Connecting Claude Code and Codex: Codex as the "Final Reviewer" Works, But... | MCP, Plugin, and codex exec Compared Hands-On
Table Of Contents
- How to call Codex from Claude Code
- How this blog put Codex into the middle of the workflow and removed it the next day
- Building the same spec under two setups
- Setup B took about a third of the time of Setup A, with identical acceptance test results
- The most serious bug was missed by Codex on every route
- Who should polish Japanese text
- Making the thumbnail three times under different conditions
- How to use Claude Code and Codex together (conclusion)
I compared three ways to call Codex from Claude Code (MCP, the official plugin, and codex exec/review) on the same diff, and built the same spec twice: once with Codex in the middle of the workflow, and once with Codex only doing the final review. I also cover how codex mcp-server disappeared in Codex CLI 0.154.0, and a thumbnail-generation comparison.
If you subscribe to both Claude Code and Codex CLI, you end up wondering two things: how to connect them, and where to use Codex.
Should you split the implementation between them, or just ask Codex for reviews?
In September 2026, this blog put Codex into the middle of its article-writing workflow, then took it out the very next day.
To check whether putting Codex in the middle of the workflow was really a bad fit, I built the same spec under two setups, and also compared three ways of calling Codex from Claude Code on the same diff.
How to call Codex from Claude Code
Online, three ways of calling Codex from Claude Code are commonly introduced: MCP, the official plugin, and calling the Codex CLI directly.
I tested them with my local Codex CLI 0.154.0 (released September 10, 2026), and here is what worked.
For MCP only, I also checked the latest version at the time of writing, 0.159.3.
| Method | Worked? | Notes |
|---|---|---|
Call codex exec from Bash | Yes | Model and effort can be pinned with arguments. This is what this article uses |
Call codex review from Bash | Yes | There is no -m, so set the model with -c |
| Official plugin codex-plugin-cc | Yes | /codex:review has no way to set effort |
MCP (register codex mcp-server) | No | 0.154.0 and 0.159.3 have no mcp-server. It connects if you specify 0.153.4 |
Write mcpServers in .claude/settings.json | No | Claude Code does not read it |
npx @openai/codex-mcp | No | The package does not exist on npm |
Note: effort is a setting that controls how much the AI thinks before it answers.
What I adopted: calling codex exec in read-only mode
For the reviews in this article, I asked Codex like this.
codex exec -m gpt-6-astra -c model_reasoning_effort=medium -s read-only "<review request>"I chose it for three reasons.
- The model and reasoning effort can be pinned on the command line
- The request can tell Codex to "check against SPEC.md"
- With
-s read-only, Codex never gets to change files
The official plugin cannot set effort for /codex:review
I installed the official plugin after adding its marketplace.
$ claude plugin marketplace add openai/codex-plugin-cc
✔ Successfully added marketplace: openai-codex (declared in user settings)
$ claude plugin install codex@openai-codex --scope project
✔ Successfully installed plugin: codex@openai-codex (scope: project)The marketplace goes into user-level settings, so I limited the plugin itself to the test directory with --scope project.
These are the 8 available commands.
| Command | What it does |
|---|---|
/codex:review | Has Codex review your local git diff |
/codex:adversarial-review | Has Codex review the implementation approach and design decisions from a deliberately critical stance |
/codex:rescue | Hands an investigation or fix request to Codex |
/codex:transfer | Moves the current Claude Code session into a Codex thread |
/codex:status | Lists running and recent Codex jobs |
/codex:result | Shows the final output of a finished job |
/codex:cancel | Stops a job running in the background |
/codex:setup | Checks whether the Codex CLI is ready to use. Also toggles the review gate |
The problem was that /codex:review has no option for setting effort.
--model gpt-6-astra was accepted, but effort simply used Codex's global setting (high in my environment).
The Codex log (~/.codex/logs_2.sqlite) also recorded model=gpt-6-astra codex.turn.reasoning_effort=high (^^;
MCP connects if you specify 0.153.4
The most widely introduced method, registering MCP, did not connect on 0.154.0 (´・ω・`)
$ claude mcp add codex -- codex mcp-server
Added stdio MCP server codex with command: codex mcp-server to local config
$ claude mcp list
codex: codex mcp-server - ✘ Failed to connect — CONNECTION_CLOSED: Connection closedmcp-server was not in the subcommand list of codex --help, and the same was true for 0.159.3.
The previous version, 0.153.4, printed a deprecation warning but worked!
$ claude mcp add codex -- npx -y @openai/[email protected] mcp-server
$ claude mcp list
codex: npx -y @openai/[email protected] mcp-server - ✔ Connectedwarning: `codex mcp-server` is deprecated and will be removed in a future release.It exposes two tools, codex and codex-reply, and passing {"model_reasoning_effort": "medium"} to the config argument of the codex tool pinned the effort.
Running once with -s workspace-write marks that directory as trusted
After the work, I compared ~/.codex/config.toml with a backup taken beforehand and found a setting I had never added.
> [projects."/Volumes/SSD-Kioxia-2TB/Projects/laratech/claude-code-codex-demo"]
> trust_level = "trusted"The file's modification time matched the time I ran codex exec -s workspace-write.
When I deleted those lines and ran -s workspace-write again, the same setting was added again.
Once a directory is trusted, the sandbox of codex exec without -s changes.
| State of the test directory | Sandbox without -s |
|---|---|
| Not trusted | read-only |
| Trusted | workspace-write [workdir, /tmp, $TMPDIR] |
If you only want reviews, it is safer to always pass -s read-only explicitly.
Also, a .codex/config.toml placed in the project was only read in trusted directories.
How this blog put Codex into the middle of the workflow and removed it the next day
On September 10, 2026, I set up a workflow with Claude Code as the director and GPT-6 Astra (via Codex CLI) as the worker.
The division of roles was that Codex edited files, while commits and pushes were always done on the Claude side.
At first, Astra wrote the first drafts of articles, and Claude checked them against the style rules and sent them back.
However, the overhead of checking and sending drafts back kept growing.
So I moved first drafts back to Claude and narrowed Astra's job to outlines and polishing, but in the end I removed Astra from the article pipeline that same night.
The reason left in the commit message is "GPT-6 malfunctioning."
I don't know whether it was a connection problem on the GPT side or something caused by processing heavy prompts, but since the quality of the generated text was not very different, I stopped the integration at that point.
Building the same spec under two setups
The language was TypeScript (Node.js), and the test framework was Vitest.
What I built was the decision logic for "scheduled publishing," which automatically publishes a blog post at a specified date and time.
No UI or database, just the following four functions as the spec.
| Function | Description |
|---|---|
isPublished(post, now) | Published if the publish time is at or before now. The exact same time counts as published |
publishedPosts(posts, now) | Published posts in descending order. Ties are ordered by id. The input array is not modified |
validateSchedule(input, now) | Validates the scheduled time input. Treats missing time zones, February 30, past times, and so on as errors |
formatPublishedAt(post, tz) | Formats as YYYY-MM-DD HH:mm in the given time zone. Defaults to Asia/Tokyo |
The development and test environment was as follows.
| Item | Version / settings |
|---|---|
| OS | macOS (Apple Silicon) |
| Claude Code | 2.1.283, --model claude-opus-5-5 --effort medium |
| Codex CLI | 0.154.0, -m gpt-6-astra -c model_reasoning_effort=medium |
| Node.js | v22.15.0 (TypeScript 5.9.3, Vitest 3.2.7) |
I built this module under the following two setups.
- Setup A: Codex in the middle of the workflow (Claude Code designs, Codex implements)
- Setup B: Codex only does the final review (Claude Code carries the implementation through to completion)
The test went in this order.
- Write the spec (SPEC.md)
- Write 22 acceptance tests and put them where neither agent can see them
- Build with Setup A (Claude Code writes a design doc → Codex implements → Claude Code reviews and sends it back → Codex fixes → Claude Code approves)
- Build with Setup B (Claude Code writes the implementation and tests → Codex reviews in read-only mode → Claude Code applies the findings)
- Run the acceptance tests on both results and compare time and pass counts
- Also show Setup B's diff to the other routes (MCP, plugin,
codex review) and to a Claude Code review subagent, and compare the findings
In Setups A and B, "Claude Code" means a separate claude -p session started in the test directory.
As a side note, when there are multiple CLAUDE.md files, you can exclude one from loading by passing claudeMdExcludes with the --settings option.
claude -p "<instructions>" --model claude-opus-5-5 --effort medium \
--settings '{"claudeMdExcludes":["/path/to/laratech/CLAUDE.md"]}'The final review request sent to Codex
This is the full review request sent to Codex in Setup B. It was written in Japanese, so here is an English translation.
Please code-review the diff between the main branch of this repository and the current branch (`git diff main...HEAD`).
Using SPEC.md as the spec, look for bugs, spec violations, and missing tests.
Give each finding a severity (P0–P3), the file and line, and the reason. If there are no findings, write "No findings."
Do not modify any files.The same request was passed when running through MCP.
For codex review and the plugin, I did not give a request and used the built-in review feature as is.
Setup B took about a third of the time of Setup A, with identical acceptance test results
The five steps of Setup A.
| Step | Owner | Time | Result |
|---|---|---|---|
| Design doc DESIGN.md (352 lines) | Claude Code | 135 s | |
| Implementation | Codex (workspace-write) | 193 s | 100 tests |
| Review, round 1 | Claude Code | 81 s | Sent back |
| Addressing the send-back | Codex | 45 s | 102 tests |
| Review, round 2 | Claude Code | 54 s | Approved |
It was sent back because "a date in year 0000 is displayed as 0001."
Nothing stopped midway, and I never had to fix code by hand or give additional instructions.
The three steps of Setup B.
| Step | Owner | Time | Result |
|---|---|---|---|
| Implementation | Claude Code | 78 s | 62 tests |
| Final review | Codex (read-only) | 53 s | 1 finding (P3) |
| Applying the finding | Claude Code | 33 s | 1 adopted, 63 tests |
Codex's finding was the same year 0000 issue that Setup A was sent back for.
Here is an English translation of Codex's (Japanese) output.
[P3] A publish date in year `0000` is displayed as `0001` — src/schedule.ts:116
Passing `publishedAt: '0000-01-01T00:00:00Z'` and `tz: 'UTC'` to `formatPublishedAt`
returns `0001-01-01 00:00` instead of the expected `0000-01-01 00:00`.
This is because it uses the BCE year returned by `Intl.DateTimeFormat` without era information.Here are the results of running the hidden acceptance tests, along with the overall time.
| Item | Setup A | Setup B |
|---|---|---|
| Acceptance tests (right after first implementation) | 22 / 22 | 22 / 22 |
| Acceptance tests (final version) | 22 / 22 | 22 / 22 |
| Total agent working time | 508 s | 164 s |
| Codex calls | 2 | 1 |
I included edge cases such as February 30, 25 o'clock, and inputs without a time zone, but both setups passed everything from the very first test run!
The results were the same when I changed the runtime time zone to America/Los_Angeles.
With equivalent output quality, Setup B, which keeps Codex out of the middle, finished in about a third of the time.
The most serious bug was missed by Codex on every route
I gave the same request about the same Setup B diff to a Claude Code review subagent (a reviewer defined with --agents).
The subagent returned 4 findings in 52 seconds.
| Severity | Finding |
|---|---|
| P2 | isPublished and formatPublishedAt rely on Date.parse, so values like "1" and "2026-02-30T00:00:00Z" are treated as valid dates |
| P3 | Passing an invalid time zone name to formatPublishedAt throws RangeError |
| P3 | There are no tests for loose string parsing or values without a time zone |
| P3 | Some boundary tests for validateSchedule (such as +23:59 and full-width spaces) are missing |
The most severe finding, the P2, was not detected by Codex on any route.
| Route | Codex | effort | Time | Findings | Files changed |
|---|---|---|---|---|---|
codex exec -s read-only | 0.154.0 | medium | 53 s | 1 (P3) | None |
codex review (run 1) | 0.154.0 | medium | 34 s | 0 | None |
codex review (run 2) | 0.154.0 | medium | 34 s | 0 | None |
Plugin /codex:review | 0.154.0 | high | 54 s | 0 | None |
| MCP | 0.153.4 | medium | 63 s | 1 (P3) | None |
On the other hand, the year 0000 issue was detected only by Codex, not by the Claude subagent.
There was also a difference in whether tests could be run.
Inside the read-only sandbox, Vitest cannot create temporary files, so on three routes, codex exec, codex review, and MCP, an EPERM error occurred and the tests could not run.
Only the Codex running through the plugin came up with the workaround npm test -- --configLoader runner --no-cache --pool threads by itself, ran it, and passed 62 tests.
Even with the same Codex, recovery behavior differed depending on the route!
Only the plugin route used high effort, so it may have reached the workaround thanks to deeper reasoning.
Note that the plugin and codex review both just pass "the diff against main" to Codex's built-in review feature, so the instructions were the same.
However, the plugin route was only tested once, so I can't be sure the difference in effort was what made the difference.
The difference in input strictness came from Claude's design doc
I ran the values from the P2 finding against the final versions from both setups.
| Input | Setup A isPublished | Setup B isPublished | Setup B formatted output |
|---|---|---|---|
"1" | false | true | 2001-01-01 00:00 |
"2026-02-30T00:00:00Z" | false | true | 2026-03-02 09:00 |
"2026-10-01T09:00" | false | true | 2026-10-01 09:00 |
"2026/09/01 09:00" | false | true | 2026-09-01 09:00 |
Setup A was strict because Claude Code wrote "do not rely on Date.parse" into the design doc, and I don't think it was a direct effect of putting Codex into the process.
Who should polish Japanese text
As a side test, I asked Codex to polish this blog article.
I ran a review on a draft of this article (excluding this section) with codex exec -s read-only, attaching the blog's style rules.
| Aspect | Findings | Adopted |
|---|---|---|
| Unnatural or hard-to-read Japanese | 6 | 6 |
| Style rule violations | 4 | 4 |
| Inconsistent numbers or statements | 3 | 3 |
All 13 findings were valid, and some were the kind my local checking script cannot catch.
However, this review took 2263 seconds (about 38 minutes), and I haven't figured out why it took so much longer than the code review (53 seconds).
Here are the actual findings. The original sentences were in Japanese, so they are translated.
Unnatural or hard-to-read Japanese (6)
| Original (gist) | Finding |
|---|---|
| "To verify that regret, ... built two versions" | "Verify that regret" makes it unclear what is being tested |
| "Whether this history was the setup's fault or just bad luck" | Connecting "the history" to "the setup's fault" is unnatural |
| "As an internal link, ... is also covered in the Laravel Boost article" | "As an internal link" is the writer's perspective and unnatural as guidance for readers |
"The method of writing in .claude/settings.json ... did not appear in the list" | What appears in the list is the server, not the "method," so subject and predicate don't match |
| "In terms of naturally polishing Japanese text, Gemini 3.8 Flash seems the most natural" | "Natural" is repeated, and "I feel it seems" is redundant |
| "Gemini had prepared polishing instructions and scripts until September 11" | Misleading subject that reads as if Gemini prepared them itself |
Style rule violations (4)
| Original (gist) | Finding |
|---|---|
"From 0.154.0 onward, mcp-server itself was gone!" | Only 0.154.0 and 0.159.3 were checked, so "onward" suggests every version was checked |
| "For the plugin, add the marketplace and then install" | It describes steps actually taken, so it should be in the past tense ("installed") |
| "Made it a 1280x720 PNG with a headless browser" | Mentions "headless browser," which the rules say need not be written, and uses the unnatural phrase "draw the instructions" |
| In a table, "would need to be redone" for Codex under "ease of editing" | Editing was never actually tested, so it's an assertion beyond what was checked |
Inconsistent numbers or statements (3)
| Original (gist) | Finding |
|---|---|
| Heading "...showing Codex the finished diff was cheaper" | What was compared was working time, not cost, so "cheaper" reads as a price comparison |
| "Testing was done with Claude Code 2.1.283... and Codex CLI 0.154.0..." | It includes the MCP test (0.153.4) and the plugin test (effort: high), so it's inaccurate as a description of the overall conditions |
| Summary: "the pass count didn't change, only the time was about 3x" | The handling of invalid dates did differ, so "only the time" overstates it |
Making the thumbnail three times under different conditions
I asked for a photorealistic thumbnail of a fictional Japanese commentator in a news studio, pointing at a monitor and explaining this article.
I changed the conditions and compared them three times.
All the people in these images are fictional, generated or drawn by AI, and have nothing to do with any real person.
Case 1: Specifying a different method for each tool
In the first round, I asked Codex to use image generation and Claude Code to draw with HTML/CSS, each with separate instructions.
| Item | Instructions to Codex | Instructions to Claude Code |
|---|---|---|
| Method | Use image generation | Build with HTML/CSS and output a PNG with the given command |
| Visual spec | Photographic, photorealistic look required | Aim for photorealistic, but if that's hard, whatever HTML/CSS can express is fine |
Codex, Case 1
Codex generated one image with its built-in image_gen tool.
The text on the monitor, "最終レビュー" (final review), and "連携" (integration) in the headline were also rendered without any garbling!
Claude Code, Case 1
Claude Code's output was an illustration, with the person drawn in SVG.
Case 1 comparison
| Aspect | Codex | Claude Code |
|---|---|---|
| Realism | Close to a real news broadcast | Illustration style |
| Time | 105 s (1 generation) | 119 s (2 renders) |
However, only Claude Code was told that an illustration was an acceptable fallback, so this was not a strictly equal comparison.
Case 2: Same prompt, method left up to each tool
In the second round, I didn't specify a method and gave both the same prompt. It was written in Japanese, so here is an English translation.
Please make one thumbnail image for a tech blog article.
Article content: A hands-on article arguing that "if you connect Claude Code and Codex CLI, it makes sense to give Codex the final review rather than a middle step of the work."
Image requirements:
- A 1280x720 PNG.
- A photographic, photorealistic look. In a news studio, a Japanese commentator (a fictional person; do not make them resemble any real person) points with their hand at a large monitor beside them while explaining this article.
- On the monitor, a simple diagram with two boxes labeled "Claude Code" on the left and "Codex" on the right, an arrow between them, and the text "最終レビュー" (final review) under the arrow.
- At the top or bottom of the image, a large Japanese headline "Claude Code×Codex 連携".
- No Anthropic or OpenAI logos, and no other product logos or brand marks.
Save the finished image as ./thumbnail/thumbnail.png.
Do not git commit or git push.
Finally, report how you made the image and which files you created.Codex, Case 2
As in Case 1, Codex used image generation and adjusted the size to 1280x720 with the sips command (time: 92 s).
Claude Code, Case 2
When Claude Code detected that the Codex CLI was available in the environment, it called codex exec on its own and handed off generating the photo part to Codex's image generation.
It then drew only the headline text on top, using Python's Pillow library with the Hiragino Kaku Gothic font (time: 136 s).
In its report, Claude Code explained that it drew the text separately because "leaving text output directly to an image generation tool tends to garble the font."
Without any explicit instruction, Claude Code ended up setting up an integration with Codex by itself.
Note that the codex exec run by Claude Code had no model or effort parameters, so it ran with Codex's global setting (effort: high).
Case 3: Same prompt, photorealism with HTML/CSS only
In the third round, I changed the Case 2 prompt as follows and gave both the same text.
- Build with HTML/CSS and output a PNG with the given command
- Aim for a photographic, photorealistic look, with no fallback clause
- No external image files, web fonts, or AI image generation (including Codex)
- Check the converted PNG and fix it yourself if anything is broken
Claude Code, Case 3
Claude Code ran the conversion twice, noticing by itself that "Claude Code" wrapped onto two lines and that the shoulders were drawn too wide in the first output, and fixed them (time: 173 s).
Codex, Case 3
Codex finished generating the HTML, but the conversion to PNG failed twice in a row (time: 254 s).
The cause was that the rendering browser (Chrome) could not be launched from inside Codex's sandbox.
Without changing the HTML Codex generated, I ran the same conversion command manually outside the sandbox to produce the PNG.
Case 3 comparison
Neither reached photographic quality, and their own reports said things like "it did not reach a photorealistic level" and "I gave up on realistic textures for hair, skin, and hands."
Codex couldn't look at the rendered image, so its conditions differed from Claude Code, which went through a process of checking the output and fixing it.
Also, both tools added text that wasn't in the instructions, such as a "検証" (verification) label and "TECH REPORT."
Summary of the three rounds
| Case | Claude Code | Codex |
|---|---|---|
| 1: Method specified separately | Illustration | Photorealistic (image generation) |
| 2: Same prompt, any method | Called Codex's image generation internally to make a photorealistic image | Photorealistic (image generation) |
| 3: Same prompt, HTML/CSS only | 3DCG-style illustration | 3DCG-style illustration |
To get photorealistic images, a dedicated image generation feature is essential.
Under an HTML/CSS-only constraint, neither tool reached photorealism.
On the other hand, when you need readable, accurate text, the combination Claude Code used in Case 2, "a background photo from image generation with text drawn by code on top," works extremely well.
The thumbnail for this article is the image Codex generated in Case 1.
How to use Claude Code and Codex together (conclusion)
Things will keep changing, but for now I think handing middle steps of the work to Codex or similar tools is questionable.
The reason is that Claude Code, acting as an agent, hands out the work, but just like with people, it turns into a game of telephone and isn't very accurate.
(The same goes for sending work to subagents.)
If you write the handoff instructions precisely yourself, accuracy may improve, so I'll count that as a gap in my testing, and apologies if someone out there is already running this perfectly.
Put Codex in the "final reviewer" role
To sum up the results of this test:
- Acceptance test pass rates were equal, and Setup A, with Codex in the middle of the workflow, took about 3.1 times as long as Setup B
- As the final reviewer, Codex finished in 34–63 seconds per run, changed no files, and found 1 issue from a different angle than Claude
- The most serious bug was found only by the Claude subagent
In conclusion: "Putting Codex in the final reviewer role is useful, but it doesn't replace Claude's own review."
That said, it's nice to have multiple reviewers giving you a list, and I don't see many downsides, so I think asking Codex for reviews makes sense.
However, I think effort should be set to high or above.
With medium, the review seems way too fast.
It feels like it's so fast that things slip through, as if it just went through the motions.
HTML images from Claude Code, photorealistic images from Codex
This is an obvious result, but if you want photorealistic images, Codex is the only choice.
However, tasks that need photorealistic images tend to be writing tasks, which I think is not really a skill engineers need for coding, and if it's about drawing architecture for documentation, Claude Code produces plenty good enough results.
For people like me who want to publish research results on a blog, having Codex make images might be fun, but I see my writing as the not-so-flashy kind (seriously examining how to use things), so slide-style images from Claude Code are good enough for me.
When Claude Code's tokens aren't enough and you want to use Codex too
Basically, doing everything with Claude Code alone is simpler, and I think the best approach is to switch models in subagents to vary their strength, but some people may run a bit short on tokens and want to split work off to Codex.
In that case, I think it's better to give Codex work that can be completely separated, or a quality-improvement role after the work is done (like the review-only role in this test).
Well, there are probably already countless people running it that way.
On having AIs work with each other 24/7
At this point, I don't think that's useful.
Having them discuss and organize things and then present it to a human is fine.
Letting AIs implement things based on their own discussions seems to have many downsides as a realistic way of working.
If they go that far on their own, the human review becomes way too much work.
I can't help thinking that the mysterious self-proclaimed AI consultants on social media who have AIs working 24/7 are probably just churning out thin, poetic posts to make pocket change on note (a Japanese blogging platform).
People are less buried in implementation now, but these days I still feel there's a need for humans to step in.
Thank you for reading to the end.




