ホーム>Other>Connecting Claude Code and Codex: Codex as the "Final Reviewer" Works, But... | MCP, Plugin, and codex exec Compared Hands-On
Other

Connecting Claude Code and Codex: Codex as the "Final Reviewer" Works, But... | MCP, Plugin, and codex exec Compared Hands-On

Thank you for your continued support.
This article contains advertisements that help fund our operations.

I compared three ways to call Codex from Claude Code (MCP, the official plugin, and codex exec/review) on the same diff, and built the same spec twice: once with Codex in the middle of the workflow, and once with Codex only doing the final review. I also cover how codex mcp-server disappeared in Codex CLI 0.154.0, and a thumbnail-generation comparison.

If you subscribe to both Claude Code and Codex CLI, you end up wondering two things: how to connect them, and where to use Codex.

Should you split the implementation between them, or just ask Codex for reviews?

In September 2026, this blog put Codex into the middle of its article-writing workflow, then took it out the very next day.

To check whether putting Codex in the middle of the workflow was really a bad fit, I built the same spec under two setups, and also compared three ways of calling Codex from Claude Code on the same diff.

How to call Codex from Claude Code

Online, three ways of calling Codex from Claude Code are commonly introduced: MCP, the official plugin, and calling the Codex CLI directly.

I tested them with my local Codex CLI 0.154.0 (released September 10, 2026), and here is what worked.

For MCP only, I also checked the latest version at the time of writing, 0.159.3.

MethodWorked?Notes
Call codex exec from BashYesModel and effort can be pinned with arguments. This is what this article uses
Call codex review from BashYesThere is no -m, so set the model with -c
Official plugin codex-plugin-ccYes/codex:review has no way to set effort
MCP (register codex mcp-server)No0.154.0 and 0.159.3 have no mcp-server. It connects if you specify 0.153.4
Write mcpServers in .claude/settings.jsonNoClaude Code does not read it
npx @openai/codex-mcpNoThe package does not exist on npm

Note: effort is a setting that controls how much the AI thinks before it answers.

What I adopted: calling codex exec in read-only mode

For the reviews in this article, I asked Codex like this.

codex exec -m gpt-6-astra -c model_reasoning_effort=medium -s read-only "<review request>"

I chose it for three reasons.

  • The model and reasoning effort can be pinned on the command line
  • The request can tell Codex to "check against SPEC.md"
  • With -s read-only, Codex never gets to change files

The official plugin cannot set effort for /codex:review

I installed the official plugin after adding its marketplace.

$ claude plugin marketplace add openai/codex-plugin-cc
✔ Successfully added marketplace: openai-codex (declared in user settings)
$ claude plugin install codex@openai-codex --scope project
✔ Successfully installed plugin: codex@openai-codex (scope: project)

The marketplace goes into user-level settings, so I limited the plugin itself to the test directory with --scope project.

These are the 8 available commands.

CommandWhat it does
/codex:reviewHas Codex review your local git diff
/codex:adversarial-reviewHas Codex review the implementation approach and design decisions from a deliberately critical stance
/codex:rescueHands an investigation or fix request to Codex
/codex:transferMoves the current Claude Code session into a Codex thread
/codex:statusLists running and recent Codex jobs
/codex:resultShows the final output of a finished job
/codex:cancelStops a job running in the background
/codex:setupChecks whether the Codex CLI is ready to use. Also toggles the review gate

The problem was that /codex:review has no option for setting effort.

--model gpt-6-astra was accepted, but effort simply used Codex's global setting (high in my environment).

The Codex log (~/.codex/logs_2.sqlite) also recorded model=gpt-6-astra codex.turn.reasoning_effort=high (^^;

MCP connects if you specify 0.153.4

The most widely introduced method, registering MCP, did not connect on 0.154.0 (´・ω・`)

$ claude mcp add codex -- codex mcp-server
Added stdio MCP server codex with command: codex mcp-server to local config
$ claude mcp list
codex: codex mcp-server - ✘ Failed to connect — CONNECTION_CLOSED: Connection closed

mcp-server was not in the subcommand list of codex --help, and the same was true for 0.159.3.

The previous version, 0.153.4, printed a deprecation warning but worked!

$ claude mcp add codex -- npx -y @openai/[email protected] mcp-server
$ claude mcp list
codex: npx -y @openai/[email protected] mcp-server - ✔ Connected
warning: `codex mcp-server` is deprecated and will be removed in a future release.

It exposes two tools, codex and codex-reply, and passing {"model_reasoning_effort": "medium"} to the config argument of the codex tool pinned the effort.

Running once with -s workspace-write marks that directory as trusted

After the work, I compared ~/.codex/config.toml with a backup taken beforehand and found a setting I had never added.

> [projects."/Volumes/SSD-Kioxia-2TB/Projects/laratech/claude-code-codex-demo"]
> trust_level = "trusted"

The file's modification time matched the time I ran codex exec -s workspace-write.

When I deleted those lines and ran -s workspace-write again, the same setting was added again.

Once a directory is trusted, the sandbox of codex exec without -s changes.

State of the test directorySandbox without -s
Not trustedread-only
Trustedworkspace-write [workdir, /tmp, $TMPDIR]

If you only want reviews, it is safer to always pass -s read-only explicitly.

Also, a .codex/config.toml placed in the project was only read in trusted directories.

How this blog put Codex into the middle of the workflow and removed it the next day

On September 10, 2026, I set up a workflow with Claude Code as the director and GPT-6 Astra (via Codex CLI) as the worker.

The division of roles was that Codex edited files, while commits and pushes were always done on the Claude side.

At first, Astra wrote the first drafts of articles, and Claude checked them against the style rules and sent them back.

However, the overhead of checking and sending drafts back kept growing.

So I moved first drafts back to Claude and narrowed Astra's job to outlines and polishing, but in the end I removed Astra from the article pipeline that same night.

The reason left in the commit message is "GPT-6 malfunctioning."

I don't know whether it was a connection problem on the GPT side or something caused by processing heavy prompts, but since the quality of the generated text was not very different, I stopped the integration at that point.

Building the same spec under two setups

The language was TypeScript (Node.js), and the test framework was Vitest.

What I built was the decision logic for "scheduled publishing," which automatically publishes a blog post at a specified date and time.

No UI or database, just the following four functions as the spec.

FunctionDescription
isPublished(post, now)Published if the publish time is at or before now. The exact same time counts as published
publishedPosts(posts, now)Published posts in descending order. Ties are ordered by id. The input array is not modified
validateSchedule(input, now)Validates the scheduled time input. Treats missing time zones, February 30, past times, and so on as errors
formatPublishedAt(post, tz)Formats as YYYY-MM-DD HH:mm in the given time zone. Defaults to Asia/Tokyo

The development and test environment was as follows.

ItemVersion / settings
OSmacOS (Apple Silicon)
Claude Code2.1.283, --model claude-opus-5-5 --effort medium
Codex CLI0.154.0, -m gpt-6-astra -c model_reasoning_effort=medium
Node.jsv22.15.0 (TypeScript 5.9.3, Vitest 3.2.7)

I built this module under the following two setups.

  • Setup A: Codex in the middle of the workflow (Claude Code designs, Codex implements)
  • Setup B: Codex only does the final review (Claude Code carries the implementation through to completion)

The test went in this order.

  1. Write the spec (SPEC.md)
  2. Write 22 acceptance tests and put them where neither agent can see them
  3. Build with Setup A (Claude Code writes a design doc → Codex implements → Claude Code reviews and sends it back → Codex fixes → Claude Code approves)
  4. Build with Setup B (Claude Code writes the implementation and tests → Codex reviews in read-only mode → Claude Code applies the findings)
  5. Run the acceptance tests on both results and compare time and pass counts
  6. Also show Setup B's diff to the other routes (MCP, plugin, codex review) and to a Claude Code review subagent, and compare the findings

In Setups A and B, "Claude Code" means a separate claude -p session started in the test directory.

As a side note, when there are multiple CLAUDE.md files, you can exclude one from loading by passing claudeMdExcludes with the --settings option.

claude -p "<instructions>" --model claude-opus-5-5 --effort medium \
  --settings '{"claudeMdExcludes":["/path/to/laratech/CLAUDE.md"]}'

The final review request sent to Codex

This is the full review request sent to Codex in Setup B. It was written in Japanese, so here is an English translation.

Please code-review the diff between the main branch of this repository and the current branch (`git diff main...HEAD`).
Using SPEC.md as the spec, look for bugs, spec violations, and missing tests.
Give each finding a severity (P0–P3), the file and line, and the reason. If there are no findings, write "No findings."
Do not modify any files.

The same request was passed when running through MCP.

For codex review and the plugin, I did not give a request and used the built-in review feature as is.

Setup B took about a third of the time of Setup A, with identical acceptance test results

The five steps of Setup A.

StepOwnerTimeResult
Design doc DESIGN.md (352 lines)Claude Code135 s
ImplementationCodex (workspace-write)193 s100 tests
Review, round 1Claude Code81 sSent back
Addressing the send-backCodex45 s102 tests
Review, round 2Claude Code54 sApproved

It was sent back because "a date in year 0000 is displayed as 0001."

Nothing stopped midway, and I never had to fix code by hand or give additional instructions.

The three steps of Setup B.

StepOwnerTimeResult
ImplementationClaude Code78 s62 tests
Final reviewCodex (read-only)53 s1 finding (P3)
Applying the findingClaude Code33 s1 adopted, 63 tests

Codex's finding was the same year 0000 issue that Setup A was sent back for. Here is an English translation of Codex's (Japanese) output.

[P3] A publish date in year `0000` is displayed as `0001` — src/schedule.ts:116
Passing `publishedAt: '0000-01-01T00:00:00Z'` and `tz: 'UTC'` to `formatPublishedAt`
returns `0001-01-01 00:00` instead of the expected `0000-01-01 00:00`.
This is because it uses the BCE year returned by `Intl.DateTimeFormat` without era information.

Here are the results of running the hidden acceptance tests, along with the overall time.

ItemSetup ASetup B
Acceptance tests (right after first implementation)22 / 2222 / 22
Acceptance tests (final version)22 / 2222 / 22
Total agent working time508 s164 s
Codex calls21

I included edge cases such as February 30, 25 o'clock, and inputs without a time zone, but both setups passed everything from the very first test run!

The results were the same when I changed the runtime time zone to America/Los_Angeles.

With equivalent output quality, Setup B, which keeps Codex out of the middle, finished in about a third of the time.

The most serious bug was missed by Codex on every route

I gave the same request about the same Setup B diff to a Claude Code review subagent (a reviewer defined with --agents).

The subagent returned 4 findings in 52 seconds.

SeverityFinding
P2isPublished and formatPublishedAt rely on Date.parse, so values like "1" and "2026-02-30T00:00:00Z" are treated as valid dates
P3Passing an invalid time zone name to formatPublishedAt throws RangeError
P3There are no tests for loose string parsing or values without a time zone
P3Some boundary tests for validateSchedule (such as +23:59 and full-width spaces) are missing

The most severe finding, the P2, was not detected by Codex on any route.

RouteCodexeffortTimeFindingsFiles changed
codex exec -s read-only0.154.0medium53 s1 (P3)None
codex review (run 1)0.154.0medium34 s0None
codex review (run 2)0.154.0medium34 s0None
Plugin /codex:review0.154.0high54 s0None
MCP0.153.4medium63 s1 (P3)None

On the other hand, the year 0000 issue was detected only by Codex, not by the Claude subagent.

There was also a difference in whether tests could be run.

Inside the read-only sandbox, Vitest cannot create temporary files, so on three routes, codex exec, codex review, and MCP, an EPERM error occurred and the tests could not run.

Only the Codex running through the plugin came up with the workaround npm test -- --configLoader runner --no-cache --pool threads by itself, ran it, and passed 62 tests.

Even with the same Codex, recovery behavior differed depending on the route!

Only the plugin route used high effort, so it may have reached the workaround thanks to deeper reasoning.

Note that the plugin and codex review both just pass "the diff against main" to Codex's built-in review feature, so the instructions were the same.

However, the plugin route was only tested once, so I can't be sure the difference in effort was what made the difference.

The difference in input strictness came from Claude's design doc

I ran the values from the P2 finding against the final versions from both setups.

InputSetup A isPublishedSetup B isPublishedSetup B formatted output
"1"falsetrue2001-01-01 00:00
"2026-02-30T00:00:00Z"falsetrue2026-03-02 09:00
"2026-10-01T09:00"falsetrue2026-10-01 09:00
"2026/09/01 09:00"falsetrue2026-09-01 09:00

Setup A was strict because Claude Code wrote "do not rely on Date.parse" into the design doc, and I don't think it was a direct effect of putting Codex into the process.

Who should polish Japanese text

As a side test, I asked Codex to polish this blog article.

I ran a review on a draft of this article (excluding this section) with codex exec -s read-only, attaching the blog's style rules.

AspectFindingsAdopted
Unnatural or hard-to-read Japanese66
Style rule violations44
Inconsistent numbers or statements33

All 13 findings were valid, and some were the kind my local checking script cannot catch.

However, this review took 2263 seconds (about 38 minutes), and I haven't figured out why it took so much longer than the code review (53 seconds).

Here are the actual findings. The original sentences were in Japanese, so they are translated.

Unnatural or hard-to-read Japanese (6)

Original (gist)Finding
"To verify that regret, ... built two versions""Verify that regret" makes it unclear what is being tested
"Whether this history was the setup's fault or just bad luck"Connecting "the history" to "the setup's fault" is unnatural
"As an internal link, ... is also covered in the Laravel Boost article""As an internal link" is the writer's perspective and unnatural as guidance for readers
"The method of writing in .claude/settings.json ... did not appear in the list"What appears in the list is the server, not the "method," so subject and predicate don't match
"In terms of naturally polishing Japanese text, Gemini 3.8 Flash seems the most natural""Natural" is repeated, and "I feel it seems" is redundant
"Gemini had prepared polishing instructions and scripts until September 11"Misleading subject that reads as if Gemini prepared them itself

Style rule violations (4)

Original (gist)Finding
"From 0.154.0 onward, mcp-server itself was gone!"Only 0.154.0 and 0.159.3 were checked, so "onward" suggests every version was checked
"For the plugin, add the marketplace and then install"It describes steps actually taken, so it should be in the past tense ("installed")
"Made it a 1280x720 PNG with a headless browser"Mentions "headless browser," which the rules say need not be written, and uses the unnatural phrase "draw the instructions"
In a table, "would need to be redone" for Codex under "ease of editing"Editing was never actually tested, so it's an assertion beyond what was checked

Inconsistent numbers or statements (3)

Original (gist)Finding
Heading "...showing Codex the finished diff was cheaper"What was compared was working time, not cost, so "cheaper" reads as a price comparison
"Testing was done with Claude Code 2.1.283... and Codex CLI 0.154.0..."It includes the MCP test (0.153.4) and the plugin test (effort: high), so it's inaccurate as a description of the overall conditions
Summary: "the pass count didn't change, only the time was about 3x"The handling of invalid dates did differ, so "only the time" overstates it

Making the thumbnail three times under different conditions

I asked for a photorealistic thumbnail of a fictional Japanese commentator in a news studio, pointing at a monitor and explaining this article.

I changed the conditions and compared them three times.

All the people in these images are fictional, generated or drawn by AI, and have nothing to do with any real person.

Case 1: Specifying a different method for each tool

In the first round, I asked Codex to use image generation and Claude Code to draw with HTML/CSS, each with separate instructions.

ItemInstructions to CodexInstructions to Claude Code
MethodUse image generationBuild with HTML/CSS and output a PNG with the given command
Visual specPhotographic, photorealistic look requiredAim for photorealistic, but if that's hard, whatever HTML/CSS can express is fine

Codex, Case 1

Codex generated one image with its built-in image_gen tool.

Thumbnail generated by Codex in Case 1. An AI-generated fictional person points at a "final review" diagram from Claude Code to Codex

The text on the monitor, "最終レビュー" (final review), and "連携" (integration) in the headline were also rendered without any garbling!

Claude Code, Case 1

Claude Code's output was an illustration, with the person drawn in SVG.

Thumbnail drawn by Claude Code in HTML/CSS in Case 1. An illustrated fictional commentator points at a monitor

Case 1 comparison

AspectCodexClaude Code
RealismClose to a real news broadcastIllustration style
Time105 s (1 generation)119 s (2 renders)

However, only Claude Code was told that an illustration was an acceptable fallback, so this was not a strictly equal comparison.

Case 2: Same prompt, method left up to each tool

In the second round, I didn't specify a method and gave both the same prompt. It was written in Japanese, so here is an English translation.

Please make one thumbnail image for a tech blog article.

Article content: A hands-on article arguing that "if you connect Claude Code and Codex CLI, it makes sense to give Codex the final review rather than a middle step of the work."

Image requirements:
- A 1280x720 PNG.
- A photographic, photorealistic look. In a news studio, a Japanese commentator (a fictional person; do not make them resemble any real person) points with their hand at a large monitor beside them while explaining this article.
- On the monitor, a simple diagram with two boxes labeled "Claude Code" on the left and "Codex" on the right, an arrow between them, and the text "最終レビュー" (final review) under the arrow.
- At the top or bottom of the image, a large Japanese headline "Claude Code×Codex 連携".
- No Anthropic or OpenAI logos, and no other product logos or brand marks.

Save the finished image as ./thumbnail/thumbnail.png.
Do not git commit or git push.
Finally, report how you made the image and which files you created.

Codex, Case 2

As in Case 1, Codex used image generation and adjusted the size to 1280x720 with the sips command (time: 92 s).

Thumbnail generated by Codex in Case 2. An AI-generated fictional male commentator points at the diagram on the monitor

Claude Code, Case 2

When Claude Code detected that the Codex CLI was available in the environment, it called codex exec on its own and handed off generating the photo part to Codex's image generation.

It then drew only the headline text on top, using Python's Pillow library with the Hiragino Kaku Gothic font (time: 136 s).

Thumbnail made by Claude Code in Case 2. The photo was generated by Codex's image generation, and the headline at the top was drawn with Pillow

In its report, Claude Code explained that it drew the text separately because "leaving text output directly to an image generation tool tends to garble the font."

Without any explicit instruction, Claude Code ended up setting up an integration with Codex by itself.

Note that the codex exec run by Claude Code had no model or effort parameters, so it ran with Codex's global setting (effort: high).

Case 3: Same prompt, photorealism with HTML/CSS only

In the third round, I changed the Case 2 prompt as follows and gave both the same text.

  • Build with HTML/CSS and output a PNG with the given command
  • Aim for a photographic, photorealistic look, with no fallback clause
  • No external image files, web fonts, or AI image generation (including Codex)
  • Check the converted PNG and fix it yourself if anything is broken

Claude Code, Case 3

Claude Code ran the conversion twice, noticing by itself that "Claude Code" wrapped onto two lines and that the shoulders were drawn too wide in the first output, and fixed them (time: 173 s).

Thumbnail drawn by Claude Code in HTML/CSS in Case 3. A 3DCG-style illustrated commentator

Codex, Case 3

Codex finished generating the HTML, but the conversion to PNG failed twice in a row (time: 254 s).

The cause was that the rendering browser (Chrome) could not be launched from inside Codex's sandbox.

Without changing the HTML Codex generated, I ran the same conversion command manually outside the sandbox to produce the PNG.

Thumbnail drawn by Codex in HTML/CSS in Case 3. A heavily shaded, 3DCG-style commentator

Case 3 comparison

Neither reached photographic quality, and their own reports said things like "it did not reach a photorealistic level" and "I gave up on realistic textures for hair, skin, and hands."

Codex couldn't look at the rendered image, so its conditions differed from Claude Code, which went through a process of checking the output and fixing it.

Also, both tools added text that wasn't in the instructions, such as a "検証" (verification) label and "TECH REPORT."

Summary of the three rounds

CaseClaude CodeCodex
1: Method specified separatelyIllustrationPhotorealistic (image generation)
2: Same prompt, any methodCalled Codex's image generation internally to make a photorealistic imagePhotorealistic (image generation)
3: Same prompt, HTML/CSS only3DCG-style illustration3DCG-style illustration

To get photorealistic images, a dedicated image generation feature is essential.

Under an HTML/CSS-only constraint, neither tool reached photorealism.

On the other hand, when you need readable, accurate text, the combination Claude Code used in Case 2, "a background photo from image generation with text drawn by code on top," works extremely well.

The thumbnail for this article is the image Codex generated in Case 1.

How to use Claude Code and Codex together (conclusion)

Things will keep changing, but for now I think handing middle steps of the work to Codex or similar tools is questionable.

The reason is that Claude Code, acting as an agent, hands out the work, but just like with people, it turns into a game of telephone and isn't very accurate.

(The same goes for sending work to subagents.)

If you write the handoff instructions precisely yourself, accuracy may improve, so I'll count that as a gap in my testing, and apologies if someone out there is already running this perfectly.

Put Codex in the "final reviewer" role

To sum up the results of this test:

  • Acceptance test pass rates were equal, and Setup A, with Codex in the middle of the workflow, took about 3.1 times as long as Setup B
  • As the final reviewer, Codex finished in 34–63 seconds per run, changed no files, and found 1 issue from a different angle than Claude
  • The most serious bug was found only by the Claude subagent

In conclusion: "Putting Codex in the final reviewer role is useful, but it doesn't replace Claude's own review."

That said, it's nice to have multiple reviewers giving you a list, and I don't see many downsides, so I think asking Codex for reviews makes sense.

However, I think effort should be set to high or above.

With medium, the review seems way too fast.

It feels like it's so fast that things slip through, as if it just went through the motions.

HTML images from Claude Code, photorealistic images from Codex

This is an obvious result, but if you want photorealistic images, Codex is the only choice.

However, tasks that need photorealistic images tend to be writing tasks, which I think is not really a skill engineers need for coding, and if it's about drawing architecture for documentation, Claude Code produces plenty good enough results.

For people like me who want to publish research results on a blog, having Codex make images might be fun, but I see my writing as the not-so-flashy kind (seriously examining how to use things), so slide-style images from Claude Code are good enough for me.

When Claude Code's tokens aren't enough and you want to use Codex too

Basically, doing everything with Claude Code alone is simpler, and I think the best approach is to switch models in subagents to vary their strength, but some people may run a bit short on tokens and want to split work off to Codex.

In that case, I think it's better to give Codex work that can be completely separated, or a quality-improvement role after the work is done (like the review-only role in this test).

Well, there are probably already countless people running it that way.

On having AIs work with each other 24/7

At this point, I don't think that's useful.

Having them discuss and organize things and then present it to a human is fine.

Letting AIs implement things based on their own discussions seems to have many downsides as a realistic way of working.

If they go that far on their own, the human review becomes way too much work.

I can't help thinking that the mysterious self-proclaimed AI consultants on social media who have AIs working 24/7 are probably just churning out thin, poetic posts to make pocket change on note (a Japanese blogging platform).

People are less buried in implementation now, but these days I still feel there's a need for humans to step in.

Thank you for reading to the end.

Please Provide Feedback
We would appreciate your feedback on this article. Feel free to leave a comment on any relevant YouTube video or reach out through the contact form. Thank you!