
Codex Subagent
io.github.parisbsv0.4.0更新於 Oct 1, 2026
Delegate coding tasks to the OpenAI Codex CLI, installed separately, with explicit model and effort.
安裝
在 SourceWeft 中
- 開啟 儀表板中的 Codex Subagent,將其新增到工作區。
- 為需要使用其工具的對話啟用該服務。
Desktop only,透過 STDIO。 STDIO 服務會啟動本機處理程序,因此需要 SourceWeft 桌面主機。
其他 MCP 客戶端
參照 儲存庫 中的啟動說明。
README
codex-subagent-mcp
[CI] [npm version] [License: MIT]
An MCP server that lets Claude Code delegate coding tasks to OpenAI's Codex CLI running on the same machine, with the model and reasoning depth chosen per task. Claude stays the orchestrator. Codex becomes a subagent it can call.
An independent project. Not affiliated with, endorsed by, or supported by OpenAI or Anthropic.
Quick start
Requires Node.js 22+ and the Codex CLI installed,
on PATH, and signed in.
Ask Claude to run codex_doctor if anything is missing. See docs/INSTALL.md
for platform setup, Claude Desktop and other installation options, and Safety before
enabling writes.
What you get
- You pick the model and the reasoning effort per task, from the catalog your Codex CLI reports live.
- Read-only by default, with a sandbox ceiling no tool call can exceed.
- Every result states what Codex actually applied (model, effort, sandbox, directory), not only what was asked.
- Optional
output_schemareturns JSON that matches your schema, for results Claude can act on directly. - Cancelling, timing out or shutting down stops Codex and every command it started.
A real read-only delegation asking for the package name, captured 2026-09-30:
Why this exists
A single model doing everything has three recurring problems, and delegation solves each one:
Your context window is finite. Having Claude read forty files to answer one question spends context you need for the actual work. Delegating the investigation returns the answer instead of the forty files.
One model has one set of blind spots. A second opinion is worth most when it comes from a different model family — different training, different failure modes. Asking the same model twice mostly gets you the same answer twice.
Not every task deserves the same reasoning budget. Renaming a variable and diagnosing a race condition are not the same job. Here they are separate dials: the model sets raw capability, the reasoning effort sets how long it deliberates. Cheap work goes to a fast model; a hard problem gets the capable one thinking for as long as it needs.
The server runs on your machine and drives the Codex CLI you already have installed; it holds no
credentials of its own. Prompts reach OpenAI through Codex, exactly as when you run codex
yourself.
See How it compares for a versioned comparison with other Codex MCP servers.
Using it
Write a bounded delegation
A delegation gets expensive when repeated commands keep adding output to the context carried into
later requests. Name the exact question, likely files, stopping condition and evidence the answer
must contain; choose higher effort for ambiguity rather than by habit.
Writing a delegation gives the measured cost model, ranked rules and
weak-versus-strong examples using the real tool parameters. Background goes in context and a
persona or extra rules in system_instructions, both layered on the built-in quality contract; see
the tool reference.
Get a second opinion from a different model family
The value here is not a second run — it is a different set of blind spots.
Ask Codex to review
src/server.tsfor correctness problems, focusing on error paths. Use a high reasoning effort and tell it to report each finding with the line and why it matters.
A delegated review like this found the terminate() defect in this repository's own runner: two
code paths could each arm a timer while only one was ever cleared.
Investigate without spending your context
Forty files go into the delegation; one answer comes back. Codex runs its own searches and reads whatever it needs; your conversation receives the conclusion.
Have Codex trace how a reasoning effort travels from the MCP tool call down to the arguments handed to the Codex CLI, and report just the call chain.
Run long work in the background while you keep going
Kick off a Codex run in the background that writes unit tests for
src/jobs.ts, then keep helping me with the API layer.
You get a job_id immediately. Ask for the status whenever you want, and read the result when it is
done. Up to eight can run at once.
Buy deep reasoning for one hard problem
Raising the reasoning effort for the whole conversation is expensive. Raising it for one delegation is not.
This intermittent test failure has beaten me twice. Ask Codex to work out the root cause at maximum reasoning effort, give it
test/runner.test.tsand the CI log, and tell it not to change anything — I want the diagnosis first.
Keep the thread going
Ask Codex to expand on its second finding.
Follow-ups reuse Codex's context, so they cost a fraction of the original. The server restates the same model, effort and directory on every follow-up because Codex itself does not keep them on resume.
Let it write, when you mean it
Have Codex apply its first two suggestions. Let it edit files, but keep it inside a git worktree so my working tree stays clean.
use_worktree sends the run's edits to ~/.codex/worktrees/; results list the files it touched and
where each landed, up to a thousand distinct files, then report the omitted count. The server does
not clean worktrees up: they may hold unapplied work. The experimental feature is enabled only for
that invocation; your Codex configuration is unchanged. See Safety for the write
boundary.
Whatever the sandbox, a delegation that writes reports what it wrote:
Safety
This server runs another program on your machine, so it is worth two minutes before you enable writes.
What protects you
Delegations are read-only by default. Writing requires an explicit sandbox: "workspace-write"
or a user-set default. use_worktree sends edits to a managed git worktree instead of your
checkout; the sandbox and add_dirs bound where it can write at all. Unsandboxed runs are
unavailable unless you explicitly opt into that ceiling.
The confinement is the operating system's own sandbox: Seatbelt on macOS, bubblewrap on Linux and
WSL2, and a native sandbox on Windows. The table was measured on macOS;
SECURITY.md records the measurement details.
Linux and Windows have not been measured here, and OpenAI's Windows documentation notes that
sandboxed commands can fail to read some directories, so reads may be stricter there:
There is no shell in the server's invocation path: the CLI is spawned with an argv array and the prompt is written to its stdin, never interpolated into a command string. Shell metacharacters in a prompt are inert.
What does not protect you
Reads are not confined. Codex can read anything your user account can, in every mode — your SSH keys, your cloud credentials. That was measured on macOS, and it is the safe assumption on every platform. Sandboxed network access is blocked so it cannot send them anywhere, but its report comes back to you, and that is a channel.
A prompt is untrusted input, and Codex acts on it. Content you did not write — an issue body, a
web page, a log, a file from someone else's repository — can carry instructions. With
workspace-write it can direct Codex to modify your repository; even read-only it can direct Codex
to read something sensitive and put it in the answer. The sandbox bounds where Codex can write. It
does not judge what it should write, or why it was asked. This is prompt injection, and it is the
risk that matters here.
The result is not sanitised. What comes back is text from a model that just read your files. Treat it as data, not as instructions; review what a delegation did rather than assuming it did what you asked.
Reducing the risk
- Leave the built-in default alone. Read-only handles investigation, review and diagnosis, which is most delegation.
- If you never want writes, cap it:
CODEX_SUBAGENT_MAX_SANDBOX=read-only. No conversation can argue past a ceiling. Register it outside the repository (Claude Code's defaultlocalscope,--scope user, or Claude Desktop's config), not in a project.mcp.jsonthat a write-enabled delegation could edit. See Configuration. - When you enable writes, add
use_worktreeso changes land somewhere you can inspect before they touch your branch. - Do not assemble delegation prompts from untrusted content when you intend to act on the answer.
- If this threat matters seriously to you, run Codex under an account or container with no access to your secrets. That solves it at the root instead of bounding it.
SECURITY.md has the full threat model, what a deny_read policy could add, and how
to report a vulnerability.
Configuration
Everything is optional, and set through environment variables on the MCP server:
A model supports a subset of efforts; an unsupported effort is adjusted to the closest supported one with a note. See docs/TOOLS.md#configuration for full semantics and Safety for the sandbox boundary.
Claude decides when to delegate; each run sends its prompt to OpenAI and spends your Codex usage, in
any conversation where the server is available. docs/CONTROL.md covers client
permission prompts, server ceilings, version ranges and your own CLAUDE.md escalation rules. Those
rules stay yours; see ADR 12 and
ADR 14 for the policy split.
Choosing a model
Use list_codex_models for the live catalog from your installed CLI. Model and reasoning effort are
independent: the model sets raw capability; the effort sets how long it deliberates. ultra
additionally delegates subtasks automatically.
The server does not choose for you. A task's model depends on your budget and how costly a wrong answer is. Without a model in the call or configuration, it refuses with a recommendation for you to decide on.
With no default, tool descriptions tell Claude to call codex_recommend first, announce the
suggested model and effort, then delegate with both explicit. This is guidance to a model, not
enforcement. For advice, ask: “Which Codex model should handle migrating this repo's tests to
vitest?” The suggestion respects your model allow-list and effort ceiling; see
the recommendation reference.
To skip that step on later delegations, set a default from list_codex_models once:
Tools
Full parameter reference: docs/TOOLS.md.
FAQ
Does this cost money?
It uses your existing Codex quota, as running codex yourself does. This server adds nothing and
has no visibility into the cost. Higher efforts consume more, and ultra also delegates subtasks;
codex_recommend helps avoid spending ultra on low work.
Why drive the CLI instead of calling the OpenAI API? Delegated coding is not a single completion — it is an agentic loop with a sandbox, an approval model, session persistence and project instruction files. All of that lives in the Codex client, not in the model endpoint. See ADR 1.
Do I need Claude Code, or does Claude Desktop work? Either. Claude Code gets a one-line install; Claude Desktop needs a manual config entry.
It says Codex is not installed, but codex works in my terminal.
Most likely Windows with a global npm install, which produces a codex.cmd batch shim that cannot
be launched without a command shell. codex_doctor reports this as unsupported-shim and offers
two fixes. On macOS and Linux, check whether a Node version manager moved codex off PATH.
Does it work on Windows and Linux?
CI builds, tests and starts the server on Windows, macOS and Linux on every change, and checks that
the Codex CLI is resolved correctly on each. A real delegation has only been verified on macOS — the
CI runners have no Codex installation or credentials. Reports from Windows and Linux are welcome.
On Windows, install the Codex CLI with the PowerShell installer rather than npm; on Linux, install
bubblewrap for Codex's sandbox. See Installation.
Documentation
- docs/INSTALL.md — installation, platform notes and troubleshooting.
- docs/CONTROL.md — permissions, ceilings and when Claude delegates.
- docs/DELEGATING.md — how to scope a delegation, with measured costs and worked prompts.
- docs/TOOLS.md — every tool and parameter.
- docs/COMPARISON.md — versioned comparisons with other Codex MCP servers.
- docs/adr/ — why the design is what it is, decision by decision.
- docs/ROADMAP.md — what is planned, and what is deliberately out of scope.
- Issues — what is actually open right now.
- docs/VERSIONING.md — what counts as a breaking change here.
- CHANGELOG.md — what changed in each release, and the Codex CLI version it was verified against.
- CONTRIBUTING.md — setup, and the rules that are not negotiable.
Disclaimer
Not an official product. This is an independent, community project. It is not affiliated with, endorsed by, sponsored by or supported by OpenAI or Anthropic. "Codex", "ChatGPT" and "OpenAI" are trademarks of OpenAI; "Claude" and "Claude Code" are trademarks of Anthropic. They are used here only to describe what this software interoperates with, which is nominative use — no claim is made to any of them. Neither company is responsible for this software, and problems with it should be reported here rather than to them.
No warranty. The software is provided "as is", without warranty of any kind, as stated in LICENSE. You use it at your own risk.
Read Safety before enabling writes; delegations spend your own Codex usage.
License
MIT. See LICENSE.
來源:README.md,提交 b1c0bbe
工具
0版本歷史
1- v0.4.0最新Oct 1, 2026


