
Pitroom
io.github.ATASTECHv0.17.0更新於 Oct 5, 2026
Hands bounded coding work to cheaper worker agents: verified answers, exact diffs, cost receipts.
概覽
Pitroom 把有界線的編碼工作(閱讀、修正、測試、審查)交給更便宜的平行 worker CLI,並回傳經驗證的答案、精確 diff 與成本回執。
- 功能
- Pitroom 提供十個 MCP 工具(pitroom_run、pitroom_wait、pitroom_show、pitroom_info、pitroom_review、pitroom_audit、pitroom_apply、pitroom_discard、pitroom_revert、pitroom_stop),把任務交給 OpenCode、Codex CLI、Claude Code、Gemini CLI 等 worker CLI。worker 平行執行,引用的 path:line 會在磁碟上核對,並回傳摘要、精確 diff 以及 token 與預估節省的回執。執行紀錄也以資源形式暴露,另有四個提示詞,本機儀表板與可搜尋歷史可查看每個 worker 做了什麼。
- 適用情境
- 當編碼代理把高價 token 花在閱讀、grep 與審查上,而你希望這些雜活由更便宜或免費的模型平行完成、決策仍由主代理掌握時,適合使用。它適用於委派調研、實作、審查與稽核任務,並讓主模型的上下文保持較小。
- 執行需求
- 需要 Node.js 22.13+ 以及至少一個 worker CLI:OpenCode v2+(預設,帶免費模型)、Codex CLI、Claude Code 或 Gemini CLI(beta,需要 Google AI Studio 的 API key)。透過 npm(pitroom)安裝,用 npx 或全域指令執行;MCP 伺服器以 stdio 在其啟動目錄或 -d DIR 指定的目錄中執行。可選設定透過環境變數,如 PITROOM_WORKER、PITROOM_MODEL、PITROOM_FALLBACK、PITROOM_TIMEOUT、PITROOM_MAX_PARALLEL、PITROOM_PRIMARY、PITROOM_PRICE、PITROOM_HOME、PITROOM_ _BIN,或 config.json 檔案。
安裝
在 SourceWeft 中
- 開啟 儀表板中的 Pitroom,將其新增到工作區。
- 為需要使用其工具的對話啟用該服務。
Desktop only,透過 STDIO。 STDIO 服務會啟動本機處理程序,因此需要 SourceWeft 桌面主機。
其他 MCP 客戶端
參照 儲存庫 中的啟動說明。
README
Pitroom
A free pit crew for your expensive coding agent
Cheaper and faster: hand the reading, fixes, tests and reviews to workers that run in parallel. Keep the decisions.
Model-agnostic · Verified answers · Receipts, not vibes · OpenCode / Codex / Claude Code / Gemini CLI workers
[Release] [Downloads] [npm] [Stars] [CI] [License]
Quick start · Dashboard · Benchmarks · Workflow · Skills · Commands · Workers · Responsibility · Write an adapter
[Pitroom — a free pit crew for your expensive coding agent. Orange Claw’d, a blue terminal-faced Codex pet and a purple-blue Gemini star work together at a coding terminal.]
Your agent decides · Cheap workers do the typing · A second model reviews
[The Pitroom dashboard: running workers with their animated mascots, a finished run opening into its task, steps, result and diff, then the History and Stats tabs (sample data)]
pitroom dash, with sample data
Why Pitroom?
Claude Code, Codex and friends spend premium tokens reading: grepping, opening files, scanning code they'll never quote.
Pitroom goes one step further:
Your agent stays in the driver's seat. A pit crew of cheap workers does the legwork in parallel, while your agent keeps going, and hands back only the answer, the exact diff and a receipt.
And more
- Receipts, not vibes. Tokens the worker burned, tokens returned to your agent, compression ratio, and an estimate of what your primary model would have charged.
pitroom savings --card card.svgmakes a shareable card. - Self-healing worker chain. Free models get rate-limited or retired. List fallback workers once (
fallback) and a run that hits "model not found", 429 or quota errors moves to the next model automatically. Still no model hardcoded: the chain is yours. - A crew, not just one worker.
pitroom crewstarts several workers at once,pitroom watch --jsonstreams one line per change so your agent follows them live,pitroom wait -gcollects every answer,pitroom apply -glands isolated patches in order. Up to 20 workers run at once by default (maxParallel, 30 at most), the rest queue, and the fallback chain absorbs free-tier rate limits; parallel writers are refused. - Faster, not only cheaper. Workers run in the background and side by side (up to 20 at once) while your agent keeps working, so slow reading and review stop blocking it. In the benchmark, 470 s of worker time finished in 156 s of wall clock, about 3× faster than one question after another, and your agent's own context stayed small.
- Free first, dearer only when needed. Runs and plan tasks start on the
cheaptier, OpenCode's free model;standard(Codex) andcapable(Claude Code) are for tasks that need them, and an optionalreviewtier names who reviews. Gemini CLI with a free Google AI Studio key is one more cheap worker (small daily quota per model).pitroom modelsshows what each worker offers, the effort levels (low…xhigh) each model accepts, what you say it costs (costs) and what it used;--effortsets the level per run. Details: Models, costs and effort. - Exact changes, even in a dirty tree. Pitroom snapshots the working tree with a throwaway git index (your index, branches and stash are never touched), so it reports only the worker's edits and can undo exactly those:
pitroom revert <id>. - Your agent sets the permissions. It creates each worker with the permissions the task and your session allow (read-only, isolated copy, in-place edits, web), and a worker never widens its own. A fixed floor stays for every run: no git history changes, no
sudo, no publishing, no secrets. A worker cannot land a deletion either:pitroom applyrefuses a patch that deletes files unless your agent passes--allow-delete. - Real isolation.
--isolateruns the worker in a private copy of your current state (its own repository, sharing your objects read-only), uncommitted and untracked files included, then hands you a patch:pitroom apply <id>(checked, refuses on conflict) orpitroom discard <id>. - Zero repo pollution. No
.pitroom/folder, no.gitignoreedits, no branches. Records live in~/.local/state/pitroom. - Works with any agent. Fourteen Agent Skills: the superpowers development workflow run by workers, plus delegation (
using-pitroom,pitroom-research,-crew,-implement), a CLI, and a Claude Code / Codex plugin whose session-start hook loads the workflow. Claude Code, Codex, Gemini CLI, Cursor, or anything that can run a shell command. - A live dashboard and a searchable history.
pitroom dashopens a local page of every run (the task, what the worker did step by step, its result and diff) with history and statistics, for the agent apps that show neither hooks nor a status line;pitroom historysearches everything Pitroom ever ran. See Dashboard. - Small. About 6,500 lines of TypeScript in the CLI, a ~210 KB CLI bundle (plus the ~1.9 MB of dashboard files it serves, with syntax highlighting for ten languages), zero runtime dependencies.
How it stays safe
Each run gets the worker's own safety mechanism set to the mode (for OpenCode, a permission profile injected through OPENCODE_CONFIG_CONTENT): no commit/push/reset/checkout/stash/clean/rebase, no bulk deletes, no sudo, a file-reading tool that refuses .env and key files (OpenCode and Claude Code workers), no web tools unless you pass --web, no subagents, no recursive delegation. Only allow/deny rules, so a headless run never stalls on a prompt. A git guard on the worker's PATH also stops sh -c "git push", env git reset, aliases and scripts from committing, pushing, resetting, stashing or touching your index. Your opencode.json is never touched. These safeguards reduce risk; they are not a security boundary (see Responsibility and SECURITY.md).
See it work
A real run on a free OpenCode Zen model: your agent reads ~554 tokens instead of 163,000.
Every run prints a receipt, and pitroom savings adds them up (--card writes this shareable card; the numbers here are sample data, and a real measurement follows in Benchmarks):
Dashboard
pitroom dash --detach prints the address of a live page on 127.0.0.1 (read-only, this machine only). It is for the places where hook messages and status lines do not reach, such as the Claude Code and Codex apps: open it in a browser or in the app's own browser pane. The skills tell your agent to start it and give you the address when it runs workers in the background.
The screenshots below show sample data (an imaginary shop-api project), not a real one.
Live. What is running now, and the latest runs. Running cards show a timer, the worker's last words and its animated pixel mascot (a finished card shows the worker's logo); finished ones show the result and, for reviews, the findings.
Click a card and it grows into a panel with the task, what the worker did step by step (files read, searches, commands, edits, with times, failed ones marked), its result, the files and diff it changed, and the details (model, tokens, cost, fallbacks, reference check).
Audits and skipped workers show on the card. A research answer that another worker re-checked (Audits) carries an audited · agrees / partly agrees / disagrees badge and, when it disagrees, the claims it disputes; a worker left out because its quota ran out is listed under the details ("Skipped … cooling down until 14:05", see pitroom cooldown).
History searches everything Pitroom ever ran (the full text of tasks, answers and steps), with filters by state, model and period; each run opens as the same card as on Live. Stats shows runs, success rate, time, tokens and savings, per worker and model and per day.
[The History tab: a search box, filters and every past run as a card] [The Stats tab: totals, runs per day, and per worker and model its logo, success rate and how many of its audited answers another worker confirmed]
The dashboard is a React app built once into dist/ui and served as two static files, so the CLI still has no runtime dependencies. It follows your system's light or dark theme (a button switches it).
Benchmarks
Measured on a real, public repository: PI-Desktop at commit c2bfe35, 2,349 tracked files and about 419,000 lines of TypeScript and Rust. Eight questions of the kind an agent asks before it changes code, run as one read-only crew (pitroom crew, at most four workers at a time) on OpenCode's free muse-spark-1.3-contributor-free model.
Workers read 5.5 million tokens of code and docs (96 steps, 179 tool calls) and handed the agent 7,684 tokens, about 960 per answer. The eight runs add up to 470 s of worker time but finished in 156 s of wall clock, because four ran at a time (the limit then; the default is now 20): about 3× faster than one after another. That compares the same workers with themselves, not with your agent doing the reading alone, which I did not measure. pitroom savings estimates $3.86 saved for these eight at Claude Sonnet list prices.
Were the answers right? Checked by hand against the repository: 46 key claims across all eight answers (the cited line and what it says) were all correct, and two completeness checks held (exactly 6 invoke and 3 event window channels; no tests in config_sync_rpc.rs). The i18n answer lists the 12 source files that import the package and leaves out 11 test and fixture files that do too, which the question did not ask about. Pitroom's own checker marks fewer references as verified than that: ¹ those two answers cite bare file names (transcripts.rs:651) that it cannot resolve to a path, and it counts them as unverified although the lines are right.
A change in an isolated copy: "add a one-line comment above resolveUpdateMode" finished in 16 s and 4 steps (142,491 tokens). The patch was one file and one line, the comment matched the code, and your working tree is untouched until you apply.
Free models against gpt-6-sol, on three large repositories. The same kind of work on React (7,252 files), Django (7,091) and Kubernetes (31,353 files), at pinned commits. Nine bounded questions, three per repository: the file and line that define a function, how many files contain a word, and which files contain it. Every answer is checked against git grep, not by opinion: an exact path:line (half a point for the right file on the wrong line), an exact count, and for the lists an F1 score that punishes both missed and invented files. Each worker ran the nine questions once, read-only, with no fallback, so a model that cannot do it fails instead of being swapped for another.
Four free models matched gpt-6-sol on these questions. Every model found every definition; the misses are count questions (a wrong number) and, for two models, list questions, one of them by adding docs/ files that are not under django/. gpt-6-sol was as fast as the quickest free models, and lowering its effort did not cost accuracy here. nemotron-3.5-lightning-free wrote an unrelated text for one list question and ran into the 10-minute limit on another. The scoring harness and every raw answer are in benchmarks/multi-repo.
Only workers that ran are in the table. Models that could not answer at all (a provider error, or the shared daily quota of OpenRouter's free tier) are left out, and so are the two gpt-oss-20b runs that ended in a provider error; a timeout is kept.
Does a review find a planted bug? A bug was changed into real code (an inverted check, a swapped &&/||, a flipped return) in React, Django, Kubernetes and Pitroom itself, committed as a bare "tidy", and pitroom review was asked to review the commit. 29 packages (14 with the bug alone, 15 with the bug and three comment-only edits around it). A review counts as a find when a Critical or Important finding names the changed file and either cites a line within 5 or names the changed identifier or its function; "exact line" is the strict version (within 3).
Almost every model found almost every bug, so this test separates them on speed and precision, not on whether they can review: muse-spark described the right bug but often quoted a wrong line number (57% exact), and gpt-6-sol and muse-spark were the fastest by a wide margin. Only three models finished all 29: the other four are scored on the reviews that ran before a provider limit (the free tier's usage limit, and the Codex plan's, which gpt-6-sol reached), and the rest are left out, not counted as misses, so their rows rest on 9 to 14 reviews. Every change is only 2 to 5 lines, which is easy; there is no false-alarm rate (the comment-only "clean" packages turned out to contain comments that were really wrong, so they cannot show one); and React, Django and Kubernetes may be in a model's training data. The harness and every raw review are in benchmarks/review-bugs.
What this does not show
- One run per question and model: there is no variance here. The questions are bounded lookups that
grepcan answer, so they do not show how a model handles design questions or large edits, and four free models andgpt-6-solall scoring 100% says the test is easy at the top, not that they are equal. - An open-ended task ("find up to five spelling mistakes in
docs/") did not finish: it was stopped after 19 minutes and 57 tool calls. Give workers bounded tasks. - "Tokens processed" is what the workers read and wrote. What your agent would have spent doing the same reading itself is an estimate, not a measurement: the savings figure assumes it would process about the same tokens. The $0 worker cost is because the model is free.
- The hand check covered key claims, not every citation.
To repeat it on your own repository: pitroom crew -d <repo> -g bench "<question 1>" "<question 2>" …, then pitroom wait -g bench, pitroom show <run> and pitroom savings --models.
Quick start
You need Node.js 22.13+ and at least one worker CLI: OpenCode v2+ (the default worker, with free models), Codex CLI, Claude Code or Gemini CLI (beta; it needs an API key from Google AI Studio, a Google account sign-in no longer works for it).
1. Install: pick your agent
Claude Code
Inside a session: /plugin marketplace add ATASTECH/pitroom, then /plugin install pitroom@pitroom.
Codex
Gemini CLI
Any agent that speaks MCP (Cursor, Claude Desktop, and the three above too)
Pitroom is also an MCP server on stdio: ten tools instead of shell commands and skills (pitroom_run, pitroom_wait, pitroom_show, pitroom_info, pitroom_review, pitroom_audit, pitroom_apply, pitroom_discard, pitroom_revert, pitroom_stop), the runs as resources, and four prompts. The tool definitions sit in the client's context for the whole session, so they are kept small: about 1.9k tokens. It is listed in the MCP Registry as io.github.ATASTECH/pitroom, so clients that browse the registry can find it. After npm i -g pitroom, let Pitroom register itself in the clients it finds (Claude Code, Codex, Gemini CLI, Cursor, Claude Desktop):
It registers the launcher by its full path (apps like Claude Desktop start without your PATH), changes Claude Code, Codex and Gemini CLI with their own mcp add command (user scope where it has one), merges the JSON of Cursor and Claude Desktop without touching their other servers (keeping a .bak-pitroom backup, and leaving a file that is not valid JSON alone), skips what is already registered, and pitroom uninstall removes it again. pitroom doctor shows where it is registered. Restart the client afterwards. Or do it by hand:
or, for Cursor (~/.cursor/mcp.json), Claude Desktop (claude_desktop_config.json) and others:
Each tool runs the matching pitroom command, so every rule of the CLI applies unchanged (permission profiles, git guard, isolation, read snapshots). A run is waited for up to waitSeconds (default 50), then comes back as "still running" with its id for pitroom_wait; keep it below your client's tool timeout. The server works in the directory it is started in: the client's for stdio, or the one -d DIR names (Claude Desktop starts it elsewhere, so give it "args": ["mcp", "-d", "/path/to/project"]); the tools' dir argument picks another per call.
- Progress. While a call waits for a run, a client that sent a progress token gets a
notifications/progressevery few seconds ("running · read · 12s · 3 steps, 2 tool calls · last: …"), and a client that resets its timeout when progress comes in (the protocol allows it; not every client does) does not give up on a long run. Cancelling apitroom_run,pitroom_revieworpitroom_auditcall stops the run it started (a cancelledpitroom_waitonly stops waiting). - Resources. The latest runs are listed as
pitroom://run/<id>(the report) and, for runs that changed files,pitroom://run/<id>/patch(the exact diff), so a client can attach one to a conversation. A client can subscribe to a run (resources/subscribe) and is told when it changes state (notifications/resources/updated), and every client is told when the newest run changes (notifications/resources/list_changed), so it need not poll; over HTTP these arrive on the session's event stream (a GET). - Prompts.
research,implement,reviewandcrewsay how to use Pitroom for that job (in Claude Code they appear as/mcp__pitroom__research, …), and every skill is a prompt of the same name (using-pitroom,pitroom-research, …), so a client without skills gets the same guidance. The server's instructions tell the agent to use the skills alongside the tools. - More than runs:
pitroom_runtakestasks(independent tasks in parallel, one worker each) andcontinue(a follow-up in the same worker session);pitroom_inforeports without changing anything, bytopic:runs,history(find an earlier answer before asking again),statsandmodels(pick a worker),savings,cooldown,config,doctor;pitroom_stopwithcooldownstries rate-limited models again. - Over HTTP.
pitroom mcp --http [--port N](default 7117) serves the same tools athttp://127.0.0.1:7117/mcpfor clients that connect to a URL, for exampleclaude mcp add --transport http pitroom http://127.0.0.1:7117/mcp --header "Authorization: Bearer $(cat ~/.local/state/pitroom/mcp-token)"(the command is printed when it starts). It is a local service: it listens on 127.0.0.1 only, needs the bearer token (made on first start, kept at<state dir>/mcp-tokenwith owner-only permissions, or set withPITROOM_MCP_TOKEN), and refuses a Host or Origin that is not local, so a web page cannot use it. Anyone who has the token can run workers as you, so keep it private. It works in the directory it was started in (printed when it starts;-d DIRfor another). Each client has its own session; a call that asks for progress is answered as an event stream. Ending a session or stopping the server ends the waiting, not the runs: collect them withpitroom_waitor the CLI.pitroom install --mcpregisters the stdio command, which needs no token and no running process.
Any other agent (a plain shell) or just the CLI
Pick one path, not several: doctor warns if the skills load twice. The plugins alone do not put pitroom on your PATH, so install it from npm too if you want to run it yourself (watch, models, savings, the status line).
2. Check the setup
It checks the worker CLIs, models, permissions and skills, and runs one live round trip.
3. Use it
Pitroom is a toolbox, not a procedure: your agent uses it when it helps, or you ask explicitly:
Use pitroom to map how sessions are created and invalidated, then propose a fix. Split this into a pitroom crew: audit auth, billing and uploads for missing input validation.
Or start a worker yourself and read its receipt:
Update and uninstall
Modes
Read snapshots. With secret-looking files in the directory, a read run reads a snapshot of your project's current state (uncommitted and untracked files included) without them and without git-ignored files, made from git objects, kept once per state and shared by every read run on it (twenty workers, one snapshot), and cleaned up after three days or by pitroom clean. Nothing is copied back: a read run has nothing to land. Its files are read-only, and the paths a worker cites are turned back into your project's. With no secret-looking files nothing changes: the worker runs in place. --in-place (or "readIn": "project" in the config) reads the directory itself, e.g. for a task that needs build output; "readIn": "snapshot" always uses one.
Answer cache. Ask the same read question on the same code and you get the earlier answer back at once, with no worker run: pitroom ⟲ cached answer · the same question on the same code as run … · no worker ran. "The same code" is exact: the commit and every working file, uncommitted and untracked ones included (git-ignored files are not code), so any edit makes it a new question. Only answers that held up are reused (the run finished, every file:line reference checked out, no audit disputed it) and only for 7 days ("cacheDays" in the config, 0 turns it off, or PITROOM_CACHE_DAYS). Follow-ups, web runs, plan tasks, reviews, audits, --verify runs, parallel work (crews) and changes are never answered from the cache, and --fresh asks a worker anyway (its answer is the one reused from then on). Not part of "the same code": git-ignored files (build output, installed dependencies) and the worker or model that answered; when a question depends on them, use --fresh.
Long jobs: --bg returns immediately; pitroom wait <id> blocks for up to 9 minutes (made for agents with 10-minute tool limits; exit code 75 means "call wait again").
Workflow
Pitroom ships a full development methodology as skills, adapted from superpowers so that its subagents are cheap workers: your agent brainstorms and plans with you, then executes the plan task by task while workers do the typing and a second model does the reviewing.
A review package holds the diff under review with 10 lines of context, and the reviewer's CLI sends it to that worker's model provider. Pitroom does not filter it: secrets committed in a reviewed range go along as they are.
Tiers map plan tasks to workers: "tiers": {"cheap": "opencode", "standard": "codex", "capable": "claude"}. pitroom plan status PLAN rebuilds where a plan stands from the run records (it survives context compaction), and pitroom plan note keeps completions and rulings outside the repo.
Audits
Pitroom checks every path:line a worker cites, which proves a line exists, not that the claim about it is true. An audit asks a second worker to verify an answer's key claims against your project and to say AGREE, PARTIAL or DISAGREE, with the claims it disputes.
or set "audit": 0.1 in the config (or PITROOM_AUDIT=0.1) to have about one read run in ten audited, in the background, without your agent asking. The same run is always in or out of the sample.
- Cost: off by default. An audit is a read run of the auditor, so about the rate times the worker's own tokens, on the
audittier (elsecheap, i.e. a free model if that is your cheap tier). It never delays the run or fails it. - Never itself: the auditor is never the worker and model that gave the answer; with no other worker configured nothing is audited (
pitroom doctorsays so). Only read runs are audited: a change haspitroom review. - Where it shows: the answer's card and
pitroom showcarry the verdict and the disputed claims; the Stats tab counts, per worker, how many audited answers were confirmed. Audits are not counted as runs and save nothing. - A sample, not a guarantee. The auditor is a model too: it can share a blind spot, and a few audits say little about a worker. Treat
DISAGREEas a reason to look, andAGREEas one more signal.
Crews
A real crew over this repository, four questions at once on a free OpenCode Zen model:
What the crew prints: the stream your agent follows and the status table
Humans get the same table live with pitroom watch -g NAME; pitroom wait -g NAME --brief prints one line per worker, pitroom show <run> the full answer.
For changes, pitroom crew -i … gives every worker its own isolated copy of your current state and pitroom apply -g NAME lands the patches in order, stopping at the first conflict. Up to maxParallel workers (default 20, at most 30) run at once, the rest queue; only one --write run per repository is allowed. pitroom stop -g NAME stops running and queued workers.
Skills
Commands
All commands and exit codes
Exit codes: 0 ok · 1 worker failed · 2 usage · 3 refused/setup · 4 timeout · 5 read-only violation · 6 verify failed · 75 still running.
History
Every finished run is also written to a SQLite database (history.db in Pitroom's state directory; Node's built-in node:sqlite, nothing to install). It keeps the task, worker and model, time, tokens, savings, the steps the worker took, its answer and the diff, and it is searchable:
While a run runs, its raw event stream is a plain file (the simplest thing that survives a crash). When it ends, the stream is compressed, and the stderr log is kept only for runs that did not succeed. pitroom clean removes old run directories but the history keeps them: pitroom show <run> still prints their report, steps, and patch. Runs from before the history existed are taken in by pitroom history import (it also runs on the first pitroom history and pitroom dash). The files remain the source: the database can be deleted and rebuilt with history import for the runs still on disk.
Seeing Pitroom at work
Two settings make every delegation visible, whether or not the agent mentions it. A status line shows running workers and this week's savings (--then keeps your own status line first); a card appears after each pitroom command the agent runs, once per phase (started, finished with its result, applied). The plugin registers the card hook itself; with pitroom install, add both to ~/.claude/settings.json:
The Claude Code and Codex apps show neither hook messages nor a status line. For them there are two things that need no setup:
pitroom dash --detachprints the address of a live dashboard with three tabs: Live (every running and recent run as an animated card; click one for the task, what the worker did step by step, its result, the diff and the details), History (search and filter everything Pitroom ever ran; every run opens as the same card) and Stats (success rate, time, tokens and savings per worker and model). It is read-only, listens on127.0.0.1only and stops itself after an hour without a request (pitroom dash --stopends it sooner). Open it in a browser or in the app's own browser pane. The skills tell the agent to start it and give you the address when it runs workers in the background.pitroom watch -g NAME --briefprints one card line when a worker starts and one when it ends. In Claude Code, the agent runs it through the Monitor tool and the lines appear in the app; in Codex the command's output block fills as it goes.
How it works
Headless problems Pitroom handles for every worker
Handled for every worker: opencode run blocks forever on an open stdin pipe; ask permissions stall headless runs; default output carries ANSI and banners; and in OpenCode v2 a plain run attaches to the shared background service, where per-run permissions and the git guard would not apply. Pitroom closes stdin, uses allow/deny-only profiles, parses --format json and always runs a private --standalone server.
Workers
A worker is named by a target: backend[:model].
Models, costs and effort
pitroom models lists what each worker offers (Codex from its own model cache, Claude Code's aliases, OpenCode's opencode models, Gemini CLI's built-in names), the reasoning-effort levels each model accepts, what you say it costs and what your own runs used:
Pitroom cannot know vendor prices and does not fetch them, so a cost is what you enter: "costs": {"codex:gpt-6-sol": 1, "codex:gpt-6.1-sol": 2} in the config, in any unit (they are only compared). pitroom doctor then prints the cost of the models in use and says when you priced a cheaper one of the same worker. A target that names only an effort (codex:#low) uses the model from models; a Codex or Claude Code worker with no pinned model gets a warning, because it would run the vendor's own default, which can change and cost more. --effort LEVEL sets the level for one run; your agent picks model and level from pitroom models (cheapest that fits: low for lookups, medium for ordinary changes, high for reviews).
Fallbacks cross backends (e.g. "fallback": ["codex:#low", "opencode"]): a worker that is rate-limited, logged out or missing its model hands the task to the next. Follow-ups (--continue) always stay on the worker that owns the session. Writing an adapter: docs/backends.md. A model that says "rate limited" (a daily quota used up, an overloaded provider) is remembered for the time its message gives, or a guess: the next runs skip it while a fallback is left, so they do not each wait for it to fail. pitroom cooldown lists what is skipped and until when, --clear tries it again, and doctor shows it too.
Configuration
Environment variables
Or put defaults in ~/.config/pitroom/config.json (flags and env still win); pitroom config shows every effective value and where it came from. pitroom init proposes a starter config from the worker CLIs and OpenCode models you have (a fallback chain of your free models, tiers when more than one CLI is installed) and writes it only with --yes; it never picks your model for you. models gives each worker a default model for targets that name none (-W codex, a "codex" fallback); a model in the target or -m still wins. tiers names workers for --tier and for plan tasks' **Worker:** lines:
workerandfallback: who runs a task when you name no one, and who takes over when it fails;models: the model each worker uses when a target names none.tiers:cheap,standardandcapablename the workers for plan tasks and--tier; an optionalreviewtier names who reviews a run (by default another worker than the implementer's,standardfirst).audit: the chance (0 to 1, default 0 = off) that a finished read run is re-checked in the background, see Audits. Theaudittier (elsecheap) names who does it.costs: your relative cost perworker:model, only compared with each other;maxParallel: workers at once (default 20, at most 30).
FAQ
Where does my code go? To whichever provider your worker's model uses. For private code, point the worker at a local model. The file-reading tools of OpenCode and Claude Code workers refuse .env files and private keys; Codex workers are not restricted that way and a shell command can still read them, so keep secrets out of the folder you delegate in.
How is "saved" computed? It assumes your primary agent would have processed about the same tokens the worker did, priced at your primary's list prices (PITROOM_PRIMARY / PITROOM_PRICE), minus the worker's cost and the cost of reading the report. It is an estimate and is labelled as one.
Can the worker break my repo? It can't touch git history, refs or your index (permission profile + git guard), can't run bulk deletes, and write-mode edits are revertible with one command. The guard is not an OS sandbox: git called by absolute path from inside a script, or non-git tools, are outside its reach, which is why --isolate exists: nothing reaches your tree until you apply it.
Can the worker read my .env? In a read run, no: when the directory holds secret-looking files (.env, prod.env, private keys and keystores, .netrc, credentials.json, secrets.json), the worker reads a clean snapshot of your project instead (see below), where they are simply not present, whatever the worker CLI would have allowed. Edits (-i, -w) and reviews are different: an isolated copy (-i) leaves out git-ignored files, so a git-ignored .env is not in it, but one that is not ignored would be copied, and Pitroom warns when it sees such files (PITROOM_NO_SECRET_WARNING=1 silences the warning). A free model may be hosted by a third party: keep real secrets out of what you delegate.
What if the worker fails? Pitroom exits non-zero with the real cause (for example a default model that no longer exists) and your agent simply continues on its own. pitroom doctor diagnoses setup problems.
Responsibility
Pitroom is provided as is, under the MIT License, and is an independent project: it is not affiliated with or endorsed by OpenAI, Anthropic, OpenCode or the other tools it can drive.
You are responsible for how you use it. You decide what you delegate and to which model provider, which permissions a worker gets, and what you apply to your projects. Workers are AI agents and can be wrong: review their changes and run your tests before you rely on them. Pitroom's permission profiles, git guard and isolated copies reduce risk, but they are not a security boundary against a determined attacker or a malicious repository. Keep secrets out of the folder you delegate in, prefer isolated mode for changes you have not reviewed, and follow the terms of the worker CLIs and model providers you connect. What reaches a provider is described in the privacy policy.
Contributing
Issues and pull requests are welcome. Start with CONTRIBUTING.md: setup, tests, the dist bundle and the commit style. Worker adapters are the easiest place to begin: docs/backends.md.
Security: please report vulnerabilities privately, as described in SECURITY.md, not in a public issue.
Report an issue · Open issues · Security policy · Code of conduct
Credits
The workflow skills (brainstorming, planning, worker-driven development, review, debugging, TDD, verification, worktrees, finishing) are adapted from superpowers by Jesse Vincent, under the MIT License; the notice is in THIRD_PARTY_NOTICES.md.
License
Pitroom is licensed under the MIT License. See LICENSE for details.
Pitroom
A free pit crew for your expensive coding agent.
Your agent decides · Cheap workers do the typing · A second model reviews
Quick start · Workflow · Skills · Write an adapter
Model-agnostic · Verified answers · Receipts, not vibes
來源:README.md,提交 515c27b
工具
0版本歷史
1- v0.17.0最新Oct 5, 2026


