
chauffeur
io.github.YashNajv0.1.0更新于 Oct 8, 2026
Drive the iOS Simulator from a coding agent, with every action verified.
概览
让编码代理驱动 iOS 模拟器:读取屏幕、点击、输入,并对照实际变化验证每一步操作。
- 功能
- 以命令行和 MCP 工具提供同一套命令:snapshot、find、wait、screenshot、tap、type、scroll、swipe、硬件按键、logs、install、launch、terminate、open、permission、location、push、appearance,以及批量执行的 do。每个操作都会通过对比操作前后的无障碍树来验证,返回 changed、NO EFFECT、INTERCEPTED、NOT DELIVERED、UNVERIFIED、APP CRASHED 或 APP EXITED 等结果。每个模拟器有一个常驻守护进程保持连接,并跟踪所启动应用的日志与崩溃报告。
- 适用场景
- 当代理能编写 iOS 应用代码却无法判断界面是否真的可用时使用;也适合希望点击、输入和导航被验证而非被假定的场景。适用于模拟器上的自动化界面检查、崩溃复现和无障碍审计。
- 运行要求
- 在 macOS 上作为本地进程运行。需要 Apple 芯片和 Xcode 26 或 27,并已启动模拟器。可通过 Homebrew 或 MCP 安装包安装;未声明任何账号、API 密钥或环境变量。目标选择使用 --udid、CHAUFFEUR_UDID 或 .chauffeur.json。
安装
在 SourceWeft 中
- 打开 控制台中的 chauffeur,将其添加到工作区。
- 为需要使用其工具的对话启用该服务。
Desktop only,通过 STDIO。 STDIO 服务会启动本地进程,因此需要 SourceWeft 桌面宿主。
其他 MCP 客户端
参照 仓库 中的启动说明。
README
chauffeur
Let your coding agent drive the iOS Simulator, and tell it the truth about what happened.
[chauffeur reporting NO EFFECT on a disabled button]
Coding agents can write your iOS app, but they can't tell whether it works. Point one at the Simulator and it will tap a disabled button, then report "Done, I submitted the form." Nothing happened, and the agent has no way of knowing.
chauffeur fixes that. It hands the agent the screen as an outline it can reason about, taps through the simulator's
own input system, and verifies every action against what actually changed. A tap that did nothing comes back as
NO EFFECT. A tap an alert swallowed comes back as INTERCEPTED. A crash comes back as APP CRASHED, with the line
that caused it. Your agent stops guessing and starts knowing.
Benchmark
Xcode 27 ships its own MCP server for driving the Simulator, so I put the two head to head. I ran 15 tasks, 3 times
each, with Claude Code on Sonnet and on Opus, giving the agent either chauffeur's tools or Xcode's (xcrun mcpbridge)
and nothing else. That's 180 runs.
Sonnet
Opus
With Opus, chauffeur verified every task. On both models it took less than half the turns, at a sixth to a seventh of
the cost. Xcode couldn't answer any of the crash or log tasks; chauffeur answered all of them. I scored strictly, even
against chauffeur, and every run is written up in docs/benchmark.md.
The same task, both ways, in real time. chauffeur is on top.
[The same task with chauffeur (top, 5 turns, $0.025) and Xcode 27's MCP (bottom, 14 turns, $0.206)]
Quick start
Requirements: Apple silicon and Xcode 26 or 27. Boot a simulator, then run:
Then ask your agent: open my app and check the login screen.
For Codex or Cursor, run chauffeur skill install --agents codex,cursor in your project. It writes the skill and an
AGENTS.md section that teach the agent the commands.
What your agent gets
The same commands are available on the command line and as MCP tools:
Every action reports what it actually did, on its first line:
How it works
- A daemon for each simulator keeps the connection open, so a command doesn't pay a startup cost each time.
- The screen is the accessibility tree, read through macOS's private accessibility translator (
AXPTranslator). It's printed as a compact outline, with refs like[e4]on everything the agent can act on. - Touches go through the simulator's HID, as real finger events.
- Every action is verified by diffing the tree before and after, and by checking the system's own record of where the touch landed.
- Logs and crash reports are followed for the app chauffeur launched, so a crash is reported with its reason.
Demos
Crash detective. The agent taps Crash and gets APP CRASHED with the fatal line. It fixes the source, rebuilds,
reinstalls, and proves the button no longer crashes.
[Claude Code finds, fixes and verifies a crash with chauffeur]
The other demos:
- It won't lie to you: the GIF at the top. A disabled button gets
NO EFFECT, and the agent says so. - The race: under Benchmark.
- Accessibility audit: a prompt-only audit of Settings › General, checked by hand.
All of them are unedited runs, recorded with demos/*.sh. Full-quality MP4s, to download: race,
crash, honest.
Limits
- Simulators only: no physical devices.
- Portrait only.
- Apple silicon only, with Xcode 26 or 27.
- It can break when Xcode updates. chauffeur uses Apple's private simulator frameworks. If an update breaks it,
please open an issue with the output of
chauffeur doctor. - Screen text is untrusted input. chauffeur quotes everything an app shows, so the agent reads it as data, not as instructions. Treat what an app says with the same care you would a web page.
What's next
- GPT-6 support. chauffeur already speaks MCP to any agent. Next, GPT-6 joins the benchmark, and the skill and tool descriptions get tuned for it, so Codex users get the same results Claude Code users do.
- A Jev feature branch. Jev, TypeSafe's typed decision model, makes fast, structured decisions. On a
jevbranch, Jev handles moment-to-moment Simulator control while your coding agent sets the goals. - Games. A game draws to the screen with Metal or SpriteKit, so there's no accessibility tree to read. chauffeur will play them through Jev or other computer vision, reading the frame and acting in real time. Then agents can test iOS games the way they test apps today.
Want one of these sooner, or have a different idea? Open an issue.
Contributing
See CONTRIBUTING.md. Report security issues as described in SECURITY.md. The design is in docs/design.md.
License
来源:README.md,提交 ab84cfb
工具
0版本历史
1- v0.1.0最新Oct 8, 2026


