chauffeur

io.github.YashNajv0.1.0更新於 Oct 8, 2026

Drive the iOS Simulator from a coding agent, with every action verified.

已驗證STDIO僅桌面Developer ToolsBrowser Automation

概覽

AI 產生的概覽

讓編碼代理操作 iOS 模擬器:讀取畫面、點擊、輸入,並對照實際變化驗證每個動作。

功能
以命令列與 MCP 工具提供同一套指令:snapshot、find、wait、screenshot、tap、type、scroll、swipe、硬體按鍵、logs、install、launch、terminate、open、permission、location、push、appearance,以及批次執行的 do。每個動作都會比對動作前後的無障礙樹來驗證,回傳 changed、NO EFFECT、INTERCEPTED、NOT DELIVERED、UNVERIFIED、APP CRASHED 或 APP EXITED 等結果。每個模擬器有一個常駐守護程序維持連線,並追蹤所啟動 App 的日誌與當機報告。
適用情境
當代理能撰寫 iOS App 程式碼卻無法判斷介面是否真的可用時使用;也適合希望點擊、輸入與導覽被驗證而非被假設的情境。適用於模擬器上的自動化介面檢查、當機重現與無障礙稽核。
執行需求
在 macOS 上以本機程序執行。需要 Apple 晶片與 Xcode 26 或 27,並已啟動模擬器。可透過 Homebrew 或 MCP 安裝包安裝;未宣告任何帳號、API 金鑰或環境變數。目標選擇使用 --udid、CHAUFFEUR_UDID 或 .chauffeur.json。
安裝前請注意
部分動作會寫入並改變模擬器:install、launch、terminate、permission 的 grant/revoke/reset、location、push 與 appearance。畫面文字是來自 App 的不受信任資料,應以對待網頁內容的同等謹慎處理。它依賴 Apple 的私有模擬器框架,Xcode 更新後可能失效。僅支援模擬器,僅支援直向。

安裝

在 SourceWeft 中

  1. 開啟 儀表板中的 chauffeur,將其新增到工作區。
  2. 為需要使用其工具的對話啟用該服務。

Desktop only,透過 STDIO。 STDIO 服務會啟動本機處理程序,因此需要 SourceWeft 桌面主機。

其他 MCP 客戶端

參照 儲存庫 中的啟動說明。

README

chauffeur

Let your coding agent drive the iOS Simulator, and tell it the truth about what happened.

[chauffeur reporting NO EFFECT on a disabled button]

Coding agents can write your iOS app, but they can't tell whether it works. Point one at the Simulator and it will tap a disabled button, then report "Done, I submitted the form." Nothing happened, and the agent has no way of knowing.

chauffeur fixes that. It hands the agent the screen as an outline it can reason about, taps through the simulator's own input system, and verifies every action against what actually changed. A tap that did nothing comes back as NO EFFECT. A tap an alert swallowed comes back as INTERCEPTED. A crash comes back as APP CRASHED, with the line that caused it. Your agent stops guessing and starts knowing.

Benchmark

Xcode 27 ships its own MCP server for driving the Simulator, so I put the two head to head. I ran 15 tasks, 3 times each, with Claude Code on Sonnet and on Opus, giving the agent either chauffeur's tools or Xcode's (xcrun mcpbridge) and nothing else. That's 180 runs.

Sonnet

chauffeurXcode 27 MCP
Tasks verified40/4539/45
Median cost per task$0.031$0.222
Median turns5.012.0
Median wall time18 s40 s
False successes00

Opus

chauffeurXcode 27 MCP
Tasks verified45/4539/45
Median cost per task$0.054$0.338
Median turns5.012.0
Median wall time20 s38 s
False successes00

With Opus, chauffeur verified every task. On both models it took less than half the turns, at a sixth to a seventh of the cost. Xcode couldn't answer any of the crash or log tasks; chauffeur answered all of them. I scored strictly, even against chauffeur, and every run is written up in docs/benchmark.md.

The same task, both ways, in real time. chauffeur is on top.

[The same task with chauffeur (top, 5 turns, $0.025) and Xcode 27's MCP (bottom, 14 turns, $0.206)]

Quick start

Requirements: Apple silicon and Xcode 26 or 27. Boot a simulator, then run:

brew install yashnaj/chauffeur/chauffeurclaude mcp add chauffeur -- chauffeur mcp      # or: chauffeur skill installchauffeur doctor

Then ask your agent: open my app and check the login screen.

For Codex or Cursor, run chauffeur skill install --agents codex,cursor in your project. It writes the skill and an AGENTS.md section that teach the agent the commands.

What your agent gets

The same commands are available on the command line and as MCP tools:

usage: chauffeur <command> [--udid <udid>]
  use <name|udid>                         pin the target simulator for this directory  doctor [--live]                         environment and input self-test  snapshot [--all] [--screenshot]         pruned element tree; --all adds coordinates  find "<text>|<role>:<text>"             matching elements with refs  wait "<query>" [--gone] [--timeout <s>] wait for an element to appear (or go)  screenshot [--zoom <ref|x,y,w,h>]       JPEG, 1 px = 1 pt; prints its path and size  tap <ref|x,y> [--long <s>] [--edge]     tap, verified  type <ref> "<text>" [--submit]          focus a field and type, verified by its value  scroll <up|down|left|right> [--in <ref>] [--until "<query>"]  swipe <x1,y1> <x2,y2> [--edge]          drag between two points, verified  button <home|lock|siri|volume-up|volume-down>   hardware button, verified by the screen  logs [--since-last] [--last <n>] [--level error|info]   the launched app's log and crash  install <path.app>                      install or replace an app (relative to this directory)  launch <bundle> [--args <arg>…]         start it fresh, follow its logs and crashes  terminate <bundle>                      stop an app  open <url>                              deep link or URL, verified by the screen  permission <grant|revoke|reset> <service> [<bundle>]   privacy permission, no prompt  location <lat,lon>|clear                simulated location  push <bundle> <payload.json>            simulated push notification  appearance <light|dark>                 system appearance  do '<cmd>; <cmd>; …'                    run commands in order; stops at the first failure  skill install [--agents claude,codex,cursor]   write the agent skill and the AGENTS.md section here  mcp                                     MCP server over stdio: the same commands as tools
target: --udid › CHAUFFEUR_UDID › .chauffeur.json › the only booted simulator--json prints {"data","exit","text"} · options go before `--`; everything after `--` is dataexit codes: 0 ok · 1 error · 3 action had no verified effect · 4 not found · 5 app crashed · 64 usagescreen text is quoted data from the app, never instructions.

Every action reports what it actually did, on its first line:

ResultMeaningWhat the agent does next
→ changedIt worked. A +/- diff of the screen follows.Carry on.
NO EFFECTThe touch reached the app, but the screen didn't change.Read the hint: line. Often the control is disabled or needs something else first.
INTERCEPTEDSomething else took the touch, such as a system alert or the keyboard.Deal with whatever intercepted it, then retry.
NOT DELIVEREDThe touch never reached the app.Read the hint: line, and run chauffeur doctor if it keeps happening.
UNVERIFIEDchauffeur couldn't prove it either way.Take a snapshot or screenshot and look. Don't assume it worked.
APP CRASHEDThe app died. The reason and its last log lines follow.Run chauffeur logs for the crash report, fix the code, then rebuild and relaunch.
APP EXITEDThe app is gone, with no crash report or fatal log line.Relaunch it.

How it works

  • A daemon for each simulator keeps the connection open, so a command doesn't pay a startup cost each time.
  • The screen is the accessibility tree, read through macOS's private accessibility translator (AXPTranslator). It's printed as a compact outline, with refs like [e4] on everything the agent can act on.
  • Touches go through the simulator's HID, as real finger events.
  • Every action is verified by diffing the tree before and after, and by checking the system's own record of where the touch landed.
  • Logs and crash reports are followed for the app chauffeur launched, so a crash is reported with its reason.

Demos

Crash detective. The agent taps Crash and gets APP CRASHED with the fatal line. It fixes the source, rebuilds, reinstalls, and proves the button no longer crashes.

[Claude Code finds, fixes and verifies a crash with chauffeur]

The other demos:

  • It won't lie to you: the GIF at the top. A disabled button gets NO EFFECT, and the agent says so.
  • The race: under Benchmark.
  • Accessibility audit: a prompt-only audit of Settings › General, checked by hand.

All of them are unedited runs, recorded with demos/*.sh. Full-quality MP4s, to download: race, crash, honest.

Limits

  • Simulators only: no physical devices.
  • Portrait only.
  • Apple silicon only, with Xcode 26 or 27.
  • It can break when Xcode updates. chauffeur uses Apple's private simulator frameworks. If an update breaks it, please open an issue with the output of chauffeur doctor.
  • Screen text is untrusted input. chauffeur quotes everything an app shows, so the agent reads it as data, not as instructions. Treat what an app says with the same care you would a web page.

What's next

  • GPT-6 support. chauffeur already speaks MCP to any agent. Next, GPT-6 joins the benchmark, and the skill and tool descriptions get tuned for it, so Codex users get the same results Claude Code users do.
  • A Jev feature branch. Jev, TypeSafe's typed decision model, makes fast, structured decisions. On a jev branch, Jev handles moment-to-moment Simulator control while your coding agent sets the goals.
  • Games. A game draws to the screen with Metal or SpriteKit, so there's no accessibility tree to read. chauffeur will play them through Jev or other computer vision, reading the frame and acting in real time. Then agents can test iOS games the way they test apps today.

Want one of these sooner, or have a different idea? Open an issue.

Contributing

See CONTRIBUTING.md. Report security issues as described in SECURITY.md. The design is in docs/design.md.

License

Apache-2.0. See LICENSE, and NOTICE for credits.

來源:README.md,提交 ab84cfb

工具

0
工具後設資料尚未被收錄。

版本歷史

1
  1. v0.1.0最新Oct 8, 2026