
SMB Sandbox
io.github.teamshift-iov0.1.0更新於 Oct 4, 2026
Mock small-business CRM, inbox, calendar, invoicing and phone tools for testing AI agents safely.
概覽
執行一套模擬的小型企業後台(CRM、收件匣、行事曆、開立發票、電話),讓助理在不碰真實帳戶的情況下完成端對端測試。
- 功能
- 以 stdio 在本機啟動 MCP 伺服器,產生一家虛構公司,並提供 CRM、收件匣、行事曆、開立發票、電話與管理等工具集,涵蓋聯絡人、潛在客戶、商機、報價、任務、電子郵件與簡訊、工單、發票與付款。所有變更都會依業務規則驗證、寫入稽核記錄,並且只留在行程內;傳送訊息只會寫入寄件匣。狀態可匯出成 JSON 檔供評分,同一套引擎也以具型別的程式庫形式提供。
- 適用情境
- 適合評估或基準測試需要操作業務軟體的 AI 助理,用固定的產業與隨機種子重現確定性的執行結果,或在不使用正式環境憑證的情況下開發助理工作流程。它不是真實的 CRM、收件匣或帳務系統。
- 執行需求
- 需要 Node.js 與 npx,以便在本機透過 stdio 執行 npm 套件 @teamshift/sandbox-mcp;需要桌面版 MCP 用戶端,例如 Claude Code、Claude Desktop、Cursor 或 VS Code。未宣告任何帳戶、API 金鑰或環境變數。選用參數可指定產業、種子、規模、資料雜亂度、工具集、資料集檔案、狀態輸出,以及在本機連接埠提供 Streamable HTTP。
安裝
在 SourceWeft 中
- 開啟 儀表板中的 SMB Sandbox,將其新增到工作區。
- 為需要使用其工具的對話啟用該服務。
Desktop only,透過 STDIO。 STDIO 服務會啟動本機處理程序,因此需要 SourceWeft 桌面主機。
其他 MCP 客戶端
參照 儲存庫 中的啟動說明。
README
sandbox-mcp
sandbox-mcp is an open-source set of mock small-business MCP servers — CRM, inbox, calendar, invoicing and phone — backed by a realistic fictional company, so you can test AI agents end-to-end without touching real customer accounts.
Point Claude, Cursor, VS Code or any Model Context Protocol client at it and your agent gets a believable business to operate: leads waiting for a reply, quotes nobody followed up on, a duplicate contact, an overdue invoice, an unmatched check. Every change is validated like a real system would, recorded in an audit log, and stays inside the process. "Sending" an email or SMS only writes to an outbox.
Quickstart
That starts an MCP server over stdio with a generated home-services company (seed 42, so it is the same company every time). Other industries: dental-clinic, marketing-agency.
Connect your MCP client
Claude Code
Claude Desktop
Add to claude_desktop_config.json:
Cursor
Add to .cursor/mcp.json:
VS Code
Add to .vscode/mcp.json:
Then try a prompt like: "Check the inbox and the call log, reply to anything urgent, and make sure every open deal has a next step."
CLI options
Logs go to stderr only, so stdout stays a clean MCP channel.
Tools
All tools are prefixed by system, take zod-validated input, return compact JSON, and carry MCP tool annotations (readOnlyHint, destructiveHint, idempotentHint, openWorldHint: false). Failures come back as tool errors with an actionable message (for example failed_precondition: Cannot move deal from "new" to "won". Allowed: qualified, quote-sent, lost.), never as a crash.
Resources: sandbox://company, sandbox://policies, and sandbox://anomalies (only with --expose-answers, so the answers never leak to an agent under test).
Business rules it enforces
- Foreign keys must exist; money is integer cents; dates are
YYYY-MM-DD. - Deal stages follow the pipeline (
new → qualified → quote-sent → negotiation → won/lost);wonis final,lostneeds a reason and can be reopened. - No payments on void, draft or fully paid invoices, and no overpayment. Invoices with payments cannot be voided.
- Unmatched payments can only be matched to an invoice of the same customer.
- Jobs cannot double-book a technician (unless
allowOverlap), and rescheduling updates the job's times. - Email and SMS need a valid address or number, so missing contact info surfaces as a real error.
- Messaging or calling a lead's contact records the first response on the lead.
- A simulated clock starts at 09:00 company time on the dataset's "today" and advances one minute per change, so timestamps are deterministic.
Grading an agent
Run your agent against the sandbox with --state-out, then score the end state, not the transcript:
run.json contains:
dataset: the final state of every record (same schema as@teamshift/fake-business), including newevents.audit: every attempted change in order:{ seq, at, tool, input, result: "ok" | "error", error?, changedIds }.outbox: every email and SMS the agent "sent".
Compare dataset against the ground-truth anomalies from the original dataset (for example: was the duplicate contact merged, was the unmatched payment applied, did every stale deal get a nextAction?), and use audit to penalize errors, destructive actions or policy violations. Leave out the admin toolset so the agent cannot reset its own run.
Use the store directly in code
The same engine is exported as a typed library, which is how reference solutions and verifiers run:
You can also embed the MCP server in your own process with createSandboxServer(store, { toolsets, exposeAnswers }) and any SDK transport.
FAQ
How do I test an MCP agent safely?
Give it a sandbox instead of production credentials. sandbox-mcp exposes the same kinds of tools a real CRM, inbox, calendar, invoicing and phone system would, over a fictional company, with no network access and no real accounts. Run the agent, save the state with --state-out, and check what it changed.
Does it send real emails?
No. Nothing leaves the process. inbox_send_email, inbox_send_sms, crm_follow_up_quote and invoicing_send_reminder append an outbound message to the thread and to the outbox, and all addresses use the reserved .example domain and 555-01xx numbers.
How do I reset the sandbox?
Call the admin_reset tool (the audit log is kept and records the reset, so graders can see it), or restart the server. The same --industry and --seed always regenerate the identical company. In code, call store.reset().
Where does the data come from?
From @teamshift/fake-business, a deterministic generator for realistic fictional small businesses, with labeled anomalies for scoring. You can also pass your own dataset with --data.
Can the agent see the answers?
Not by default. The ground-truth anomalies are only exposed as a resource with --expose-answers; they are always included in --state-out files for graders.
Does it support Streamable HTTP?
Yes: --http 8787 serves stateless Streamable HTTP at http://127.0.0.1:8787/mcp. All requests share one sandbox state.
License
Apache-2.0
Built by TeamShift — AI workers for small-business operations.
來源:packages/sandbox-mcp/README.md,提交 fca9085
工具
0版本歷史
1- v0.1.0最新Oct 4, 2026
