SMB Sandbox

io.github.teamshift-iov0.1.0更新于 Oct 4, 2026

Mock small-business CRM, inbox, calendar, invoicing and phone tools for testing AI agents safely.

概览

AI 生成的概览

运行一套模拟的小型企业后台(CRM、收件箱、日历、开票、电话),让助手在不接触真实账户的情况下完成端到端测试。

功能
以 stdio 在本地启动 MCP 服务器,生成一家虚构公司,并提供 CRM、收件箱、日历、开票、电话和管理等工具集,涵盖联系人、线索、商机、报价、任务、邮件与短信、工单、发票与付款。所有变更都会按业务规则校验、记入审计日志,并只保留在进程内;发送消息只会写入发件箱。状态可导出为 JSON 文件用于评分,同一引擎也以类型化库的形式提供。
适用场景
适合评估或基准测试需要操作业务软件的 AI 助手,用固定的行业与随机种子复现确定性的运行结果,或在不使用生产凭据的情况下开发助手工作流。它不是真实的 CRM、收件箱或计费系统。
运行要求
需要 Node.js 与 npx,以便在本地通过 stdio 运行 npm 包 @teamshift/sandbox-mcp;需要桌面版 MCP 客户端,例如 Claude Code、Claude Desktop、Cursor 或 VS Code。未声明任何账户、API 密钥或环境变量。可选参数用于选择行业、种子、规模、数据杂乱度、工具集、数据集文件、状态输出,以及在本地端口提供 Streamable HTTP。
安装前请注意
这是模拟环境:不会真正发送邮件、短信或拨打电话,地址使用保留的示例域名,也不涉及网络访问或真实账户。但工具仍会在沙箱内写入、更新和删除记录,管理类工具还能重置状态或推进模拟时钟。除非启用暴露答案的参数,否则标准答案异常是隐藏的,测试助手时应保持关闭。

安装

在 SourceWeft 中

  1. 打开 控制台中的 SMB Sandbox,将其添加到工作区。
  2. 为需要使用其工具的对话启用该服务。

Desktop only,通过 STDIO。 STDIO 服务会启动本地进程,因此需要 SourceWeft 桌面宿主。

其他 MCP 客户端

参照 仓库 中的启动说明。

README

sandbox-mcp

sandbox-mcp is an open-source set of mock small-business MCP servers — CRM, inbox, calendar, invoicing and phone — backed by a realistic fictional company, so you can test AI agents end-to-end without touching real customer accounts.

Point Claude, Cursor, VS Code or any Model Context Protocol client at it and your agent gets a believable business to operate: leads waiting for a reply, quotes nobody followed up on, a duplicate contact, an overdue invoice, an unmatched check. Every change is validated like a real system would, recorded in an audit log, and stays inside the process. "Sending" an email or SMS only writes to an outbox.

[npm] [license] [node]

Quickstart

bash
npx @teamshift/sandbox-mcp --industry home-services

That starts an MCP server over stdio with a generated home-services company (seed 42, so it is the same company every time). Other industries: dental-clinic, marketing-agency.

bash
npx @teamshift/sandbox-mcp --industry dental-clinic --seed 7 --toolsets crm,inbox --state-out run.json

Connect your MCP client

Claude Code

bash
claude mcp add sandbox -- npx -y @teamshift/sandbox-mcp

Claude Desktop

Add to claude_desktop_config.json:

json
{  "mcpServers": {    "sandbox": {      "command": "npx",      "args": ["-y", "@teamshift/sandbox-mcp", "--industry", "home-services"]    }  }}

Cursor

Add to .cursor/mcp.json:

json
{  "mcpServers": {    "sandbox": {      "command": "npx",      "args": ["-y", "@teamshift/sandbox-mcp", "--industry", "home-services"]    }  }}

VS Code

Add to .vscode/mcp.json:

json
{  "servers": {    "sandbox": {      "type": "stdio",      "command": "npx",      "args": ["-y", "@teamshift/sandbox-mcp", "--industry", "home-services"]    }  }}

Then try a prompt like: "Check the inbox and the call log, reply to anything urgent, and make sure every open deal has a next step."

CLI options

OptionDefaultDescription
--industry <id>home-serviceshome-services, dental-clinic or marketing-agency
--seed <n>42Same seed, same company, byte for byte
--size <s>mediumsmall, medium or large
--messiness <x>1Multiplier for injected data problems; 0 gives clean data
--data <file>Load a dataset JSON (e.g. from @teamshift/fake-business) instead of generating
--toolsets <list>allComma list of crm,inbox,calendar,invoicing,phone,admin
--expose-answersoffExpose ground-truth anomalies as the sandbox://anomalies resource
--state-out <file>Write final dataset + audit log + outbox JSON on exit and on admin_save_state
--http <port>Serve Streamable HTTP at http://127.0.0.1:<port>/mcp instead of stdio

Logs go to stderr only, so stdout stays a clean MCP channel.

Tools

All tools are prefixed by system, take zod-validated input, return compact JSON, and carry MCP tool annotations (readOnlyHint, destructiveHint, idempotentHint, openWorldHint: false). Failures come back as tool errors with an actionable message (for example failed_precondition: Cannot move deal from "new" to "won". Allowed: qualified, quote-sent, lost.), never as a crash.

ToolsetTools
crmcrm_search_contacts, crm_get_contact, crm_create_contact, crm_update_contact, crm_merge_contacts, crm_list_customers, crm_get_customer, crm_list_employees, crm_list_leads, crm_update_lead, crm_list_deals, crm_get_deal, crm_update_deal, crm_list_quotes, crm_send_quote, crm_follow_up_quote, crm_list_tasks, crm_create_task, crm_complete_task
inboxinbox_list_threads, inbox_get_thread, inbox_send_email, inbox_send_sms, inbox_mark_thread_read
calendarcalendar_list_jobs, calendar_get_job, calendar_schedule_job, calendar_reschedule_job, calendar_cancel_job, calendar_complete_job
invoicinginvoicing_list_invoices, invoicing_get_invoice, invoicing_create_invoice, invoicing_send_reminder, invoicing_void_invoice, invoicing_record_payment, invoicing_list_payments, invoicing_match_payment
phonephone_list_calls, phone_log_call
adminadmin_get_company_policies, admin_get_company, admin_get_audit_log, admin_reset, admin_save_state, admin_advance_clock

Resources: sandbox://company, sandbox://policies, and sandbox://anomalies (only with --expose-answers, so the answers never leak to an agent under test).

Business rules it enforces

  • Foreign keys must exist; money is integer cents; dates are YYYY-MM-DD.
  • Deal stages follow the pipeline (new → qualified → quote-sent → negotiation → won/lost); won is final, lost needs a reason and can be reopened.
  • No payments on void, draft or fully paid invoices, and no overpayment. Invoices with payments cannot be voided.
  • Unmatched payments can only be matched to an invoice of the same customer.
  • Jobs cannot double-book a technician (unless allowOverlap), and rescheduling updates the job's times.
  • Email and SMS need a valid address or number, so missing contact info surfaces as a real error.
  • Messaging or calling a lead's contact records the first response on the lead.
  • A simulated clock starts at 09:00 company time on the dataset's "today" and advances one minute per change, so timestamps are deterministic.

Grading an agent

Run your agent against the sandbox with --state-out, then score the end state, not the transcript:

bash
npx @teamshift/sandbox-mcp --industry home-services --seed 42 --toolsets crm,inbox,calendar,invoicing,phone --state-out run.json

run.json contains:

  • dataset: the final state of every record (same schema as @teamshift/fake-business), including new events.
  • audit: every attempted change in order: { seq, at, tool, input, result: "ok" | "error", error?, changedIds }.
  • outbox: every email and SMS the agent "sent".

Compare dataset against the ground-truth anomalies from the original dataset (for example: was the duplicate contact merged, was the unmatched payment applied, did every stale deal get a nextAction?), and use audit to penalize errors, destructive actions or policy violations. Leave out the admin toolset so the agent cannot reset its own run.

Use the store directly in code

The same engine is exported as a typed library, which is how reference solutions and verifiers run:

ts
import { generate } from "@teamshift/fake-business";import { SandboxStore } from "@teamshift/sandbox-mcp";
const store = new SandboxStore(generate({ industry: "home-services", seed: 42 }));for (const quote of store.listQuotes({ notFollowedUp: true, limit: 100 }).items) {  store.followUpQuote({ quoteId: quote.id, body: "Just checking in on your quote. Any questions?" });}console.log(store.audit().length, store.outbox().length);const finalState = store.snapshot();

You can also embed the MCP server in your own process with createSandboxServer(store, { toolsets, exposeAnswers }) and any SDK transport.

FAQ

How do I test an MCP agent safely?

Give it a sandbox instead of production credentials. sandbox-mcp exposes the same kinds of tools a real CRM, inbox, calendar, invoicing and phone system would, over a fictional company, with no network access and no real accounts. Run the agent, save the state with --state-out, and check what it changed.

Does it send real emails?

No. Nothing leaves the process. inbox_send_email, inbox_send_sms, crm_follow_up_quote and invoicing_send_reminder append an outbound message to the thread and to the outbox, and all addresses use the reserved .example domain and 555-01xx numbers.

How do I reset the sandbox?

Call the admin_reset tool (the audit log is kept and records the reset, so graders can see it), or restart the server. The same --industry and --seed always regenerate the identical company. In code, call store.reset().

Where does the data come from?

From @teamshift/fake-business, a deterministic generator for realistic fictional small businesses, with labeled anomalies for scoring. You can also pass your own dataset with --data.

Can the agent see the answers?

Not by default. The ground-truth anomalies are only exposed as a resource with --expose-answers; they are always included in --state-out files for graders.

Does it support Streamable HTTP?

Yes: --http 8787 serves stateless Streamable HTTP at http://127.0.0.1:8787/mcp. All requests share one sandbox state.

License

Apache-2.0


Built by TeamShift — AI workers for small-business operations.

来源:packages/sandbox-mcp/README.md,提交 fca9085

工具

0
工具元数据尚未被收录。

版本历史

1
  1. v0.1.0最新Oct 4, 2026