
SMB Sandbox
io.github.teamshift-iov0.1.0Updated Oct 4, 2026
Mock small-business CRM, inbox, calendar, invoicing and phone tools for testing AI agents safely.
Overview
Runs a mock small-business backend (CRM, inbox, calendar, invoicing, phone) so an assistant can be tested end-to-end without touching real accounts.
- What it does
- Starts a local MCP server over stdio with a generated fictional company and exposes tool sets for CRM, inbox, calendar, invoicing, phone and admin, covering contacts, leads, deals, quotes, tasks, email and SMS, jobs, invoices and payments. Changes are validated against business rules, recorded in an audit log, and kept inside the process; sending a message only appends to an outbox. State can be written to a JSON file for grading, and the same engine is available as a typed library.
- When to use it
- Use it to evaluate or benchmark an AI agent that operates business software, to reproduce agent runs deterministically with a fixed industry and seed, or to develop agent workflows without production credentials. It is not a real CRM, inbox or billing system.
- Requirements
- Node.js with npx to run the npm package @teamshift/sandbox-mcp locally over stdio; a desktop MCP client such as Claude Code, Claude Desktop, Cursor or VS Code. No accounts, API keys or environment variables are declared. Optional flags select industry, seed, size, messiness, tool sets, a dataset file, state output, and Streamable HTTP on a local port.
Installation
In SourceWeft
- Open SMB Sandbox in the dashboard and add it to a workspace.
- Enable the server for the chats that should use its tools.
Desktop only via STDIO. STDIO servers start a local process, so they need the SourceWeft desktop host.
Other MCP clients
Follow the launch instructions in the repository.
README
sandbox-mcp
sandbox-mcp is an open-source set of mock small-business MCP servers — CRM, inbox, calendar, invoicing and phone — backed by a realistic fictional company, so you can test AI agents end-to-end without touching real customer accounts.
Point Claude, Cursor, VS Code or any Model Context Protocol client at it and your agent gets a believable business to operate: leads waiting for a reply, quotes nobody followed up on, a duplicate contact, an overdue invoice, an unmatched check. Every change is validated like a real system would, recorded in an audit log, and stays inside the process. "Sending" an email or SMS only writes to an outbox.
Quickstart
That starts an MCP server over stdio with a generated home-services company (seed 42, so it is the same company every time). Other industries: dental-clinic, marketing-agency.
Connect your MCP client
Claude Code
Claude Desktop
Add to claude_desktop_config.json:
Cursor
Add to .cursor/mcp.json:
VS Code
Add to .vscode/mcp.json:
Then try a prompt like: "Check the inbox and the call log, reply to anything urgent, and make sure every open deal has a next step."
CLI options
Logs go to stderr only, so stdout stays a clean MCP channel.
Tools
All tools are prefixed by system, take zod-validated input, return compact JSON, and carry MCP tool annotations (readOnlyHint, destructiveHint, idempotentHint, openWorldHint: false). Failures come back as tool errors with an actionable message (for example failed_precondition: Cannot move deal from "new" to "won". Allowed: qualified, quote-sent, lost.), never as a crash.
Resources: sandbox://company, sandbox://policies, and sandbox://anomalies (only with --expose-answers, so the answers never leak to an agent under test).
Business rules it enforces
- Foreign keys must exist; money is integer cents; dates are
YYYY-MM-DD. - Deal stages follow the pipeline (
new → qualified → quote-sent → negotiation → won/lost);wonis final,lostneeds a reason and can be reopened. - No payments on void, draft or fully paid invoices, and no overpayment. Invoices with payments cannot be voided.
- Unmatched payments can only be matched to an invoice of the same customer.
- Jobs cannot double-book a technician (unless
allowOverlap), and rescheduling updates the job's times. - Email and SMS need a valid address or number, so missing contact info surfaces as a real error.
- Messaging or calling a lead's contact records the first response on the lead.
- A simulated clock starts at 09:00 company time on the dataset's "today" and advances one minute per change, so timestamps are deterministic.
Grading an agent
Run your agent against the sandbox with --state-out, then score the end state, not the transcript:
run.json contains:
dataset: the final state of every record (same schema as@teamshift/fake-business), including newevents.audit: every attempted change in order:{ seq, at, tool, input, result: "ok" | "error", error?, changedIds }.outbox: every email and SMS the agent "sent".
Compare dataset against the ground-truth anomalies from the original dataset (for example: was the duplicate contact merged, was the unmatched payment applied, did every stale deal get a nextAction?), and use audit to penalize errors, destructive actions or policy violations. Leave out the admin toolset so the agent cannot reset its own run.
Use the store directly in code
The same engine is exported as a typed library, which is how reference solutions and verifiers run:
You can also embed the MCP server in your own process with createSandboxServer(store, { toolsets, exposeAnswers }) and any SDK transport.
FAQ
How do I test an MCP agent safely?
Give it a sandbox instead of production credentials. sandbox-mcp exposes the same kinds of tools a real CRM, inbox, calendar, invoicing and phone system would, over a fictional company, with no network access and no real accounts. Run the agent, save the state with --state-out, and check what it changed.
Does it send real emails?
No. Nothing leaves the process. inbox_send_email, inbox_send_sms, crm_follow_up_quote and invoicing_send_reminder append an outbound message to the thread and to the outbox, and all addresses use the reserved .example domain and 555-01xx numbers.
How do I reset the sandbox?
Call the admin_reset tool (the audit log is kept and records the reset, so graders can see it), or restart the server. The same --industry and --seed always regenerate the identical company. In code, call store.reset().
Where does the data come from?
From @teamshift/fake-business, a deterministic generator for realistic fictional small businesses, with labeled anomalies for scoring. You can also pass your own dataset with --data.
Can the agent see the answers?
Not by default. The ground-truth anomalies are only exposed as a resource with --expose-answers; they are always included in --state-out files for graders.
Does it support Streamable HTTP?
Yes: --http 8787 serves stateless Streamable HTTP at http://127.0.0.1:8787/mcp. All requests share one sandbox state.
License
Apache-2.0
Built by TeamShift — AI workers for small-business operations.
Source: packages/sandbox-mcp/README.md at commit fca9085
Tools
0Version history
1- v0.1.0LatestOct 4, 2026
