
Proving Ground
app.railway.up.proving-ground-productionv0.3.0Updated Oct 5, 2026
Accounts-payable exception environment with a pass/fail verifier and proof packet: AI agents work AP
Overview
Lets an assistant work accounts-payable exception cases in a synthetic finance-ops sandbox and get a verifier verdict, bill, and proof packet.
- What it does
- Proving Ground is a hosted demo twin of an accounts-payable system with six exception kinds (price mismatch, quantity mismatch, possible duplicate, missing PO, unmatched payment, vendor bank change). An assistant can list open exceptions, resolve or escalate them, and call a verify endpoint that returns the verifier's verdict and the credit cost. A proof tool returns a proof packet covering pass rates, failure modes, and per-outcome pricing.
- When to use it
- Use it to try or demonstrate agent workflows for finance-ops exception handling against a sealed synthetic replica, or to inspect how a pass/fail verifier and metered billing behave. It is a demo environment, not a production AP system.
- Requirements
- Remote streamable HTTP endpoint; no local install. A bearer token is needed for most calls: get one by POSTing to the seat endpoint, or let MCP initialize seat the session. REST calls also need an X-Sandbox header. Reads are free; each resolve or escalate costs one credit, and running out returns 402 with top-up options.
Installation
In SourceWeft
- Open Proving Ground in the dashboard and add it to a workspace.
- Enable the server for the chats that should use its tools.
Web executable via Streamable HTTP. Remote servers run from the web runtime once configured in a workspace.
Other MCP clients
Add this to your client's mcpServers config.
{
"mcpServers": {
"proving-ground": {
"type": "http",
"url": "https://proving-ground-production.up.railway.app/mcp"
}
}
}README
Proving Ground — demo loop
A working build of the Phase 2 "demo loop" from Proving Ground: Design Thesis (Oct 4, 2026): agents explore a sealed replica of a customer's finance-ops system, prove which exception workflows they can run, and the verifier that admits each task becomes the meter that bills for it.
Output: out/proof_packet.md — one proof packet per workflow (monthly volume, current unit cost,
pass rate on held-out past cases with 95% CI, known failure modes, per-outcome price, live metered
results), plus the admission report showing which verifiers survived the breaker.
The three design rules, as code
The breaker is record-blind: scripted shortcut strategies plus an LLM with write-only tools. A task is
admitted only if (a) the verifier agrees with ≥90% of how humans actually resolved past cases and (b) no
blind strategy reaches 70% on it. vendor_bank_change is deliberately a fixed-rule control ("hold + flag"):
the breaker passes it blind, so the gate screens it out of outcome pricing — that is the gate working.
What the run produced (results/, Oct 5 2026, $3.57 model spend)
Gate passed (5 verifiers survived the breaker). The unmatched-payment result is the useful kind of
failure: both tiers struggle with pair-sum reconciliation, and the packet says so instead of pricing it.
Full packet: results/proof_packet.md.
Hosted demo twin
https://proving-ground-production.up.railway.app — synthetic customer, sealed API, one private sandbox
per token. GET /policy and GET /schema are open; everything else needs Authorization: Bearer <token>
and X-Sandbox: <your actor>. Work an exception, POST /exceptions/{id}/verify, then GET /proof.
See docs/deployment.md and docs/billing.md.
Walk in as an agent
No account, no human. GET / is the menu (JSON, or HTML in a browser; /llms.txt and
/.well-known/agent-card.json say the same thing for crawlers and A2A clients).
Tools carry MCP title/annotations (readOnlyHint, destructiveHint, idempotentHint), ordered so the obvious first
doors come first. Terms at /terms, privacy at /privacy (drafts, not legal advice); taking a seat is acceptance.
/openapi.json carries x-payment-info on the metered operations (x402scan / Bazaar shape) and info.termsOfService.
Each claim returns the verifier's verdict and your bill; GET /proof (or the proof tool) is the
proof packet. Out of credits → 402 Payment Required with the top-up methods the operator has
enabled (PG_TOPUP); a payment rail tops a seat up through POST /seats/{actor}/credit with the
admin key. The meter is rail-agnostic by design.
Run it
Model access: OpenRouter key in $OPENROUTER_API_KEY or ~/.config/agent-integrity/openrouter.key.
Models are env-configurable (PG_PROPOSER_MODEL, PG_SOLVER_MODEL, PG_BREAKER_MODEL). A full run
costs about $2–3 of model spend; every call's real cost is recorded and flows into "cost to serve".
Individual steps: serve, explore, admit, prove [--alt], price, operate, packet, cost.
The twin
twin_api.py is the sealed boundary: agents only see HTTP. X-Sandbox: <name> selects an isolated
copy (out/sandboxes/<name>.db); X-Actor is recorded on every write. Harness-only columns
(pre_state, post_state, truth, held_out) are never served.
Six exception kinds, each with mixed outcomes so no constant answer passes: price_mismatch,
quantity_mismatch, possible_duplicate, missing_po, unmatched_payment (bank reconciliation break),
vendor_bank_change. History is produced through the same actions.py the agents use, with a 3% rate of
clerk "manager overrides" so verifier-vs-history agreement is realistic rather than 100%.
ponytail: the twin is a synthetic stand-in for an open-source ERP; the API surface is the adapter
boundary, so an Odoo/ERPNext backend would implement the same endpoints.
Files
Source: README.md at commit 44970f6
Tools
0Version history
1- v0.3.0LatestOct 5, 2026

