Proving Ground

app.railway.up.proving-ground-productionv0.3.0Updated Oct 5, 2026

Accounts-payable exception environment with a pass/fail verifier and proof packet: AI agents work AP

VerifiedStreamable HTTPWeb executableBusiness & CommerceFinanceData & Analytics

Overview

AI-generated overview

Lets an assistant work accounts-payable exception cases in a synthetic finance-ops sandbox and get a verifier verdict, bill, and proof packet.

What it does
Proving Ground is a hosted demo twin of an accounts-payable system with six exception kinds (price mismatch, quantity mismatch, possible duplicate, missing PO, unmatched payment, vendor bank change). An assistant can list open exceptions, resolve or escalate them, and call a verify endpoint that returns the verifier's verdict and the credit cost. A proof tool returns a proof packet covering pass rates, failure modes, and per-outcome pricing.
When to use it
Use it to try or demonstrate agent workflows for finance-ops exception handling against a sealed synthetic replica, or to inspect how a pass/fail verifier and metered billing behave. It is a demo environment, not a production AP system.
Requirements
Remote streamable HTTP endpoint; no local install. A bearer token is needed for most calls: get one by POSTing to the seat endpoint, or let MCP initialize seat the session. REST calls also need an X-Sandbox header. Reads are free; each resolve or escalate costs one credit, and running out returns 402 with top-up options.
Before you install
The Authorization bearer token is shown once and is the seat credential; treat it as a secret. Metered actions spend credits, and top-up can involve a payment rail. Resolve and escalate actions write to the sandbox and are recorded. The data is synthetic, and the service is a demo with draft terms and privacy pages.

Installation

In SourceWeft

  1. Open Proving Ground in the dashboard and add it to a workspace.
  2. Enable the server for the chats that should use its tools.

Web executable via Streamable HTTP. Remote servers run from the web runtime once configured in a workspace.

Other MCP clients

Add this to your client's mcpServers config.

{
  "mcpServers": {
    "proving-ground": {
      "type": "http",
      "url": "https://proving-ground-production.up.railway.app/mcp"
    }
  }
}

README

Proving Ground — demo loop

A working build of the Phase 2 "demo loop" from Proving Ground: Design Thesis (Oct 4, 2026): agents explore a sealed replica of a customer's finance-ops system, prove which exception workflows they can run, and the verifier that admits each task becomes the meter that bills for it.

1. Twin ─────▶ 2. Explore ─────▶ 3. Propose                                      │  exceptions become new tasks4. Admit ─────▶ 5. Prove & price ──▶ 6. Operate & bill  ◀──────────┘

Output: out/proof_packet.md — one proof packet per workflow (monthly volume, current unit cost, pass rate on held-out past cases with 95% CI, known failure modes, per-outcome price, live metered results), plus the admission report showing which verifiers survived the breaker.

The three design rules, as code

Rule (thesis)Where
Tasks come from the customer's records. A candidate is admitted only with evidence of real volume.agents.propose must cite volume + unit cost; loop.explore checks every number against /stats/exceptions.
No agent grades its own work. Proposer, solver, verifier, breaker are separate, from different model families.Proposer google/gemini-2.5-flash, solver openai/gpt-4.1-mini (tier openai/gpt-4.1 compared), breaker deepseek/deepseek-chat-v3-0324, verifier = code (policy.py).
Proof is replay, not a judge's opinion.loop.prove reopens 20 held-out historical cases per workflow in sandbox copies; policy.verify recomputes the required end state from the system of record and checks the audit log for out-of-scope writes. No LLM judges anything.

The breaker is record-blind: scripted shortcut strategies plus an LLM with write-only tools. A task is admitted only if (a) the verifier agrees with ≥90% of how humans actually resolved past cases and (b) no blind strategy reaches 70% on it. vendor_bank_change is deliberately a fixed-rule control ("hold + flag"): the breaker passes it blind, so the gate screens it out of outcome pricing — that is the gate working.

What the run produced (results/, Oct 5 2026, $3.57 model spend)

WorkflowAdmittedPriced tierReplay pass (n=20, 95% CI)Price/outcomeLive metered
quantity_mismatchyesgpt-4.1-mini100% [84–100]$4.587/8
price_mismatchyesgpt-4.1-mini95% [76–99]$3.177/8
missing_poyesgpt-4.1-mini90% [70–97]$5.228/8
possible_duplicateyesgpt-4.1 (mini: 60%)100% [84–100]$2.028/8
unmatched_paymentyesgpt-4.160% [39–78] — below contract bar$5.726/8
vendor_bank_changescreened—blind "hold + flag" passes 100%——

Gate passed (5 verifiers survived the breaker). The unmatched-payment result is the useful kind of failure: both tiers struggle with pair-sum reconciliation, and the packet says so instead of pricing it. Full packet: results/proof_packet.md.

Hosted demo twin

https://proving-ground-production.up.railway.app — synthetic customer, sealed API, one private sandbox per token. GET /policy and GET /schema are open; everything else needs Authorization: Bearer <token> and X-Sandbox: <your actor>. Work an exception, POST /exceptions/{id}/verify, then GET /proof. See docs/deployment.md and docs/billing.md.

Walk in as an agent

No account, no human. GET / is the menu (JSON, or HTML in a browser; /llms.txt and /.well-known/agent-card.json say the same thing for crawlers and A2A clients).

# 1. take a seat: a bearer token (shown once), a private sandbox, free creditscurl -s -X POST https://proving-ground-production.up.railway.app/seat \     -H 'Content-Type: application/json' -d '{"name":"my-agent"}'
# 2a. REST: send both headers on every call; reads are free, each resolve/escalate costs 1 creditcurl -s -H 'Authorization: Bearer <token>' -H 'X-Sandbox: my-agent' \     'https://proving-ground-production.up.railway.app/exceptions?status=open&kind=price_mismatch'
# 2b. MCP: point any MCP client at /mcp with the same token (streamable HTTP, no X-Sandbox needed)#     — or with no auth at all: `initialize` seats you and returns Mcp-Session-Id; that session IS a seat#     (same free allowance, same x402 top-up), which is what directory listings that forbid static tokens need

Tools carry MCP title/annotations (readOnlyHint, destructiveHint, idempotentHint), ordered so the obvious first doors come first. Terms at /terms, privacy at /privacy (drafts, not legal advice); taking a seat is acceptance. /openapi.json carries x-payment-info on the metered operations (x402scan / Bazaar shape) and info.termsOfService.

Each claim returns the verifier's verdict and your bill; GET /proof (or the proof tool) is the proof packet. Out of credits → 402 Payment Required with the top-up methods the operator has enabled (PG_TOPUP); a payment rail tops a seat up through POST /seats/{actor}/credit with the admin key. The meter is rail-agnostic by design.

Run it

pip install -r requirements.txt          # fastapi, uvicorn (everything else is stdlib)python -m src.cli seed                   # deterministic synthetic AP twin: out/twin.db (12 months, ~8.7k invoices, 3,960 exceptions)python -m src.cli selfcheck              # asserts: verifier agrees with history, truth stable, reopen/replay round-tripspython -m src.cli run-all                # explore → admit → prove → price → operate → packet (starts the API if needed)python -m src.cli packet                 # re-render out/proof_packet.md

Model access: OpenRouter key in $OPENROUTER_API_KEY or ~/.config/agent-integrity/openrouter.key. Models are env-configurable (PG_PROPOSER_MODEL, PG_SOLVER_MODEL, PG_BREAKER_MODEL). A full run costs about $2–3 of model spend; every call's real cost is recorded and flows into "cost to serve".

Individual steps: serve, explore, admit, prove [--alt], price, operate, packet, cost.

The twin

twin_api.py is the sealed boundary: agents only see HTTP. X-Sandbox: <name> selects an isolated copy (out/sandboxes/<name>.db); X-Actor is recorded on every write. Harness-only columns (pre_state, post_state, truth, held_out) are never served.

Six exception kinds, each with mixed outcomes so no constant answer passes: price_mismatch, quantity_mismatch, possible_duplicate, missing_po, unmatched_payment (bank reconciliation break), vendor_bank_change. History is produced through the same actions.py the agents use, with a 3% rate of clerk "manager overrides" so verifier-vs-history agreement is realistic rather than 100%.

ponytail: the twin is a synthetic stand-in for an open-source ERP; the API surface is the adapter boundary, so an Odoo/ERPNext backend would implement the same endpoints.

Files

src/schema.sql    tables            src/policy.md   the customer's AP policy (served at /policy)src/seed.py       synthetic twin    src/policy.py   expected outcomes + verifiers (code keyed to the system of record)src/actions.py    all state changes src/twin_api.py sealed APIsrc/llm.py        OpenRouter + tool loop + $ tally     src/agents.py  proposer / solver / breakersrc/loop.py       admit / prove / price / operate / packet            src/cli.py  entry pointsrc/selfcheck.py  the one runnable check

Source: README.md at commit 44970f6

Tools

0
Tool metadata has not been indexed yet.

Version history

1
  1. v0.3.0LatestOct 5, 2026