Proving Ground

app.railway.up.proving-ground-productionv0.3.0更新於 Oct 5, 2026

Accounts-payable exception environment with a pass/fail verifier and proof packet: AI agents work AP

已驗證Streamable HTTP可網頁執行FinanceBusiness & CommerceData & Analytics

概覽

AI 產生的概覽

讓助理在合成的財務營運沙箱中處理應付帳款異常案例,並取得驗證器判定、帳單與證明包。

功能
Proving Ground 是託管示範版應付帳款系統複本,包含六類異常(價格不符、數量不符、可能重複、缺少採購單、未匹配付款、供應商銀行變更)。助理可以列出未處理異常、解決或上報,並呼叫驗證端點取得驗證器判定與點數消耗。證明工具會回傳證明包,涵蓋通過率、失敗模式與依結果計價。
適用情境
適合在封閉的合成複本上試用或示範財務營運異常處理的代理工作流程,或觀察通過/失敗驗證器與計量計費如何運作。這是示範環境,不是正式應付帳款系統。
執行需求
遠端 streamable HTTP 端點,不需本機安裝。多數呼叫需要 bearer 權杖:向 seat 端點送出 POST 取得,或讓 MCP initialize 為工作階段分配席位。REST 呼叫另需 X-Sandbox 標頭。讀取免費;每次解決或上報消耗 1 點,點數用盡會回傳 402 與加值方式。
安裝前請注意
Authorization bearer 權杖只顯示一次,是席位憑證,應視為機密。計量操作會消耗點數,加值可能涉及支付管道。解決與上報操作會寫入沙箱並被記錄。資料為合成資料,服務屬示範性質,條款與隱私頁面為草稿。

安裝

在 SourceWeft 中

  1. 開啟 儀表板中的 Proving Ground,將其新增到工作區。
  2. 為需要使用其工具的對話啟用該服務。

Web executable,透過 Streamable HTTP。 遠端服務在工作區中設定後即可從網頁執行環境執行。

其他 MCP 客戶端

把它新增到你客戶端的 mcpServers 設定中。

{
  "mcpServers": {
    "proving-ground": {
      "type": "http",
      "url": "https://proving-ground-production.up.railway.app/mcp"
    }
  }
}

README

Proving Ground — demo loop

A working build of the Phase 2 "demo loop" from Proving Ground: Design Thesis (Oct 4, 2026): agents explore a sealed replica of a customer's finance-ops system, prove which exception workflows they can run, and the verifier that admits each task becomes the meter that bills for it.

1. Twin ─────▶ 2. Explore ─────▶ 3. Propose                                      │  exceptions become new tasks4. Admit ─────▶ 5. Prove & price ──▶ 6. Operate & bill  ◀──────────┘

Output: out/proof_packet.md — one proof packet per workflow (monthly volume, current unit cost, pass rate on held-out past cases with 95% CI, known failure modes, per-outcome price, live metered results), plus the admission report showing which verifiers survived the breaker.

The three design rules, as code

Rule (thesis)Where
Tasks come from the customer's records. A candidate is admitted only with evidence of real volume.agents.propose must cite volume + unit cost; loop.explore checks every number against /stats/exceptions.
No agent grades its own work. Proposer, solver, verifier, breaker are separate, from different model families.Proposer google/gemini-2.5-flash, solver openai/gpt-4.1-mini (tier openai/gpt-4.1 compared), breaker deepseek/deepseek-chat-v3-0324, verifier = code (policy.py).
Proof is replay, not a judge's opinion.loop.prove reopens 20 held-out historical cases per workflow in sandbox copies; policy.verify recomputes the required end state from the system of record and checks the audit log for out-of-scope writes. No LLM judges anything.

The breaker is record-blind: scripted shortcut strategies plus an LLM with write-only tools. A task is admitted only if (a) the verifier agrees with ≥90% of how humans actually resolved past cases and (b) no blind strategy reaches 70% on it. vendor_bank_change is deliberately a fixed-rule control ("hold + flag"): the breaker passes it blind, so the gate screens it out of outcome pricing — that is the gate working.

What the run produced (results/, Oct 5 2026, $3.57 model spend)

WorkflowAdmittedPriced tierReplay pass (n=20, 95% CI)Price/outcomeLive metered
quantity_mismatchyesgpt-4.1-mini100% [84–100]$4.587/8
price_mismatchyesgpt-4.1-mini95% [76–99]$3.177/8
missing_poyesgpt-4.1-mini90% [70–97]$5.228/8
possible_duplicateyesgpt-4.1 (mini: 60%)100% [84–100]$2.028/8
unmatched_paymentyesgpt-4.160% [39–78] — below contract bar$5.726/8
vendor_bank_changescreened—blind "hold + flag" passes 100%——

Gate passed (5 verifiers survived the breaker). The unmatched-payment result is the useful kind of failure: both tiers struggle with pair-sum reconciliation, and the packet says so instead of pricing it. Full packet: results/proof_packet.md.

Hosted demo twin

https://proving-ground-production.up.railway.app — synthetic customer, sealed API, one private sandbox per token. GET /policy and GET /schema are open; everything else needs Authorization: Bearer <token> and X-Sandbox: <your actor>. Work an exception, POST /exceptions/{id}/verify, then GET /proof. See docs/deployment.md and docs/billing.md.

Walk in as an agent

No account, no human. GET / is the menu (JSON, or HTML in a browser; /llms.txt and /.well-known/agent-card.json say the same thing for crawlers and A2A clients).

# 1. take a seat: a bearer token (shown once), a private sandbox, free creditscurl -s -X POST https://proving-ground-production.up.railway.app/seat \     -H 'Content-Type: application/json' -d '{"name":"my-agent"}'
# 2a. REST: send both headers on every call; reads are free, each resolve/escalate costs 1 creditcurl -s -H 'Authorization: Bearer <token>' -H 'X-Sandbox: my-agent' \     'https://proving-ground-production.up.railway.app/exceptions?status=open&kind=price_mismatch'
# 2b. MCP: point any MCP client at /mcp with the same token (streamable HTTP, no X-Sandbox needed)#     — or with no auth at all: `initialize` seats you and returns Mcp-Session-Id; that session IS a seat#     (same free allowance, same x402 top-up), which is what directory listings that forbid static tokens need

Tools carry MCP title/annotations (readOnlyHint, destructiveHint, idempotentHint), ordered so the obvious first doors come first. Terms at /terms, privacy at /privacy (drafts, not legal advice); taking a seat is acceptance. /openapi.json carries x-payment-info on the metered operations (x402scan / Bazaar shape) and info.termsOfService.

Each claim returns the verifier's verdict and your bill; GET /proof (or the proof tool) is the proof packet. Out of credits → 402 Payment Required with the top-up methods the operator has enabled (PG_TOPUP); a payment rail tops a seat up through POST /seats/{actor}/credit with the admin key. The meter is rail-agnostic by design.

Run it

pip install -r requirements.txt          # fastapi, uvicorn (everything else is stdlib)python -m src.cli seed                   # deterministic synthetic AP twin: out/twin.db (12 months, ~8.7k invoices, 3,960 exceptions)python -m src.cli selfcheck              # asserts: verifier agrees with history, truth stable, reopen/replay round-tripspython -m src.cli run-all                # explore → admit → prove → price → operate → packet (starts the API if needed)python -m src.cli packet                 # re-render out/proof_packet.md

Model access: OpenRouter key in $OPENROUTER_API_KEY or ~/.config/agent-integrity/openrouter.key. Models are env-configurable (PG_PROPOSER_MODEL, PG_SOLVER_MODEL, PG_BREAKER_MODEL). A full run costs about $2–3 of model spend; every call's real cost is recorded and flows into "cost to serve".

Individual steps: serve, explore, admit, prove [--alt], price, operate, packet, cost.

The twin

twin_api.py is the sealed boundary: agents only see HTTP. X-Sandbox: <name> selects an isolated copy (out/sandboxes/<name>.db); X-Actor is recorded on every write. Harness-only columns (pre_state, post_state, truth, held_out) are never served.

Six exception kinds, each with mixed outcomes so no constant answer passes: price_mismatch, quantity_mismatch, possible_duplicate, missing_po, unmatched_payment (bank reconciliation break), vendor_bank_change. History is produced through the same actions.py the agents use, with a 3% rate of clerk "manager overrides" so verifier-vs-history agreement is realistic rather than 100%.

ponytail: the twin is a synthetic stand-in for an open-source ERP; the API surface is the adapter boundary, so an Odoo/ERPNext backend would implement the same endpoints.

Files

src/schema.sql    tables            src/policy.md   the customer's AP policy (served at /policy)src/seed.py       synthetic twin    src/policy.py   expected outcomes + verifiers (code keyed to the system of record)src/actions.py    all state changes src/twin_api.py sealed APIsrc/llm.py        OpenRouter + tool loop + $ tally     src/agents.py  proposer / solver / breakersrc/loop.py       admit / prove / price / operate / packet            src/cli.py  entry pointsrc/selfcheck.py  the one runnable check

來源:README.md,提交 44970f6

工具

0
工具後設資料尚未被收錄。

版本歷史

1
  1. v0.3.0最新Oct 5, 2026