Ponytail Gain
Display this scoreboard when invoked. One-shot: do NOT change mode, write flag files, or persist anything.
The figures are the published agentic benchmark of Ponytail 5: headless Claude
Code (Opus 5.5, default effort) on 39 tasks (feature tickets in a real FastAPI +
React repo, bug fixes, security and privacy cases, small apps), 5 runs each,
against the same agent without the skill. 18 of the tasks have hidden checks
for correctness and safety. Each figure is the geometric mean of the per-task
medians. They are measured, not computed from the current repo.
Source: benchmarks/results/2026-10-07-agentic.md and the README.
Scoreboard
Render plain ASCII bars. The bar length shows ponytail as a share of the no-skill baseline; the label carries the exact figure:
Honesty boundary
These are benchmark averages, not this repo. NEVER print a per-repo savings
number ("you saved X lines/tokens here"): the unbuilt version was never
written, so there is no real baseline to subtract from in a live repo. The
only real per-repo figures come from /ponytail-debt (a counted ledger), and
this card points there instead of inventing one.
Boundaries
One-shot display. Edits nothing, changes no mode. "stop ponytail" or "normal mode": revert.
