Ponytail Gain

dietrichgebert/ponytail/.openclaw/skills/ponytail-gain

by dietrichgebertb088b2df6e08MIT158K starsListed Oct 8, 2026Updated Oct 8, 2026Repository updated today

Show ponytail measured impact as a scoreboard: less code, less cost, more speed, from the agentic benchmark averages. One-shot display.

Instructions onlyProductivity & Workflow
AI-generated overview

Displays a one-shot ASCII scoreboard of Ponytail's published agentic benchmark results.

What it does
Renders a plain ASCII scoreboard showing Ponytail's measured benchmark impact: lines of code, output tokens, cost, time, hidden checks passed, and tests where the logic needs one, each compared against a no-skill baseline. It presents benchmark averages from a published agentic benchmark of 39 tasks with 5 runs each, and points to related per-repo commands. It is a one-shot display that changes no mode and persists nothing.
When to use it
Use when someone wants to see Ponytail's measured benchmark gains as a scoreboard. It is intended for a quick one-shot display rather than ongoing tracking or per-repository measurement.
Requirements
No scripts, tools, packages, or credentials; it only renders text. It cites benchmark source files and a homepage but does not require network access to display the scoreboard.

Ponytail Gain

Display this scoreboard when invoked. One-shot: do NOT change mode, write flag files, or persist anything.

The figures are the published agentic benchmark of Ponytail 5: headless Claude Code (Opus 5.5, default effort) on 39 tasks (feature tickets in a real FastAPI + React repo, bug fixes, security and privacy cases, small apps), 5 runs each, against the same agent without the skill. 18 of the tasks have hidden checks for correctness and safety. Each figure is the geometric mean of the per-task medians. They are measured, not computed from the current repo. Source: benchmarks/results/2026-10-07-agentic.md and the README.

Scoreboard

Render plain ASCII bars. The bar length shows ponytail as a share of the no-skill baseline; the label carries the exact figure:

  ponytail gain           benchmark · 39 tasks × 5 runs · Opus 5.5
  no-skill        ████████████████████  100%  Lines of code   █████████···········   47%   ▼ 53%  Output tokens   ███████████·········   55%   ▼ 45%  Cost            ███████████████·····   74%   ▼ 26%  Time            ████████████········   59%   ▼ 41%  Hidden checks passed   97%  (no-skill 96%)  Tests where the logic needs one   98%  (no-skill 68%)
  This repo:  /ponytail-debt  (shortcuts you deferred)              /ponytail-audit (what's still cuttable)

Honesty boundary

These are benchmark averages, not this repo. NEVER print a per-repo savings number ("you saved X lines/tokens here"): the unbuilt version was never written, so there is no real baseline to subtract from in a live repo. The only real per-repo figures come from /ponytail-debt (a counted ledger), and this card points there instead of inventing one.

Boundaries

One-shot display. Edits nothing, changes no mode. "stop ponytail" or "normal mode": revert.

Source and attribution

Source:dietrichgebert/ponytailin.openclaw/skills/ponytail-gainat commitb088b2d

License: MIT

Content belongs to its original authors. SourceWeft indexes it from a public repository.

Report or request removal

More from dietrichgebert/ponytail

Ponytail Review

DietrichGebert

Reviews a code change for correctness, safety, scale, tests, speed and unnecessary code, and reports findings.

Software Development158Kupdated today

Ponytail Help

DietrichGebert

Quick reference for ponytail levels, skills and commands. One-shot display. Use for /ponytail-help, "ponytail help", "how do I use ponytail".

Awaiting classification158Kupdated today

Ponytail Debt

DietrichGebert

Scans a repository for shortcut: and ponytail: comment markers and reports them as a debt ledger.

Software Development158Kupdated today

Ponytail Audit

DietrichGebert

Quality audit of a whole repo: bugs, security holes, what breaks under real load, risky code without tests, slow paths, and what to delete, merge or split. Ranked, each finding explained in plain English. One-shot report, changes nothing. Use for "audit this codebase", "review the whole repo", "find bloat", "what can I delete", /ponytail-audit.

Awaiting classification158Kupdated today

Ponytail

DietrichGebert

Lazy senior dev mode: the smallest change that fully solves the task, and a reply a busy human understands in one read. Use on any coding task (writing, fixing, refactoring, reviewing, choosing dependencies) and when the user says "ponytail", "be lazy", "simplest solution", "yagni", or complains about over-engineering or bloat. Levels: lite, full (default), ultra.

Awaiting classification158Kupdated today

Ponytail Review

dietrichgebert

Reviews a code diff for bugs, security, scale, missing tests, speed and lean code, and reports findings.

Software Development158Kupdated today