Benchmark Optimization Loop

by affaan-mef648e01899bMIT275K starsListed Oct 8, 2026Updated Oct 8, 2026Repository updated 3 days ago

Convert 'make it faster' requests into a bounded measured optimization loop — baseline first, generate one-hypothesis variants, benchmark each against a correctness gate, and promote the fastest safe variant with reproducible commands. Use when asked to speed something up, try many variants, run recursive optimization, benchmark latency/throughput/cost, or pick the best implementation by repeated measured tests.

AI-generated overview

Turns speed-up requests into a bounded, measured optimization loop with baselines, variants, benchmarks and a promotion gate.

What it does
This skill defines a structured process for performance optimization: it requires a baseline, a correctness gate, a metric and a search budget before any work starts. It then generates one-hypothesis variants, benchmarks them on the same input shape, rejects unsafe or non-reproducible ones, and promotes the fastest safe variant. It also covers recursive or hyperparameter search with a run ledger and holdout checks, and produces a variant table plus exact commands and measurements.
When to use it
Use it when asked to make something faster, to try many variants, to run recursive optimization, or to benchmark latency, throughput or cost. It fits requests to pick the best implementation through repeated measured tests.
Requirements
No scripts ship with the skill; it is instructions only. It assumes an agent with file read/write, edit, shell and search tools, plus a runnable operation, a correctness test and a way to measure the chosen metric.

Benchmark Optimization Loop

Use this skill to convert "make it 20x faster" or "try 50 recursive optimizations" into a bounded measured loop that can actually improve a system.

Required Baseline

Do not optimize until these exist:

  • the operation being optimized;
  • the correctness gate that must stay green;
  • the metric: wall time, p95 latency, rows/sec, cost/run, memory, error rate;
  • the current baseline;
  • the search budget: max variants, max time, max spend, max data impact.

If the user asks for an unrealistic target, keep the ambition but make the loop bounded and measurable.

Loop

  1. Measure the baseline.
  2. Identify bottlenecks from evidence.
  3. Generate variants that test one hypothesis each.
  4. Run variants with the same input shape.
  5. Reject variants that fail correctness, safety, or reproducibility.
  6. Promote the fastest safe variant.
  7. Codify the winning path in a script, command, test, config, or doc.
  8. Rerun the baseline and winner to confirm the delta.

Variant Table

Track variants like this:

text
Variant | Hypothesis | Command | Time | Correct? | Notesbaseline | current path | npm run job | 120s | yes | stablebatch-500 | fewer round trips | npm run job -- --batch 500 | 42s | yes | winnerparallel-8 | more workers | npm run job -- --workers 8 | 31s | no | rate limited

Recursive Search

For recursive or hyperparameter work:

  • persist every run to a ledger;
  • compare against the prior accepted winner, not only the previous run;
  • keep a holdout or replay check;
  • stop when improvement is within noise, correctness fails, cost exceeds the budget, or the search starts changing more variables than it can explain.

Use phrases like "best measured safe variant" instead of "global optimum" unless the search space was actually exhaustive.

Promotion Gate

A variant cannot become the new default until:

  • correctness tests pass;
  • the performance delta is repeated or explained;
  • rollback is obvious;
  • the change is encoded in source control or a durable runbook;
  • the final summary includes exact commands and measurements.

Source and attribution

Source:affaan-m/eccinskills/benchmark-optimization-loopat commitef648e0

License: MIT

Content belongs to its original authors. SourceWeft indexes it from a public repository.

Report or request removal