Benchmark Optimization Loop

作者 affaan-mef648e01899bMIT275K 个星标收录于 2026年10月8日更新于 2026年10月8日仓库3天前更新

Convert 'make it faster' requests into a bounded measured optimization loop — baseline first, generate one-hypothesis variants, benchmark each against a correctness gate, and promote the fastest safe variant with reproducible commands. Use when asked to speed something up, try many variants, run recursive optimization, benchmark latency/throughput/cost, or pick the best implementation by repeated measured tests.

AI 生成的概览

把提速请求转化为有边界、可度量的优化循环,包含基线、变体、基准测试与晋级门槛。

功能
该技能定义了一套结构化的性能优化流程:在开始前必须先有基线、正确性门槛、度量指标和搜索预算。随后生成每次只验证一个假设的变体,在相同输入形态下进行基准测试,剔除不安全或不可复现的变体,并晋级最快且安全的变体。它还涵盖带运行台账和留出验证的递归或超参数搜索,并产出变体表格以及确切的命令与测量结果。
适用场景
当被要求让某物更快、尝试大量变体、执行递归优化,或对延迟、吞吐量、成本做基准测试时使用。也适用于通过反复实测挑选最佳实现的请求。
运行要求
该技能不附带脚本,仅为说明性指令。它假定智能体具备文件读写、编辑、命令行和搜索工具,并有一个可运行的操作、一个正确性测试以及度量所选指标的方法。

Benchmark Optimization Loop

Use this skill to convert "make it 20x faster" or "try 50 recursive optimizations" into a bounded measured loop that can actually improve a system.

Required Baseline

Do not optimize until these exist:

  • the operation being optimized;
  • the correctness gate that must stay green;
  • the metric: wall time, p95 latency, rows/sec, cost/run, memory, error rate;
  • the current baseline;
  • the search budget: max variants, max time, max spend, max data impact.

If the user asks for an unrealistic target, keep the ambition but make the loop bounded and measurable.

Loop

  1. Measure the baseline.
  2. Identify bottlenecks from evidence.
  3. Generate variants that test one hypothesis each.
  4. Run variants with the same input shape.
  5. Reject variants that fail correctness, safety, or reproducibility.
  6. Promote the fastest safe variant.
  7. Codify the winning path in a script, command, test, config, or doc.
  8. Rerun the baseline and winner to confirm the delta.

Variant Table

Track variants like this:

text
Variant | Hypothesis | Command | Time | Correct? | Notesbaseline | current path | npm run job | 120s | yes | stablebatch-500 | fewer round trips | npm run job -- --batch 500 | 42s | yes | winnerparallel-8 | more workers | npm run job -- --workers 8 | 31s | no | rate limited

Recursive Search

For recursive or hyperparameter work:

  • persist every run to a ledger;
  • compare against the prior accepted winner, not only the previous run;
  • keep a holdout or replay check;
  • stop when improvement is within noise, correctness fails, cost exceeds the budget, or the search starts changing more variables than it can explain.

Use phrases like "best measured safe variant" instead of "global optimum" unless the search space was actually exhaustive.

Promotion Gate

A variant cannot become the new default until:

  • correctness tests pass;
  • the performance delta is repeated or explained;
  • rollback is obvious;
  • the change is encoded in source control or a durable runbook;
  • the final summary includes exact commands and measurements.

来源与署名

来源:affaan-m/ecc位于skills/benchmark-optimization-loop提交ef648e0

许可证: MIT

内容归原作者所有。SourceWeft 从公开仓库中收录这些内容。

举报或申请下架