Benchmark Optimization Loop

作者 affaan-mef648e01899bMIT275K 個星標收錄於 2026年10月8日更新於 2026年10月8日儲存庫3 天前更新

Convert 'make it faster' requests into a bounded measured optimization loop — baseline first, generate one-hypothesis variants, benchmark each against a correctness gate, and promote the fastest safe variant with reproducible commands. Use when asked to speed something up, try many variants, run recursive optimization, benchmark latency/throughput/cost, or pick the best implementation by repeated measured tests.

AI 產生的概覽

把加速需求轉化為有界、可量測的最佳化迴圈,包含基準、變體、效能測試與晉升門檻。

功能
此技能定義一套結構化的效能最佳化流程:開始前必須先有基準、正確性門檻、量測指標與搜尋預算。接著產生每次只驗證一個假設的變體,在相同輸入形態下進行效能測試,剔除不安全或無法重現的變體,並晉升最快且安全的變體。它也涵蓋搭配執行紀錄與留出驗證的遞迴或超參數搜尋,並產出變體表格以及確切的指令與量測結果。
適用情境
當被要求讓某個東西更快、嘗試大量變體、執行遞迴最佳化,或對延遲、輸送量、成本做效能測試時使用。也適用於透過反覆實測挑選最佳實作的請求。
執行需求
此技能不隨附指令碼,僅為說明性指示。它假定代理具備檔案讀寫、編輯、命令列與搜尋工具,並有一個可執行的操作、一個正確性測試以及量測所選指標的方法。

Benchmark Optimization Loop

Use this skill to convert "make it 20x faster" or "try 50 recursive optimizations" into a bounded measured loop that can actually improve a system.

Required Baseline

Do not optimize until these exist:

  • the operation being optimized;
  • the correctness gate that must stay green;
  • the metric: wall time, p95 latency, rows/sec, cost/run, memory, error rate;
  • the current baseline;
  • the search budget: max variants, max time, max spend, max data impact.

If the user asks for an unrealistic target, keep the ambition but make the loop bounded and measurable.

Loop

  1. Measure the baseline.
  2. Identify bottlenecks from evidence.
  3. Generate variants that test one hypothesis each.
  4. Run variants with the same input shape.
  5. Reject variants that fail correctness, safety, or reproducibility.
  6. Promote the fastest safe variant.
  7. Codify the winning path in a script, command, test, config, or doc.
  8. Rerun the baseline and winner to confirm the delta.

Variant Table

Track variants like this:

text
Variant | Hypothesis | Command | Time | Correct? | Notesbaseline | current path | npm run job | 120s | yes | stablebatch-500 | fewer round trips | npm run job -- --batch 500 | 42s | yes | winnerparallel-8 | more workers | npm run job -- --workers 8 | 31s | no | rate limited

Recursive Search

For recursive or hyperparameter work:

  • persist every run to a ledger;
  • compare against the prior accepted winner, not only the previous run;
  • keep a holdout or replay check;
  • stop when improvement is within noise, correctness fails, cost exceeds the budget, or the search starts changing more variables than it can explain.

Use phrases like "best measured safe variant" instead of "global optimum" unless the search space was actually exhaustive.

Promotion Gate

A variant cannot become the new default until:

  • correctness tests pass;
  • the performance delta is repeated or explained;
  • rollback is obvious;
  • the change is encoded in source control or a durable runbook;
  • the final summary includes exact commands and measurements.

來源與署名

來源:affaan-m/ecc位於skills/benchmark-optimization-loop提交ef648e0

授權條款: MIT

內容歸原作者所有。SourceWeft 從公開儲存庫中收錄這些內容。

檢舉或申請下架