Ce Retune

everyinc/compound-engineering-plugin/skills/ce-retune

作者 everyinc67035e931c5c無授權條款25K 個星標收錄於 2026年10月8日更新於 2026年10月8日儲存庫今天更新

Retune a skill corpus for a new model, measurement-first: mine the run archive for a baseline, establish a noise floor, audit the corpus adversarially, then cut in measured passes until a pre-registered bar clears. Requires a benchmark harness that can A/B two builds of the corpus; refuses without one.

僅含說明AI & Agents
AI 產生的概覽

以量測為先,為新模型重新調校技能語料庫:挖掘基準、建立雜訊底線、對抗式稽核並依量測結果裁剪。

功能
此技能指導以量測為先的方式,為目標模型重新調校技能語料庫。它先挖掘執行封存檔以建立基準,用兩份相同的語料庫副本執行以確定雜訊底線,再由獨立代理對語料庫進行對抗式稽核(一方提出裁剪、一方辯護),接著進行外科式裁剪並持續量測,直到預先登記的達標門檻被滿足。產出是帶有可歸因刪除項的重新調校語料庫,或一份說明無法支持的具體主張以及未被量測內容的報告。
適用情境
適用於技能語料庫在新模型上表現退化,且希望依據實測行為而非猜測來修復的情況。適合擁有可對兩份語料庫組建進行 A/B 對比的基準測試框架,並有可重複端到端任務的團隊。不適用於靜態稽核或單純減少字數。
執行需求
需要執行封存檔或能產生逐次執行記錄的測試框架,記錄包含工具呼叫軌跡、終止標記、token 計數和最終訊息;需要組建選擇器,可將執行指向特定語料庫簽出;還需要語料庫能端到端執行的可重複任務。若沒有可對兩份組建進行 A/B 對比的基準測試框架,此技能會拒絕執行。不附帶指令碼,僅含說明與參考文件。

Retune a Corpus for a New Model

A corpus that degrades on a new model is a measurement problem before it is a writing problem: rewriting what looks wrong produces a plausible fix list and no way to know whether any item mattered.

Outcome: a corpus whose measured behavior on the target model clears a bar registered before any change, with the regression classes removed and each removal attributable.

Done: the bar is cleared, or the run reports the specific claim it could not support. A green test suite is not done: it proves nothing broke, not that behavior improved.

Non-goal: word reduction. Leanness and performance are separate programs that share a corpus; only one of them is the result here. Report completion, not word count.

Phase 0: the measurement gate — check this first

This skill cannot run without a way to observe behavior. Check for all three, and name whichever is missing:

  1. A run archive or a harness that produces one — per-run logs carrying the tool-call trace, a terminal marker, token counts, and the final message.
  2. A build selector — the harness can point a run at a specific source checkout of the corpus (a --plugin-dir-style override, a configurable skills path, an env var), so two builds are comparable under one runner.
  3. A repeatable task the corpus actually executes end to end.

If any is missing, stop and say so, naming what to build. Do not fall back to a static audit and present it as retuning: an audit can say what looks cuttable and never whether cutting helped. An audit-only pass is a legitimate thing to want; it is a different request.

State the target model and the harness you found before continuing.

The phases

They run in order, and each names the reference it cannot start without. Read references/workflow-shapes.md before dispatching any phase: the wrong orchestration shape is the common failure. Fan out by disjoint file ownership, never by item. Items cross files, and agents that share a file lose each other's edits.

Before assessing whether the registered bar is met or interpreting its results, read references/noise-floor.md.

  1. Mine the archive before spending a run — references/baseline-mining.md. Historical runs are a free baseline, usually larger than any experiment affordable now.
  2. Establish the noise floor — references/noise-floor.md. Run the harness against two identical copies of the corpus, same commit on both sides; whatever difference appears is the floor every later claim must clear. Register the bar now, in writing, before any change exists. A bar chosen after seeing results is not a bar.
  3. Audit the corpus adversarially — references/corpus-audit.md. One agent per skill proposes cuts; a second per skill does the opposite and defends the existing prose. The two passes require independent contexts. If the host exposes no way to run them as separate agents, report that as a blocker and stop the audit — do not argue both sides in one context and present the result as an audit.
  4. Cut in surgical passes, one problem per agent — references/cut-passes.md, and references/halt-taxonomy.md when the symptom is stalling, halting, or a run that ends while naming work it did not do. Two rules bound every pass, whatever class it is cutting. Never edit a test to make a suite green: a removed string a test pins is a finding to report, not a test to weaken. And not every stop is the enemy. Some workflows exist to stop and ask; that is the product. Sort every stop by who is actually on the other side before touching it. references/halt-taxonomy.md carries the screens that decide, so read them before cutting any stop.
  5. Measure, then let the failure choose the next fix — references/cut-passes.md again for what each failure site means and for auditing the phases the instrument never enters. Loop 4 and 5 until the registered bar clears. Then stop; a bar cleared is done. Also report what stayed unmeasured: a cleared bar never implies coverage it does not have.
  6. Ship — references/cut-passes.md carries what the commits and the write-up must preserve.

來源與署名

來源:everyinc/compound-engineering-plugin位於skills/ce-retune提交67035e9

授權條款: 無授權條款

內容歸原作者所有。SourceWeft 從公開儲存庫中收錄這些內容。

檢舉或申請下架