Testing Changes

作者 riekelte67b7af9ac74無授權條款5 個星標收錄於 2026年10月8日更新於 2026年10月8日儲存庫3 週前更新

Use when deciding what tests a change needs - a feature, a bug fix, a refactor, any behavior change - or when reviewing whether a diff's tests are sufficient. Encodes tests-change-with-behavior, the bug-regression pattern, and assertion discrimination. Use whenever behavior changes and the test diff is empty, especially when the change is "too small to test".

AI 產生的概覽

指導行為變更需要哪些測試,包括回歸測試、可區分斷言與紅燈測試處理。

功能
此技能提出判斷一次變更需要哪些測試的規則,涵蓋新功能、缺陷修正、重構及其他行為變更。它定義變更應承擔的測試,例如缺陷修正的回歸測試、包含已浮現邊界情況的具體情境、每個任務的驗證指令、可區分斷言,以及彙總資料的不變量測試。它也說明紅燈測試規則,列出變更不必承擔的測試,並列舉常見錯誤。
適用情境
適用於判斷一次變更需要哪些測試,或審查某個差異的測試是否充分。也適用於行為改變但測試差異為空的情況,尤其是當該變更被認為太小而不值得測試時。
執行需求
僅為說明性內容,不附帶指令碼。它依名稱引用其他技能,包括 principal-engineering、writing-unit-tests、handling-failures、scoping-changes、verifying-before-done,以及一份技術寫作外掛參考。

Testing changes

REQUIRED BACKGROUND: the principal-engineering skill. Test craft lives in writing-unit-tests; this skill governs which tests a change owes.

Overview

A test is the executable form of a claim about behavior. A change that alters behavior without touching tests is a claim nobody wrote down. Two failure modes follow: the green suite over code that could not work, and the test or gate that never ran.

Tests a change owes

  1. Tests change with behavior, in the same change. An empty test diff on a behavior change is a review finding, not a style preference. A pure refactor owes the opposite proof: the existing tests still pass unmodified, which is what makes it a refactor.
  2. A bug fix ships its regression test. Named after the failure mode, not the ticket. The author shows both halves: red against the unfixed code, green against the fix. A regression test that never went red proves only that it compiles.
  3. Scenarios are concrete and include the surfaced edges. The edge cases that grounding and review turned up go into tests by name; the happy path alone tests the demo, not the change. When the change's risk is in the failure path, the failure path gets the tests (see handling-failures: it requires a logged, typed failure, and that surfacing is behavior a test owes).
  4. Every task carries its targeted verify command. The task names the verify command that proves this change, runnable alone and stated where the reviewer can run it. "The suite passed" vouches for nothing the suite never covered.
  5. Assertions must discriminate. A test that passes regardless of the change proves nothing. Break the code once, confirm the test goes red, then restore it. Non-discriminating assertions are how suites stay green over broken behavior.
  6. Aggregates that must reconcile get invariant tests. Anything on the project's declared critical paths (see the risk tiers in principal-engineering) that sums, derives, or mirrors other data gets more than point examples: the test asserts the reconciliation itself (the conservation pattern), the aggregate equals what the raw records imply.

The red-test rule

A red test in a gate you own gets fixed, never silenced: weakening the assertion, deleting the test, or marking it skipped to ship is converting a detected defect into an undetected one. Changing the test is legitimate exactly when the test asserted the old, wrong behavior, and the change says so explicitly. Attribute the origin first, then fix the test regardless of whose it is; "pre-existing" is a footnote, never an excuse (see verifying-before-done).

Tests a change does not owe

  • Tests for unreachable edges (see scoping-changes: fencing what cannot happen is dead code with good intentions).
  • Tests of framework internals or generated code. Test your use of them at the boundary you own.
  • A test-first process: whether tests come first is workflow (a test-first workflow skill, where the project installs one, governs that); this skill governs what must exist when the change ships, whichever order produced it.

Common mistakes

  • "Too small to test." A one-line change reaches behavior no test covers, and the regression stays silent until someone audits it.
  • Testing the fix without reproducing the bug. Red-before-green is the half that proves the test sees the defect.
  • Counting coverage by feel. Count the changed behaviors against the tests naming them (the same counting rule as the documentation-coverage count under "Keeping documents true" in the technical-writer plugin's technical-writing/references/truth.md).
  • Adding the test that discriminates against nothing, ever: an assertion no plausible defect could fail. Distinct from the legitimate pinning test that deliberately passes against both old and new code to guard unchanged adjacent behavior from overcorrection; a pinning test says that is what it is for.

來源與署名

來源:riekelt/principal-engineer位於plugins/principal-engineer/skills/testing-changes提交e67b7af

授權條款: 無授權條款

內容歸原作者所有。SourceWeft 從公開儲存庫中收錄這些內容。

檢舉或申請下架