Verify This

作者 cursorccb5507cec15无许可证10K 个星标收录于 2026年10月8日更新于 2026年10月8日仓库今天更新

Verify a claim with fresh local evidence: restate it falsifiably, capture baseline and treatment, compare artifacts, and return VERIFIED, NOT VERIFIED, or INCONCLUSIVE.

AI 生成的概览

通过采集基线与处理后的证据来验证具体主张,并返回 VERIFIED、NOT VERIFIED 或 INCONCLUSIVE。

功能
引导智能体把主张改写为可证伪的形式,选择能够证伪它的最小本地验证面,并在相同命令、数据和环境下采集基线与处理后证据。随后比较原始产物,例如数值、截图、终端记录、HTTP 响应、性能剖析或堆快照。最终只给出一个判定——VERIFIED、NOT VERIFIED 或 INCONCLUSIVE——并附上证据、差值、阈值和可能的干扰因素。
适用场景
适用于用户要求验证某个主张、证明某功能可用、确认修复是否生效或展示证据的场景。也适合缺陷修复的前后复现,以及 UI、CLI、API、性能或内存类主张的测量。不适用于缺少可测量条件、指标和阈值的模糊主张。
运行要求
不需要脚本或特殊工具,仅为操作说明。根据主张不同,可能使用本地测试运行器、CLI 或 UI 控制工具、HTTP 客户端、性能剖析器或堆快照工具。向磁盘写入产物是可选的;涉及敏感内容时应避免写入,除非用户同意。

Verify This

Verification is not a recap. It proves or disproves a specific claim with repeatable evidence.

When To Use

  • The user asks "verify this", "prove it works", "did this fix it", or "show me the evidence".
  • A bug fix needs a before/after repro.
  • A UI, CLI, API, performance, or memory claim needs measurement.
  • A test passes but the user-visible behavior still needs confirmation.

Do not use this for vague claims like "the code is cleaner". Ask for a measurable claim first.

Workflow

  1. Restate the claim in falsifiable form: condition, metric, and threshold.
  2. Pick the smallest local surface that can disprove it.
  3. Capture a baseline from the old state: merge base, parent commit, failing branch, or current broken repro.
  4. Capture treatment from the changed state with the same command, data, warmup, and environment.
  5. Compare raw artifacts: numbers, screenshots, terminal transcripts, HTTP responses, profiles, heap snapshots, or test output.
  6. Return exactly one verdict: VERIFIED, NOT VERIFIED, or INCONCLUSIVE.

Local Surfaces

  • Code behavior: focused unit/integration tests or a minimal repro script.
  • CLI/TUI behavior: control-cli, terminal transcript, or demo recording.
  • UI behavior: control-ui, screenshots, accessibility snapshots, or browser traces.
  • API behavior: local HTTP/RPC request and response diff.
  • Performance: same-machine baseline/treatment timings or CPU profiles.
  • Memory: heap snapshots before and after the suspected operation.

Artifact Layout

When safe to write artifacts:

text
/tmp/verify-this/<claim-slug>/├── claim.md├── timeline.md├── baseline/├── treatment/├── diff/└── verdict.md

If artifacts may contain sensitive code, prompts, screenshots, HTTP bodies, or heap data, keep only the minimal inline evidence unless the user agrees to disk storage.

Verdict Rules

  • VERIFIED: baseline and treatment differ in the predicted direction, by the claimed threshold, with no obvious confound.
  • NOT VERIFIED: the behavior is unchanged, moves the wrong way, or misses the threshold.
  • INCONCLUSIVE: no valid baseline, noisy signal, failed measurement, or an environment difference invalidates the comparison.

Output

Use this shape:

text
VERIFIED | NOT VERIFIED | INCONCLUSIVEClaim: <falsifiable claim>
Evidence:<metric/artifact>: baseline=<...>, treatment=<...>, delta=<...>, threshold=<...>
Reasoning:<one tight paragraph naming the evidence and any confounds>

Do not soften a negative result. A clear NOT VERIFIED is useful.

来源与署名

来源:cursor/plugins位于cursor-team-kit/skills/verify-this提交ccb5507

许可证: 无许可证

内容归原作者所有。SourceWeft 从公开仓库中收录这些内容。

举报或申请下架