Autoresearch

Yeachan-Heo/oh-my-claudecode/skills/autoresearch

作者 Yeachan-Heo454bae017356ed1d056182cd929b6eb6cf6143ad无许可证收录于 2026年10月9日更新于 2026年10月9日

Stateful single-mission improvement loop with strict evaluator contract, markdown decision logs, and max-runtime stop behavior

仅含说明AI & Agents
AI 生成的概览

为单一任务运行有界、由评估器驱动的改进循环,记录每次迭代直至触发停止条件。

功能
Autoresearch 是一个有状态技能,一次只负责一个任务,并按实验与评估循环反复迭代。每次迭代都会运行评估器、保存机器可读的评估 JSON,并在任务目录下追加人类可读的 markdown 决策日志。即使评估未通过也会继续,仅在达到明确的 max-runtime 上限、用户取消或记录到其他终止条件时停止。
适用场景
适用于任务与评估器已存在、且需要以严格评估进行持续单任务改进的场景。也适合需要持久实验日志以及通过原生 cron 定期重跑的情况。不适用于在运行时生成评估器或编排多个任务。
运行要求
需要先前 deep-interview --autoresearch 步骤产出的任务与评估器,以及可写入的任务目录用于存放产物。评估器输出必须是包含布尔 pass 字段的结构化 JSON。该技能不附带脚本,仅为指令。

<Purpose>

Autoresearch is a stateful skill for bounded, evaluator-driven iterative improvement. It owns one mission at a time, keeps iterating through non-passing results, records each evaluation and decision as durable artifacts, and stops only when an explicit max-runtime ceiling or another explicit terminal condition is reached.

</Purpose>

<Use_When>

  • You already have a mission and evaluator from /deep-interview --autoresearch
  • You want persistent single-mission improvement with strict evaluation
  • You need durable experiment logs under .omc/autoresearch/
  • You want a supported path for periodic reruns via Claude Code native cron

</Use_When>

<Do_Not_Use_When>

  • You need evaluator generation at runtime — use /deep-interview --autoresearch first
  • You need multiple missions orchestrated together — v1 forbids that
  • You want the deprecated omc autoresearch CLI flow — it is no longer authoritative

</Do_Not_Use_When>

<Contract>

  • Single-mission only in v1
  • Mission setup/evaluator generation stays in deep-interview --autoresearch
  • Evaluator output must be structured JSON with required boolean pass and optional numeric score
  • Non-passing iterations do not stop the run
  • Stop conditions are explicit and bounded, with max-runtime as the primary strict stop hook

</Contract>

<Required_Artifacts>

Canonical persistent storage lives under .omc/autoresearch/<mission-slug>/ and/or .omc/logs/autoresearch/<run-id>/.

Minimum required artifacts:

  • mission spec
  • evaluator script or command reference
  • per-iteration evaluation JSON
  • markdown decision logs

Recommended canonical shape:

text
.omc/autoresearch/<mission-slug>/  mission.md  evaluator.json  runs/<run-id>/    evaluations/      iteration-0001.json      iteration-0002.json    decision-log.md

Reuse existing runtime artifacts when available rather than duplicating them unnecessarily.

</Required_Artifacts>

<Workflow>

  1. Confirm a single mission exists and evaluator setup is already available.
  2. Ensure mode/state is active for autoresearch and records:
    • mission slug/dir
    • evaluator reference
    • iteration count
    • started/updated timestamps
    • explicit max-runtime or deadline
  3. On every iteration:
    • run exactly one experiment/change cycle
    • run the evaluator
    • persist machine-readable evaluation JSON
    • append a human-readable markdown decision log entry
    • continue even when evaluation does not pass
  4. Stop when:
    • max-runtime ceiling is reached
    • user explicitly cancels
    • another explicit terminal condition is recorded by the runtime

</Workflow>

<Cron_Integration>

Claude Code native cron is a supported integration point for periodic mission enhancement. In v1, prefer documenting/configuring cron inputs over building a large scheduler UI.

If cron is used:

  • keep one mission per scheduled job
  • preserve the same mission/evaluator contract
  • append new run artifacts rather than overwriting prior experiments

</Cron_Integration>

<Execution_Policy>

  • Do not hand execution back to omc autoresearch
  • Do not create multi-mission orchestration
  • Prefer reusing src/autoresearch/* runtime/schema helpers where they already match the stricter contract
  • Keep logs useful to humans, not only machines

</Execution_Policy>

来源与署名

来源:Yeachan-Heo/oh-my-claudecode位于skills/autoresearch提交454bae0

许可证: 无许可证

内容归原作者所有。SourceWeft 从公开仓库中收录这些内容。

举报或申请下架