Experiment Design

作者 lingzhi2279e6c085d65e3無授權條款386 個星標收錄於 2026年10月8日更新於 2026年10月8日儲存庫7 個月前更新

Design experiment plans with progressive stages — initial implementation, baseline tuning, creative research, and ablation studies. Plan baselines, datasets, hyperparameter sweeps, and evaluation metrics. Use when planning experiments for a research paper.

AI 產生的概覽

規劃分階段的機器學習實驗設計,涵蓋基準、資料集、超參數網格、指標與消融實驗。

功能
把研究想法、計畫或方法描述轉換為結構化的實驗設計。它將工作組織成四個遞進階段:初始實作、基準調校、創意研究與消融研究。輸出為 JSON 或 Markdown 設計,包含基準、資料集、指標、消融元件、超參數網格與隨機種子數量,並附上一個僅用標準函式庫的 Python 指令碼來產生這些內容。
適用情境
在為研究論文規劃實驗、需要包含基準、資料集、參數掃描與評估指標的分階段方案時使用。適合在實作與程式碼執行之前的實驗範圍界定階段。
執行需求
執行隨附指令碼 scripts/design_experiments.py 需要 Python 3 環境,該指令碼僅依賴標準函式庫。輸入為研究想法、計畫或方法描述,可選提供 JSON 研究計畫檔案。未說明需要憑證或網路存取。

Experiment Design

Design structured, progressive experiment plans for research papers.

Input

  • $0 — Research idea, plan, or method description

References

  • 4-stage progressive experiment prompts: ~/.claude/skills/experiment-design/references/stage-prompts.md

Scripts

Generate experiment design

bash
python ~/.claude/skills/experiment-design/scripts/design_experiments.py --plan research_plan.json --output experiment_design.jsonpython ~/.claude/skills/experiment-design/scripts/design_experiments.py --method "contrastive learning" --task classification --format markdown

Generates baselines, ablation matrix, hyperparameter grid, metric selection. Stdlib-only.

4-Stage Progressive Framework (from AI-Scientist-v2)

Stage 1: Initial Implementation

  • Focus on getting a basic working implementation
  • Use a simple dataset
  • Aim for basic functional correctness
  • Completion: at least one working (non-buggy) implementation

Stage 2: Baseline Tuning

  • Tune hyperparameters (learning rate, epochs, batch size)
  • Do NOT change model architecture
  • Test on at least TWO datasets
  • Completion: stable training curves, improvement over Stage 1

Stage 3: Creative Research

  • Explore novel improvements and insights
  • Be creative and think outside the box
  • Test on at least THREE datasets
  • Completion: demonstrated novel improvement

Stage 4: Ablation Studies

  • Systematic component analysis
  • Each ablation tests a different aspect
  • Use same datasets as Stage 3
  • Completion: all planned ablations done

Output Format

json
{  "stages": [    {      "name": "initial_implementation",      "goals": ["Basic working baseline", "Simple dataset"],      "max_iterations": 5,      "completion_criteria": "Working implementation with non-zero accuracy"    }  ],  "baselines": ["Method A", "Method B"],  "datasets": ["Dataset1", "Dataset2", "Dataset3"],  "metrics": ["accuracy", "F1", "inference_time"],  "ablation_components": ["component_A", "component_B"],  "hyperparameter_grid": {    "lr": [1e-4, 1e-3, 1e-2],    "batch_size": [32, 64, 128]  },  "num_seeds": 3}

Rules

  • Always start simple (Stage 1) before complex experiments
  • Each stage builds on the best result from the previous stage
  • Multi-seed evaluation for statistical significance
  • Document every experiment run in notes.txt
  • Generate figures for training curves and comparisons

Related Skills

  • Upstream: research-planning, idea-generation
  • Downstream: experiment-code, data-analysis
  • See also: paper-assembly

來源與署名

來源:lingzhi227/agent-research-skills位於skills/experiment-design提交9e6c085

授權條款: 無授權條款

內容歸原作者所有。SourceWeft 從公開儲存庫中收錄這些內容。

檢舉或申請下架