Experiment Designer

alirezarezvani/claude-skills/product-team/skills/experiment-designer

作者 alirezarezvani19392f7a0826無授權條款27K 個星標收錄於 2026年10月8日更新於 2026年10月8日儲存庫5 週前更新

Use when planning product experiments, writing testable hypotheses, estimating sample size, prioritizing tests, or interpreting A/B outcomes with practical statistical rigor.

AI 產生的概覽

規劃產品 A/B 實驗:假設、指標、樣本量、ICE 優先順序排序與統計解讀。

功能
指導產品實驗的設計與評估,從撰寫 If/Then/Because 假設、定義主要指標、護欄指標與次要指標,到設定停止規則並解讀結果。其中包含 ICE 優先順序計算公式,以及關於 p 值、信賴區間與實務顯著性的統計護欄。隨附的 Python 指令碼可依基準轉換率、最小可偵測效果、顯著水準與統計檢定力,計算每個變體及整體所需樣本量。
適用情境
適用於規劃 A/B 或多變數測試、定義成功標準、估算樣本量或最小可偵測效果、為測試待辦事項排序,或為產品決策解讀統計輸出。
執行需求
需要 Python 3 來執行 scripts/sample_size_calculator.py;未說明需要憑證或網路存取。隨附兩份參考文件。

Experiment Designer

Design, prioritize, and evaluate product experiments with clear hypotheses and defensible decisions.

When To Use

Use this skill for:

  • A/B and multivariate experiment planning
  • Hypothesis writing and success criteria definition
  • Sample size and minimum detectable effect planning
  • Experiment prioritization with ICE scoring
  • Reading statistical output for product decisions

Core Workflow

  1. Write hypothesis in If/Then/Because format
  • If we change [intervention]
  • Then [metric] will change by [expected direction/magnitude]
  • Because [behavioral mechanism]
  1. Define metrics before running test
  • Primary metric: single decision metric
  • Guardrail metrics: quality/risk protection
  • Secondary metrics: diagnostics only
  1. Estimate sample size
  • Baseline conversion or baseline mean
  • Minimum detectable effect (MDE)
  • Significance level (alpha) and power

Use:

bash
python3 scripts/sample_size_calculator.py --baseline-rate 0.12 --mde 0.02 --mde-type absolute
  1. Prioritize experiments with ICE
  • Impact: potential upside
  • Confidence: evidence quality
  • Ease: cost/speed/complexity

ICE Score = (Impact * Confidence * Ease) / 10

  1. Launch with stopping rules
  • Decide fixed sample size or fixed duration in advance
  • Avoid repeated peeking without proper method
  • Monitor guardrails continuously
  1. Interpret results
  • Statistical significance is not business significance
  • Compare point estimate + confidence interval to decision threshold
  • Investigate novelty effects and segment heterogeneity

Hypothesis Quality Checklist

  • Contains explicit intervention and audience
  • Specifies measurable metric change
  • States plausible causal reason
  • Includes expected minimum effect
  • Defines failure condition

Common Experiment Pitfalls

  • Underpowered tests leading to false negatives
  • Running too many simultaneous changes without isolation
  • Changing targeting or implementation mid-test
  • Stopping early on random spikes
  • Ignoring sample ratio mismatch and instrumentation drift
  • Declaring success from p-value without effect-size context

Statistical Interpretation Guardrails

  • p-value < alpha indicates evidence against null, not guaranteed truth.
  • Confidence interval crossing zero/no-effect means uncertain directional claim.
  • Wide intervals imply low precision even when significant.
  • Use practical significance thresholds tied to business impact.

See:

  • references/experiment-playbook.md
  • references/statistics-reference.md

Tooling

scripts/sample_size_calculator.py

Computes required sample size (per variant and total) from:

  • baseline rate
  • MDE (absolute or relative)
  • significance level (alpha)
  • statistical power

Example:

bash
python3 scripts/sample_size_calculator.py \  --baseline-rate 0.10 \  --mde 0.015 \  --mde-type absolute \  --alpha 0.05 \  --power 0.8

來源與署名

來源:alirezarezvani/claude-skills位於product-team/skills/experiment-designer提交19392f7

授權條款: 無授權條款

內容歸原作者所有。SourceWeft 從公開儲存庫中收錄這些內容。

檢舉或申請下架