A B Test Design

owl-listener/designer-skills/prototyping-testing/skills/a-b-test-design

作者 owl-listener9a6930cf84a8無授權條款2.8K 個星標收錄於 2026年10月8日更新於 2026年10月8日儲存庫4 週前更新

Design an A/B experiment — hypothesis, variants, primary metric, and sample size. Use when a change can be measured quantitatively at scale. For observing behaviour qualitatively, use `test-scenario`.

僅含說明Data & Analytics
AI 產生的概覽

設計嚴謹的 A/B 實驗,涵蓋假設、變體、指標、樣本數與實驗期間。

功能
引導 A/B 實驗的設計:結構化假設、對照組與實驗組變體、單一主要指標、次要指標與護欄指標、依最小可偵測效果、基準轉換率、顯著水準與統計檢定力計算的樣本數,以及實驗期間。同時列出常見陷阱、不適合做 A/B 測試的情況,以及文件記錄與分析的最佳實務。產出是一份書面的實驗設計方案,而非程式碼或資料檔案。
適用情境
當某項變更能在較大規模上以量化方式衡量,且需要在執行前擬定有依據的實驗方案時使用。也適合用來判斷是否該採用 A/B 測試,例如流量極低或涉及基礎性變更的情況。
執行需求
除代理本身外無需其他條件;僅為指示性內容,不含指令碼、工具或憑證。

A/B Test Design

You are an expert in designing rigorous A/B experiments that produce actionable results.

What You Do

You design A/B tests with clear hypotheses, controlled variants, appropriate metrics, and statistical rigor.

Test Structure

1. Hypothesis

Structured as: 'If we [change], then [outcome] will [improve/decrease] because [rationale].'

2. Variants

  • Control (A): current design
  • Treatment (B): proposed change
  • Keep changes isolated — test one variable at a time

3. Primary Metric

The single most important measure of success. Must be measurable, relevant, and sensitive to the change.

4. Secondary Metrics

Supporting measures and guardrail metrics to detect unintended consequences.

5. Sample Size

Based on: minimum detectable effect, baseline conversion rate, statistical significance level (typically 95%), and power (typically 80%).

6. Duration

Run until sample size is reached. Account for weekly cycles (run in full weeks). Minimum 1-2 weeks typically.

Common Pitfalls

  • Peeking at results before completion
  • Too many variants at once
  • Metric not sensitive enough to detect change
  • Sample size too small
  • Not accounting for novelty effects
  • Ignoring segmentation effects

When Not to A/B Test

  • Very low traffic (insufficient sample)
  • Ethical concerns with withholding improvement
  • Foundational changes that affect everything
  • When qualitative insight is more valuable

Best Practices

  • One hypothesis per test
  • Document everything before starting
  • Don't stop early on positive results
  • Analyze segments after overall results
  • Share learnings broadly regardless of outcome

來源與署名

來源:owl-listener/designer-skills位於prototyping-testing/skills/a-b-test-design提交9a6930c

授權條款: 無授權條款

內容歸原作者所有。SourceWeft 從公開儲存庫中收錄這些內容。

檢舉或申請下架