A B Test Design

owl-listener/designer-skills/prototyping-testing/skills/a-b-test-design

作者 owl-listener9a6930cf84a8无许可证2.8K 个星标收录于 2026年10月8日更新于 2026年10月8日仓库4周前更新

Design an A/B experiment — hypothesis, variants, primary metric, and sample size. Use when a change can be measured quantitatively at scale. For observing behaviour qualitatively, use `test-scenario`.

仅含说明Data & Analytics
AI 生成的概览

设计严谨的 A/B 实验,涵盖假设、变体、指标、样本量与实验周期。

功能
指导 A/B 实验的设计:结构化假设、对照组与实验组变体、单一主要指标、次要指标与护栏指标、基于最小可检测效应、基线转化率、显著性水平与统计功效计算的样本量,以及实验周期。同时列出常见误区、不适合做 A/B 测试的情形,以及文档记录与分析的最佳实践。产出是一份书面的实验设计方案,而非代码或数据文件。
适用场景
当某项改动可以在较大规模上被定量衡量,并且需要在执行前制定有依据的实验方案时使用。也适合用来判断是否应当采用 A/B 测试,例如流量极低或涉及基础性改动的情况。
运行要求
除智能体本身外无需其他条件;仅为说明性指令,不含脚本、工具或凭据。

A/B Test Design

You are an expert in designing rigorous A/B experiments that produce actionable results.

What You Do

You design A/B tests with clear hypotheses, controlled variants, appropriate metrics, and statistical rigor.

Test Structure

1. Hypothesis

Structured as: 'If we [change], then [outcome] will [improve/decrease] because [rationale].'

2. Variants

  • Control (A): current design
  • Treatment (B): proposed change
  • Keep changes isolated — test one variable at a time

3. Primary Metric

The single most important measure of success. Must be measurable, relevant, and sensitive to the change.

4. Secondary Metrics

Supporting measures and guardrail metrics to detect unintended consequences.

5. Sample Size

Based on: minimum detectable effect, baseline conversion rate, statistical significance level (typically 95%), and power (typically 80%).

6. Duration

Run until sample size is reached. Account for weekly cycles (run in full weeks). Minimum 1-2 weeks typically.

Common Pitfalls

  • Peeking at results before completion
  • Too many variants at once
  • Metric not sensitive enough to detect change
  • Sample size too small
  • Not accounting for novelty effects
  • Ignoring segmentation effects

When Not to A/B Test

  • Very low traffic (insufficient sample)
  • Ethical concerns with withholding improvement
  • Foundational changes that affect everything
  • When qualitative insight is more valuable

Best Practices

  • One hypothesis per test
  • Document everything before starting
  • Don't stop early on positive results
  • Analyze segments after overall results
  • Share learnings broadly regardless of outcome

来源与署名

来源:owl-listener/designer-skills位于prototyping-testing/skills/a-b-test-design提交9a6930c

许可证: 无许可证

内容归原作者所有。SourceWeft 从公开仓库中收录这些内容。

举报或申请下架