A B Test Design

owl-listener/designer-skills/prototyping-testing/skills/a-b-test-design

by owl-listener9a6930cf84a8No license2.8K starsListed Oct 8, 2026Updated Oct 8, 2026Repository updated 4 weeks ago

Design an A/B experiment — hypothesis, variants, primary metric, and sample size. Use when a change can be measured quantitatively at scale. For observing behaviour qualitatively, use `test-scenario`.

Instructions onlyData & Analytics
AI-generated overview

Designs rigorous A/B experiments with hypotheses, variants, metrics, sample size and duration.

What it does
Guides the design of an A/B experiment: a structured hypothesis, control and treatment variants, a single primary metric, secondary and guardrail metrics, sample size based on detectable effect, baseline rate, significance and power, and a run duration. It also lists common pitfalls, cases where A/B testing is inappropriate, and best practices for documentation and analysis. The output is a written experiment design rather than code or a data file.
When to use it
Use it when a change can be measured quantitatively at scale and you need a defensible experiment plan before running it. It is also useful for checking whether an A/B test is the right approach at all, for example with very low traffic or foundational changes.
Requirements
None beyond the agent; instructions only, no scripts, tools or credentials.

A/B Test Design

You are an expert in designing rigorous A/B experiments that produce actionable results.

What You Do

You design A/B tests with clear hypotheses, controlled variants, appropriate metrics, and statistical rigor.

Test Structure

1. Hypothesis

Structured as: 'If we [change], then [outcome] will [improve/decrease] because [rationale].'

2. Variants

  • Control (A): current design
  • Treatment (B): proposed change
  • Keep changes isolated — test one variable at a time

3. Primary Metric

The single most important measure of success. Must be measurable, relevant, and sensitive to the change.

4. Secondary Metrics

Supporting measures and guardrail metrics to detect unintended consequences.

5. Sample Size

Based on: minimum detectable effect, baseline conversion rate, statistical significance level (typically 95%), and power (typically 80%).

6. Duration

Run until sample size is reached. Account for weekly cycles (run in full weeks). Minimum 1-2 weeks typically.

Common Pitfalls

  • Peeking at results before completion
  • Too many variants at once
  • Metric not sensitive enough to detect change
  • Sample size too small
  • Not accounting for novelty effects
  • Ignoring segmentation effects

When Not to A/B Test

  • Very low traffic (insufficient sample)
  • Ethical concerns with withholding improvement
  • Foundational changes that affect everything
  • When qualitative insight is more valuable

Best Practices

  • One hypothesis per test
  • Document everything before starting
  • Don't stop early on positive results
  • Analyze segments after overall results
  • Share learnings broadly regardless of outcome

Source and attribution

Source:owl-listener/designer-skillsinprototyping-testing/skills/a-b-test-designat commit9a6930c

License: No license

Content belongs to its original authors. SourceWeft indexes it from a public repository.

Report or request removal