Ab Test Analysis

作者 phuryn8607e3b07781無授權條款26K 個星標收錄於 2026年10月8日更新於 2026年10月8日儲存庫3 週前更新

Analyze A/B test results with statistical significance, sample size validation, confidence intervals, and ship/extend/stop recommendations. Use when evaluating experiment results, checking if a test reached significance, interpreting split test data, or deciding whether to ship a variant.

AI 產生的概覽

分析 A/B 測試結果的統計顯著性,並給出上線、延長、停止或調查的建議。

功能
引導代理評估 A/B 測試結果:釐清假設、變體、主要指標與護欄指標以及流量分配;驗證樣本量、測試期間、隨機化和新奇效應;並計算轉換率、相對提升、p 值和信賴區間。它會檢查護欄指標,並將結果對應到上線、延長、停止或調查的建議。它會產出 Markdown 總結報告,包含指標表格、建議、理由和後續步驟;在提供原始資料時,還可能產生用於計算的 Python 腳本。
適用情境
適用於評估實驗結果、判斷測試是否達到統計顯著性、解讀分流測試資料,或決定是否上線某個變體。它適合依賴 A/B 測試結果的產品或成長決策。
執行需求
僅為說明文件,不附帶腳本。它可以讀取使用者提供的資料檔案,如 CSV、Excel 或分析匯出檔,並可能產生用於統計計算的 Python 腳本,因此在分析原始資料時需要 Python 執行環境。

A/B Test Analysis

Evaluate A/B test results with statistical rigor and translate findings into clear product decisions.

Context

You are analyzing A/B test results for $ARGUMENTS.

If the user provides data files (CSV, Excel, or analytics exports), read and analyze them directly. Generate Python scripts for statistical calculations when needed.

Instructions

  1. Understand the experiment:

    • What was the hypothesis?
    • What was changed (the variant)?
    • What is the primary metric? Any guardrail metrics?
    • How long did the test run?
    • What is the traffic split?
  2. Validate the test setup:

    • Sample size: Is the sample large enough for the expected effect size?
      • Use the formula: n = (Z²α/2 × 2 × p × (1-p)) / MDE²
      • Flag if the test is underpowered (<80% power)
    • Duration: Did the test run for at least 1-2 full business cycles?
    • Randomization: Any evidence of sample ratio mismatch (SRM)?
    • Novelty/primacy effects: Was there enough time to wash out initial behavior changes?
  3. Calculate statistical significance:

    • Conversion rate for control and variant
    • Relative lift: (variant - control) / control × 100
    • p-value: Using a two-tailed z-test or chi-squared test
    • Confidence interval: 95% CI for the difference
    • Statistical significance: Is p < 0.05?
    • Practical significance: Is the lift meaningful for the business?

    If the user provides raw data, generate and run a Python script to calculate these.

  4. Check guardrail metrics:

    • Did any guardrail metrics (revenue, engagement, page load time) degrade?
    • A winning primary metric with degraded guardrails may not be a true win
  5. Interpret results:

    OutcomeRecommendation
    Significant positive lift, no guardrail issuesShip it — roll out to 100%
    Significant positive lift, guardrail concernsInvestigate — understand trade-offs before shipping
    Not significant, positive trendExtend the test — need more data or larger effect
    Not significant, flatStop the test — no meaningful difference detected
    Significant negative liftDon't ship — revert to control, analyze why
  6. Provide the analysis summary:

    ## A/B Test Results: [Test Name]
    **Hypothesis**: [What we expected]**Duration**: [X days] | **Sample**: [N control / M variant]
    | Metric | Control | Variant | Lift | p-value | Significant? ||---|---|---|---|---|---|| [Primary] | X% | Y% | +Z% | 0.0X | Yes/No || [Guardrail] | ... | ... | ... | ... | ... |
    **Recommendation**: [Ship / Extend / Stop / Investigate]**Reasoning**: [Why]**Next steps**: [What to do]

Think step by step. Save as markdown. Generate Python scripts for calculations if raw data is provided.


Further Reading

來源與署名

來源:phuryn/pm-skills位於pm-data-analytics/skills/ab-test-analysis提交8607e3b

授權條款: 無授權條款

內容歸原作者所有。SourceWeft 從公開儲存庫中收錄這些內容。

檢舉或申請下架