Analyzing Experiment Session Replays

作者 PostHog469d1773e9cb无许可证收录于 2026年10月8日更新于 2026年10月8日

Analyze session replay patterns across experiment variants to understand user behavior differences. Use when the user wants to see how users interact with different experiment variants, identify usability issues, compare behavior patterns between control and test groups, or get qualitative insights to complement quantitative experiment results. Also covers pairing the observed behavior with a linked survey when the user wants qualitative feedback beyond what recordings show.

AI 生成的概览

分析实验各变体的会话回放,比较对照组与测试组的用户行为差异。

功能
引导智能体获取实验元数据和功能开关变体,并为每个变体构建会话录制筛选条件。它会分别检索各变体的录制,比较时长、活跃度和控制台错误等量化模式,以及困惑、导航路径差异等定性信号。最终产出行为差异、可用性问题和建议的书面总结,并可能建议配合简短问卷。
适用场景
当用户想了解用户在不同实验变体中的行为差异、比较对照组与测试组行为,或需要定性洞察来补充实验指标时使用。也适用于查看某个特定变体的会话回放。
运行要求
需要已启动且有开始日期的实验、项目已启用会话回放,以及已接触实验的用户。依赖 experiment-get、query-session-recordings-list 和 execute-sql 等工具,并能访问实验与功能开关数据。仅为说明文档,不附带脚本。

Analyzing experiment session replays

This skill guides you through analyzing session recordings for experiment variants to understand behavioral differences between control and test groups.

When to use this skill

Use this skill when:

  • The user asks to analyze session replays for an experiment
  • The user wants to understand how users behave differently across experiment variants
  • The user asks to compare user behavior between control and test variants
  • The user wants qualitative insights to complement experiment metrics
  • The user asks questions like "How are users behaving in my experiment?" or "Show me session replays for variant X"

Prerequisites

Before analyzing session replays:

  1. The experiment must be launched (not in draft state)
  2. Session replay must be enabled for the project
  3. Users must have been exposed to the experiment variants
  4. The experiment must have a start date

Workflow

1. Get experiment details and feature flag variants

First, retrieve the experiment information and the feature flag variants (source of truth).

Step 1a: Get experiment metadata

You can either:

  • Option A: Use the experiment-get tool if you already have the experiment ID from context
  • Option B: Query the experiments table via HogQL:
sql
SELECT    e.id,    e.name,    f.key AS feature_flag_key,    e.start_date,    e.end_dateFROM system.experiments eJOIN system.feature_flags f ON f.id = e.feature_flag_idWHERE e.id = <experiment_id>

From the experiment data, extract:

  • feature_flag_key: The feature flag controlling the experiment
  • start_date and end_date: The experiment's time range

Step 1b: Get variants from the feature flag

IMPORTANT: Always get variants from the feature flag, NOT from experiment.parameters.feature_flag_variants. The parameters can be out of sync or deprecated. The feature flag is the source of truth.

Query the feature flag to get the current variants:

sql
SELECT filters.multivariate.variants AS variantsFROM system.feature_flagsWHERE key = '<feature_flag_key>'

Select the variants path directly — selecting the whole filters object gets truncated in results for flags with large targeting configs. Example structure: [{"key": "control", "name": "Control", "rollout_percentage": 50}, {"key": "test", ...}]

The variant key values (e.g., "control", "test", "variant_a") are what you'll use to filter session recordings.

2. Build session recording filters for each variant

For each variant in the experiment, construct recording filters that match users exposed to that variant.

Filter structure for a variant (input to query-session-recordings-list):

json
{  "date_from": "<experiment.start_date>",  "date_to": "<experiment.end_date or current time>",  "filter_test_accounts": true,  "properties": [    {      "type": "event",      "key": "$feature/<feature_flag_key>",      "operator": "exact",      "value": ["<variant_key>"]    }  ]}

Key points:

  • The $feature/<flag_key> event property records the flag's value on each event — filtering on it matches recordings where the flag was active with that variant. This is an approximation of exposure, broader than the experiment's exposure event ($feature_flag_called, or $experiment_exposure on the new rollout — both deduped per identity): right for browsing behavior across variants, but not an exact mirror of the analysis population — the scanning-experiments-with-replay-vision skill derives that exact filter when you need it
  • value is an array of variant key strings (e.g. ["control"]); for boolean flags use ["true"] or ["false"]
  • Avoid the type: "flag" / flag_evaluates_to property filter for variant scoping — the recordings query accepts it but silently ignores it, returning unfiltered results (last verified 2026-06-10). If you want to try it anyway, verify it actually filters first: a query with a nonexistent flag key should return zero recordings
  • Set the date range to the experiment's start and end dates
  • Enable filter_test_accounts: true to exclude test users

3. Retrieve recordings for each variant

Use the query-session-recordings-list tool with the filters constructed in step 2.

Call the tool once per variant to get recordings for each group:

  • Variant "control" → recordings for control group
  • Variant "test" → recordings for test variant
  • Additional variants if the experiment has more than 2

The tool returns a list of recordings with metadata including:

  • distinct_id — the person's distinct ID
  • recording_duration, active_seconds, inactive_seconds
  • click_count, keypress_count, mouse_activity_count
  • console_log_count, console_warn_count, console_error_count
  • start_url — first page URL visited
  • start_time / end_time, activity_score

4. Compare and analyze

Compare the recordings between variants by looking for:

Quantitative patterns:

  • Session duration differences
  • Activity levels (clicks, keypresses)
  • Console error rates
  • Bounce rates

Qualitative insights:

  • User confusion or frustration indicators
  • Different navigation paths
  • Feature discovery patterns
  • Error recovery behavior

5. Present findings

Summarize the behavioral differences between variants, highlighting:

  • Total recordings per variant
  • Notable behavior patterns unique to each variant
  • Usability issues or friction points observed
  • Recommendations based on the qualitative data

6. Observing shows behavior; asking adds what users think of it

Watching sessions and asking users are different instruments, not substitutes. Recordings show what people did with the change; a short survey, shown when they finish the experimented flow, captures what they thought of it — a rating and an optional comment, readable per variant. For a user-facing change of real size, the two together make a fuller qualitative read than either alone, so mention the option when the behavioral comparison in step 4 leaves opinion unaccounted for, or when a pattern in the recordings is a hypothesis worth checking with the people who produced it. Once per conversation at most; drop it if declined.

Default to asking every exposed user rather than one variant: a popover shown to only one arm is itself a difference between the arms, and the response event carries the variant anyway, so the split survives.

→ See references/qualitative-feedback.md in [[diagnosing-experiment-results]]

Example interaction

text
User: "How are users behaving in my checkout experiment?"Agent steps:1. Query experiment details (ID: 123, feature_flag_key: "checkout-flow-test", date range: 2025-01-01 to 2025-01-31)2. Query feature flag "checkout-flow-test" to get variants from filters.multivariate.variants3. Extract variant keys: "control" and "new-checkout"4. Build filters for control variant:   - Property filter: { type: "event", key: "$feature/checkout-flow-test", operator: "exact", value: ["control"] }   - Date range: 2025-01-01 to 2025-01-315. Call query-session-recordings-list with control filters → 147 recordings found6. Build filters for new-checkout variant and call query-session-recordings-list → 152 recordings found7. Compare patterns:   - Control: Average 3m 45s session duration, 12% console errors   - New-checkout: Average 2m 30s session duration, 5% console errors8. Present findings:   "I analyzed session replays for your checkout experiment. The new checkout flow shows:   - 33% faster completion (2m 30s vs 3m 45s)   - 58% fewer console errors (5% vs 12%)   - Users in the new variant navigate directly to payment, while control users often backtrack to review cart   - Recommendation: The new checkout flow reduces friction and errors"

Important notes

Do not make assumptions:

  • Always verify the experiment has recordings before analyzing
  • Check that the experiment is launched (has a start_date)
  • If no recordings are found, inform the user clearly

Filter construction:

  • The $feature/<flag_key> event property is how you scope recordings to a variant
  • One filter per variant — call the tool once per variant with its own filter
  • For boolean flags, use ["true"]/["false"] as the value instead of a variant key

Error handling:

  • If the experiment is in draft state, tell the user it hasn't started yet
  • If no recordings exist, suggest enabling session replay or waiting for user traffic
  • If the variant count is unexpected, double-check the experiment configuration

Related tools

  • query-session-recordings-list: Core tool for retrieving session recordings with filters
  • experiment-get: Get experiment metadata; experiment-results-get for statistical results
  • execute-sql: Query experiments table for details via HogQL

Related skills

  • diagnosing-experiment-results — the quantitative side: bias checks and significance on the same experiment
  • investigating-replay — deep-dive a single session from either variant
  • finding-sessions-to-watch — general session shortlisting outside the experiment context

来源与署名

来源:PostHog/ai-plugin位于skills/analyzing-experiment-session-replays提交469d177

许可证: 无许可证

内容归原作者所有。SourceWeft 从公开仓库中收录这些内容。

举报或申请下架

更多来自 PostHog/ai-plugin 的技能

Writing Simplified Technical English

PostHog

应用 ASD-STE100 简化技术英语规则,让智能体撰写的文字含义明确、便于执行。

Writing & Content2026年10月8日

Working With Task Comments

PostHog

通过 PostHog MCP exec 调度器读取并解读 PostHog 任务、产物和画布上的评论。

Productivity & Workflow2026年10月8日

Working With Skills

PostHog

指导智能体使用 PostHog 的 skill-* MCP 工具来发现、读取、创建、更新和重构技能。

AI & Agents2026年10月8日

Working With Scouts

PostHog

关于如何把监控任务委派给 PostHog Signals 侦察代理、处理其报告并长期调校整个代理集群的操作手册。

AI & Agents2026年10月8日

Validating And Publishing Canvases

PostHog

Validate and publish a canvas source project safely: the source-project shape, declared capabilities, reading the current version pointer, iterating on validation diagnostics, guarded publishing with expected_current_version_id, staging a draft build and promoting it, waiting out the queued build, and recovering from a 409 version_conflict or a 429 capacity limit without overwriting concurrent work. Use whenever a canvas edit is ready to save, a draft build is wanted, a canvas publish or build returns diagnostics or a conflict, or a task needs to understand canvas version history.

待分类2026年10月8日

Understanding Billing Usage

PostHog

Explains PostHog billing usage and spend from the customer's visible Billing MCP tools. Use when the user asks why usage or spend is high, which product or project is driving usage, what a usage type means, how to reduce usage, what changed over time, why they got a usage change alert, or whether a spike/drop alert was real or noisy. Also use before product-specific analytics skills when the user names a billable PostHog product metric such as events, recordings, feature flag requests, exceptions, survey responses, synced rows, logs, AI events, AI credits, or Inbox credits. Starts from Billing usage/spend tools, then routes to customer-visible product MCP surfaces for deeper investigation.

待分类2026年10月8日