Analyzing Experiment Session Replays

by PostHog469d1773e9cbNo licenseListed Oct 8, 2026Updated Oct 8, 2026

Analyze session replay patterns across experiment variants to understand user behavior differences. Use when the user wants to see how users interact with different experiment variants, identify usability issues, compare behavior patterns between control and test groups, or get qualitative insights to complement quantitative experiment results. Also covers pairing the observed behavior with a linked survey when the user wants qualitative feedback beyond what recordings show.

AI-generated overview

Analyzes session replays across experiment variants to compare user behavior between control and test groups.

What it does
Guides an agent through retrieving experiment metadata and feature flag variants, then building per-variant session recording filters. It retrieves recordings for each variant and compares quantitative patterns such as duration, activity and console errors, plus qualitative signals like confusion or navigation differences. It produces a written summary of behavioral differences, usability issues and recommendations, and may suggest pairing findings with a short survey.
When to use it
Use it when someone wants to see how users behave differently across experiment variants, compare control and test behavior, or get qualitative insight to complement experiment metrics. It also fits requests to review session replays for a specific variant.
Requirements
Requires a launched experiment with a start date, session replay enabled for the project, and exposed users. It relies on tools such as experiment-get, query-session-recordings-list and execute-sql, plus access to experiment and feature flag data. Instructions only; no scripts are shipped.

Analyzing experiment session replays

This skill guides you through analyzing session recordings for experiment variants to understand behavioral differences between control and test groups.

When to use this skill

Use this skill when:

  • The user asks to analyze session replays for an experiment
  • The user wants to understand how users behave differently across experiment variants
  • The user asks to compare user behavior between control and test variants
  • The user wants qualitative insights to complement experiment metrics
  • The user asks questions like "How are users behaving in my experiment?" or "Show me session replays for variant X"

Prerequisites

Before analyzing session replays:

  1. The experiment must be launched (not in draft state)
  2. Session replay must be enabled for the project
  3. Users must have been exposed to the experiment variants
  4. The experiment must have a start date

Workflow

1. Get experiment details and feature flag variants

First, retrieve the experiment information and the feature flag variants (source of truth).

Step 1a: Get experiment metadata

You can either:

  • Option A: Use the experiment-get tool if you already have the experiment ID from context
  • Option B: Query the experiments table via HogQL:
sql
SELECT    e.id,    e.name,    f.key AS feature_flag_key,    e.start_date,    e.end_dateFROM system.experiments eJOIN system.feature_flags f ON f.id = e.feature_flag_idWHERE e.id = <experiment_id>

From the experiment data, extract:

  • feature_flag_key: The feature flag controlling the experiment
  • start_date and end_date: The experiment's time range

Step 1b: Get variants from the feature flag

IMPORTANT: Always get variants from the feature flag, NOT from experiment.parameters.feature_flag_variants. The parameters can be out of sync or deprecated. The feature flag is the source of truth.

Query the feature flag to get the current variants:

sql
SELECT filters.multivariate.variants AS variantsFROM system.feature_flagsWHERE key = '<feature_flag_key>'

Select the variants path directly — selecting the whole filters object gets truncated in results for flags with large targeting configs. Example structure: [{"key": "control", "name": "Control", "rollout_percentage": 50}, {"key": "test", ...}]

The variant key values (e.g., "control", "test", "variant_a") are what you'll use to filter session recordings.

2. Build session recording filters for each variant

For each variant in the experiment, construct recording filters that match users exposed to that variant.

Filter structure for a variant (input to query-session-recordings-list):

json
{  "date_from": "<experiment.start_date>",  "date_to": "<experiment.end_date or current time>",  "filter_test_accounts": true,  "properties": [    {      "type": "event",      "key": "$feature/<feature_flag_key>",      "operator": "exact",      "value": ["<variant_key>"]    }  ]}

Key points:

  • The $feature/<flag_key> event property records the flag's value on each event — filtering on it matches recordings where the flag was active with that variant. This is an approximation of exposure, broader than the experiment's exposure event ($feature_flag_called, or $experiment_exposure on the new rollout — both deduped per identity): right for browsing behavior across variants, but not an exact mirror of the analysis population — the scanning-experiments-with-replay-vision skill derives that exact filter when you need it
  • value is an array of variant key strings (e.g. ["control"]); for boolean flags use ["true"] or ["false"]
  • Avoid the type: "flag" / flag_evaluates_to property filter for variant scoping — the recordings query accepts it but silently ignores it, returning unfiltered results (last verified 2026-06-10). If you want to try it anyway, verify it actually filters first: a query with a nonexistent flag key should return zero recordings
  • Set the date range to the experiment's start and end dates
  • Enable filter_test_accounts: true to exclude test users

3. Retrieve recordings for each variant

Use the query-session-recordings-list tool with the filters constructed in step 2.

Call the tool once per variant to get recordings for each group:

  • Variant "control" → recordings for control group
  • Variant "test" → recordings for test variant
  • Additional variants if the experiment has more than 2

The tool returns a list of recordings with metadata including:

  • distinct_id — the person's distinct ID
  • recording_duration, active_seconds, inactive_seconds
  • click_count, keypress_count, mouse_activity_count
  • console_log_count, console_warn_count, console_error_count
  • start_url — first page URL visited
  • start_time / end_time, activity_score

4. Compare and analyze

Compare the recordings between variants by looking for:

Quantitative patterns:

  • Session duration differences
  • Activity levels (clicks, keypresses)
  • Console error rates
  • Bounce rates

Qualitative insights:

  • User confusion or frustration indicators
  • Different navigation paths
  • Feature discovery patterns
  • Error recovery behavior

5. Present findings

Summarize the behavioral differences between variants, highlighting:

  • Total recordings per variant
  • Notable behavior patterns unique to each variant
  • Usability issues or friction points observed
  • Recommendations based on the qualitative data

6. Observing shows behavior; asking adds what users think of it

Watching sessions and asking users are different instruments, not substitutes. Recordings show what people did with the change; a short survey, shown when they finish the experimented flow, captures what they thought of it — a rating and an optional comment, readable per variant. For a user-facing change of real size, the two together make a fuller qualitative read than either alone, so mention the option when the behavioral comparison in step 4 leaves opinion unaccounted for, or when a pattern in the recordings is a hypothesis worth checking with the people who produced it. Once per conversation at most; drop it if declined.

Default to asking every exposed user rather than one variant: a popover shown to only one arm is itself a difference between the arms, and the response event carries the variant anyway, so the split survives.

→ See references/qualitative-feedback.md in [[diagnosing-experiment-results]]

Example interaction

text
User: "How are users behaving in my checkout experiment?"Agent steps:1. Query experiment details (ID: 123, feature_flag_key: "checkout-flow-test", date range: 2025-01-01 to 2025-01-31)2. Query feature flag "checkout-flow-test" to get variants from filters.multivariate.variants3. Extract variant keys: "control" and "new-checkout"4. Build filters for control variant:   - Property filter: { type: "event", key: "$feature/checkout-flow-test", operator: "exact", value: ["control"] }   - Date range: 2025-01-01 to 2025-01-315. Call query-session-recordings-list with control filters → 147 recordings found6. Build filters for new-checkout variant and call query-session-recordings-list → 152 recordings found7. Compare patterns:   - Control: Average 3m 45s session duration, 12% console errors   - New-checkout: Average 2m 30s session duration, 5% console errors8. Present findings:   "I analyzed session replays for your checkout experiment. The new checkout flow shows:   - 33% faster completion (2m 30s vs 3m 45s)   - 58% fewer console errors (5% vs 12%)   - Users in the new variant navigate directly to payment, while control users often backtrack to review cart   - Recommendation: The new checkout flow reduces friction and errors"

Important notes

Do not make assumptions:

  • Always verify the experiment has recordings before analyzing
  • Check that the experiment is launched (has a start_date)
  • If no recordings are found, inform the user clearly

Filter construction:

  • The $feature/<flag_key> event property is how you scope recordings to a variant
  • One filter per variant — call the tool once per variant with its own filter
  • For boolean flags, use ["true"]/["false"] as the value instead of a variant key

Error handling:

  • If the experiment is in draft state, tell the user it hasn't started yet
  • If no recordings exist, suggest enabling session replay or waiting for user traffic
  • If the variant count is unexpected, double-check the experiment configuration

Related tools

  • query-session-recordings-list: Core tool for retrieving session recordings with filters
  • experiment-get: Get experiment metadata; experiment-results-get for statistical results
  • execute-sql: Query experiments table for details via HogQL

Related skills

  • diagnosing-experiment-results — the quantitative side: bias checks and significance on the same experiment
  • investigating-replay — deep-dive a single session from either variant
  • finding-sessions-to-watch — general session shortlisting outside the experiment context

Source and attribution

Source:PostHog/ai-plugininskills/analyzing-experiment-session-replaysat commit469d177

License: No license

Content belongs to its original authors. SourceWeft indexes it from a public repository.

Report or request removal

More from PostHog/ai-plugin

Writing Simplified Technical English

PostHog

Applies ASD-STE100 simplified technical English rules to make agent-written prose unambiguous and actionable.

Writing & ContentOct 8, 2026

Working With Task Comments

PostHog

Reads and interprets comments on PostHog tasks, artifacts, and canvases through the PostHog MCP exec dispatcher.

Productivity & WorkflowOct 8, 2026

Working With Skills

PostHog

Guides agents in using PostHog's skill-* MCP tools to discover, read, create, update, and refactor skills.

AI & AgentsOct 8, 2026

Working With Scouts

PostHog

Operating manual for delegating watching jobs to PostHog Signals scouts, acting on their reports, and steering the fleet over time.

AI & AgentsOct 8, 2026

Validating And Publishing Canvases

PostHog

Validate and publish a canvas source project safely: the source-project shape, declared capabilities, reading the current version pointer, iterating on validation diagnostics, guarded publishing with expected_current_version_id, staging a draft build and promoting it, waiting out the queued build, and recovering from a 409 version_conflict or a 429 capacity limit without overwriting concurrent work. Use whenever a canvas edit is ready to save, a draft build is wanted, a canvas publish or build returns diagnostics or a conflict, or a task needs to understand canvas version history.

Awaiting classificationOct 8, 2026

Understanding Billing Usage

PostHog

Explains PostHog billing usage and spend from the customer's visible Billing MCP tools. Use when the user asks why usage or spend is high, which product or project is driving usage, what a usage type means, how to reduce usage, what changed over time, why they got a usage change alert, or whether a spike/drop alert was real or noisy. Also use before product-specific analytics skills when the user names a billable PostHog product metric such as events, recordings, feature flag requests, exceptions, survey responses, synced rows, logs, AI events, AI credits, or Inbox credits. Starts from Billing usage/spend tools, then routes to customer-visible product MCP surfaces for deeper investigation.

Awaiting classificationOct 8, 2026