Analyzing experiment session replays
This skill guides you through analyzing session recordings for experiment variants to understand behavioral differences between control and test groups.
When to use this skill
Use this skill when:
- The user asks to analyze session replays for an experiment
- The user wants to understand how users behave differently across experiment variants
- The user asks to compare user behavior between control and test variants
- The user wants qualitative insights to complement experiment metrics
- The user asks questions like "How are users behaving in my experiment?" or "Show me session replays for variant X"
Prerequisites
Before analyzing session replays:
- The experiment must be launched (not in draft state)
- Session replay must be enabled for the project
- Users must have been exposed to the experiment variants
- The experiment must have a start date
Workflow
1. Get experiment details and feature flag variants
First, retrieve the experiment information and the feature flag variants (source of truth).
Step 1a: Get experiment metadata
You can either:
- Option A: Use the
experiment-gettool if you already have the experiment ID from context - Option B: Query the experiments table via HogQL:
From the experiment data, extract:
feature_flag_key: The feature flag controlling the experimentstart_dateandend_date: The experiment's time range
Step 1b: Get variants from the feature flag
IMPORTANT: Always get variants from the feature flag, NOT from experiment.parameters.feature_flag_variants.
The parameters can be out of sync or deprecated. The feature flag is the source of truth.
Query the feature flag to get the current variants:
Select the variants path directly — selecting the whole filters object gets truncated in results for flags with large targeting configs.
Example structure: [{"key": "control", "name": "Control", "rollout_percentage": 50}, {"key": "test", ...}]
The variant key values (e.g., "control", "test", "variant_a") are what you'll use to filter session recordings.
2. Build session recording filters for each variant
For each variant in the experiment, construct recording filters that match users exposed to that variant.
Filter structure for a variant (input to query-session-recordings-list):
Key points:
- The
$feature/<flag_key>event property records the flag's value on each event — filtering on it matches recordings where the flag was active with that variant. This is an approximation of exposure, broader than the experiment's exposure event ($feature_flag_called, or$experiment_exposureon the new rollout — both deduped per identity): right for browsing behavior across variants, but not an exact mirror of the analysis population — thescanning-experiments-with-replay-visionskill derives that exact filter when you need it valueis an array of variant key strings (e.g.["control"]); for boolean flags use["true"]or["false"]- Avoid the
type: "flag"/flag_evaluates_toproperty filter for variant scoping — the recordings query accepts it but silently ignores it, returning unfiltered results (last verified 2026-06-10). If you want to try it anyway, verify it actually filters first: a query with a nonexistent flag key should return zero recordings - Set the date range to the experiment's start and end dates
- Enable
filter_test_accounts: trueto exclude test users
3. Retrieve recordings for each variant
Use the query-session-recordings-list tool with the filters constructed in step 2.
Call the tool once per variant to get recordings for each group:
- Variant "control" → recordings for control group
- Variant "test" → recordings for test variant
- Additional variants if the experiment has more than 2
The tool returns a list of recordings with metadata including:
distinct_id— the person's distinct IDrecording_duration,active_seconds,inactive_secondsclick_count,keypress_count,mouse_activity_countconsole_log_count,console_warn_count,console_error_countstart_url— first page URL visitedstart_time/end_time,activity_score
4. Compare and analyze
Compare the recordings between variants by looking for:
Quantitative patterns:
- Session duration differences
- Activity levels (clicks, keypresses)
- Console error rates
- Bounce rates
Qualitative insights:
- User confusion or frustration indicators
- Different navigation paths
- Feature discovery patterns
- Error recovery behavior
5. Present findings
Summarize the behavioral differences between variants, highlighting:
- Total recordings per variant
- Notable behavior patterns unique to each variant
- Usability issues or friction points observed
- Recommendations based on the qualitative data
6. Observing shows behavior; asking adds what users think of it
Watching sessions and asking users are different instruments, not substitutes. Recordings show what people did with the change; a short survey, shown when they finish the experimented flow, captures what they thought of it — a rating and an optional comment, readable per variant. For a user-facing change of real size, the two together make a fuller qualitative read than either alone, so mention the option when the behavioral comparison in step 4 leaves opinion unaccounted for, or when a pattern in the recordings is a hypothesis worth checking with the people who produced it. Once per conversation at most; drop it if declined.
Default to asking every exposed user rather than one variant: a popover shown to only one arm is itself a difference between the arms, and the response event carries the variant anyway, so the split survives.
→ See references/qualitative-feedback.md in [[diagnosing-experiment-results]]
Example interaction
Important notes
Do not make assumptions:
- Always verify the experiment has recordings before analyzing
- Check that the experiment is launched (has a start_date)
- If no recordings are found, inform the user clearly
Filter construction:
- The
$feature/<flag_key>event property is how you scope recordings to a variant - One filter per variant — call the tool once per variant with its own filter
- For boolean flags, use
["true"]/["false"]as the value instead of a variant key
Error handling:
- If the experiment is in draft state, tell the user it hasn't started yet
- If no recordings exist, suggest enabling session replay or waiting for user traffic
- If the variant count is unexpected, double-check the experiment configuration
Related tools
query-session-recordings-list: Core tool for retrieving session recordings with filtersexperiment-get: Get experiment metadata;experiment-results-getfor statistical resultsexecute-sql: Query experiments table for details via HogQL
Related skills
diagnosing-experiment-results— the quantitative side: bias checks and significance on the same experimentinvestigating-replay— deep-dive a single session from either variantfinding-sessions-to-watch— general session shortlisting outside the experiment context

