Managing Experiment Lifecycle

作者 PostHog469d1773e9cb无许可证收录于 2026年10月8日更新于 2026年10月8日

Guides experiment state transitions: launching, pausing, resuming, freezing/unfreezing exposure, ending, shipping variants, archiving, resetting, duplicating, and copying to another project. Covers preconditions, implications for variant assignment and analysis, and the decision framework for when to use each action. TRIGGER when: user asks to launch, pause, resume, end, ship, archive, reset, duplicate, or copy an experiment to another project, or to freeze/unfreeze exposure (stop enrolling new users while metrics keep flowing, or reopen enrollment). DO NOT TRIGGER when: user is creating an experiment (use creating-experiments), configuring rollout (use configuring-experiment-rollout), or setting up metrics (use configuring-experiment-analytics).

AI 生成的概览

指导实验生命周期的状态转换,如启动、暂停、冻结曝光、结束、发布、归档、重置和复制。

功能
该技能说明实验生命周期中的各项状态转换,逐一解释每个操作的作用、前置条件,以及对变体分配和数据分析的影响。它包含状态图、将情境映射到操作与工具的决策框架表、错误处理表,以及操作之间相互影响的说明。还涵盖开启功能标志清理拉取请求、将实验复制到其他项目等选项。
适用场景
当用户要求启动、暂停、恢复、结束、发布、归档、重置、复制实验,或冻结、解冻曝光时使用。在需要判断某种情境适合哪种生命周期操作时也适用,例如结果显著或结论不明确的情况。
运行要求
仅为说明性指令,不附带脚本。需要访问实验平台的生命周期工具,并需要实验 ID,可通过另一个技能解析获得。

Managing experiment lifecycle

This skill covers experiment state transitions — what each action does, when to use it, and how it affects variant assignment and analysis.

State diagram

text
draft ──launch──▶ running ──end──▶ stopped ──archive──▶ archived                  │ │   ▲              │                  │ pause resume  ship_variant                  │ │   │         (also ends if running)                  │ ▼   │                  │ paused (flag inactive, still "running" status)                  │                  ├─freeze_exposure──▶ exposure_frozen ──unfreeze_exposure──▶ running                  │                    (enrollment closed, metrics keep flowing)
Any non-draft state ──reset──▶ draft

Actions and their implications

For each action, the two key questions:

  1. Who sees what variant? (user perspective)
  2. Who is in my analysis? (statistical perspective)

Launch (experiment-launch)

Transitions draft → running. Activates the feature flag and sets start_date.

  • Preconditions: must be in draft, flag needs 2-20 multivariate variants (no specific key required; the baseline defaults to "control" when present, else the first variant)
  • Pre-launch checklist: has at least one metric? Variants correct? Flag implemented in code?
  • Variants: users start being bucketed into variants based on the configured split
  • Analysis: data collection begins from start_date

No request body needed.

One optional item worth a single mention at launch, when the change is user-facing and substantial: a short survey, shown when users finish the experimented flow (e.g. triggered by the form's submit event), collects qualitative feedback (a rating, an optional comment) alongside the metrics, from day one. Offer it once as setup advice, drop it if declined, and never let it delay the launch. Do not raise it at end or ship-variant time — there it reads as a gate on rolling out. → See references/qualitative-feedback.md in [[diagnosing-experiment-results]]

Pause (experiment-pause)

Deactivates the feature flag. Users fall back to the default experience (typically control).

  • Preconditions: must be running and not already paused
  • Variants: flag is not returned by /decide — no new exposure events recorded
  • Analysis: no new data while paused, but existing data is preserved. Experiment stays "running".

No request body. Use experiment-resume to reactivate.

Resume (experiment-resume)

Reactivates the feature flag after a pause. Users are re-bucketed deterministically into the same variants.

  • Preconditions: must be paused
  • Variants: same assignment as before pause — deterministic bucketing
  • Analysis: exposure tracking resumes

No request body.

Freeze exposure (experiment-freeze-exposure)

Stops enrolling new users while everything else keeps going: already-enrolled users keep their variant, metrics keep flowing, and end_date stays null. Snapshots the already-exposed users into a static cohort and narrows every release condition on the feature flag to that cohort. Status becomes exposure_frozen.

Use for long-horizon metrics (revenue, LTV, retention, renewals) when the sample is big enough and you want to stop adding users without stopping measurement. Neither end nor pause fits that job: end stops measurement at end_date, and pause deactivates the flag for everyone.

  • Preconditions: must be running (not draft, stopped, paused, or already frozen), flag linked and not deleted, at least one release condition
  • Variants: enrolled users keep their variant (deterministic bucketing); new users no longer match the flag
  • Analysis: exposures stop growing (a flat exposure curve is expected), but metric data keeps accumulating for enrolled users

Timing: the exposure scan and cohort snapshot run synchronously inside the API call, and duration scales with the number of exposed persons — an experiment with tens of thousands of exposed users can take on the order of tens of seconds. Set expectations with the user, wait for the response, and don't treat a slow call as a failure or retry it.

Not applicable (400) for:

  • Group-aggregated experiments — the flag targets groups, not persons, and a person cohort can't freeze group-based matching
  • Experiments in a holdout — holdout assignment is evaluated before release conditions, so new users would keep entering the holdout
  • Flags with early access conditions — also evaluated before release conditions, so freezing can't stop new enrollment
  • Mostly-anonymous exposure (e.g. experiments on logged-out surfaces) — anonymous "personless" users can never match a person cohort and would silently lose their variant, so freezes with more than a small unresolved share are rejected
  • Very large exposed sets — the exposure scan is bounded by a person cap and a timeout; over either bound the API returns a clean 400 rather than freezing

When a freeze is rejected, explain which limitation applies rather than retrying — these are structural, not transient.

Interactions with other actions: ship-variant and reset strip the freeze (both also delete the snapshot cohort); end does NOT touch the flag, so ending a frozen experiment leaves the flag narrowed to the snapshot cohort. SDKs using local evaluation can't resolve static cohorts, so a frozen flag evaluates via the /decide endpoint (standard static-cohort behavior). Exposures ingested in the final moments before freezing may miss the snapshot (ingestion lag).

No request body. Use experiment-unfreeze-exposure to reopen enrollment.

Unfreeze exposure (experiment-unfreeze-exposure)

Reopens enrollment on an exposure-frozen experiment. Removes the snapshot-cohort condition and freeze markers from every release group, restoring the flag's original targeting, and deletes the snapshot cohort. Status returns to running.

  • Preconditions: exposure must be frozen (and the experiment not ended)
  • Variants: enrolled users keep their variant; new users can enroll again under the original release conditions
  • Analysis: exposures resume growing

Can introduce bias: reopening enrollment re-exposes the flag to a potentially new population. Users who enrolled before the freeze and those who enroll after the unfreeze joined at different times, and possibly under different conditions — mixing the two cohorts in one analysis can bias the results. Warn the user before unfreezing, especially after a long freeze or if the audience or product changed in between. If they only wanted to sanity-check the frozen results, they may not need to unfreeze at all.

No request body.

End (experiment-end)

Sets end_date and transitions to stopped. The feature flag is NOT modified.

  • Preconditions: must be running (launched, not already stopped)
  • Variants: users continue seeing assigned variants (flag stays active)
  • Analysis: results frozen to data up to end_date

Optional body: conclusion ("won", "lost", "inconclusive", "stopped_early", "invalid") and conclusion_comment.

Use this when you want to freeze results without changing what users see. If the experiment's exposure was frozen, ending does not strip the freeze — the flag stays narrowed to the snapshot cohort (unfreeze first, or ship a variant, if that's not desired).

Ship variant (experiment-ship-variant)

Rewrites the feature flag so the selected variant is served to 100% of users.

  • Preconditions: must be launched (running or stopped). Cannot ship from draft.
  • Variants: ALL users see the shipped variant. The flag is rewritten with a catch-all group.
  • Analysis: if still running, the experiment is also ended (end_date set)

Always confirm with the user before shipping — this permanently rewrites the feature flag.

Required: variant_key (e.g. "test"). Optional: conclusion, conclusion_comment.

Returns 409 if an approval policy requires review before the flag change.

Flag cleanup PR (option on end and ship variant)

Both experiment-end and experiment-ship-variant accept open_cleanup_pr: true. A background PostHog Code task then removes the experiment's feature flag code and opens a draft pull request in the team's connected GitHub repository.

  • Only set this when the user asks for it or confirms it.
  • The key must carry the task:write scope, or the whole request is rejected with a 403 and the experiment is not ended or shipped.
  • The cleanup runs only when the call actually ends the experiment and a conclusion is set — shipping an already-stopped experiment, or ending without a conclusion, skips it. It also requires the team to have the flag cleanup feature enabled; silently skipped when it isn't.
  • repository ("organization/repository") picks the target when several repositories are connected. Omit it to fall back to the experiment's saved repository, the team default, or the only connected repository. With several candidates and no default, the cleanup is skipped unless provided.
  • Track progress with experiment-cleanup-task — the PR URL appears there once opened; a cleanup typically takes several minutes.

Archive (experiment-archive)

Hides a stopped experiment from the default list view.

  • Preconditions: must be stopped (end_date set)
  • Variants: no change — flag is unaffected
  • Analysis: no change — results remain accessible

No request body. Can be restored by setting archived=false via experiment-update.

Reset (experiment-reset)

Returns an experiment to draft state. Clears start_date, end_date, conclusion, and archived.

  • Preconditions: must not already be in draft
  • Variants: flag is left unchanged — users continue seeing assigned variants
  • Analysis: previously collected data still exists but won't be included in results unless start_date is adjusted after re-launch

No request body.

Duplicate (experiment-duplicate)

Creates a copy as a new draft with fresh dates and no results.

Important: always provide a unique feature_flag_key different from the original. If the same key is used, both experiments share a flag — changes to one affect both.

Optional: custom name (defaults to "Original Name (Copy)").

Copy to project (experiment-copy-to-project)

Copies an experiment into a different project in the same organization as a new draft. Use this instead of experiment-duplicate when the copy should land in another project; use duplicate when it stays in the same project.

  • Preconditions: source must not use legacy metrics; target project must be in the same organization and you must have write access to it. Cannot copy across organizations or regions.
  • What's copied: name, description, type, parameters, filters, primary/secondary metrics (fresh uuids), stats and scheduling config, exposure criteria. Not copied: saved-metric references (project-scoped), holdout, exposure cohort, dates, results, conclusion.
  • Feature flag: target_team_id is required; feature_flag_key is optional. The resolved key is then looked up in the target project, and the lookup result — not whether you passed the key — decides what happens:
    • If feature_flag_key is omitted: it defaults to the source experiment's flag key. That key normally doesn't exist in the target project, so a new flag with it is created there. (The default can still collide — see the next point — so to be safe, pass an explicit key.)
    • If the resolved key already exists as a flag in the target project: the copy shares that existing flag instead of creating one. Both experiments then point at the same flag, so lifecycle ops (ship, pause) on either affect both. The existing flag must be multivariate with 2-20 variants, otherwise the call returns 400.
    • If the resolved key does not exist in the target project: a new, independent flag is created with that key. To guarantee independence, pass a feature_flag_key that doesn't already exist in the target.

Confirm the source experiment and target project by name before calling — this writes into a project the user isn't looking at. The returned experiment (and its id) belongs to the target project.

Decision framework

SituationActionTool
Draft ready, flag implemented, metrics setLaunchexperiment-launch
Clear winner, significant resultsShip the winning variantexperiment-ship-variant
No significant difference after sufficient timeEnd as inconclusiveexperiment-end
Something wrong, need to stop exposure temporarilyPauseexperiment-pause
Resume after pauseResumeexperiment-resume
Stop enrolling new users, keep measuring enrolledFreeze exposureexperiment-freeze-exposure
Reopen enrollment after a freezeUnfreeze exposureexperiment-unfreeze-exposure
Experiment ended, ready to clean upArchiveexperiment-archive
Need to start over with same configReset to draftexperiment-reset
Want a similar experiment with a fresh startDuplicateexperiment-duplicate
Want the same experiment in a different projectCopy to another projectexperiment-copy-to-project

Resolving experiments

All lifecycle actions require an experiment ID. If you don't have one, load the finding-experiments skill to resolve the user's reference (name, description, "latest", etc.) to a concrete ID before proceeding.

Error handling

Error messageMeaning
"Experiment has already been launched."Can't launch a non-draft experiment
"Experiment has not been launched yet."Can't end/pause/ship a draft
"Experiment has already ended."Can't end/pause a stopped experiment
"Experiment is already paused."Use resume instead
"Experiment is not paused."It's already active
"Experiment is already in draft state."Nothing to reset
"Experiment is already archived."Already done
"Experiment exposure is already frozen."Nothing to freeze
"Experiment exposure is not frozen."Nothing to unfreeze
"Cannot freeze a paused experiment. Resume it first."Resume, then freeze
"Group-aggregated experiments cannot have their exposure frozen."Structural limitation — don't retry

When you get a 400, explain the situation to the user rather than retrying.

Related skills

  • creating-experiments — create the next experiment from scratch
  • diagnosing-experiment-results — sanity-check results before a ship or end decision
  • configuring-experiment-rollout — split and rollout changes, which are config edits rather than lifecycle operations

来源与署名

来源:PostHog/ai-plugin位于skills/managing-experiment-lifecycle提交469d177

许可证: 无许可证

内容归原作者所有。SourceWeft 从公开仓库中收录这些内容。

举报或申请下架

更多来自 PostHog/ai-plugin 的技能

Writing Simplified Technical English

PostHog

应用 ASD-STE100 简化技术英语规则,让智能体撰写的文字含义明确、便于执行。

Writing & Content2026年10月8日

Working With Task Comments

PostHog

通过 PostHog MCP exec 调度器读取并解读 PostHog 任务、产物和画布上的评论。

Productivity & Workflow2026年10月8日

Working With Skills

PostHog

指导智能体使用 PostHog 的 skill-* MCP 工具来发现、读取、创建、更新和重构技能。

AI & Agents2026年10月8日

Working With Scouts

PostHog

关于如何把监控任务委派给 PostHog Signals 侦察代理、处理其报告并长期调校整个代理集群的操作手册。

AI & Agents2026年10月8日

Validating And Publishing Canvases

PostHog

Validate and publish a canvas source project safely: the source-project shape, declared capabilities, reading the current version pointer, iterating on validation diagnostics, guarded publishing with expected_current_version_id, staging a draft build and promoting it, waiting out the queued build, and recovering from a 409 version_conflict or a 429 capacity limit without overwriting concurrent work. Use whenever a canvas edit is ready to save, a draft build is wanted, a canvas publish or build returns diagnostics or a conflict, or a task needs to understand canvas version history.

待分类2026年10月8日

Understanding Billing Usage

PostHog

Explains PostHog billing usage and spend from the customer's visible Billing MCP tools. Use when the user asks why usage or spend is high, which product or project is driving usage, what a usage type means, how to reduce usage, what changed over time, why they got a usage change alert, or whether a spike/drop alert was real or noisy. Also use before product-specific analytics skills when the user names a billable PostHog product metric such as events, recordings, feature flag requests, exceptions, survey responses, synced rows, logs, AI events, AI credits, or Inbox credits. Starts from Billing usage/spend tools, then routes to customer-visible product MCP surfaces for deeper investigation.

待分类2026年10月8日