Auditing Experiments Flags

by PostHog469d1773e9cbNo licenseListed Oct 8, 2026Updated Oct 8, 2026

Audit PostHog experiments and feature flags for configuration issues, staleness, and best-practice violations. Read when the user asks to audit, health-check, or review experiments or feature flags, check flag hygiene, or verify experiment setup.

Instructions onlySoftware Development
AI-generated overview

Audits PostHog experiments and feature flags for configuration issues, staleness, and best-practice violations.

What it does
Guides an agent through configuration audits of PostHog experiments and feature flags using read tools such as experiment-get, experiment-list, feature-flag-get-definition, and feature-flag-get-all. It supports quick single-entity checks, scoped audits of one domain, and full audits of both, applying checks from bundled reference files. Findings are reported grouped by severity (CRITICAL, WARNING, INFO) with entity links, a one-sentence problem description, and a remediation action, either inline as markdown or in a notebook when more than five entities have findings.
When to use it
Use when the user asks to audit, health-check, or review PostHog experiments or feature flags, check flag hygiene, or verify experiment setup. It fits requests for a single entity check, a domain-wide audit, or a comprehensive audit of both experiments and flags.
Requirements
Requires access to PostHog experiment and feature flag read tools (experiment-get, experiment-list, feature-flag-get-definition, feature-flag-get-all), and optionally feature-flags-activity-retrieve for activity history checks and notebook tools for larger reports. Ships no scripts; it is instructions plus reference markdown files.

Auditing experiments and feature flags

This skill teaches you how to run configuration audits on experiments and feature flags. All checks use the experiment and feature flag read tools (experiment-get, experiment-list, feature-flag-get-definition, feature-flag-get-all) — no SQL queries are needed for Phase 1 checks.

Usage modes

Quick check (single entity)

When the user asks about a specific experiment or flag:

  1. Fetch the entity via experiment-get (experiment ID) or feature-flag-get-definition (numeric flag ID).
  2. Apply the relevant checks from experiment checks or flag checks.
  3. Report findings inline as markdown, grouped by severity (CRITICAL first, then WARNING, then INFO).
  4. Include entity links as [Experiment: name](/experiments/id) or [Flag: key](/feature_flags/id).

Scoped audit (one domain)

When the user asks to audit all experiments or all flags:

  1. Bulk-fetch via experiment-list or feature-flag-get-all.
  2. Run all checks for that domain against each entity.
  3. Group findings by severity, then by entity.
  4. Report as inline markdown.

Full audit (comprehensive)

When the user asks for a comprehensive audit of both experiments and flags:

  1. Fetch all experiments via experiment-list and all flags via feature-flag-get-all.
  2. Run all experiment checks and all flag checks.
  3. Apply recurring patterns to identify patterns across multiple findings.
  4. If there are more than 5 entities with findings, write them to a notebook for easier navigation. Otherwise report inline. Create the notebook from the project's own notebook tools. Run search notebooks?- to load them and read the titles.

Output format

For each finding, include:

  • Severity badge: 🔴 CRITICAL, 🟡 WARNING, or 🔵 INFO
  • Check name: Which check produced this finding
  • Entity link: Markdown link to the entity
  • What's wrong: One-sentence description
  • Action: What to do about it (see remediation actions)

Example:

🟡 WARNING — Flag integration · Experiment: checkout-redesign The linked feature flag is inactive (paused). Traffic is not being split. Action: Re-enable the flag or end the experiment.

Handling unavailable data

Some checks require activity logs (feature-flags-activity-retrieve for flags), which may not be available in every session. If activity log data is unavailable:

  • Skip checkActivityHistory (experiment check) entirely.
  • Skip the "toggle instability" and "never activated" sub-checks in flag lifecycle checks.
  • In your report, note which checks were skipped and why:

    Skipped: Activity history checks (activity logs not available via current tools)

Partial failures

If a fetch call fails for some entities:

  • Continue with the entities you could fetch.
  • Report which entities could not be assessed and why.
  • Do not silently omit entities from the audit.

Reference files

  • Experiment checks — experiment configuration checks
  • Flag checks — feature flag checks
  • Finding types — severity and category definitions
  • Recurring patterns — patterns across multiple findings
  • Remediation actions — what to do about each finding

Source and attribution

Source:PostHog/ai-plugininskills/auditing-experiments-flagsat commit469d177

License: No license

Content belongs to its original authors. SourceWeft indexes it from a public repository.

Report or request removal

More from PostHog/ai-plugin

Writing Simplified Technical English

PostHog

Applies ASD-STE100 simplified technical English rules to make agent-written prose unambiguous and actionable.

Writing & ContentOct 8, 2026

Working With Task Comments

PostHog

Reads and interprets comments on PostHog tasks, artifacts, and canvases through the PostHog MCP exec dispatcher.

Productivity & WorkflowOct 8, 2026

Working With Skills

PostHog

Guides agents in using PostHog's skill-* MCP tools to discover, read, create, update, and refactor skills.

AI & AgentsOct 8, 2026

Working With Scouts

PostHog

Operating manual for delegating watching jobs to PostHog Signals scouts, acting on their reports, and steering the fleet over time.

AI & AgentsOct 8, 2026

Validating And Publishing Canvases

PostHog

Validate and publish a canvas source project safely: the source-project shape, declared capabilities, reading the current version pointer, iterating on validation diagnostics, guarded publishing with expected_current_version_id, staging a draft build and promoting it, waiting out the queued build, and recovering from a 409 version_conflict or a 429 capacity limit without overwriting concurrent work. Use whenever a canvas edit is ready to save, a draft build is wanted, a canvas publish or build returns diagnostics or a conflict, or a task needs to understand canvas version history.

Awaiting classificationOct 8, 2026

Understanding Billing Usage

PostHog

Explains PostHog billing usage and spend from the customer's visible Billing MCP tools. Use when the user asks why usage or spend is high, which product or project is driving usage, what a usage type means, how to reduce usage, what changed over time, why they got a usage change alert, or whether a spike/drop alert was real or noisy. Also use before product-specific analytics skills when the user names a billable PostHog product metric such as events, recordings, feature flag requests, exceptions, survey responses, synced rows, logs, AI events, AI credits, or Inbox credits. Starts from Billing usage/spend tools, then routes to customer-visible product MCP surfaces for deeper investigation.

Awaiting classificationOct 8, 2026
Auditing Experiments Flags Agent Skill | SourceWeft