
Creating Online Evaluations
by PostHog469d1773e9cbNo licenseListed Oct 8, 2026Updated Oct 8, 2026
Author continuously-running online evaluations in PostHog AI observability, grounded in real failure modes you've identified. Use when the user wants evaluations that automatically score new generations or whole traces going forward — "create an eval to catch X", "continuously check that responses do Y", "turn these failures into evals". Covers letting the explored data decide how many evals to create, proposing that set for the user to pick, choosing the target and eval type (hog / llm_judge / sentiment), configuring a provider and model for an llm_judge eval (a provider key gates enabling, not creation), scoping which generations trigger it via conditions, creating disabled, verifying scope, and enabling. Proposes a sentiment eval when no failure mode is worth catching. Finding and ranking the failure modes worth evaluating is its own job — use exploring-ai-failures first. To debug or manage evaluations that already exist, use exploring-llm-evaluations.
Only the file list is public. File contents are available once the skill is installed in a workspace.
| Path | Size | Type |
|---|---|---|
| references/evaluation-payload.md | 11.5 KB | text/markdown |
| SKILL.md | 23.9 KB | text/markdown |
Source and attribution
Source:PostHog/ai-plugininskills/creating-online-evaluationsat commit469d177
License: No license
Content belongs to its original authors. SourceWeft indexes it from a public repository.
More from PostHog/ai-plugin

Writing Simplified Technical English
PostHog
Applies ASD-STE100 simplified technical English rules to make agent-written prose unambiguous and actionable.

Working With Task Comments
PostHog
Reads and interprets comments on PostHog tasks, artifacts, and canvases through the PostHog MCP exec dispatcher.

Working With Skills
PostHog
Guides agents in using PostHog's skill-* MCP tools to discover, read, create, update, and refactor skills.

Working With Scouts
PostHog
Operating manual for delegating watching jobs to PostHog Signals scouts, acting on their reports, and steering the fleet over time.

Validating And Publishing Canvases
PostHog
Validate and publish a canvas source project safely: the source-project shape, declared capabilities, reading the current version pointer, iterating on validation diagnostics, guarded publishing with expected_current_version_id, staging a draft build and promoting it, waiting out the queued build, and recovering from a 409 version_conflict or a 429 capacity limit without overwriting concurrent work. Use whenever a canvas edit is ready to save, a draft build is wanted, a canvas publish or build returns diagnostics or a conflict, or a task needs to understand canvas version history.

Understanding Billing Usage
PostHog
Explains PostHog billing usage and spend from the customer's visible Billing MCP tools. Use when the user asks why usage or spend is high, which product or project is driving usage, what a usage type means, how to reduce usage, what changed over time, why they got a usage change alert, or whether a spike/drop alert was real or noisy. Also use before product-specific analytics skills when the user names a billable PostHog product metric such as events, recordings, feature flag requests, exceptions, survey responses, synced rows, logs, AI events, AI credits, or Inbox credits. Starts from Billing usage/spend tools, then routes to customer-visible product MCP surfaces for deeper investigation.
More in AI & Agents

Skill Development
anthropics
Guides creation of Claude Code plugin skills, covering structure, descriptions, progressive disclosure and validation.

Plugin Structure
anthropics
Guides the structure, manifest, and component layout of Claude Code plugins.

Command Development
anthropics
Guides creation of Claude Code slash commands, covering structure, YAML frontmatter, arguments and plugin features.

Claude Md Improver
anthropics
Audits CLAUDE.md files in a repository, scores their quality, and applies approved targeted improvements.

Google Cloud Solution Agentic Ai Data Science Workflow
Guides design of a multi-product agentic data science architecture on Google Cloud, from requirements to deployment and validation.

Google Cloud Solution Agentic Ai Borderless Data Lakehouse
Guides design of a borderless multicloud open data lakehouse on Google Cloud, from requirements discovery to deployment and validation.