
Exploring Llm Evaluations
by PostHog469d1773e9cbNo licenseListed Oct 8, 2026Updated Oct 8, 2026
Investigate AI observability evaluations — `hog` (deterministic code-based), `llm_judge` (LLM-prompt-based), and `sentiment` (user-message sentiment). Find existing evaluations, inspect their configuration, run them against specific generations, query individual results, and set up scheduled reports on an evaluation. Use when the user asks to debug why an evaluation is failing, surface common failure modes, compare results across filters, dry-run a Hog evaluator, prototype a new LLM-judge prompt, inspect sentiment classifications, or manage the evaluation lifecycle.
- 469d1773e9cbCurrentcommit 469d177Published Oct 8, 2026
Source and attribution
Source:PostHog/ai-plugininskills/exploring-llm-evaluationsat commit469d177
License: No license
Content belongs to its original authors. SourceWeft indexes it from a public repository.
More from PostHog/ai-plugin

Writing Simplified Technical English
PostHog
Applies ASD-STE100 simplified technical English rules to make agent-written prose unambiguous and actionable.

Working With Task Comments
PostHog
Reads and interprets comments on PostHog tasks, artifacts, and canvases through the PostHog MCP exec dispatcher.

Working With Skills
PostHog
Guides agents in using PostHog's skill-* MCP tools to discover, read, create, update, and refactor skills.

Working With Scouts
PostHog
Operating manual for delegating watching jobs to PostHog Signals scouts, acting on their reports, and steering the fleet over time.

Validating And Publishing Canvases
PostHog
Validate and publish a canvas source project safely: the source-project shape, declared capabilities, reading the current version pointer, iterating on validation diagnostics, guarded publishing with expected_current_version_id, staging a draft build and promoting it, waiting out the queued build, and recovering from a 409 version_conflict or a 429 capacity limit without overwriting concurrent work. Use whenever a canvas edit is ready to save, a draft build is wanted, a canvas publish or build returns diagnostics or a conflict, or a task needs to understand canvas version history.

Understanding Billing Usage
PostHog
Explains PostHog billing usage and spend from the customer's visible Billing MCP tools. Use when the user asks why usage or spend is high, which product or project is driving usage, what a usage type means, how to reduce usage, what changed over time, why they got a usage change alert, or whether a spike/drop alert was real or noisy. Also use before product-specific analytics skills when the user names a billable PostHog product metric such as events, recordings, feature flag requests, exceptions, survey responses, synced rows, logs, AI events, AI credits, or Inbox credits. Starts from Billing usage/spend tools, then routes to customer-visible product MCP surfaces for deeper investigation.
More in AI & Agents

Skill Development
anthropics
Guides creation of Claude Code plugin skills, covering structure, descriptions, progressive disclosure and validation.

Plugin Structure
anthropics
Guides the structure, manifest, and component layout of Claude Code plugins.

Command Development
anthropics
Guides creation of Claude Code slash commands, covering structure, YAML frontmatter, arguments and plugin features.

Claude Md Improver
anthropics
Audits CLAUDE.md files in a repository, scores their quality, and applies approved targeted improvements.

Smb Onboard
anthropics
Guides a small-business owner through first-time setup: connecting tools, running a value-proof recipe, capturing business context, and setting a weekly…

Smb Router
anthropics
Routes a small-business owner's request to the right plugin skill or command and explains what is available.