
Creating Online Evaluations
作者 PostHog469d1773e9cb无许可证收录于 2026年10月8日更新于 2026年10月8日
Author continuously-running online evaluations in PostHog AI observability, grounded in real failure modes you've identified. Use when the user wants evaluations that automatically score new generations or whole traces going forward — "create an eval to catch X", "continuously check that responses do Y", "turn these failures into evals". Covers letting the explored data decide how many evals to create, proposing that set for the user to pick, choosing the target and eval type (hog / llm_judge / sentiment), configuring a provider and model for an llm_judge eval (a provider key gates enabling, not creation), scoping which generations trigger it via conditions, creating disabled, verifying scope, and enabling. Proposes a sentiment eval when no failure mode is worth catching. Finding and ranking the failure modes worth evaluating is its own job — use exploring-ai-failures first. To debug or manage evaluations that already exist, use exploring-llm-evaluations.
- 469d1773e9cb当前提交 469d177发布于 2026年10月8日
来源与署名
来源:PostHog/ai-plugin位于skills/creating-online-evaluations提交469d177
许可证: 无许可证
内容归原作者所有。SourceWeft 从公开仓库中收录这些内容。
更多来自 PostHog/ai-plugin 的技能

Writing Simplified Technical English
PostHog
应用 ASD-STE100 简化技术英语规则,让智能体撰写的文字含义明确、便于执行。

Working With Task Comments
PostHog
通过 PostHog MCP exec 调度器读取并解读 PostHog 任务、产物和画布上的评论。

Working With Skills
PostHog
指导智能体使用 PostHog 的 skill-* MCP 工具来发现、读取、创建、更新和重构技能。

Working With Scouts
PostHog
关于如何把监控任务委派给 PostHog Signals 侦察代理、处理其报告并长期调校整个代理集群的操作手册。

Validating And Publishing Canvases
PostHog
Validate and publish a canvas source project safely: the source-project shape, declared capabilities, reading the current version pointer, iterating on validation diagnostics, guarded publishing with expected_current_version_id, staging a draft build and promoting it, waiting out the queued build, and recovering from a 409 version_conflict or a 429 capacity limit without overwriting concurrent work. Use whenever a canvas edit is ready to save, a draft build is wanted, a canvas publish or build returns diagnostics or a conflict, or a task needs to understand canvas version history.

Understanding Billing Usage
PostHog
Explains PostHog billing usage and spend from the customer's visible Billing MCP tools. Use when the user asks why usage or spend is high, which product or project is driving usage, what a usage type means, how to reduce usage, what changed over time, why they got a usage change alert, or whether a spike/drop alert was real or noisy. Also use before product-specific analytics skills when the user names a billable PostHog product metric such as events, recordings, feature flag requests, exceptions, survey responses, synced rows, logs, AI events, AI credits, or Inbox credits. Starts from Billing usage/spend tools, then routes to customer-visible product MCP surfaces for deeper investigation.
更多AI & Agents技能

Skill Development
anthropics
指导创建 Claude Code 插件技能,涵盖结构、描述、渐进式披露与验证。

Plugin Structure
anthropics
指导 Claude Code 插件的结构、清单与组件布局。

Command Development
anthropics
指导创建 Claude Code 斜杠命令,涵盖结构、YAML frontmatter、参数与插件功能。

Claude Md Improver
anthropics
审查仓库中的 CLAUDE.md 文件,评估其质量,并在获得批准后应用有针对性的改进。

Google Cloud Solution Agentic Ai Data Science Workflow
指导在 Google Cloud 上设计多产品智能体数据科学架构,涵盖需求、部署与验证。

Google Cloud Solution Agentic Ai Borderless Data Lakehouse
指导在 Google Cloud 上设计无边界多云开放数据湖仓,涵盖需求发现、部署与验证。