Authoring log alerts
Authoring an alert is a measurement problem, not a guessing problem. You are not trying to be exhaustive — you are trying to land thresholds that fire 0–3 times per week on real production patterns, on services that matter.
When to use this skill
- The user asks to "set up alerts" / "suggest alerts" for their project.
- The user wants to evaluate whether a service is producing alertable signal.
- The user has just enabled log alerting and wants a starter set.
When not to use this skill
- Tuning an alert that already exists — that's a different job (use
posthog:logs-alerts-events-listto inspect fire/resolve cadence andposthog:logs-alerts-partial-updateto adjust). - Investigating an active incident — pull rows with
posthog:query-logs, don't author an alert mid-incident.
Tools
Do not call posthog:query-logs during authoring. You need distributions, not rows. Reserve posthog:query-logs for
the very end if the user asks "show me a sample of what would have fired" — limit: 10 is plenty.
Workflow
1. Triage — pick candidate services
Call posthog:logs-services for the last 24h with no filters. The response is capped at 25 services and includes a
sparkline, so it is small and bounded.
A service is a candidate when both are true:
log_countis non-trivial (≥ ~1k in 24h — quieter services produce too little signal to alert on).error_rateis non-zero, or the user has named the service explicitly.
Skip services with high volume but error_rate == 0 unless the user wants a volume-shape alert (e.g. "warn me
if api-gateway suddenly stops producing logs"). Volume-floor alerts use threshold_operator: below and need
different reasoning — see references/volume-floor-alerts.md.
If the user names a service, treat it as a candidate even without error signal.
2. (Optional) Narrow the filter
If a service has many error sub-types, an alert on "all errors" is usually too broad. Use
posthog:logs-attributes-list (try attribute_type: log) and posthog:logs-attribute-values-list to find a discriminator —
common ones are http.status_code, error.type, k8s.container.name. Add the narrowing filter to your draft.
Keep it simple: one severity filter + one or two attribute filters is plenty. Multi-clause filters are harder to reason about and rarely improve precision.
3. Baseline — characterise the candidate over 7 days
Call posthog:logs-count-ranges with the candidate's filters, dateRange: { date_from: "-7d" }, and
targetBuckets: 24 (one bucket ≈ 7h). The response gives you bucket counts.
Do not eyeball the percentiles or scale the threshold to the alert window manually. Pipe the count-ranges response into the helper script:
The script returns:
Use suggested_threshold_count as your starting threshold. Read health:
4. Draft and simulate
Pick a starter draft from these defaults — see references/threshold-defaults.md for the reasoning:
Call posthog:logs-alerts-simulate-create with these settings and date_from: "-7d". The response gives you fire_count
and resolve_count.
5. Iterate — three rounds, then ship or skip
Target: fire_count between 0 and ~3 over -7d. If outside the band:
When adjusting the threshold, read values from the script's stats block — never recompute percentiles
by hand.
Cap iteration at 3 simulate calls per candidate. If you can't land in the band in 3 rounds, the metric is wrong — either the filter is too broad, the window is wrong, or the service genuinely doesn't have a threshold-shape signal. Note it and move on.
6. Ship — create + attach destination
Once a draft simulates cleanly:
-
Call
posthog:logs-alerts-createwith the validated config. Use a name like<service> error rate (auto)so the user can see at a glance which alerts came from this skill. -
Call
posthog:logs-alerts-destinations-createto wire it to a notification target. An alert with no destination is silent. Supported destination fields:- Slack:
type: "slack",slack_workspace_id, andslack_channel_id.slack_channel_nameis optional. - Webhook:
type: "webhook"andwebhook_url. - Microsoft Teams:
type: "teams"andwebhook_url.
Always confirm the channel name or webhook URL with the user before attaching. Never wire an auto-generated alert to a production channel without explicit confirmation. If the user is unsure, suggest a low-traffic testing channel for the first few alerts.
- Slack:
If the user wants alerts created in enabled: false state for review-then-flip, pass enabled: false to
-create and tell them how many drafts you produced.
Filter shape — required
The filters field on posthog:logs-alerts-create takes a subset of LogsViewerFilters and must contain at
least one of:
severityLevels— list of["trace","debug","info","warn","error","fatal"]serviceNames— list of service name stringsfilterGroup— property filter group
The same shape goes into posthog:logs-alerts-simulate-create's filters field. Match the simulate filters to the alert filters
exactly — otherwise the simulation is testing a different alert than the one you ship.
Example minimum:
Token-economy rules
- One
posthog:logs-servicescall at the start, not per-candidate. - One
posthog:logs-count-rangescall per candidate attargetBuckets: 24. Don't go above 30 during authoring. - ≤ 3
posthog:logs-alerts-simulate-createcalls per candidate. - Zero
posthog:query-logscalls during the authoring loop. - Prefer reporting a small set of well-validated alerts over a long list of unvalidated drafts.
Output
Report what you did, in this shape:
- For each shipped alert: name, filters, threshold, simulated fire_count over 7d, destination.
- For each skipped candidate: service name + why (flat baseline, can't land threshold, low volume).
- Total simulate calls made, total alerts created.
The user should be able to read this and decide whether to disable any drafts before they go live.
Related skills
investigating-logs— characterize a service's baseline before alerting on it, and investigate firings afterauthoring-error-tracking-alerts— alert on exceptions rather than log lines
