Filtering Bot Traffic

by PostHog469d1773e9cbNo licenseListed Oct 8, 2026Updated Oct 8, 2026

Identify, measure, and exclude bot / crawler / AI-agent traffic in PostHog web and product analytics using the traffic classification surface (the isLikelyBot / getTrafficType HogQL functions and the $virt_* virtual properties). Use when the user asks to "exclude bots", "filter out crawlers", "remove bot traffic from my numbers", "how much of my traffic is bots / AI crawlers", "is GPTBot / ChatGPT / Claude hitting my site", "break down traffic by human vs bot", or wants clean human-only counts in an insight or dashboard. For the real-time Live tab bot tiles, use exploring-live-traffic instead.

Instructions onlyData & Analytics
AI-generated overview

Guides filtering and measuring bot, crawler and AI-agent traffic in PostHog analytics using its traffic classification surface.

What it does
This skill explains how to identify, quantify and exclude automated traffic in PostHog web and product analytics. It documents the classification surface — the isLikelyBot and getTrafficType HogQL functions and the $virt_* virtual properties — and gives recipes for human-only counts, traffic-type breakdowns, bot-name and operator breakdowns, and AI-agent measurement. It also covers adding custom bot rules and why non-JavaScript crawlers need server-side $http_log ingestion.
When to use it
Use it when a user wants to exclude bots from analytics, measure what share of traffic is automated, find which crawlers or AI agents hit a site, or break a trend down by traffic type. It targets historical windows, saved insights and dashboards rather than the real-time Live tab.
Requirements
Requires PostHog with captured user agents on events; queries run in HogQL or the insight builder. Custom bot rules need project admin rights and can be managed via the settings UI, MCP tools or the REST API. No scripts are shipped.

Filtering and measuring bot traffic

PostHog classifies every request by user agent so you can tell humans apart from bots, crawlers, and AI agents anywhere HogQL runs — the SQL editor, insights, trends, and Web analytics breakdowns. This skill teaches you (the agent) how to use that classification to:

  • exclude bots so analytics reflect human traffic only
  • measure how much traffic is automated, and which bots / operators are responsible
  • separate AI-agent traffic (worth measuring) from noise (worth dropping)
  • pick the right surface — virtual properties for the insight builder, functions for raw SQL

For real-time ("right now", last 30 min) bot questions and the Live tab tiles, use the exploring-live-traffic skill instead. This skill is for historical windows, saved insights, dashboards, and filtering.

When to use this skill

Use it when the user wants to:

  • exclude or filter out bots ("remove bots from my pageviews", "humans only")
  • quantify automated traffic ("what % of traffic is bots?", "how much is AI crawlers?")
  • find which bots hit them ("which crawlers visit us?", "is ChatGPT reading our docs?")
  • break a trend down by traffic type or bot name
  • measure AI-agent / AI-search traffic specifically (AEO / answer-engine visibility)

Do not use it for the Live tab, real-time numbers, or the per-minute bot charts — that is exploring-live-traffic.

The classification surface

Two equivalent ways to reach the same classification. Prefer virtual properties in the insight builder and filters; use functions in hand-written SQL or when you need a value the virtual properties don't expose.

Virtual properties (insight builder, filters, breakdowns)

These read the user agent for you (falling back from $raw_user_agent to $user_agent), so you don't pass anything in. Available wherever you pick an event property.

PropertyValue
$virt_is_botboolean — true for bots / crawlers / automation
$virt_traffic_typeRegular, AI Agent, Bot, or Automation
$virt_traffic_categoryfiner category, e.g. ai_crawler, ai_search, ai_assistant, search_crawler, seo_crawler, social_crawler, monitoring, http_client, headless_browser, no_user_agent, regular
$virt_bot_namedisplay name, e.g. Googlebot, GPTBot, ClaudeBot
$virt_bot_operatorcompany behind the bot, e.g. Google, OpenAI, Anthropic

HogQL functions (raw SQL)

Pass the user agent explicitly. Use coalesce(nullIf(properties.$raw_user_agent, ''), properties.$user_agent) to cover both server-side ($raw_user_agent) and JS SDK ($user_agent) captures. The nullIf keeps an empty $raw_user_agent from shadowing a real $user_agent and being misread as a bot — this mirrors the expression the virtual properties use internally.

FunctionReturns
isLikelyBot(ua)true if the UA matches a bot/automation pattern (empty UA counts as a bot)
getTrafficType(ua)AI Agent / Bot / Automation / Regular
getTrafficCategory(ua)subcategory; regular for humans
getBotType(ua)same subcategory but empty string for humans — handy for filtering
getBotName(ua)bot name; empty for humans
getBotOperator(ua)operator/company; empty for humans

Traffic types — what to keep vs drop

getTrafficType / $virt_traffic_type sorts every request into four buckets. The default move differs per bucket — don't treat them all as noise:

TypeWhat it isDefault move
RegularHuman visitorsKeep
AI AgentAI crawlers, AI search, AI assistants (GPTBot, ClaudeBot, PerplexityBot, ChatGPT-User)Often measure, don't drop — these are how AI tools find and cite content
BotSearch crawlers, SEO tools, social previews, monitoring (Googlebot, AhrefsBot, Pingdom)Exclude from human metrics; track separately for SEO
AutomationHTTP clients and headless browsers (curl, python-requests, Puppeteer)Usually noise — exclude

Recipes

Exclude bots from an insight (humans only)

Add a property filter $virt_is_bot exact false:

json
{ "key": "$virt_is_bot", "value": ["false"], "operator": "exact", "type": "event" }

Drop it into any TrendsQuery / FunnelsQuery / etc. properties. Visitor, session, and pageview counts then reflect human traffic only, without changing stored data.

To exclude a narrower slice (e.g. keep AI agents but drop monitoring + automation), filter on $virt_traffic_type or $virt_traffic_category with operator: is_not instead.

What share of traffic is automated

Break a pageview trend down by $virt_traffic_type:

json
{  "kind": "TrendsQuery",  "dateRange": { "date_from": "-30d" },  "series": [{ "kind": "EventsNode", "event": "$pageview", "math": "total" }],  "breakdownFilter": { "breakdown": "$virt_traffic_type", "breakdown_type": "event" },  "trendsFilter": { "display": "ActionsBarValue" }}

Which bots / operators are hitting us

Filter to bots and break down by name (or $virt_bot_operator for company-level):

json
{  "kind": "TrendsQuery",  "dateRange": { "date_from": "-30d" },  "series": [{ "kind": "EventsNode", "event": "$pageview", "math": "total" }],  "properties": [{ "key": "$virt_is_bot", "value": ["true"], "operator": "exact", "type": "event" }],  "breakdownFilter": { "breakdown": "$virt_bot_name", "breakdown_type": "event", "breakdown_limit": 25 },  "trendsFilter": { "display": "ActionsBarValue" }}

Measure AI-agent traffic specifically

Filter $virt_traffic_type exact AI Agent, break down by $virt_bot_operator to see which tools (OpenAI, Anthropic, Perplexity, …) read your site and which pages they hit.

Raw SQL equivalents

sql
-- human pageviews onlySELECT count() AS human_pageviewsFROM eventsWHERE event = '$pageview'    AND NOT isLikelyBot(coalesce(nullIf(properties.$raw_user_agent, ''), properties.$user_agent))
-- top bots by hitsSELECT    getBotName(coalesce(nullIf(properties.$raw_user_agent, ''), properties.$user_agent)) AS bot,    getBotOperator(coalesce(nullIf(properties.$raw_user_agent, ''), properties.$user_agent)) AS operator,    count() AS hitsFROM eventsWHERE event = '$pageview'    AND isLikelyBot(coalesce(nullIf(properties.$raw_user_agent, ''), properties.$user_agent))GROUP BY bot, operatorORDER BY hits DESC

Adding a bot PostHog doesn't know yet

The built-in list only covers self-declared user agents PostHog already knows. When a scraper matters to a project but isn't detected — an internal load test, a partner integration, a niche crawler — add a custom bot rule instead of waiting for the built-in list to catch up. Rules extend the same classification surface, so Is bot (isLikelyBot), Bot name, and Traffic category reflect them everywhere HogQL runs.

A rule matches one event property — the user agent by default, but also $ip, $lib, $host, $pathname, $current_url, $browser, $os, $browser_language, $screen_width, $screen_height, $geoip_country_code, $referrer, or $referring_domain — using contains (case-insensitive substring), regex (RE2), or cidr (an IP range, only valid with $ip). Set name to the label reported by Bot name, and optionally category to a built-in category like ai_crawler to relabel the traffic type.

Three ways to manage rules:

  • Settings UI — Settings → Environment → Custom bots.
  • MCP tools — web-analytics-bot-rules-list, web-analytics-bot-rules-create, web-analytics-bot-rules-destroy. Prefer these when driving PostHog through an agent.
  • REST API — GET/POST /api/projects/{project_id}/web_analytics_bot_rules/ and DELETE /api/projects/{project_id}/web_analytics_bot_rules/{id}/.

Listing is open to project members; creating and deleting require a project admin (they mutate the admin-only modifiers team setting). A rule whose pattern can't run is rejected on save, so a bad rule can never take down the project's classification queries.

Seeing bots that don't run JavaScript

Most crawlers and AI agents never execute JS, so posthog-js never fires a $pageview for them — they're invisible to client-side analytics. To measure them, the project must forward server access logs as $http_log events carrying $raw_user_agent. If a user asks "why don't I see GPTBot when I know it's crawling us?", the answer is almost always: no $http_log ingestion. Point them at server-side capture (the Vercel logs source, an edge worker, or the capture API) before building bot insights.

Gotchas

  • Needs a captured user agent. Classification is computed at query time from the event's $raw_user_agent / $user_agent, so it works on any historical event — there's no need to restrict dateRange.date_from. The one requirement is that a user agent was captured; events from sources that never set one can't be classified (and empty UAs fall through to Automation / no_user_agent, below).
  • isLikelyBot is "likely". Detection is a user-agent heuristic — some bots spoof real browser UAs, and some legit tools use bot-like ones. Treat it as best-effort, not ground truth.
  • Empty user agent = bot. Requests with no UA (server-to-server, misconfigured SDKs) classify as Automation / no_user_agent, so isLikelyBot returns true.
  • Don't silently drop the host filter. If the user is scoped to one domain, inherit $host in properties — leaving it out changes the answer.
  • Bot definitions evolve. The detected-bot list changes over time, so re-running the same query later can classify older events differently.

Source and attribution

Source:PostHog/ai-plugininskills/filtering-bot-trafficat commit469d177

License: No license

Content belongs to its original authors. SourceWeft indexes it from a public repository.

Report or request removal

More from PostHog/ai-plugin

Writing Simplified Technical English

PostHog

Applies ASD-STE100 simplified technical English rules to make agent-written prose unambiguous and actionable.

Writing & ContentOct 8, 2026

Working With Task Comments

PostHog

Reads and interprets comments on PostHog tasks, artifacts, and canvases through the PostHog MCP exec dispatcher.

Productivity & WorkflowOct 8, 2026

Working With Skills

PostHog

Guides agents in using PostHog's skill-* MCP tools to discover, read, create, update, and refactor skills.

AI & AgentsOct 8, 2026

Working With Scouts

PostHog

Operating manual for delegating watching jobs to PostHog Signals scouts, acting on their reports, and steering the fleet over time.

AI & AgentsOct 8, 2026

Validating And Publishing Canvases

PostHog

Validate and publish a canvas source project safely: the source-project shape, declared capabilities, reading the current version pointer, iterating on validation diagnostics, guarded publishing with expected_current_version_id, staging a draft build and promoting it, waiting out the queued build, and recovering from a 409 version_conflict or a 429 capacity limit without overwriting concurrent work. Use whenever a canvas edit is ready to save, a draft build is wanted, a canvas publish or build returns diagnostics or a conflict, or a task needs to understand canvas version history.

Awaiting classificationOct 8, 2026

Understanding Billing Usage

PostHog

Explains PostHog billing usage and spend from the customer's visible Billing MCP tools. Use when the user asks why usage or spend is high, which product or project is driving usage, what a usage type means, how to reduce usage, what changed over time, why they got a usage change alert, or whether a spike/drop alert was real or noisy. Also use before product-specific analytics skills when the user names a billable PostHog product metric such as events, recordings, feature flag requests, exceptions, survey responses, synced rows, logs, AI events, AI credits, or Inbox credits. Starts from Billing usage/spend tools, then routes to customer-visible product MCP surfaces for deeper investigation.

Awaiting classificationOct 8, 2026