CrawlCheck

io.github.emmanuelortav1.0.0更新於 Oct 10, 2026

Verification layer for the agentic web: a signed answer about any site before an agent acts.

已驗證Streamable HTTP可網頁執行Developer ToolsSecurity & MonitoringWeb Search & Scraping

概覽

AI 產生的概覽

CrawlCheck 讓助理取得網站如何對待 AI 爬蟲的簽章讀數:robots.txt 政策、爬蟲身分、可引用內容與代理介面。

功能
CrawlCheck 是託管驗證層,回傳關於網站允許並提供給 AI 爬蟲與代理什麼內容的簽章答案。其讀取器涵蓋依具名代理解析的 robots.txt 政策(含被遮蔽規則)、依營運方 IP 來源驗證爬蟲身分、可引用內容訊號、JSON-LD 實體解析、圖片訊號,以及 markdown 協商與 .well-known 路徑等代理介面探測。開源核心套件只公開這些讀取器;評分、門檻、監控與計費留在託管服務中。
適用情境
當助理在對網站採取行動前需要確認爬蟲是否被允許、聲稱的爬蟲是否真實,或頁面是否含有答案引擎可直接引用的句子時使用。適合 AI 爬蟲存取稽核、日誌偽造檢查與代理就緒度評估。
執行需求
透過 streamable HTTP 連線 crawlcheck.io 遠端端點;未宣告任何套件、環境變數或標頭。開源核心套件需從程式碼倉庫安裝(尚未發布到 npm 登錄),並要求 Node 20+;其爬蟲身分讀取器會連網取得十個營運方 IP 來源。
安裝前請注意
託管服務保留評分、監控、資料集與計費,因此僅有讀數得不到評級。爬蟲身分讀取器是套件中唯一的網路呼叫,會存取營運方 IP 來源;未提供來源的營運方會被標為無法驗證而非偽造。將私有或內部網站交給託管端點前,應先確認其對提交 URL 的處理方式。

安裝

在 SourceWeft 中

  1. 開啟 儀表板中的 CrawlCheck,將其新增到工作區。
  2. 為需要使用其工具的對話啟用該服務。

Web executable,透過 Streamable HTTP。 遠端服務在工作區中設定後即可從網頁執行環境執行。

其他 MCP 客戶端

把它新增到你客戶端的 mcpServers 設定中。

{
  "mcpServers": {
    "crawlcheck": {
      "type": "http",
      "url": "https://crawlcheck.io/mcp"
    }
  }
}

README

crawlcheck-core

[tests]

The dependency-free modules of CrawlCheck, the verification layer for the agentic web, extracted from the production Worker. These are the readers behind its signed answer about what a site permits and serves to AI crawlers and agents. No build step, no dependencies, ESM, Node 20+.

The hosted service fetches a site as 15 crawler identities and grades 22 sections across three stages — reach (can the crawler get in), read (what it actually receives), quote (can an answer engine lift it). This repository is the open core: the readers and checks that need no network, no store and no key, published so the measurements can be reproduced.

What is here today: the robots.txt policy reader, the crawler identifier / log-line forgery verifier, the quotable-content reader, the image-signal reader, and the markdown-negotiation and agent-surface probes. What stays in the service: the scoring — the thresholds each reading is judged against, the weight every section carries in the grade, and the corpus calibration those numbers were set from — plus fetching as fifteen crawler identities, the record over time, monitoring, the dataset, and billing. The rule is simple: a reading of your page is a fact you are entitled to check, and the calibration is the product.

Install

bash
npm install github:emmanuelorta/crawlcheck-core

Not on the npm registry yet, so install it from the repository. Every subpath below resolves from that install.

robots — which crawlers your robots.txt actually admits

js
import { crawlerPolicy } from "crawlcheck-core/robots";
const policy = crawlerPolicy(robotsTxt, /* unreadable */ false, "/wp-admin/");policy.agents.find(a => a.ua === "GPTBot");// { ua:"GPTBot", label:"OpenAI training corpus", role:"train",//   allowed:true, allowed_root:true, named:true, rules_shadowed:true }policy.shadowed;// [{ uas:["gptbot"], missing:["/wp-admin/","/feed/"], missing_count:2 }]

For each of 92 named agents (answer engines, search indexes, training crawlers, SEO tools, social unfurlers, regional engines) it reports: which user-agent group applies (named), whether the given path and the root are allowed under that group, and whether the group shadows the * rules.

The shadowing defect

A crawler obeys only its most specific group. The moment you write

User-agent: *Disallow: /wp-admin/
User-agent: GPTBotAllow: /

GPTBot no longer sees Disallow: /wp-admin/ — the * group stops applying to it entirely. Sites that add an allowlist for AI agents routinely void every rule they thought still held. shadowedGroups() lists each named group and the * disallows it lost. This was found on our own sites first: how to block AI crawlers and the glossary entry for user-agent group.

API

ExportWhat it does
crawlerPolicy(txt, unreadable, path)Full per-agent policy for one file. measured:false with a reason when the file is empty or unreadable — an absence is never reported as a policy
parseGroups(txt)RFC 9309 groups: consecutive User-agent lines share one group; Allow/Disallow rules in order
groupFor(groups, ua)The group that applies to an agent: longest-prefix match on the agent token, * as fallback
pathAllowed(group, path) / rootAllowed(group)Longest rule wins; Allow beats Disallow at equal length
shadowedGroups(groups)Named groups missing any * disallow
AGENTSThe 92-agent table: [token, label, role]

What it does not do

* inside a rule path is stripped and the remainder matched as a prefix; $ end anchors are not honoured; Crawl-delay, Sitemap and unknown fields are ignored. These are the production service's readings too — the tests include byte-for-byte parity against the live /api/tool/robots endpoint on real files, so this module answers exactly what the service answers.

agents — who fetched, and whether they are who they claim

js
import { identifyBot, loadRanges, verifyLines } from "crawlcheck-core/agents";
identifyBot("Mozilla/5.0 (compatible; GPTBot/1.2; +https://openai.com/gptbot)"); // "GPTBot"
const ranges = await loadRanges(fetch);          // the operators' published IP feeds (cache this for a day)const r = verifyLines(accessLogText, ranges);r.tally;            // { verified, spoofed, unverifiable, no_bot }r.forged_rate_pct;  // spoofed / (verified + spoofed), or null when nothing is checkabler.rows[0];          // { line, bot, ip, verdict, why }

Three verdicts, never two. Verified: the address is inside a range the claimed operator publishes. Spoofed: it is outside every range that operator publishes. Unverifiable: the operator publishes no feed (Common Crawl, ByteDance, Meta, Amazon) or the feed could not be read — and that is reported as unverifiable, not as forgery. The forged rate divides by verified + spoofed only; counting unverifiable hits as forged would inflate it and counting them as genuine would understate it.

ExportWhat it does
identifyBot(ua)The named crawler token in a user-agent, longest match wins; null for browsers
verifyLines(lines, ranges)The verifier behind /api/tool/verify: string or array of log lines → tally, rate, per-line verdicts
verifyBotIP(bot, ip, ranges)true / false / null for one hit
loadRanges(fetch)Fetches the 10 operator feeds (OpenAI ×3, Anthropic, Perplexity ×2, Google ×2, Bing, Apple) into the shape the verifier reads. The only network call in the package; absent feeds make their operator unverifiable
ipInCidr(ip, cidr), ip4ToInt, ip6ToBigRange matching, IPv4 and IPv6 (including IPv4-mapped)
BOT_TOKENS, BOT_FEED_OF, BOT_IP_FEEDS, CRAWLERSThe tables: tokens, token → feed family, feed URLs, and the 18-row operator/role/purpose table

Why this exists: CrawlCheck's own telemetry once counted 78 of 85 "GPTBot" hits as forged — they were our own test requests. Reverse DNS is the older check the feeds replace; a user-agent is a self-assertion and proves nothing on its own.

quotable — what an answer engine can lift from the page

js
import { quotable } from "crawlcheck-core/quotable";
const { signals } = quotable(html, "Acme Fencing LLC");
signals.lead_defines;    // does the first sentence say what this ISsignals.answer_units;    // paragraphs of 15-70 words that carry a factsignals.quotable_share;  // % of sentences an engine could attribute uneditedsignals.faq_questions;   // declared in FAQPage markupsignals.faq_visible;     // ...and actually readable on the page

The question is not whether the copy is good. It is whether, when an answer engine wants to state a fact about this page, there is a sentence on it that can be lifted as it stands — or whether the engine has to assemble one, in which case the answer carries the engine's wording rather than yours.

This module publishes the reading, not the scoring. Every count and share the production service records for a page is produced here and is reproducible byte-for-byte against a live scan — that is what the parity test does, and it refuses to compare unless the page bytes still match the ones the fixture was captured from. What the service keeps is what those readings are judged against: the thresholds, the weight the section carries in the overall grade, and the multi-hundred-site corpus pass the numbers were calibrated from.

The distinction is not a marketing line, it is the one that keeps the product honest. Anyone can check that we read their page correctly; nobody has to take our word for a number. What they cannot lift is the calibration, which is the part that took a corpus to earn.

Two details that are easy to get wrong and are settled here. FAQ parity is checked against the whole visible page, not the stripped body — an FAQ accordion inside a <footer> is still text a reader can see. And counts come from body paragraphs and list items with navigation, header, footer and forms removed, because a nav menu is not prose and counting it inflates every ratio on the page.

ExportWhat it does
quotable(html, nameHint)The production call site in one call: parses the JSON-LD graph and returns { signals }
quoteSignals(html, nodes, nameHint)The reader. nodes comes from ldGraphNodes(html); null when the input has no <body>

Scoring the reading — quoteRows, quoteBlock, the section weight and the score version — is in the service, not this package. Point the free scan at a URL to see the rows: https://crawlcheck.io/.

rows — the three-state row every reader returns

js
import { tgt, pctScore } from "crawlcheck-core/rows";

A row is { k, cur, opt, ok, why }: the check, what this page currently does, what it should do, the verdict, and why it matters. ok: null is the third state and it is load-bearing — it means measured but not scored. A signal too rare to judge on, one that does not apply to this kind of site, or one that is purely informational is shown to the reader and deliberately left out of the arithmetic. pctScore() divides by the scored rows only and returns null, never 0, for a section with none: an unmeasured section is not a failing one.

That rule is why a service-area business is not marked down for having no street address, and why a brand-new site with no field performance data is not reported as slow. Most audit tools have two states, and every gap in their own measurement becomes your zero.

surfaces — markdown negotiation, and the Accept header that made a site an F

js
import { agentSurfaces, grab } from "crawlcheck-core/surfaces";
const root = await grab("https://example.com/");const out = await agentSurfaces("https://example.com", root);
out.found.markdown;    // Accept: text/markdown got a text/markdown answerout.found.agents_md;   // an authored /agents.md, not a 200 of HTMLout.found.link_header; // read off the root response — no extra requestout.evidence.markdown; // { content_type: "text/markdown; charset=utf-8" }

A scanner's Accept header is part of its identity, and getting it wrong makes the whole grade a measurement of the wrong document. This fetcher sent Accept: */* until 2026-09-01. vercel.com answered that with 3,025 bytes of text/markdown, while every named-crawler row — which sends text/html — received the 521 KB page. The scan then scored mobile 0 (there is no viewport meta in a markdown file), entity-map parity 0, and graded the site F: a grade about a document no crawler is ever handed. Googlebot, GPTBot and ClaudeBot all send a browser-shaped Accept, so CRAWLER_ACCEPT does too.

Markdown negotiation is therefore measured on purpose, once, in its own request — never as a side effect of how the page was fetched. Conflating "this site offers markdown to agents" with "our fetcher asked for the wrong thing" is the failure mode, and one test in this package exists solely to pin it: exactly one request carries Accept: text/markdown, and every other request carries CRAWLER_ACCEPT.

The other rule here is that a 200 answering with HTML is a soft 404. /SKILL.md, /agents.md and the three .well-known paths are absent on most sites, and a catch-all route returns the SPA shell at 200 for all of them; a naive check reports every one of those as present. Nine of the ten rows this module renders are informational and none is scored — adoption is early, absence is not a defect, and a row that says NOT APPLICABLE: no Product, Offer or checkout found is measuring the site rather than the checklist.

ExportWhat it does
agentSurfaces(base, root)The four .well-known probes fired together, the Link header off root, and the markdown negotiation request
grab(url)CrawlCheck's fetcher: crawler UA, browser-shaped Accept, 10s abort, body read to 60,000 chars
skillMd(host) / agentsMd(host)/SKILL.md with its YAML frontmatter, /agents.md — both with the soft-404 rule
agentSurfaceRows(result)The ten rows, including the applicability wording
looksHtml(body, ctype), UA, CRAWLER_ACCEPT, SURFACE_PATHSThe pieces, exported so a caller can reproduce one probe

jsonld — the entities a page actually declares

js
import { ldGraphNodes } from "crawlcheck-core/jsonld";ldGraphNodes(html); // root nodes and @graph children, merged by @id

Two rules, both easy to get wrong. A node nested inside another node is a value — a PostalAddress, an Offer, a ListItem, an Answer — not a subject the page is asserting, and it is not held to the standards an entity is. And the same @id twice is one subject: RDF merges statements, so a reader that keeps the first node and drops the second reads a site declaring a typeless fingerprint stub before its full Organization as having an untyped business. That was measured on a live site, which is why the merge is here.

images — what an engine can learn from the pictures

js
import { imageSignals, imageRows } from "crawlcheck-core/images";
const sig = imageSignals(html);// { total:9, no_alt:0, empty_alt:0, no_dimensions:0, lazy:7, lazy_above_fold:0,//   modern_format:8, svg:1, og_image:true, ld_image:true, ld_imageobject:true, ... }imageRows({ images: sig });   // eight rows, five scored

Five rows are scored: a missing alt attribute, width and height declared, nothing lazy-loaded above the fold, an entity image in the JSON-LD, an og:image. Three are reported and deliberately never scored, and that is the point of the module.

alt="" is not a missing alt. It is the correct markup for a decorative image, and most real sites are nearly all decorative alt, so scoring it fails almost every site on a signal that is usually right. Only a missing attribute counts. File format and hero priority are preferences rather than defects. And a page with no images returns one unscored row, so the section scores null instead of zero.

The reader carries a fix that came out of writing these tests: the attribute pattern excluded whitespace, so any quoted value containing a space read back as an empty string and alt="a cedar fence" was indistinguishable from alt="". Measured on six live homepages, that reported 15 of 15, 9 of 9 and 10 of 11 images as decorative when almost all carried real alt text.

Tests

bash
npm test            # 48 tests, no network (4 skipped)CC_LIVE=1 npm test  # 48: adds the four live parity checks

robots — RFC group parsing, precedence, the shadowing case both ways, unmeasured inputs, table integrity, and parity with the production service on two real robots.txt files.

agents — token matching, IPv4/IPv6 range logic, the three-state verdict, and parity with the production verifier on 11 real log lines against a stored snapshot of the five operator feeds those lines touch (test/fixtures/ranges.json, dated inside the file).

quotable — the reader (what counts as a paragraph, the lead test, question headings, FAQ parity, what makes a sentence quotable), the rows (which are scored, which are deliberately not, what happens under five paragraphs), the JSON-LD graph rules, and parity with a live production scan.

surfaces — the Accept headers actually sent, one negotiation request and only one, soft-404 handling, and the row wording.

How the parity fixtures work

Each .quotable.json holds the output of one real scan at crawlcheck.io — the service's own numbers, copied unaltered, with the scan id, score version and build sha they came from. The page itself is not committed. Instead the fixture stores page_sha256, the digest of the exact response body that scan read, and the live test re-fetches the page with the same user-agent and Accept, refuses to compare unless the bytes are identical, and only then checks that this module reproduces all 21 fields and the section score.

That makes drift loud instead of silent. A stored HTML blob would keep passing forever against a page that no longer exists; a digest mismatch says plainly that the site changed, not the reader, and the fixture needs recapturing.

One of the two fixtures is deliberately not a re-fetch target. emmanuelorta.com renders live figures on an hourly refresh — measured on 2026-09-03, an identical-length response returned a different digest three hours after capture — so its fixture carries reproducible: false and the reason. It stays because the two recorded readings disagree (100 against 67; nine FAQ questions against none; markdown negotiated against not), which is what proves the reader does not flatten one site into another.

images — the three distinctions above pinned as tests (missing vs empty alt, a page with no images, format and priority unscored), the row contract and the three-state verdict, malformed input, and parity with a live production scan reproducing every field the service recorded from the same bytes.

Roadmap

Next: the machine-file readers (llms.txt, entity map, the soft-404 and challenge verdicts behind them). Each ships only with a parity test against the live service.

Licence

MIT. Built by Emmanuel Orta. The service itself is at crawlcheck.io: a free A–F audit of what AI crawlers and agents are served, no account.

來源:README.md,提交 8816e34

工具

0
工具後設資料尚未被收錄。

版本歷史

1
  1. v1.0.0最新Oct 10, 2026