
Octocrawl
io.github.77777R7v0.3.1更新於 Oct 9, 2026
Web scraper for agents: blocked, empty and wrong pages reported as such, with an Evidence Record.
概覽
Octocrawl 將網頁擷取成 Markdown 與結構化欄位供助理使用,並把被封鎖、空白或錯誤的頁面如實回報為失敗,附上證據記錄。
- 功能
- Octocrawl 是為 RAG 與代理工作流程設計的網頁轉 LLM 內容擷取服務。它支援單頁擷取、批次擷取、爬取與網站地圖,回傳 Markdown、連結、中繼資料,以及最多 20 個不需模型即可讀取的頁面欄位。它會從 HTTP 逐級升級到瀏覽器、已儲存的登入或代理,並把結果標記為 empty_verified、blocked 或 failed 及原因,而不是默默回報成功。每筆結果都帶有證據記錄,包含最終 URL、重新導向鏈、狀態、通道、robots.txt 決定與雜湊。本機工具也可匯入、列出與移除已儲存的網站登入。
- 適用情境
- 當助理需要可靠的頁面內容用於檢索或分析,而且你在意擷取是否真的成功時,適合使用。它適用於大量 URL 的批次或爬取工作、帶驗證或需要登入的網站,以及需要可稽核擷取證據的工作流程。
- 執行需求
- 可作為託管遠端端點執行,也可透過 npm 套件 @octocrawl/mcp 以 stdio 在本機執行。本機使用需要 Node.js 22.13+ 或 24+。本機 API 服務以 npx octocrawl serve 啟動(預設 API 為 API 需要 W2L_API_TOKEN 權杖,本機 API 不需要;W2L_API_URL 用來選擇 API。登入匯入與 my-browser 通道需要開啟遠端偵錯的 Google Chrome。
安裝
在 SourceWeft 中
- 開啟 儀表板中的 Octocrawl,將其新增到工作區。
- 為需要使用其工具的對話啟用該服務。
Web executable,透過 Streamable HTTP。 遠端服務在工作區中設定後即可從網頁執行環境執行。
其他 MCP 客戶端
把它新增到你客戶端的 mcpServers 設定中。
{
"mcpServers": {
"octocrawl": {
"type": "http",
"url": "https://mcp.octocrawl.dev/mcp"
}
}
}README
Octocrawl — Web-to-LLM Context Extraction
A transparent, verifiable web extraction system built for RAG and Agent workflows.
Try it without installing anything at octocrawl.dev, read the documentation, or connect your agent to hosted Octocrawl in one line, no key needed to start (how, for Claude Code, Cursor, OpenCode and Codex):
Why This Exists
Most crawlers report "success" when they return empty pages, challenge screens, or the wrong content. Octocrawl makes failure visible and fixable:
Before (typical crawler):
After (Octocrawl):
What's Different
- Failure is a first-class outcome —
empty_verified,blocked,failedwith reasons, not silent empties; a page answered with an error status keeps itshttpStatusand Markdown as evidence, never as success, and so does a page on which Octocrawl finds no main content (failed/empty_unverifiedwith the whole page's Markdown) - Five false-success checks — challenge text, wrong-page content, missing facts, truncation, yield-below-floor
- Execution ladder — HTTP → browser → user auth → proxy, with automatic routing, per-attempt trace, and task-level cost accounting
- Ground-truth benchmark — a 56-case fixture suite (soft 404s, challenge pages, SPAs, timeouts, zip bombs, tables) with verified false-success rates
- Honest evidence —
artifacts: []is an explicit empty artifact list, not a promise that every failed page has a screenshot or DOM snapshot; browserbytesWire: nullmeans wire bytes were not measured - One Evidence Record — every scrape result, batch item and crawl page carries
evidenceRecord(final URL, redirect chain, fetch time, status and reason, lane, robots.txt decision, raw and output hashes, field evidence), stated the same way in every lane and described by a versioned JSON Schema; see the reference
Quick Start
The no-install, single-page web preview runs at octocrawl.dev
(five previews a day; see Public preview for how it is deployed). Besides Markdown it
returns a page's links and metadata, and up to 20 fields read from the page without a model. The same site serves
the documentation: MCP
connection steps for four clients, four task guides, and limits. The pages are generated from
Markdown in this repository during the public web build;
npm run public:preview:local serves them at http://127.0.0.1:8798/docs/.
Install nothing first: with Node.js 22.13+ or 24+,
The packages are published from this repository (0.3.0 on 2026-10-05); the Install check workflow runs the npx line from an empty cache on macOS, Windows and Linux each week. To work on the code, use Node.js 22.13+ or 24+ and npm (the PDF text engine, pdf.js, needs 22.13 or later) and build from this checkout. For the full Monitor → result → HTTPS event → restart workflow, follow the onboarding guide and independent developer acceptance checklist.
octocrawl (the @octocrawl/cli package, also published as octocrawl; npm run w2l -- <command> in a checkout) runs the API's engine in its own process, so every option of the REST API works on the command line. Each option is a flag under its kebab-case name: maxAge is --max-age, onlyMainContent: false is --no-only-main-content, includeTags takes a,b and may repeat, formats takes markdown,tables or a JSON array for an entry with options, parsers takes pdf, none or JSON, headers takes JSON or --header name=value (repeatable), and the URLs are arguments. The request is checked by the REST API's own parser, so a value the API refuses is refused here with the same message (exit code 2).
octocrawl scrape <url>prints the scrape response as JSON (compact, as MCP gets it;--debugfor the full one), or the Markdown alone with--markdown.octocrawl batch <url>...(or--urls-file <file>, one URL per line) andoctocrawl crawl <url>run to the end and print{ report, items }. Ctrl-C leaves the job paused:octocrawl crawl --resume <taskId>continues a crawl (a crawl recorded as pending or running is refused, as another process may be running it), and a batch resumes when the API starts on the same task root.--webhookis not offered, since a command runs no delivery worker: send the job tooctocrawl serveinstead.octocrawl map <url>prints the map.--out <dir>writesresults.jsonl(one result per line),results.csv(one row per page with its evidence, failed pages included: the columns of the Python client'sto_pandas()without the Markdown, thenmarkdown_file),report.jsonfor a job, and per page<n>-<host-path>.mdwith its Markdown and<n>-<host-path>.table-<i>.csvfor each table of thetablesformat.octocrawl serveruns the local API, asnpm run apidoes, with the same flags (--port,--host,--hosted,--token, ...).octocrawl login import <site>saves your login to a site from the Chrome you already use, so--mode authedreads its pages signed in as you;octocrawl login listandoctocrawl login remove <site>show and forget saved logins without printing a cookie. See Your own login (mode authed).
Exit codes: 0 for a page read as content or a completed job, 1 for anything else Octocrawl answered, 2 for a refused command line, 130 when interrupted. The task root, where tasks, saved files and the page cache live, is --task-root, else W2L_TASK_ROOT, else .w2l/cli, apart from the API's .w2l/api; never point a command at the task root of a running API server, which could run the same job twice. A command never resumes the task root's earlier jobs, as the API does when it starts. The earlier in-process ladder tool is w2l-ladder (and w2l-fetch) in @w2l/bench.
Your own login (mode authed)
Mode authed reads a page with the login you saved for its site, in Octocrawl's own browser. To save one, sign in to the site in Google Chrome (144 or later), in the default profile and a normal window; open chrome://inspect/#remote-debugging and turn on "Allow remote debugging for this browser instance" once; then run octocrawl login import example.com (a domain or a page URL). Chrome asks "Allow remote debugging?" for every connection: click Allow. Octocrawl connects once, reads that site's cookies (its own, a parent domain's, and its subdomains' when the site sets cookies on the name itself) and the localStorage of the site's tabs you have open, saves them and disconnects; it never reads Chrome's files on disk, and reads a tab's storage without loading anything or running script in it. Logins are kept in W2L_SESSIONS_FILE, else ~/.w2l/sessions.json, readable by you alone, and the command line, the local API and the local MCP service read the same file. Records carry the login's SHA-256, cookie count and localStorage origins and item count, never a value. The same import works without the command line on a server running on your machine: POST /v1/logins/import with { "site": "example.com" } (and approveTimeoutMs, how long to wait for your Allow, 10 s to 10 min, default 2 min) answers { domain, savedAt, cookieCount, localStorage, localStorageRead, sessionSha256 } once you click Allow (localStorage: { origins, itemCount } saved, or null; localStorageRead false when no tab of the site was open, so none was read; localStorageUnread, the origins of open tabs Chrome did not give the storage of, a crashed or discarded tab, which a reload brings back). GET /v1/logins lists the saved logins and DELETE /v1/logins/:site forgets one. The SDK's importLogin, listLogins and removeLogin, and the local MCP tools import_login, list_logins and remove_login, call these routes. A local MCP server offers all three. import_login lets an agent choose which site's login to save, so its description tells the agent to ask you first, and Chrome's Allow is still yours to click. No route answers a cookie. A server that is not on loopback, the hosted service, and a server without your Chrome refuse an import with 409.
- A login for
example.comcoverswww.example.comand other subdomains. When one applies, the authed rung goes first, before the public rungs, which would take a logged-out page as the answer; if the site refuses the login (login_wall, for example when it expired), the run ends there instead of returning the logged-out page. A refusal is alogin_wallpage, a redirect to the site's login page, or a page that asks for a sign-in in place: a heading or plain line among its first eight that begins with the request, such as "Log in to view your wishlists", "Please sign in", "You must be logged in" or "Login required". A header's Sign in link, a table row, a list item, a line with a link, a line where the words come after others ("Step 2: Sign in to continue") and prose further down are the page's content, not a refusal; so is a heading the page's title names (an issue or a question whose subject begins with those words, "Please log in again #1411"). The rule matches English only.ladder_session_rejectednames which:redirectedToorsignInPrompt. Import the site again to replace an expired login; a running server uses the new one at once. - Mode
authedworks on scrape and batch, not on crawl: a crawl follows every link, and a sign-out link would end your session in Chrome too, since the saved cookies are that session. Send the pages as a batch. - A site that keeps its login in
localStorage(a token its script reads) needs a tab of it open when you import: Chrome reads an origin's storage only through a page that shows it, so with none open the import saves the cookies alone and sayslocalStorageRead: false. A tab whose storage Chrome does not give (it crashed, was discarded or closed meanwhile) is saved without it and named inlocalStorageUnread, with the request to Chrome that failed and Chrome's answer inlocalStorageUnreadReasons. Only tabs in the profile the cookies come from are read, never an Incognito window's, and only each tab's own origin, not a frame of another origin inside it;sessionStorageand IndexedDB are not saved. Chrome's remote debugging reaches the default profile only. - A site may tie its login to more than the cookies (a server-side session, the browser, the network address); Octocrawl does not imitate your browser, so such a site answers as if you were signed out (see the I1 run).
Enhanced access (an access grant)
For people who do not want the technical switches, one option chooses how pages are reached: "access": "standard", "enhanced" or "my-browser" on POST /v1/scrape and POST /v1/batches (standard or enhanced on POST /v1/crawl), the MCP scrape, batch_scrape and crawl tools, and --access on the command line.
standard: Octocrawl's own fetching and its local browser. No rung that costs a third party runs; the routing audit says which it dropped (ladder_channels_filtered, reasonaccess standard).enhanced: also what the server's access grant of tier enhanced approves: the paid providers below, in any mode, mode standard included. The grant's run budget caps a batch or a crawl; a single scrape calls each provider the server names at most once. A server without such a grant refuses it by name (unsupported_parameter).my-browser: your own Chrome, the same as"lane": "my-browser"(below). A crawl does not take it.
Without access, a request runs as the server is configured. A batch or crawl keeps its choice, so a resumed run makes the same one. The per-route options below stay for developers.
By default a server uses none of the capabilities ADR 0005 puts behind a grant: no provider browser, no challenge solving, no provider stealth. An operator who wants them starts the server with a grant, a JSON file the server checks at startup; a grant with any problem stops startup with every problem listed.
W2L_ACCESS_GRANTtakes the same, as a file path or the JSON itself.- A capability ADR 0005 refuses (identity rotation, patching your own Chrome, and the others it lists) or defers (an own browser engine, Camoufox, and the others) is an error, not ignored; a deferred one names the ROADMAP row that says when it restarts.
- A
standardormy_browsergrant may name onlycompatible_transportandegress_sessions. A capability that can cost money needs a positiveperRunUsd, and anenhancedone needs an attestation. - Provider rungs (
W2L_VENDORSwith its key, in moderesearchorauthed) are built only when the grant namesvendor_remote_browser; the provider's challenge solving and stealth followvendor_captcha_solvingandvendor_stealth. - The ladder goes on to its next rung (the browser, then a provider) for a block or check it recognises, and also for a 403 or 405 answered without one (a site's bot defence often answers so), for a page a browser rendered with no main content it could verify, and for the HTTP rung's refused connection (its next rung is the browser, whose own network failure ends the run); the trace says why (
ladder_stepwithescalate). When the stronger rungs then fail without a page (a network error, a provider's own error, the deadline), the page it stepped past is the answer (ladder_evidence_kept). It stops at a timeout (a slow or dead site would hold the scrape for the browser's wait too), at a rate limit (429, whoseRetry-Aftera batch or crawl honours), at an error status that is the page's own answer (404, 410, 5xx), at your saved login's rung (the rungs after it do not carry the login), and after an earlier rung found content, which stays the answer. - A provider is called only at a known price ceiling (ROADMAP PA item 4).
tariffsnames, per provider, the prices you accepted from its pricing page (perCallUsd,perHourUsd) and what bounds a call's time: a per-hour price needsmaxSessionMs, the longest session one call may hold;minBilledMsis the shortest time the provider bills a session, andbillingIncrementMsthe step it bills in (a minute, rounded up). A call's ceiling isperCallUsdplusperHourUsdfor its longest session, at leastminBilledMsand at least the shortest timeout the provider takes (60 s for Browserbase, 15 s for Steel), rounded up to the step. Under a tariff each call opens a connection and a session of its own, never shared with another call, ends it atmaxSessionMsand releases it; the session is also created with the provider's own timeout and its proxies off, so it ends on the provider's side if the release fails, and bills no bandwidth. A per-GB price is refused: the bytes a provider's browser receives across every target it opens cannot be counted from here, so a bandwidth cost has no ceiling. A provider with no tariff is not called (ladder_channel_skipped,no price ceiling). - Spend is reserved before each call and settled after it, in one ledger per task that every page, retry and provider of a batch or crawl shares, and that a resumed or appended run opens with what the task was already charged: a call is made only when its ceiling fits what
perRunUsd(and the task's own cap) has left, so concurrent pages cannot pass it, and it is settled at the price the provider reported or, when it reported none (Browserbase and Steel state no price per request), at its ceiling, never below the cost; a call that threw or that the deadline cut is charged at its ceiling.perRequestUsdcaps one page's calls within the run, and a single scrape's (elseperRunUsddoes). A call whose ceiling does not fit is skipped and the trace says so (budget); a task whose cap is used up stops (budget_exceededwithcost). The answer keepsusage.externalCostUsdfor the exact cost (null when unknown) and addsexternalCostChargedUsd, what the ledger charged all the run's calls, with aspend_settledtrace event per call. A provider called with no budget at all is reserved and settled in an uncapped ledger of the run's own, so its calls are on the record too. Each page's Evidence Record lists them inaccess.paidCalls, in order: the provider (provider), the rung, the ADR 0005 capabilities its session was created with (capabilities), the ceiling reserved (ceilingUsd), what the ledger charged (chargedUsd), the price the provider stated (reportedCostUsd, null when it stated none), andoutcomeandreason, what Octocrawl made of the page the call returned by its own checks (a block page, an empty or unverified read, an identity it did not send), never the provider's word that it succeeded, null when the call returned no page;answermarks the call whose page is the record's.access.grantnames the grant they were made under by the SHA-256 of its text (shasum -a 256 grant.jsongives the same), its tier and its attestation time, never the attestation's principal or statement. A page read again on another egress keeps the calls of the read it gave up, a page you read in your Chrome after a check keeps those of the run the check stopped, and a batch or crawl page whose run threw after a paid call keeps it (none of these is itsanswer); a page from the cache lists those of the fetch it reuses. Both arenullwhen no provider was called. A model's cost for JSON extraction is reported apart (modelUsage) and not counted. A grant that namesscope.hostsis refused, since nothing limits the routes to those hosts yet. W2L_BROWSER_ENGINE=patchrightruns the public browser rung on Patchright, a maintained Playwright fork, when the grant namesenhanced_browser; a hosted server refuses it, and a saved login's rung and a managed session keep stock Playwright. Patchright is not installed with Octocrawl: install the two together (npm install octocrawl patchright, ornpx -p octocrawl -p patchright octocrawl ...), thennpx patchright install chromium. AnexecuteJavascriptstep still runs in the page's own JavaScript world on it. A page fetched on it says so in its trace (browser_engine). It stays an experiment: on 2026-10-06, over one exit, it reached no more pages than stock Playwright and lost none (record), so Octocrawl does not turn it on by default.W2L_COMPAT_HOSTS=example.com,shop.examplesends those hosts' pages, and their subdomains', over a browser-compatible HTTP transport (impit) when the grant namescompatible_transport: thehttprung becomeshttp_compat, in standard mode on a local server; a hosted server refuses it. The request carries Chrome's own headers and TLS handshake, so the page records that Chrome identity (identity_sent) and the transport (transport). A request with customheadersormobilekeeps thehttprung, since the transport cannot send them without changing Chrome's header set, and so does a map's start page, read under the identity the map reports. impit decodes compressed bodies itself, so such a page reports its wire size as unknown. WithoutW2L_COMPAT_HOSTS, a grant that namescompatible_transportuses it for the hosts its acceptance showed it helps: five hosts (research/access/benefit-hosts.v1.json) whose blocked task it verified in both G1 windows with no regression (research/access/runs/2026-10-07-g1-acceptance-5afe577.md). It is not for every host: across the whole task set it also turned refusals into answers whose data was wrong.W2L_COMPAT_HOSTS=noneturns it off; naming hosts replaces the list.egress_sessionsgives each batch and crawl (not a single scrape, nor modeauthed, which has your saved login) a cookie session: the cookies its pages set are sent again to their site on its later pages, by the HTTP rung, the compatible transport and the browser alike, so a page whose check the browser cleared lets the task's next page of that site go over HTTP. Cookies are matched to domain and path as a browser does, kept in the task's directory (cookie-session.<route>.json, one per egress route, readable by you alone) so a batch or crawl resumed after a restart goes on with them, deleted when the task ends, and never recorded: a page's trace names the session's random id and counts (session_cookies). A page read with a session is not cached.W2L_EGRESS_PROXIES=http://user:pass@proxy-a:8080,http://proxy-b:3128(withegress_sessions; a hosted server refuses it) sends every fetch through your own proxies. A batch or crawl keeps one for its run, and its cookie session belongs to that one; it moves to the next healthy proxy only when its proxy itself fails: after a page got no HTTP answer from any rung, Octocrawl asks the proxy for a tunnel to a name that does not exist, and moves on only if the proxy does not answer or refuses its credentials (407), at most twice a run, with a new cookie session, and the page is read again there (egress_switchedin its trace). It never moves because of what a site did: a block, a challenge, a 429 or a connection the site reset stays on its proxy, since moving to another address to get past one is identity rotation, which Octocrawl does not do. A proxy that failed is set aside for 10 minutes; a scrape takes the next healthy one. A task resumed on another route (the pool or the proxy changed) starts a new cookie session; modeauthednever uses the pool. Records name a proxy byhost:port(proxy), never its credentials. WithW2L_EGRESS_ECHO_URLset to a service that answers with the caller's address (https://ipinfo.io/json,https://api.ipify.org,https://httpbin.org/ip), each proxy is asked it through itself once, again after it failed or after 10 minutes, and every page read through that proxy records where it left from asaccess.egress.exit({ ip, country, observedAt }, the country when the service gives one); without it, or when the echo did not answer,exitisnull. The echo request goes to the service you name, through your proxy, on a connection of its own: a gateway that gives each connection or session a new exit may have sent the page from another address, soexitis where that proxy was seen to leave from, not proof of the page's own address. A page from the cache keeps the exit its original read recorded, and a page never waits for the echo past itstimeout(the exit is thennull).vendor_unlock_htmlandthird_party_captcha_solvercan be granted, but the routes that use them are not built yet (ROADMAP PA items 4 and 6).- The
octocrawlcommands readW2L_ACCESS_GRANTtoo, as their engine runs in their own process. - The grant is part of the page cache key, so a page fetched under one grant is not reused under another.
A check you get through yourself (handoff)
By default Octocrawl does not solve captchas or challenges and does not disguise itself (see the access grant above for what a grant changes). When a page of a batch stops at one (blocked with captcha, cloudflare_challenge, bot_detected_generic or login_wall), Octocrawl on your own machine can hand it to you in the Chrome you already use: octocrawl batch <urls> --handoff, POST /v1/batches/:id/handoff on a local server (octocrawl serve on loopback), or the local MCP tool hand_off_batch. Remote debugging must be on, as for octocrawl login import, and Chrome asks "Allow remote debugging?" once per handoff. While remote debugging is on, every page Chrome opens sees navigator.webdriver as true, whether Octocrawl is connected or not (seen 2026-10-07 on Chrome 153 with the chrome://inspect switch, and on Chrome 154 started with --remote-debugging-port): a site that looks for it, as bot checks may, takes your Chrome for an automated one, and Chrome shows "Chrome is being controlled by automated test software". Octocrawl does not hide it, since it never changes your Chrome. Turn remote debugging off at chrome://inspect/#remote-debugging when you are done.
For one page, ask for it in the request: octocrawl scrape <url> --handoff, or "handoff": true (or { "waitMs": 60000 }) on POST /v1/scrape, the SDK's scrape and the MCP scrape tool. When Octocrawl's own fetch stops at such a check, the page opens in your Chrome in the same way, and the scrape answers with the page you get through to. That answer is recorded as below, and its routing audit is the stopped run's. A page that is not read answers as stopped, with a handoff_not_through warning that says why. The scrape's timeout bounds Octocrawl's fetch, not your time; the SDK waits for the answer as long as the handoff takes. Without the option, a stopped page on a server that offers the handoff carries handoff: { reason, liveViewUrl: null, rationale }, saying how to ask for it. A server that does not offer the handoff refuses the option (unsupported_parameter), as it does beside actions or a screenshot.
-
While the batch waits. It runs to the end as usual. The batch's status counts the stopped items in
waitingForPerson. Each such item carrieshandoff: { reason, liveViewUrl: null, rationale }, wherereasoniscaptcha_required,bot_gateorlogin_required. A rate limit or a region block is not handed over. -
What happens when you hand off. Each stopped page opens in a new tab of your Chrome, one at a time. You get through the check there, as you would on your own. Octocrawl then reads the page and closes the tab. A page counts as through when, on three reads a second apart, all of these hold:
- it has loaded and its document answered 2xx;
- Octocrawl's gate, given the document's own status and headers, finds no check on it;
- it is on the site asked for (that host, a subdomain or a parent domain) and not on a login path;
- you are not at a step of your own: no password or one-time-code field showing on the page, and no form field whose value you are changing (a search box that only has the focus is not one).
The page's address is the one Chrome shows, not what the page's script says. A page that reloads itself (a challenge that runs a script, then reloads) is waited for, not taken for a closed tab. A way through that ends elsewhere on the site (a sign-in that lands on the home page) is followed by Octocrawl taking the tab back to the page asked for, twice at most; a tab still elsewhere after that is not read. The page asked for is the URL itself, or where the URL leads when Octocrawl takes the tab back to it, or that page with its address rewritten by its own script once it came, with no new document and before you clicked or typed on it (Indeed drops its paging token and names the job it shows): the same path, the same value for every parameter both addresses name, and no number dropped (a
start=10orpage=2that disappears may mean the site fell back to its first page). An address your own click or key moved in place (Next, a sort) is not the page asked for, nor is another page of a list; the tab is taken back.Octocrawl reads a page in your Chrome only after you acted in its tab: a click or a key press that Chrome itself counts as a user's (the document's user activation, read in a world of Octocrawl's own that the page's script cannot reach, or a navigation Chrome marks as made with a user gesture). Nothing the page does by itself counts: not a reload, a redirect, a check that passes on its own, or a script filling a field. A page that shows no check in your Chrome (you are already signed in there, say) is read only once you click on it;
octocrawl batch --handofftells you so, and until you do the item keeps its stopped result. A click counts only in the tab Octocrawl opened: when that tab stays out of sight for 3 s (another tab or window in front of it),octocrawl scrape --handoffandoctocrawl batch --handofftell you to switch to it. So one Allow never lets a caller read the sites you are signed into without you; the my-browser lane below reads without a click only on the sites you allowed in Octocrawl's own page in your Chrome. To read pages with your login and no handoff, import it for the site (octocrawl login import) and run the batch in modeauthed. -
What replaces the stopped result. Only a read that is the page (
successorpartial) replaces the item's stopped result, in the formats the batch asked for, under the item's own id; a read that still shows a check, or is an error or empty, leaves the stopped result standing. The page is recorded as what it is: lanebrowser_local_authed, modeauthed,compliance: null(Octocrawl sent nothing, so it signs nothing),usage.requestCount: 0, the User-Agent unknown (identity_unobserved: your browser sent the request), the document's status andContent-Typeas your browser received them, and the robots.txt decision of the stopped fetch. The trace records the stopped result inhandoff_fromand the read inuser_browser_read, with how you acted (act:user_activationorgesture_navigation) and the check Octocrawl saw (sawGate); the stopped run's routing audit is dropped with it. -
When a page is not read. You may not get through in time (
waitMs, 10 s to 30 min, default 10 min per page; a page left on another site, or one you did not click on, is given up when that time ends), close the tab or quit Chrome; or the caller may go away (the request's connection closes, Ctrl-C on the CLI), or Octocrawl may shut down; each ends the wait and closes the tab. A page Chrome refuses to open or answer for is not read, and the others are still handed over. That item keeps its stopped result, and the answer says why:{ id, handedOff, through, notThrough, items: [{ id, url, through, status, reason? }] }. -
A list that stopped at a check. For a batch whose only step is
paginatewith anitemSelector(the items are what tells a page of the list from another page you open in that tab), the page the check was on opens in your Chrome; for a pager whose pages have no address of their own, that is the list's own page, and the pages read before the check are shown again on the way without counting againstmaxPages. You get through it there, then page on yourself by clicking Next: Octocrawl only reads that tab (a read-only script; it clicks nothing and sends nothing), and stops when Next has been gone, hidden or disabled on the last page it read for 5 s, at the step'smaxPages, after 60 s without a new page, or at the handoff'swaitMs(default 10 minutes in all); wait for the terminal's or the tool's word that the check's page was read before you click Next, and do not close the tab (that, or quitting Chrome, gives up the item, pages read included); if you got through but showed no page after the check's within that time, the item keeps its stopped result and a later handoff goes on from there. The pages Octocrawl's own browser read before the check and the pages you showed it are merged into one list, each page once (actions.scrapes[].byisuser_browserfor yours); the item becomes the whole list,actions.lists[].continuedis{ from, pages, by: "user_browser" }, andstoppedBysays how the reading ended (end,max, ordeadlinewith alist_not_exhaustedwarning). -
Where it is offered. Only on a server that runs on your machine and answers you alone: a server not on loopback, and the hosted service, answer 409. A batch that asked for page
actionsor ascreenshotis not handed over either (409, nohandoffon its items): those are Octocrawl's browser's to take, and a page read in your Chrome cannot give them. Cancelling while Chrome still asks "Allow remote debugging?" drops the connection, so an Allow clicked later attaches to nothing. The handoff runs on a finished batch, not while it runs. It sends no webhook event for an item it replaces: read the items again. -
octocrawl serve(andnpm run api) reads saved logins only when it listens on loopback, and then answers only requests addressed to127.0.0.1,localhostor[::1]with no foreignOrigin, so another machine or a web page using a rebound DNS name cannot read pages as you. Listening on another address, it reads none and says so when it starts. A hosted server never reads them.
Your own Chrome (lane my-browser)
Remote debugging, which this lane needs, makes every page see navigator.webdriver as true while it is on (see the handoff above). The one exception is a batch whose only step is paginate with an itemSelector: an item whose list stopped at a check (actions.lists[].stoppedBy is challenge) is handed over as above; its other stopped items are not. A page that only needs your login, your address or a real browser is read; a site whose bot check looks at navigator.webdriver may refuse your Chrome as it refuses Octocrawl's own lanes. In the 2026-10-07 acceptance run studylib.net and imf.org were read this way, while crunchbase.com (a Cloudflare block page) and stackoverflow.com (a Cloudflare challenge that did not clear) were not; why those two refused is not isolated.
On a server running on your machine, a scrape can read its page in your own Chrome instead of fetching it: octocrawl scrape <url> --lane my-browser, or "lane": "my-browser" on POST /v1/scrape, the SDK's scrape and the MCP scrape tool. Octocrawl fetches nothing itself; your Chrome loads the page, signed in as you where you are.
- You allow it twice per connection. Chrome asks "Allow remote debugging?" (remote debugging on, as for
octocrawl login import). Octocrawl then opens a page of its own in your Chrome that lists the site and the task, and waits for you to click Allow reading these sites there. Only that click, which Chrome counts as yours, allows the site: a script cannot. Closing that page, or clicking Revoke, stops it, and a page being read then ends ascancelled. Not allowing the site in time (10 minutes), closing the page or clicking Revoke first refuses the request (409), and the refusal says what the page last answered (not clicked, or allowed without a click Chrome counted). A batch waiting for you to allow its sites says so in its status (waitingForApproval: true), and the server's log shows each step of the approval (my_browser_approval), never a page's content. - The page is read without a click, on that host only. It is read when it has loaded, answered 2xx, shows no check and is on the host you allowed, exactly: a page that leads to a subdomain or a parent domain of it is not read without you, however the site links them. Otherwise it is judged as a handoff judges a page (three reads in a row). A page that shows a check waits for you to get through it (
handoff.waitMs, default 10 min). A page not read answerscancelled,blockedwith the check it still showed,failed/connection_errorwhen Chrome refuses a command (a tab it will not open),failed/redirect_limitwhen the site leads the tab to another page of it each time Octocrawl takes it back (twice), orfailed/timeout, with amy_browser_not_readwarning that says why. - How it is recorded. Lane
my_browser, never cached;compliance: nullandusage.requestCount: 0(Octocrawl sent nothing), no robots.txt decision (it fetched nothing), and the Evidence Record'saccesssaysroute: user_browser, the browser that read it, andcompletion:user_browser, orhanded_to_personwhen the page showed a check you got through. - Where it is offered. As the handoff: only on a server on your machine, answering you alone; elsewhere, and with page
actions, ascreenshot,lockdownor a mode other than standard, the request is refused by name (unsupported_parameter). - Batches.
"lane": "my-browser"onPOST /v1/batches, the MCPbatch_scrapetool oroctocrawl batch <urls> --lane my-browserreads every page that way, one at a time. Octocrawl's page in your Chrome lists every site of the batch (host and port) and you allow them once for the run; a page on a site not among them is not read, nor any page after you revoke them. Not allowing them answers every pagecancelled, and Chrome not reached answersfailed/connection_error, each with the reason. A run resumed later (a restart, a pause) asks you again: a server restarted while such a batch ran opens Chrome's prompt as it starts, and pages left when you do not answer endcancelled. A URL appended while the batch runs, on a site not in that run's list, endscancelledtoo, and is not read later. Cancelling the batch, or its time budget ending, closes the tab being read and the connection. A page here has notimeoutof its own (your time is yours), only the batch's. A batch on this lane takes no webhook (pages read as you are not sent elsewhere) and nomaxConcurrencyabove 1.
Every Evidence Record's access.completion counts how a page was read: unattended (Octocrawl's own lanes), authorized_session (your saved login, mode authed), user_browser (your Chrome, on a site you allowed, without a step of yours) or handed_to_person (your Chrome, after you got through a check). It is null when no page was read.
For researchers, two guides walk through a real run: From a URL list to a CSV with evidence (the command line and the Python client, every evidence column, and why failed rows stay) and Citing web data in a paper (a methods section, a reference with its access date and hash, and personal data).
The ports the local services listen on, all on 127.0.0.1:
For MCP use there are three ways, from least to most setup; the configs for Cursor, OpenCode and Codex, and a first task, are on Connect MCP.
Hosted (scrape and map; keyless within a daily allowance over HTTP, a key for more pages and the browser lane; see docs/hosted-api.md):
On your computer (everything: scrape, map, crawl, batch, the Amazon.sg product tool and the Monitor tools; no limit; nothing from this checkout is needed):
Self-hosted for others: npx octocrawl serve --hosted --token <token> listens on all interfaces behind a bearer token, with private addresses, robots overrides, saved logins, handoff and non-HTTPS webhooks refused; point @octocrawl/mcp at it with --base-url and --token (or W2L_API_URL and W2L_API_TOKEN).
The checkout's managed local service is for the Monitor → HTTPS delivery flow: one background service runs the API,
Monitor scheduler, delivery worker and a Streamable HTTP MCP endpoint at http://127.0.0.1:8791/mcp. On macOS,
install it as a LaunchAgent and connect Codex to its loopback URL:
It restarts after a process crash and at login. No hosting or sign-in account is
needed for this local path. npm run local:mcp:uninstall removes the agent;
codex mcp remove w2l-local removes the client entry. On other systems, run
npm run local:mcp in one terminal. The state stays in .w2l/api by default.
See the MCP first-use walkthrough for the actual
Monitor and HTTPS delivery flow and secret setup. Keep this checkout while
the LaunchAgent points to it.
To receive signed events on the same Mac with a fixed HTTPS loopback URL,
run npm run local:receiver:install, then reinstall the MCP service with
W2L_LOCAL_DELIVERY_LOOPBACK=1 npm run local:mcp:install. The option only
permits loopback delivery and pins trust to the generated local certificate.
The receiver and its SQLite inbox run as a separate LaunchAgent; neither
service becomes reachable from another machine.
The legacy standalone REST API remains available for SDK and Firecrawl-shim clients:
To connect a standalone stdio MCP process to that API, run:
npm run api binds 127.0.0.1 and allows loopback/RFC1918 so fixture servers work. Hosted mode is explicit: npm run api -- --hosted --token $W2L_API_TOKEN. That binds 0.0.0.0, requires Authorization: Bearer, denies private/metadata IPs, and limits a crawl to 100 pages: an omitted or null maxPages takes 100, and a larger one is refused with invalid_request. It obeys robots.txt for every URL and offers no way past it: robotsOverride, robotsOverrides and ignoreRobotsTxt are refused with unsupported_parameter (see below).
A server started with tokens, hosted or local, accepts any one of them: repeat --token, or set W2L_API_TOKEN and the comma-separated W2L_API_TOKENS. Tokens on the command line replace those in the environment. A --token without a value (the last argument, followed by another flag, or blank) stops the server at startup, and the error never repeats a token. Give each client its own token; restarting the server without a token revokes it. Tokens are compared as fixed-length SHA-256 digests in constant time, and a missing or unknown token gets HTTP 401 with { "error": "unauthorized", "code": "unauthorized" }. The SDK sends its token option, or W2L_API_TOKEN from the environment when none is passed; token: '' sends none.
An operator can cap how many requests that start work each caller may make: W2L_RATE_LIMIT_PER_MINUTE=<n> or --rate-limit-per-minute <n> (an integer from 1 to 100,000; unset or empty means no limit, anything else stops startup with W2L_RATE_LIMIT_PER_MINUTE must be an integer between 1 and 100000) counts POST /v1/scrape, /v1/crawl, /v1/batches, /v1/map, /fc/v1/scrape, /fc/v1/crawl and /fc/v1/map in a sliding 60-second window per bearer token (for the one local caller when the server takes no token); status reads are free. Over the limit the answer is HTTP 429 with a Retry-After header (whole seconds, at least 1) and { "error": "rate limit exceeded: <n> requests per minute", "code": "rate_limited", "retryAfterSeconds": <s>, "agentHints": ["wait <s> s before the next request"] }; /fc answers { success: false, error, code: "rate_limited", agent_hints } with the same header. rate_limited is not one of the request-error codes: the request was well formed, the caller's budget was spent. The SDK throws W2LError with status 429, code rate_limited, retryAfterMs read from the header (delta-seconds or an HTTP date) and agentHints, and retries nothing, as Firecrawl's SDKs do not; MCP tool calls fail with rate limited: retry after <s> s (rate_limited). The window is in memory and per process, so a restart resets it, and it keys on the token's digest, not the client address, so rotated tokens have separate budgets. Nothing about outbound politeness changes: the per-origin gate and the Retry-After cooldowns toward sites stay as they are.
Behind a proxy, local mode (npm run api, the local MCP service, npm run scrape/crawl) sends its outbound requests, including robots.txt and the local browser, through HTTPS_PROXY for https: URLs and HTTP_PROXY for http: URLs (lower-case names too), with curl's rules: NO_PROXY hosts and their subdomains, host:port, IP and CIDR entries go direct, * disables the proxy, and loopback is always direct. The proxy must be http:// or https://, and both variables must name the same one. The proxy resolves the names it fetches, so for proxied requests Octocrawl trusts it for resolution and checks only the URL itself (scheme, credentials, IP literals, metadata names); direct requests are still resolved, validated and pinned. Results name the proxy's host:port in evidence.envProxy and an egress_proxy trace event, never its credentials. W2L_PROXY=off ignores the variables; hosted mode never uses them. Without the environment proxy, the local browser connects directly like the HTTP lane: it never falls back to the operating system's proxy settings, a route no result would record. The macOS LaunchAgent does not inherit your shell, so put these variables in .w2l/local-mcp.env. Octocrawl verifies certificates by default, through the proxy too (the proxy tunnels TLS end to end); skipTlsVerification turns it off for one local request, is recorded, and is refused in hosted mode (see the scrape options below).
robots.txt is read for every URL Octocrawl fetches, and its verdict is recorded; what it decides depends on who chose the URL (decided 2026-10-05). robots.txt addresses crawlers that discover links, so on a local server a URL the request names (a scrape, a batch entry, the CLI's URL list, MCP scrape and batch_scrape, /fc/v1/scrape) is fetched whatever robots.txt says, as a browser visit would be: the result keeps the verdict (robotsDecision.decision: "disallowed", userOverride: true, overrideBasis: "user_named_url"), a robots_overridden warning and the trace events below. The links a crawl or map discovers obey robots.txt, unless the crawl or map was started with ignoreRobotsTxt on a local server (overrideBasis: "ignore_robots_txt"; a crawl then also reads the sitemap files robots.txt disallows, and a map returns the URLs it disallows, each link's robots saying disallowed or unreachable). A Monitor's scheduled re-reads obey it too. A hosted server obeys robots.txt for every URL, since it fetches from the operator's addresses. A site owner can address Octocrawl itself: robots.txt User-agent lines are matched against the request's User-Agent with the product token Octocrawl added, in every mode, so a group for Octocrawl (or research mode's w2l-research) governs it whatever header was sent. Such a rule is the owner's targeted opt-out: a named URL and ignoreRobotsTxt do not set it aside, and only a robotsOverride with your recorded reason does, on a local server. Pacing is unchanged by any of this: a batch and a crawl space a host's pages by its Crawl-delay, whether robots.txt allows the page or not, and a 429 cools the host down for every request. A 4xx robots.txt means no restrictions. A robots.txt that cannot be fetched (a 5xx, a network error, or no answer within 5 seconds) is a complete disallow, as RFC 9309 §2.3.1.4 requires: where robots.txt is obeyed the page is not fetched, the result is failed with policy_denied, and the robots_checked and robots_disallowed trace events and the compliance record's robots decision (browser and provider lanes) carry unreachable: "server_error", "network_error" or "timeout", so it never reads like a rule the publisher wrote. Octocrawl asks for that robots.txt again after five minutes (robotsUnreachableTtlMs in the network policy). On a direct connection a name that does not resolve is still dns_error; through the environment proxy, which resolves names itself, its robots.txt request fails first, so the page is policy_denied with unreachable: "network_error". Where a URL is fetched past it (a named URL, ignoreRobotsTxt, robotsOverride), the unreachable reason stays in the trace and the robots_overridden warning says the file could not be read. On a local server a scrape or batch entry may also carry your own recorded reason (robotsOverride, described with the scrape options below); ignoreRobotsTxt on a scrape or batch is refused with HTTP 400 naming it.
Research mode (mode: "research", --mode research on the command line) declares Octocrawl as a bot in its User-Agent. Set W2L_CONTACT to say who runs it, as a name and email address or a URL, for example W2L_CONTACT="Jane Doe [email protected]": printable ASCII, at most 200 characters, no parentheses or backslashes. The research User-Agent then ends ; contact: Jane Doe [email protected]). To sec.gov and its subdomains it takes the format SEC's fair-access policy prescribes, <Company or name> <email>, instead: W2L Research Jane Doe [email protected] (so give an email address). Their robots.txt is requested with it too, a robots.txt group for w2l-research still applies there, and the Evidence Record's identity.contact reads the contact from either format. Results keep the User-Agent that was sent (the identity_sent trace event on the HTTP lane, the compliance record's sentHeaders in the browser lane). Standard mode sends a browser User-Agent and declares no contact. SEC.gov answers 403 to automated clients that declare no contact; on the HTTP lane such a 403 from an SEC host carries a declared_contact_hint trace event saying to use mode: "research" with W2L_CONTACT set. npm run api, the local MCP service (in .w2l/local-mcp.env) and npm run scrape/crawl read the variable.
The unified local MCP covers scrape, map, Crawl, persistent URL-array batches, and
Monitor/Delivery without separate worker terminals. A unified service also
implements authenticated Streamable HTTP for the reviewed public-document
Monitor and anonymous Amazon.sg product JSON/batch flows (experimental; its setup is
archived in docs/archive/hosted-mcp-pilot.md); hosting is paused on the roadmap
(ROADMAP.md). For both flows on one Mac, run
npm run first-use:local after npm ci; see the
two-flow first-use guide,
MCP first-use walkthrough and
C2/C3 status. Advanced clients may
still launch the legacy stdio adapter from this repository:
MCP scrape is compact by default: it returns the selected content, document/product metadata, aggregate usage and errors without repeating the body under summary.attempts, with the response's status and content-type header in snapshot (httpStatus, contentType) and the final URL and every redirect hop in evidenceRecord. Pass debug: true when you need the full route, trace and per-attempt audit. REST and SDK calls that omit formats and debug keep the legacy full Markdown response.
For many known URLs, use batch_scrape in MCP, then get_batch, get_batch_items, or wait_batch. REST and SDK support the same durable task and paginated items. REST streams a crawl or a batch as it runs (GET /v1/crawl/:id/events and GET /v1/batches/:id/events as server-sent events, /ws on each for a WebSocket), and the SDK's watcher(jobId, { kind }) follows a job over those routes or, where they are off, by polling (the watcher paragraph below). See batch scraping. The per-origin concurrency ceiling is configurable up to four, with a shared Retry-After cooldown and minimum request interval. The controlled 1/2/4 comparison is local fixture evidence, not an Amazon speed claim. A batch also takes maxConcurrency (an integer from 1 to 4): the most of its pages in flight at once, across all its hosts. It only lowers the service's worker count (4 locally, 2 on the hosted MCP host; GET /v1/batches/:id reports the cap in force as maxConcurrency), the per-host ceiling and minimum interval still apply, and the cap is stored with the task, so a batch resumed after a restart runs under it. ignoreInvalidURLs: true starts the batch with the entries of urls that are http(s) URLs and reports the rest as invalidURLs on the 202 and on the status, where without it one such entry refuses the whole request by its index (urls[2] must be http(s)); an entry that is not a string, or a duplicate, is refused either way. Firecrawl's extract scope flags allowExternalLinks and includeSubdomains are taken on a batch as false only, which already holds (a batch fetches the URLs given and follows no link); true is HTTP 400 naming the crawl option that does it (allowExternalLinks: true is not offered on a batch: a batch fetches only the URLs given; a crawl takes allowExternalLinks, and extraction across links is the M5 multi-URL extract; includeSubdomains points at the crawl's allowSubdomains). GET /v1/batches/:id reports succeeded (items recorded success, partial or empty_verified) and failed (the items the errors report lists) beside completed. GET /v1/batches/:id/errors (SDK getBatchErrors, MCP get_batch_errors) lists the items that did not succeed across every attempt of the batch, so a batch interrupted and resumed keeps its earlier failures on the record, each as { id, timestamp, url, status, code, error, httpStatus } (Firecrawl's names; code is the item's failureReason, blockReason or budgetExceeded), in pages of up to 1000, with robotsBlocked, the URLs robots.txt refused: a policy_denied item with a robots_disallowed trace event that no recorded override set aside. A governance or SSRF refusal is policy_denied too, but not robots.txt, and stays in errors only. A batch also takes idempotencyKey (1 to 200 characters without control characters; also the x-idempotency-key or Idempotency-Key header Firecrawl's clients send, on POST /v1/batches, POST /v1/crawl and /fc/v1/crawl; a crawl start takes the same field): a submission sent again with the same key and the same body answers the first submission's taskId with replayed: true and starts nothing, so a retry after a dropped connection does not fetch every URL twice; the same key with another body is HTTP 409 conflict (idempotency key was used for a different request), a body key that differs from the header is HTTP 400, and a key lives 24 hours, in an index beside the task directories (<task root>/idempotency.sqlite) kept by the one API process that runs that task root. appendToId: "<taskId>" adds urls to an existing batch instead of starting a new job (SDK appendToBatch(id, urls, options), MCP batch_scrape with appendToId): the job keeps its mode, formats, includeLinks, maxConcurrency and page options (sending one is HTTP 400 appendToId keeps the job's options; formats cannot be changed; ignoreInvalidURLs, idempotencyKey and robotsOverrides for the new URLs may come along), the URLs go to the end of its list (at most 1000 in all; a URL already in the batch is refused by name, appended url is already in the batch: <url>; a URL on a new host is fetched like the others, under its own robots.txt and SSRF checks), a running batch picks them up in the same attempt (no job is added, so an active-batch limit does not count the append), a completed one runs again for them in a new attempt (active again, it counts against such a limit as a new batch does: on a server that runs one batch at a time the append is HTTP 400 active batch limit reached while another batch is active, and the batch is left as it was), and a cancelled or failed one is HTTP 409 conflict (batch is cancelled); the 202 carries requested (the job's URLs now) and appended, and GET /v1/batches/:id counts the longer list. For a list longer than 1000 URLs the SDK's batchScrapeChunked(urls, options, { chunkSize, itemLimit, pollIntervalMs, timeoutMs }) runs one batch per chunk of chunkSize URLs (default 100), each started, waited for and listed before the next starts (a caller's idempotencyKey becomes <key>:<chunk index> per job), and returns { jobs, items, invalidURLs } with the items in the order the URLs were submitted; chunkUrls(urls, chunkSize) splits a list alone. The hosted MCP host's batch_scrape refuses idempotencyKey and appendToId as it refuses the other batch options.
A crawl (POST /v1/crawl, MCP crawl) takes the same formats and includeLinks as scrape, plus includePaths / excludePaths: regular expressions matched against the URL path of each discovered link. The start URL is always fetched and an excludePaths match wins. A pattern that can backtrack catastrophically on a crafted link path, such as ^/(a+)+$ or .*a.*b, is refused with invalid_request: a repeated group with a repeated part and no separator, a repeated choice whose alternatives can start alike, or three or more overlapping repeated parts in a row. The rest run on V8's linear-time regular expression engine where it can run them; one with a lookaround, a backreference or a counted repetition above 16 (such as {3,40}) runs on the backtracking engine, only on paths of up to 2,048 characters and with a 100 ms limit per link. A link a filter cannot decide is not followed, and a filter that ran out of time decides no later link either. Scrape, batch and crawl reject an unknown field or an unsupported format with HTTP 400 naming it.
A crawl follows links on the start URL's host, that host's www. twin (example.com and www.example.com) and the host the start URL redirects to, and only inside the start URL's path subtree: a start URL ending in / scopes the crawl to that directory, one naming a file (/3/tutorial/index.html) to the file's directory, any other to itself and the paths under it (/search admits /search?page=2 and /search/x, not /searching); a redirect of the start URL adds the final URL's subtree. crawlEntireDomain: true follows links anywhere on that host, as crawls did before the option existed. allowSubdomains: true adds every host under the start URL's apex (the host with one leading www. removed; there is no public-suffix list, so a start URL on www.gov.uk admits every *.gov.uk host), allowExternalLinks: true adds every host and cannot be combined with allowlistedDomains, and a non-empty allowlistedDomains adds the hosts it names (exact or *.domain) to the start URL's own. Hosts other than the start URL's are not path-scoped, and every page on a new host gets its own robots.txt read, SSRF check and identity record on whichever lane serves it. Governance follows the same rule: with allowlistedDomains the ladder may fetch the start URL's host, its twin, the listed hosts and *.apex under allowSubdomains; otherwise the frontier alone bounds the crawl. A crawl does not follow links whose path ends in an image, font, stylesheet, script, audio, video or program extension (.png, .woff2, .css, .js, .mp4, .exe and the like); documents and data files such as PDF, CSV, XLSX, JSON, XML and ZIP are followed.
includePaths / excludePaths match a link's pathname; with regexOnFullURL: true they match its canonical URL (scheme://host/path?query, with the host lower-cased, the default port and tracking parameters dropped and the query sorted), so a pattern can name a host or a query. ignoreQueryParameters: true treats URLs that differ only in their query string as one page: the first variant seen is fetched (its url keeps the query, its canonicalUrl does not) and later ones are reported as collapsed. deduplicateSimilarURLs (default true) does the same for /a and /a/, / and /index.html (also .htm, .php), www. and the apex, and http and https; a page's canonicalUrl stays a real URL, and a site that serves different pages at /a and /a/ needs the option off. A page fetched and then found to repeat an earlier page's body (the same rawBodySha256) keeps status duplicate in the checkpoint, is left out of GET /v1/crawl/:id/pages unless includeDuplicates=true (MCP get_crawl_pages, SDK getCrawlPages) and out of a /fc crawl status's data, and still counts in that status's total. Every crawl reports what became of the links it found: GET /v1/crawl/:id carries discovery (offered, enqueued, duplicate, collapsed, hostDenied, subtreeDenied, pathDenied, depthDenied, duplicateContent; null for a batch), written after every page, and each page's trace (pages?debug=true) carries a discovered event (via: seed or link, from: the linking page) and a links_offered event with that page's counters and up to 20 collapsed (url, into) and host-refused links as samples. A crawl task stores all of these options; one stored before they existed resumes with its original rule, the whole host and exact canonical URLs. The octocrawl crawl CLI has no flags for these options yet: it keeps following the whole host (crawlEntireDomain) and takes the other defaults.
A crawl reads the site's sitemap by default (sitemap: "include", as Firecrawl does): the files the start URL's robots.txt names in Sitemap: lines, or /sitemap.xml when it names none. A <sitemapindex> is followed one level, its children in listed order; a .gz file, or one whose bytes start with the gzip magic number, is inflated under the 50 MiB decompression cap; at most 20 files are read per attempt, and the load stops once it holds as many entries as maxPages (50 000 when the crawl is unbounded). The entries go to the frontier at depth 1, after the start URL and ahead of its links, under the same host, subtree, includePaths / excludePaths and maxDepth rules as a link (so maxDepth: 0 drops them all), so a bounded crawl of a site with a sitemap returns different pages than it did before the option. sitemap: "only" follows no page link: the pages are the start URL and the sitemap's entries, and each page's links are still returned when asked for. sitemap: "skip" reads none. A sitemap file is an auxiliary fetch like robots.txt, never a page: it is requested with the crawl mode's own http identity (its User-Agent and client hints, the mobile ones when the crawl's pages declare them, and nothing a caller added), after the SSRF check on its URL and every redirect hop, through the same DNS-pinned route or operator proxy, paced by the origin scheduler, within the policy's redirect limit and 10 MiB wire cap, and only after its own URL passed its host's robots.txt under that identity (a disallowed or unreachable robots.txt refuses it, as it would a page, unless the crawl was started with ignoreRobotsTxt). It never runs through the ladder or the extractor, and it has no signed compliance record: the record of these fetches is the report's discovery.sitemap (mode, sources: robots or guess, files, listed, enqueued, truncated, error), where each file carries its url, finalUrl, status, contentType, bytes, sha256, kind (index, urlset, absent for a 4xx, not_sitemap for a 2xx body that is neither, unreadable with the reason, such as body_too_large or decompressed_too_large, or refused), entries, robots verdict and proxyUsed. A 4xx or an HTML body is never an error; an unreadable child is recorded and the load moves to the next. Sitemap fetches are spaced by the scheduler's minimum interval; the robots.txt Crawl-delay governs the crawl's page starts from the first page on. A page found through the sitemap carries discovered { via: "sitemap", from: <the file's URL> } in its trace, and the sitemap's entries count in discovery beside the links. The octocrawl crawl CLI reads no sitemap yet, and a crawl task stored before the option resumes without one.
maxConcurrency caps the pages one crawl fetches at once: an integer from 1 to the service's worker count (W2L_WORKER_COUNT, 4 by default on a local API; 2 on the hosted MCP host; a larger value is refused with maxConcurrency must be at most N on this service). It only lowers a crawl's parallelism: the per-origin ceiling (W2L_PER_HOST_CONCURRENCY, at most 4) and the minimum interval still apply, so it never raises the host ceiling. The evidence of the setting is in the pages themselves: each page's createdAt and usage.wallMs give its fetch interval. A resume keeps the value.
GET /v1/crawl/active (SDK getActiveCrawls(), MCP list_active_crawls, local only) lists the crawls this API process is running, those it started and those it resumed at startup, oldest start first: { crawls: [{ id, url, status, startedAt, pagesFetched, options }] }, where options are the crawl's stored options (maxPages, maxDepth, allowlistedDomains, includePaths, excludePaths, useCached, sitemap, the URL-scope options, maxConcurrency and scrapeOptions, the per-page options with formats and includeLinks). It is always 200, { "crawls": [] } when nothing runs; a batch is never listed (it has GET /v1/batches/:id), and there is no team id, since Octocrawl has no teams. The list is this process's own crawls: a crawl another process runs on the same task root is not in it.
A crawl or batch takes webhook (REST, SDK, MCP crawl and batch_scrape): where the job posts its events as durable, retried deliveries, a URL string or { url, headers, metadata, events, secretEnv }. Five events: started (sequence 0) when the job is accepted, one page per page recorded (every outcome: success, partial, failed, blocked, duplicate; the payload's page is the page as GET /v1/crawl/:id/pages or /v1/batches/:id/items lists it, trace empty, no audit, its json included when a json format was asked for), then completed, failed or cancelled with the job's status as GET /v1/crawl/:id or GET /v1/batches/:id reports it then (report, and error on failed). events narrows them (default all five; a filtered event is never enqueued, so nothing pending or dead-lettered appears for it; a cancelled job is cancelled, never failed). Each payload is { schemaVersion: "w2l.job-event/v1", eventId, sequence, jobId, jobKind: "crawl" | "batch", event, at, metadata, page?, report?, error? }: eventId is <taskId>:started, <taskId>:page:<stepId> or <taskId>:<status> (a job that runs again, a resumed crawl or a batch appended after completion, suffixes its later terminal event with the attempt id), or <taskId>:handoff:<stepId> for a batch item a handoff replaced with the page the person got through to, on a batch recorded before 2026-10-05 (a batch with a webhook is no longer handed over: a page read in the person's Chrome is read signed in as them), and sequence counts the job's events in order, 0 and then one per page, terminal and handoff event, so a job of n pages ends at n+1 and each item handed over after it adds one; metadata is the request's (at most 32 strings of at most 1000 characters, 8 KiB in all), {} when none, copied into each payload when the event is enqueued, so a retry resends the identical body. Every request carries content-type: application/json, x-w2l-event-id, x-w2l-event-version (the sequence) and x-w2l-delivery-id, plus x-w2l-timestamp and x-w2l-signature (sha256= HMAC over <timestamp>.<body>) when secretEnv names an operator W2L_WEBHOOK_SECRET_* variable (never a literal secret), and the request's headers (at most 32, 8 KiB in all, RFC 7230 token names lower-cased, no line breaks; content-type, content-length, host, connection, transfer-encoding and every x-w2l-* name are refused by name, webhook.headers: content-length is reserved), sent on every attempt, retries included, after Octocrawl's own, which they never override. Header values are stored in the control database (section-b-control.sqlite, mode 0600) alone: the task, the status and every delivery route show their names only (headerNames), so a bearer token for the receiver belongs there and nowhere in a response; secretEnv remains the recommended signing path. A terminal event is enqueued only after the task row is written terminal, so a completed delivery arrives after GET /v1/crawl/:id already says completed. The destination is job:<taskId> (GET /v1/delivery/destinations?jobId=<taskId>); GET /v1/deliveries?jobId=<taskId> and GET /v1/deliveries/page?jobId= list the deliveries, each with its attempts at GET /v1/deliveries/:id, retried with backoff and Retry-After and dead-lettered after the attempt budget as a Monitor's are (POST /v1/deliveries/:id/retry replays one), and the job's status reports webhook: { destinationId, url (origin and path, no query), events, pending, delivered, deadLetter }. Event ids are deterministic and a delivery is unique per event, so a resume or a restart offers every persisted page again and sends none twice, and a finished job whose events a crash cut off is completed when the API starts, a batch item a handoff replaced included; an event offered again takes the number it would have had, or, when another event holds that number, the next one after every number sent. The receiver must be https; a local server (npm run api, the local MCP service) also takes plain http to a loopback receiver (127.0.0.1, ::1, localhost, sent direct over node:http), as fixture servers are allowed, and refuses any other http URL with HTTP 400 webhook.url must be https (http is accepted only for a loopback receiver of a local service); a hosted server (--hosted) refuses every http receiver the same way and a private or metadata address with webhook.url must be a public address, and dead-letters a name that resolves privately (webhook egress denied: <violation>). w2l-api delivers them itself: it runs the delivery worker the MCP runtime runs, under the delivery policy it prints at start (TLS always verified, the shell's HTTPS_PROXY never used, W2L_DELIVERY_PROXY_URL for an explicit proxy, W2L_DELIVERY_CA_FILE trusted, W2L_DELIVERY_PRIVATE_ALLOWLIST for private https receivers in hosted mode), so npm run delivery:worker is no longer needed beside it, and a second worker on the same control database is safe, since a delivery is leased and fenced. The hosted MCP host's batch_scrape refuses webhook (unsupported remote tool option) and offers no crawl. On /fc/v1/crawl Firecrawl's webhook (a string or { url, headers, metadata, events }) is mapped onto the native option and its receiver gets Firecrawl's shape, { success, type: "crawl.started" | "crawl.page" | "crawl.completed" | "crawl.failed", id, data: [page], metadata, error? } (a cancelled crawl is crawl.failed with error: "cancelled"), the Octocrawl event identity on the headers (docs/firecrawl-shim.md). A delivery is bookkeeping about the job: nothing about a page's fetch, trace or compliance record changes with a webhook.
The SDK follows a listing's cursors for you: listCrawlPages(id, options) and listBatchItems(id, options) take, beside limit, attemptId, debug and includeDuplicates, the caps maxPages (pages read after the first), maxResults (items in all) and maxWaitMs (no further page after that long), and stop quietly when one is reached; the generator's return value says where (nextCursor, stoppedBy: end, maxPages, maxResults or maxWait). With maxResults each page is requested no larger than what is still wanted, so the cursor the listing stops at continues exactly after the last item returned. getCrawlDocuments(id, options) returns { report, pages, nextCursor, stoppedBy } and getBatchDocuments(id, options) { report, items, nextCursor, stoppedBy }, the status and the documents in one answer, every document unless a cap stops the listing; collectCrawlPages and collectBatchItems return the documents alone. A listing is the latest attempt's unless attemptId names another (a resume with useCached re-records the pages it reuses in its new attempt, so that attempt normally holds every page). The server's page sizes are unchanged: crawl pages 1 to 1 000 per request (default 50), batch items at most 50. MCP get_crawl_pages and get_batch_items take maxResults (1 to 200) and then follow the cursors themselves, answering { items, nextCursor, hasMore, stoppedBy }; get_crawl_pages also takes includeDuplicates. The /fc shim's crawl status still pages with next alone.
A crawl task stores every option it was started with. A crawl paused by a shutdown or left running by a crash resumes when the API starts again, POST /v1/crawl/:id/resume (SDK resumeCrawl, MCP resume_crawl) restarts a paused or failed one, and octocrawl crawl --resume <taskId> continues one from the command line; all three run with the stored options. A resume refetches the pages the crawl already has, or reuses them when the crawl was started with useCached: true (the crawl's own pages only; maxAge reuses the cache across requests). maxPages counts the task's distinct pages across resumes, so a resumed crawl never exceeds it. GET /v1/crawl/:id counts pages while the crawl runs; /v1/crawl/:id/pages items carry the routing audit and trace only with debug=true, like batch items. Between two page starts on a host the crawl waits that host's robots.txt Crawl-delay or the minimum interval (W2L_PER_HOST_MIN_DELAY_MS), whichever is longer, and until a page on a host has reported its robots.txt the crawl starts one page at a time there. Each fetched page records the wait in a crawl_delay trace event: startedAt, previousStartedAt, observedDelayMs, requiredDelayMs and robotsCrawlDelayMs.
To wait for a crawl or a batch from the SDK, waitCrawl(taskId) and waitBatch(taskId) poll it until it is completed, failed or cancelled (a paused task is still waited on) and return its status; crawlAndWait(url, options, wait) and batchAndWait(urls, options, wait) start the task, wait, and also return every page and error of the crawl, or every item of the batch. The wait options are pollIntervalMs (default 500), timeoutMs (no limit by default), maxRetries (default 5) and signal. When timeoutMs runs out, a status request still in flight included, the wait throws WaitTimeoutError with taskId, timeoutMs and last, the last status read (null when none answered in time), and the task keeps running. A status request that fails with a network error, HTTP 408, 429 or 5xx is retried after 1, 2, 4, 8, then 10 s, or after its Retry-After when that asks for 60 s or less; the wait throws any other error at once, and a transient one once maxRetries retries in a row have failed.
To follow a job as it runs instead of waiting for it, GET /v1/crawl/:id/events and GET /v1/batches/:id/events stream it as server-sent events: catchup with the report as the stream opens, one document per page recorded, whatever its outcome (the compact page GET /v1/crawl/:id/pages or /v1/batches/:id/items lists, trace empty, no audit, every attempt of the job, each step once; its id: is the step cursor the listing routes take), a snapshot with the report after each page, then done with the terminal report, after which the stream closes; error carries { code, message }. ?after=<cursor> or a Last-Event-ID header resumes after a document (a cursor the API did not issue is HTTP 400 cursor is not one this API issued); an id that is not a job is 404. GET /v1/crawl/:id/ws and GET /v1/batches/:id/ws upgrade to a WebSocket carrying the same frames as JSON ({ type, data | error, cursor? }), close with 1000 after done, with 4404 for a missing job and 4400 for a bad cursor; on a server with tokens the upgrade presents the bearer token in the Authorization header or, where the WebSocket API gives no header, as the subprotocol w2l.token.<token>, echoed back as the selected protocol (never in the URL; a token with characters a subprotocol cannot carry uses the SSE route instead). The server reads the checkpoint once per subscriber as the stream opens and then once per page for every subscriber of that job together; a stream is a view of the job, records nothing, and W2L_JOB_STREAMS=off turns all four routes into 404s. The SDK's watcher(jobId, { kind: 'crawl' | 'batch', transport, pollIntervalMs, timeoutMs, after, signal, WebSocket }) returns a JobWatcher (an EventTarget): events document (CustomEvent<CrawlPage>), snapshot, done (the report) and error ({ code, message }), also yielded by for await (const event of watcher); fields jobId, kind, status, data (every document, each once) and transport; close() stops watching and leaves the job running. transport: 'auto' (default) tries the WebSocket route, then server-sent events, then polling (getCrawlPages with includeDuplicates, getCrawlErrors and getBatchItems with the cursor of the last document, the status every pollIntervalMs, default 2000, at least 250: a smaller value is a TypeError), switching once per level when a transport is unavailable (a 404 on the stream routes, no WebSocket constructor) or ends before done, and continuing from the last cursor, so a document is emitted once per step id across catch-up, live delivery and any switch; a 401 or 403 is final (error unauthorized), and timeoutMs ends the watch with error { code: "watcher_timeout", message: "job <id> did not finish within <ms> ms (last status: <status>)" } while the job keeps running. crawlAndWatch(url, options, watch) and batchScrapeAndWatch(urls, options, watch) start the job and return its watcher. MCP has no streaming surface (request and response only): wait_batch and get_batch_items are the way there. The frozen v1 Firecrawl shim adds no WebSocket path.
A map (POST /v1/map, SDK map(url, opts)) lists a site's URLs without fetching each page: it reads the start host's robots.txt, at most one page body (the start URL, on the http rung alone; no browser is started) and the sitemaps the site declares (the start URL's robots.txt Sitemap: lines, else /sitemap.xml; an index is followed one level, .gz files are inflated, at most 50 files), inside one deadline. It takes url, mode (standard or research; authed is refused: a map reads public sitemaps and one public page), limit (an integer from 1 to 100,000, default 5,000, Firecrawl's documented default and maximum), timeout (milliseconds, 1,000 to 300,000, default 60,000, for the whole map), the options below, ignoreRobotsTxt (above; a local server only), origin and integration; any other key is refused by name with unsupported_parameter (useIndex with the hint that Octocrawl keeps no URL index; a page option such as headers, mobile or skipTlsVerification, so a map has nothing to loosen; the crawl's allowSubdomains and allowExternalLinks, which a map calls includeSubdomains and does not offer). Every candidate goes through the crawl's scope rules (the start host, its www. twin and where the start URL redirects; the start URL's path subtree; assets left out; similar URLs folded, as deduplicateSimilarURLs does on a crawl, except that a link seen first over http: gives way to its https: variant when that variant comes too, on an origin whose robots.txt the map read anyway and which allows it (no further robots.txt is read for it); the start URL stays as given) and its host's robots.txt under the map's declared identity: a disallowed URL, or one whose robots.txt cannot be read, is not returned unless the map was started with ignoreRobotsTxt, and robots.txt is read for the start host and at most 20 others. The answer (HTTP 200) is { id, url, status, stoppedBy, links, sources, refused, identity, warnings, agentHints?, elapsedMs }: links hold the start URL, then the start page's links in document order, then sitemap entries not already found, each { url, title?, description?, titleSource?, via, sitemapFile?, lastmod?, robots }. A title is never invented: the start URL's is its page's <title> (and only it has a description), a page link's is its anchor text (whitespace collapsed, at most 300 characters; else aria-label, title or an inner image's alt), a sitemap entry's is its <news:title>; no other URL is fetched for one. via says how a URL was found (start, link, sitemap; a URL found both ways is one link), lastmod is the sitemap's as written. sources records the start page read (httpStatus, status, robots, rawBodySha256, linksFound) and every sitemap file (as on a crawl), and refused counts what was left out and why (duplicate, collapsed, hostDenied, subtreeDenied, pathDenied, assetDenied, robots, robotsUnchecked, searchFiltered, overLimit, with samples; once the start page's links alone fill limit the map answers at once and reads no sitemap, so overLimit then counts only the candidates seen before it stopped). status is completed when every source was read or is definitively absent (reaching limit is completed, with stoppedBy: "limit"), partial when links came back but the deadline cut the run or a source failed, and failed when nothing came back and a source failed or the deadline fired; at the deadline the answer is still HTTP 200 with what was found, stoppedBy: "timeout" and a map_timeout warning that names the links found and the sitemap files not read, never a 408 and never a complete-looking list. Other warnings: start_page_unreadable, start_page_client_rendered (the http lane found the page filled by script; the hint is to scrape it with formats: ["links"]), sitemap_unreadable, sitemap_files_capped, robots_host_cap, robots_unreachable (a robots.txt that could not be read, which counts as a complete disallow; the warning names the host and the reason, and a start URL on such a host has sources.startPage.robots: "unreachable" with robotsUnreachable, never disallowed, which is kept for a rule the publisher wrote). Each map is recorded at <taskRoot>/maps/<id>.json before it is answered and read back with GET /v1/maps/:id (SDK getMap(id)); a client that disconnects cancels the map and leaves no record. A hosted server takes limit up to 5,000 and timeout up to 60,000, and refuses a URL its channel policy serves with the browser lane only. The SDK waits timeout + 5000 ms for the answer. What a map does not do, against Firecrawl's: no URL index, so a site without a sitemap maps only its start page's links (docs.python.org/3/ gave 24 links on 2026-10-03: its sitemap lists 8 version roots); search filters and does not rank; a title is anchor text, not the target's <title>; no description except the start URL's; no location.
A map's options: sitemap is include (default: the start page's links and the sitemaps), skip (no sitemap file is requested; the links are the start URL and its page's links, sources.sitemap is null) or only (no page body is read at all, sources.startPage is null; the links are the sitemap entries scope, robots.txt and search admit, in listed order, up to limit, and the start URL is among them only when a sitemap lists it; no sitemap read, absent included, is failed with sitemap_unreadable). With only the operator's browser-only channel policy does not refuse the URL, since no page is read. search (a string of 1 to 200 characters with at most 10 words, after trimming; else HTTP 400 search must be a string of 1 to 200 characters with at most 10 words) keeps a URL when every word appears, case-insensitively, in its percent-decoded URL or its title in hand (the start page's own, an anchor's text, a sitemap's <news:title>); it runs before robots.txt and limit, so limit counts matches, keeps the discovery order (Firecrawl orders by relevance; Octocrawl filters), fetches nothing more, and counts what it left out in refused.searchFiltered. includeSubdomains (default false; Firecrawl v2 documents true) admits every host under the start URL's apex, the crawl's allowSubdomains rule: the start host with one leading www. removed and no public-suffix list; each new host's robots.txt is read once under the map's identity, for at most 20 hosts beside the start host, and the URLs on further hosts are counted in refused.robotsUnchecked with a robots_host_cap warning. Without it every link is on the start host, its www. twin (python.org and www.python.org are one site) or the host the start URL redirected to. ignoreQueryParameters (default false; Firecrawl v2 documents true) folds URLs that differ only in their query string into the first one seen and returns it without its query; each fold is counted in refused.collapsed with up to 20 { url, into } samples, never merged silently. Without it query variants stay apart, with tracking parameters (utm_*, gclid, ...) dropped and the rest sorted. The crawl's includePaths, excludePaths (at most 1,000 regexes of 1 to 2,000 characters, exclude winning; a pattern that can backtrack catastrophically is refused), regexOnFullURL, crawlEntireDomain (lifts the start URL's path subtree) and deduplicateSimilarURLs (default true) apply with their crawl names, messages and rules. MCP has the same map as the local map tool (annotated read-only, idempotent, open-world, with an output schema): compact by default, { id, status, stoppedBy, links: [{ url, title?, description? }], warning?, agentHints?, counts: { returned, refused } } with the warnings' messages joined, the native response with debug: true, each answered as text and as structuredContent; the hosted MCP host does not offer it, since it takes no arbitrary URL. /fc/v1/map takes Firecrawl's map request (shim).
Scrape, batch and crawl also take twelve page options and the four cache options (below); batch and crawl apply them to every page and store them with the task, so a resumed task keeps them:
onlyMainContent(defaulttrue).falsereturns the Markdown of the whole page: the document body with scripts, styles, form controls and embedded media left out, and the header, navigation and footer kept, through the same converter and base URL. The evidence (hashes, status) is the same in both modes, lane routing still reads the main content, and theextracttrace event recordsonlyMainContent: false. One exception: on a page where the extractor finds no main block, the default isfailed/empty_unverifiedwith the whole page's Markdown kept as evidence, whilefalsereturns that Markdown assuccess; either way the browser rung is still tried and answers when it renders more.linksalways come from the whole page.includeTags(up to 100 CSS selectors of 1 to 200 characters, with at most 100 selector parts in all; see below). The content is the elements the selectors name, in document order, an element inside another named one once. They are taken from the page as it was received, before main-content selection and cleaning, so a named navigation, header or footer stays, andonlyMainContentno longer chooses the content; scripts, styles, form controls and embedded media are left out as always. The page's type, title,metadataandlinksare still read from the whole page, JSON extraction still reads the page's own facts and the label/value pairs of its main content, anddocument.confidenceis 1 when the named elements hold any text or image, since what you named is the content. When nothing is named, the answer issuccesswith empty Markdown from the rung that read the page, not a failure; the browser rung is tried only when the page itself reads as thin or script-filled, as for any such page; when that rung then finds the page blocked, the block is the answer and the empty one is given up (the ladder audit recordsladder_empty_answer_dropped), and when it names nothing either, the empty answer is confirmed (confirmsEmptyon itsladder_step) and no further rung, a vendor's included, is asked. A page that is blocked staysblockedwhatever the selectors name, on whichever rung the block is found: a sign-in wall's heading is not the page that was asked for. Namingbodynames the whole page.excludeTags(the same kind of list). The elements the selectors name are removed, with everything inside them, before the content is taken: from the main content, from the whole page (onlyMainContent: false), from anincludeTagsselection, and from the page a failed or blocked result keeps as evidence. Both lists are matched against the whole page as it was received, sofooter pinexcludeTagstakes the footer's paragraphs out of anincludeTags: ["p"]selection. Theextracttrace event records both lists.waitFor(milliseconds, an integer from 0 to 60 000, default 0). The browser rung waits this long after the page has loaded and settled, then captures it; a document the page moved on to meanwhile (a script, a meta refresh) settles before the capture, and the result reports that document's URL and status. The HTTP rung cannot run scripts, so a request withwaitForstarts at the browser rung, and the ladder audit records the skipped rung (ladder_channel_skipped). Where no browser rung is configured, the result isfailedwithpolicy_deniedand await_for_unavailabletrace event, never an answer that ignored the wait.timeout(milliseconds, an integer from 1 000 to 300 000, default 300 000). The deadline for the whole scrape,waitForincluded. When it fires, the API still answers HTTP 200:partialwith the best content a rung produced so far (for example the HTTP content while the browser rung was still loading), orfailedwithfailureReason: "timeout"when nothing usable exists. Both carryusage.deadlineExceeded: trueand adeadline_exceededtrace event. When awaitForwould run past the deadline, the browser stops waiting about one second before it and captures the page as it is then:partialwhen that page has content, otherwisefailed/timeout. The lanes' waits for a slow server follow thetimeoutyou set: the HTTP lane waits for the response headers and for each chunk of the body, and the browser waits for navigation, until the deadline. Without atimeoutthey keep their defaults inside the 300 000 ms deadline: 10 s for headers, 30 s between body chunks and 20 s for navigation. A lane that stops at one of these fails withtimeout, withoutusage.deadlineExceeded, and the ladder does not move on to the next rung for it: the answer is the content an earlier rung produced, or thatfailed/timeout. The browser lane'snavigatetrace event records the wait it allowed (timeoutMs). A client that disconnects still cancels the scrape. An MCP client that cancels itsscrapecall (notifications/cancelled), or closes the call's HTTP request before the result, cancels it too, over stdio and over the local and hosted HTTP services; over HTTP only the same client can cancel a call (the same session id, and on the hosted service the same bearer token), see cancelling an MCP call. The SDK'sscrapewaits for the answer untiltimeoutplus 30 s; on Node that replaces fetch's own 300 s wait for the response headers, which an answer at a 300 000 ms deadline can outlast. JSON extraction reads fields from apartialpage but reports itincompletewith apage_partialissue, and never calls the model for it.maxFileBytes(bytes, an integer from 1 up to the server'sW2L_MAX_FILE_BYTES). A lower size cap for a file (PDF, CSV, XLSX, ZIP, JSON, text) than the server's; see Files.headers(an object of at most 32 header names with string values of at most 4,096 characters, no line breaks). Sent to the requested URL, its same-origin redirect hops and, on the browser rung, the files the page loads from that origin, after Octocrawl's declared identity, which they can never override:User-Agent, anysec-ch-*orsec-fetch-*header, the credentialsAuthorization,Proxy-AuthorizationandCookie, and the transport headers (Host,Accept-Encoding,Connection,Content-Length, ...) are refused with HTTP 400 naming them (headers.user-agent is refused: the User-Agent and client hints are Octocrawl's declared identity;headers.cookie is refused: credentials are not sent as headers; mode 'authed' carries your own session on the record;headers.accept-encoding is refused: transport headers are set by the lane). Names are lower-cased and a name given twice is refused;Accept,Accept-Language,Referer,Cache-Control,If-None-MatchandX-*headers are the common uses. A redirect to another origin is fetched with the identity alone on both rungs and the trace says so (custom_headers_withheldwith the URL and the names); on the browser rung the headers are added per request through Chromium's own request interception, which judges every hop and every file the page loads by its origin, so a navigation the page itself makes to another origin (a script, a meta refresh) and a same-origin file that redirects elsewhere get none either. robots.txt is always fetched with the identity alone. Everything inheadersis on the record: the trace (request_headers_added, values included), the HTTP rung'sidentity_sentlist and the browser rung's signed compliancesentHeaders(which never carry aCookieorAuthorizationheader: a session is on the record as a hash), so secrets belong to the authed session path, not to headers. On the browser rung the context locale staysen-US, so anAccept-Languageyou set can differ from the page'snavigator.language; that is visible to the page, not hidden. A request withheadersnever goes on to a vendor rung (ladder_channels_filteredin the ladder audit names the dropped rungs).mobile(defaultfalse).truefetches as Octocrawl's second declared browser identity: Android Chrome (Mozilla/5.0 (Linux; Android 14; Pixel 7) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/<major>.0.0.0 Mobile Safari/537.36, the same Chrome major as the desktop identity), aligned client hints (sec-ch-ua-mobile: ?1,sec-ch-ua-platform: "Android"), a 412x915 viewport at 2.625 device pixels per CSS pixel, and touch. It passes the same coherence and honesty checks as the desktop identity, robots.txt is evaluated against its User-Agent, and the trace records it (identity_sentwithdevice: "mobile"on the HTTP rung,identity_declaredon the browser rung); the browser rung's compliance record carries the mobile User-Agent and hints in its as-sentsentHeaders, and has no device field (schemaVersion 2; adding one would be schemaVersion 3 and a new hash for every record, an owner decision). On the browser rung the identity is set as Chromium's own user-agent metadata too, so the client hints Chromium generates itself (on a redirect hop, on the page's own requests) andnavigator.userAgentDatacarry the same brands, platform and mobile flag as the headers, for the desktop identity as for the mobile one. The page is whatever the site serves to that identity: a redirect to its mobile host, a responsive layout at 412 CSS pixels, or a page without a viewport meta laid out at 980 CSS pixels and scaled, as on a phone; Octocrawl rewrites nothing. The mobile identity claims Android on desktop Chromium the way the desktop identity claims macOS on any host: internally coherent, not a statement about the machine. It is refused withmode: "research"(mobile is not available in research mode: the research identity declares a bot, not a device) and never sent to a vendor rung. The Evidence Record says which identity answered:identity.deviceismobileordesktopas the answering lane declared it (null in research mode, which declares no device, and when no request was sent), andidentity.requestHeaderslists theheadersthat lane sent, names lower-cased and sorted, each with the SHA-256 of its value (valueSha256), never the value ([]when none).skipTlsVerification(defaultfalse). By default a certificate that does not verify (self-signed, expired, wrong name) isfailedwithfailureReason: "tls_error"on both rungs, with the error code in the trace (request_failedwithreason: "tls_error"andcode, such asDEPTH_ZERO_SELF_SIGNED_CERTorCERT_HAS_EXPIRED), also when it is the robots.txt request that fails on it, and the ladder does not escalate for it.true, on a local server, loads the page anyway: the HTTP rung uses routes of its own that do not verify, for this one fetch and its robots.txt lookup, and closes them after it (the shared routes, the robots cache's own and the delivery worker keep verifying); the browser rung opens its context withignoreHTTPSErrors. Every result of such a fetch carries atls_verification_skippedtrace event and atls_unverifiedwarning ("The certificate of was not verified at the caller's request; the content cannot be attributed to that host with certainty."), after arobots_overriddenwarning when there is one, on batch items and crawl pages too; the browser rung's compliance record (schemaVersion 2) has no TLS field, so the trace and the warning are the record. A hosted server (--hosted, the hosted MCP host) refuses the option with HTTP 400skipTlsVerification is not available in hosted modebefore anything is fetched, and runs a stored task without it. Never sent to a vendor rung.fastMode(defaultfalse).truekeeps the HTTP rung alone: no Chromium is launched,channelsTriedis["http"],summary.browserMsis 0, and the ladder audit records the rungs it dropped (ladder_channels_filteredwithreason: "fastMode"). The HTTP rung's own verdict is the answer: a page whose content its scripts write isfailed/empty_unverifiedwith the page as evidence, never rendered, and a thin or client-rendered-looking success keeps its warnings. When that rung asked for the browser rung, the response (full and compact) carriesagentHints: ["the http lane asked for the browser lane; fastMode declined it; retry without fastMode"]. It skips the browser rung; it does not make a page the HTTP rung already serves any faster, andwaitForhas no effect under it. A URL the server binds to the browser lane (the hosted Amazon flow) refuses it with HTTP 400fastMode is not available for this URL: it is served by the browser lane only.blockAds(defaulttrue). By default the extractor removes ad containers (elements whose id or class has one of the tokensad,ads,advert,advertisement,sponsored,promo) and cookie-consent banners (ids withcookie,consent,gdpr,cmp,onetrust,didomi,usercentricsand the like, and elements whosearia-labelnames cookies or consent) before extraction on every rung, and the local browser rung aborts, before any connection, every request to a bundled list of about fifty reviewed ad-serving hosts (doubleclick.net,googlesyndication.com,adnxs.com,criteo.com,taboola.com,outbrain.com,amazon-adsystem.com, ...;packages/bench/src/subjects/adHosts.ts, matched by host or dot-suffix), recording their number and the first 20 hosts in anads_blockedtrace event. The HTTP rung fetches no page resources, so only the extraction switch applies there. The token list is heuristic: an element whose class happens to bepromois removed too, which is why the switch exists; the host list is curated, not EasyList, and is not downloaded or updated, so ads from hosts outside it load (false negatives are expected).falsekeeps ads and cookie banners inmarkdownandhtmland loads every host;rawHtmlis always the page as received. The hosted browser's host allowlist and image, font and media rule stay in force whateverblockAdssays, sofalsecannot widen them.removeBase64Images(defaulttrue). By default an<img>whosesrcis adata:URI is left out of the Markdown and its alt text kept, on every rung and on the page a failed result keeps as evidence: Octocrawl has always done this, which is also what Firecrawl'sremoveBase64Imagesdoes by default (Octocrawl keeps the alt text where Firecrawl writes a(<Base64-Image-Removed>)placeholder).falsekeeps the image as, andusage.contentTokensthen counts it.htmlandrawHtmlare never rewritten: they are the page. A link whose target is adata:URI is always written as its text. A rendering choice, not a fetch fact: no trace event, warning or compliance change; the hosted MCP service refuses the option like every option outside its allowlist.
includeTags and excludeTags take tag, class, id and attribute selectors, the descendant and child combinators, :root, :empty, and :not(), :is() and :where() around selectors without combinators, for example table.wikitable, main > article p:not(.note) or a[href$=".pdf"]. A selector that does not parse is HTTP 400 invalid_request (includeTags entry is not a valid CSS selector: div[[). One that parses and uses anything else is refused with unsupported_parameter, naming what it uses and its place (excludeTags entry uses :nth-child, which Octocrawl does not match: li:nth-child(2), then the supported list; details.parameters: ["excludeTags[1]"]): the sibling combinators + and ~, positional pseudo-classes such as :nth-child and :first-child, :has(), :contains() and every other pseudo-class. Octocrawl does not run those because the DOM library it uses matches them at a cost the size of the page does not bound: on a page of 200 paragraphs x ~ p ~ p ~ p took 4.7 s and one more ~ p did not finish in 20 s, and a scrape's timeout cannot stop a selector while it is matching. The selectors Octocrawl accepts are matched in time proportional to the size of the page, however many combinators they chain, and a list is limited to 100 selector parts in all so that the proportion is bounded too: a tag name, *, a class, an id, an attribute test and a pseudo-class each count as one, those inside :not(), :is() and :where() included, so main > article p:not(.note) has five; a longer list is HTTP 400 invalid_request naming the count (includeTags must hold at most 100 selector parts in all, and holds 102). Each part costs one test of every element and each compound selector one pass over the page: on a page of 1.2 MB and 55,000 elements, a list of 50 descendant chains with 100 distinct compound selectors took 0.53 s, div:not() with 98 alternatives 0.06 s, and extraction with such a list in both includeTags and excludeTags 1.6 s against 0.3 s without them. Selectors are read with the DOM library's own parser, so an escaped name is what it is to the library: #\31 23, which CSS.escape writes for the id 123, names that element, and .\32 xl\:grid the class 2xl:grid.
The cache options. Octocrawl stores the latest successful result of each page under the task root (page-cache.sqlite) and reuses it only when a request asks. One stored result per page and set of options: the key is the URL without its fragment, the mode, every option the lanes receive except timeout (formats such as html or screenshot, onlyMainContent, includeTags, excludeTags, waitFor, headers, mobile, ...), the rungs the request may use (so the server's channel policy is never crossed), fastMode, a recorded robots override and the build (the extractor versions and W2L_SOURCE_COMMIT), so a result is reused only for a request that would have shaped it the same way, and its Evidence Record names the build that produced it; after an upgrade, pages are fetched again. Only a success with a recorded fetch time is stored, never a partial, failed or blocked page; a later fetch of the same page and options replaces it, an older one never does. A request with custom headers is stored only when it says storeInCache: true, since the stored trace keeps the header values. The cache has no size limit or expiry of its own: delete page-cache.sqlite to empty it.
maxAge(milliseconds, 0 to 315 360 000 000, default 0). Reuse a stored result fetched at most this long ago instead of fetching the page. 0 looks nothing up: the page is fetched live, as it always was.minAge(milliseconds, same range, at mostmaxAge). Reuse only a stored result at least this old; withoutmaxAge, of any age from this one on.storeInCache(defaulttrue, except for a request with customheaders, which stores only withtrue).falsestores nothing from this request.lockdown(defaultfalse). Answer from a stored result only, never fetch: a page without one isfailedwithfailureReason: "cache_miss"and anagentHintsentry, and nothing is requested, robots.txt included.maxAgeandminAgestill bound the age when given;lockdownwithmaxAge: 0is HTTP 400. A crawl in lockdown reads no sitemap, so it needssitemap: "skip"(HTTP 400 otherwise).
A reused result is the original fetch's, unchanged: its content, its evidenceRecord (fetchedAt, hashes, robots.txt decision, identity) and its trace, with a cache_hit event added at the end. The response says so in metadata.cacheState: "hit" with metadata.cachedAt, the reused fetch's fetchedAt; channelsTried is [] and usage counts no request, attempt, byte or browser time. A request that looked a page up and found nothing that fits says cacheState: "miss" and fetches it (a cache_miss event, and cache_stored when the result was stored). A request that looked nothing up (no maxAge above 0, no minAge, no lockdown) carries no cacheState at all: it is never reported as a miss. Batch items and crawl pages carry cacheState and cachedAt the same way, and a reused page has cached: true and counts in the crawl report's cachedPages; it waits for its host's pacing like any page but leaves the host's robots.txt Crawl-delay as it was. A URL the request's allowlistedDomains refuses is never answered from the cache. Mode authed neither stores nor reuses (a page read with your session stays yours): a lookup is HTTP 400 there. A Monitor's captures neither read nor fill the cache. A crawl's useCached is a different thing: a resume reuses the pages that crawl already fetched (above); maxAge reaches the results of any request on the same server. /fc maps the four options under the same names (see docs/firecrawl-shim.md).
formats takes markdown, links, json, html, rawHtml, images, tables and screenshot, one { "type": "attributes", "selectors": [{ "selector", "attribute" }] } entry and one { "type": "screenshot", "fullPage", "quality", "viewport" } entry, on scrape, batch and crawl. images is every image URL of the whole document as received, like links: img src and every srcset candidate (so a 1.5x/2x or 480w variant is its own entry), <picture> <source srcset> candidates, the lazy-loading attributes data-src, data-srcset, data-lazy-src and data-original on img and source, video[poster], <link rel="image_src">, og:image (and its :url / :secure_url forms) and twitter:image; each resolved against the document base (<base href> or the final URL), absolute http(s) only, the fragment stripped, each URL once, in document order, with data: URIs left out and counted. There is no file-extension filter, so an extensionless CDN URL stays, and includeTags, excludeTags and onlyMainContent do not narrow the list. On the browser rung the list is read from the rendered DOM, so an image a lazy loader has already moved from data-src into src appears once. The trace records images_collected with count, srcsetCandidates, lazy and dataUrisDropped. attributes reads, for each selector (1 to 50 entries, a selector of 1 to 200 characters under the same rules and the same 100-part bound as includeTags, an attribute name matching ^[A-Za-z_][A-Za-z0-9_:.-]*$ of at most 100 characters), the named attribute's values on every element the selector matches in the document as received, as written in the HTML (not resolved; links and images carry the resolved forms), elements without the attribute skipped, in document order, [] when nothing matches: [{ "selector", "attribute", "values" }] in request order, with an attributes_extracted trace event (selectors, counts). A selector that does not parse is HTTP 400 attributes selectors[i].selector is not a valid CSS selector: …, one Octocrawl does not match unsupported_parameter naming formats[j].selectors[i], and a request may carry one attributes entry (formats must contain at most one attributes entry). Both are returned only when asked for, by the full and compact scrape responses (whose formats list names them), batch items, crawl pages and /fc (data.images, data.attributes), and only for a page read as content: a file, a failed or blocked page and a result without HTML carry neither; a page without images gives images: []. summary.attempts and stored audits never repeat them. The vendor (provider) rung carries them too, read from the page the vendor returned, as it carries html and rawHtml. html is the cleaned HTML the Markdown is written from: the main content region; with onlyMainContent: false the page's <body> with its header, navigation and footer, without scripts, styles, form controls, embedded media and the excludeTags elements; with includeTags a <body> holding the named elements. It is the page's own markup: link and image targets stay as the page wrote them (the Markdown resolves them), and the browser rung's layout markers, which shape the Markdown, are not in it. rawHtml is the page as the answering rung received it, scripts and all: the response body on the HTTP rung, the rendered DOM on a browser rung. Its UTF-8 bytes hash to snapshot.rawBodySha256 (the Evidence Record's rawSha256). Both are returned only when asked for, by the full and compact scrape responses, batch items, crawl pages and /fc, and summary.attempts repeats them only with debug: true (stored batch and crawl audits never do). They are null for a file and for a result that is not success or partial: a failed or blocked page keeps its Markdown as evidence, not its HTML, and a crawl's duplicate page gives up both with its Markdown. The Evidence Record's outputSha256 covers Markdown and JSON, not html. screenshot (the string, Firecrawl v1's screenshot@fullPage, or one { "type": "screenshot", "fullPage", "quality", "viewport" } entry per request: formats must contain at most one screenshot entry) captures the rendered page as an image on the local browser rung, which such a request then selects alone: channelsTried is ["browser_local"], no HTTP attempt is made, the ladder audit names the dropped rungs (ladder_channels_filtered, reason screenshot), a server without a browser rung refuses the format with HTTP 400 screenshot requires the browser lane, which this deployment does not offer, and fastMode beside it is refused the same way (which fastMode declines). The capture is taken after load, stability and waitFor and before the DOM is read, so the image and the Markdown show the same page, with Playwright's scale: "css": the image is CSS-pixel sized, 1280x800 for the declared desktop viewport (the declared device scale factor 2 is reported, not baked into the image) or the viewport asked for (integers 320..1920 by 240..1080, within the declared 1920x1080 screen; with mobile, within the declared 412x915), which is a window size and not a change of identity: the User-Agent, client hints, locale, time zone, screen and scale factor stay as declared, and the trace records screenshot_viewport. fullPage: true captures the document's whole height at the viewport's width without scrolling first, so sections a page loads on scroll may show unloaded (the Chromium Playwright 1.62.1 installs captured 100,000 px tall fixtures whole, with no clipping). quality (1 to 100) gives a JPEG at that quality, otherwise a PNG. The response's screenshot is { contentType, width, height, fullPage, viewport, deviceScaleFactor, quality, bytes, sha256, path, base64 }: the bytes inline, hashing to sha256; path is <sha256>.png or .jpg under W2L_CAPTURE_RAW_DIR when that is set (listed in snapshot.artifacts too, and in the Evidence Record as kind: "screenshot" with its size and type), else null; the trace records screenshot_captured with the size, the bytes, the hash and captureMs. The capture goes on whatever the rendered page turns out to be, a success, an error page or a gate kept as evidence; it is null when the browser could not take it (screenshot_failed in the trace, a screenshot_unavailable warning and an agentHints entry, the page result kept) and when no page rendered (a file, a robots.txt denial). summary.attempts and stored batch and crawl audits carry screenshot: null, so the image travels once per response; a batch or crawl with the format takes every page on the browser rung and stores each capture inline in its checkpoint, so viewport captures are the lighter choice there. /fc returns it as Firecrawl's data.screenshot data URI string. The vendor (provider) rung never carries it: a screenshot request keeps the local browser rungs alone, so the vendor rungs are dropped with the http rung (ladder_channels_filtered, reason screenshot).
The list format ({ "type": "list", "itemSelector": "article.product", "fields": [{ "name": "title", "selector": "h3 a" }, { "name": "url", "selector": "h3 a", "attribute": "href" }, { "name": "price", "selector": ".price" }] }, one per request, 1 to 50 fields, names unique) turns a page of repeated items into records: every element itemSelector matches is one (an element inside another matched one is part of it, not a record of its own), and each field is read from it: the text of its first match within the record as a reader sees it (scripts and styles left out, blocks kept apart, whitespace collapsed); the record element itself when the field has no selector; or the named attribute instead of the text, a link or source (href, src, data-src, ...) made absolute against the page, its <base href> included. Field names are unique, and source_url, page and index are the CSV's own columns, so no field takes them. The selectors follow the includeTags rules and are checked before anything is fetched. The response's list is { itemSelector, fields, records: [{ values, missing, source: { url, page, index } }], pages, incomplete, truncated, csv, csvSha256 }: a value the record does not have is null and named in its missing, never filled in or guessed, and incomplete counts the records with one; csv has the fields, then source_url, page and index, so each row can be traced to the page and the place it came from. Read from the page as received (the rendered DOM on a browser rung), with no model. With a paginate action, the records are those of every page it read, each with its page number, also when a later step failed; a page whose records repeat a page already merged is not counted again. Otherwise they are those of the page as it stands. A list stops at 10,000 records or 5,000,000 characters of values: truncated is then true and a list_truncated warning says so. A page with at least one record holding a value is not failed as having no main content (a list of products or quotes is not an article): it answers success, its Markdown the whole page; elements that match but hold nothing (a loading skeleton) do not count. A crawl's duplicate page carries no records. Batch items carry their page's list. /fc does not offer it.
{ "type": "list" } alone finds the page's list itself, and { "type": "list", "itemSelector": "..." } the fields of the items named (fields without itemSelector is refused). The list is the elements that repeat beside each other with the same tag and classes under parents of the same tag and classes (a grid's rows together), scored by how many they are, how much text they hold and how alike they are inside. A list in nav, header, footer or aside, under a menu's role, or hidden (hidden, aria-hidden, a closed <details>) is never chosen and none of its items is read as a record (items hidden by their own hidden or aria-hidden attribute, such as loading skeletons, are kept out with :not([hidden]):not([aria-hidden="true"]) on the itemSelector); a menu by its class (menu, dropdown, tabs, pagination), or items that are one short link each, count for little. Its fields are what at least 60% of the items hold at the same place inside them: an element's text, a link's text and href, an image's src; a text every item has the same (a label, a button) is not a field. Each field is named after its class (price for a sum of money), never source_url, page or index; an item that lacks a field has it missing, never another element's value (each field's selector is checked on every item, and one the bounded work cannot check is left out). The selectors chosen stay within what a request may send (200 characters each, 100 selector parts in all), and the work is bounded on any page. The response's itemSelector is the one chosen and list.detected is { fields, alternatives: [{ itemSelector, count }] }: the fields as a request names them, to send back as they are or edited, and the other lists found, best first. A page with no list answers itemSelector: null, no records and a list_not_detected warning. With paginate, the list is found on the first page and every page is read the same way. It is a guess from the page's structure, with no model: check detected before relying on it.
tables gives every data table of the content the Markdown was written from (the main content, the whole page with onlyMainContent: false, or the includeTags selection) as data: one entry per GFM table of that Markdown, in its order, so tableIndex N is the Nth GFM table. Layout tables and single-row tables are not data tables, and a table nested in a cell is that cell's text, as in the Markdown. Each entry is { tableIndex, caption, sourceUrl, headerRows, columns, rows, csv, csvSha256 }: caption is the <caption> as plain text (null when there is none; a title written above the table is not a caption), sourceUrl the page's final URL, headerRows the leading rows in <thead> or made of <th> cells alone, and rows the cells as plain text (a link is its text, an image its alt text, whitespace collapsed, nothing escaped) with a cell that spans rows or columns repeated in every slot it covers, so every row has columns cells and none is shifted. csv is the rows as RFC 4180 CSV (CRLF line ends, a field quoted when it holds a comma, a quote or a line break, quotes doubled) and csvSha256 its SHA-256. Spans are read and capped as browsers read them: the attribute's leading digits (2.5 is 2), 1 when it has none, is negative or is a colspan of 0, a rowspan of 0 to the end of its <thead>, <tbody> or <tfoot> (or of the run of rows directly in the table), at most colspan 1,000 and rowspan 65,534. A rowspan covers every row it spans, also one whose cells end before its column, and never goes past the end of its row group. Rows are in the order browsers lay them out: the first <thead> first and the first <tfoot> last (an empty one included), wherever they are written; a later <thead> or <tfoot> stays where it is. A page is parsed by the HTML standard's tree construction (parse5), so a table holds the rows and cells a browser builds from the same tags, misnested ones included: a row or row group written inside a cell closes the cell, text written between rows comes before the table, an end tag a browser ignores closes nothing, and an svg's own <tr> or <td> is not a row or cell. A fragment (such as the main content) is read as a <template>'s content, so one row or cell of a table stays one. A table whose cells, spans repeated and every row padded to the widest, would exceed 2,000,000 characters, or what is left of 5,000,000 for the page's tables together (a cell counts its text and its CSV and JSON escaping), is given as { tableIndex, …, rows: [], csv: "", omitted: "too_large" }, so a small page cannot make a huge CSV. The trace records tables_extracted with the count, each table's rows and columns, and the omitted tables' indexes. Returned only when asked for, by the full and compact scrape responses, batch items and crawl pages, and only for a page read as content ([] for a page without tables); /fc refuses the format, which Firecrawl does not have. A PDF's tables are not reconstructed (see PDF text).
A scrape or a batch also takes actions: steps the local browser runs on the page after load, stability and waitFor, and before the screenshot format and the content are read, so a page whose data appears only after an interaction can be read (a "load more" button, a search form, an infinite list, a tab). The steps are Firecrawl's: wait (milliseconds up to 60,000, or a selector waited for up to 60 s), click (selector; all: true clicks every match; a control something else keeps covering, a modal or a consent banner, fails the step within about 10 s, naming what covers it, once it has stayed covered through a whole 5 s try of the click, as do loadMore's and paginate's clicks; a cover that goes sooner is waited out), write (text, typed into the element that has focus, so click it first), press (key: Enter, Tab, ArrowDown, ...), scroll (direction up or down, one screen of the page or of the element a selector names), screenshot (fullPage, quality, viewport), scrape (the page's HTML at that point), executeJavascript (script, a function body whose return value is kept; it runs through the DevTools protocol, so the page's Content-Security-Policy does not stop it) and pdf (format A0 to A6, Letter, Legal, Tabloid or Ledger, landscape, scale 0.1 to 2); at most 50, checked before anything is fetched. They run in order, each recorded in the trace as an action event with its index, type, outcome and time (and navigatedTo when it moved the page). The response's actions holds what they produced: screenshots (as the screenshot format's evidence), scrapes ({ url, html }), javascriptReturns ({ type, value }) and pdfs ({ format, landscape, scale, bytes, sha256, base64 }), each in step order. A step that fails stops the steps after it: actions.failed names it (index, type, code, message; the codes are selector_not_found, selector_timeout, script_error, navigation_refused, deadline_exceeded and action_error), and a page read as content is failed with action_failed, its Markdown the page as it stood, since it is not the page the steps were to reach. Every step is bounded by the scrape's timeout less the second kept back to read the page: a step that does not finish in time fails as deadline_exceeded, and what the steps before it produced is kept. From the first step until the page is read, a navigation of the page (a click on a link, a script's redirect, a form) goes through robots.txt and the egress policy as the requested URL did, before its request is sent: one Octocrawl does not fetch is answered 204 No Content inside the browser, so it is never requested and the page stays where it was, with a navigation_refused trace event and the step that led to it failing as navigation_refused (a page the steps left at such a URL through a server redirect is not read). A page a step's navigation reaches through a server redirect is checked when it loads, and a step that lands on one Octocrawl does not fetch fails there, keeping nothing it read of it; a redirect of the requested URL itself is the fetch's, not a step's, and the requested URL, when robots.txt was set aside for it (a URL the request named on a local server, or a robotsOverride), is fetched as the request said; any other URL a step navigates to is checked against robots.txt. A window a step opens is guarded the same way and closed at once (popup_closed in the trace). A change of the URL within the page (history.pushState) requests nothing and is not a navigation. A URL that answers a file has no page: its steps are reported as not run (action_error on the first), never skipped silently. The Evidence Record's pageActions lists the steps that ran with their outcome and says whether a script ran (scriptRan): the hashes are of the page as the steps left it, so a page a script rewrote is on the record as such. A request with actions selects the local browser rungs alone (ladder_channels_filtered names the others), is refused with fastMode and with the cache options (a page after actions is never stored or reused), and is refused by a hosted server and the hosted MCP endpoint; a crawl and a map take none. /fc maps a scrape's actions (see docs/firecrawl-shim.md).
Three more steps are Octocrawl's own, for lists that end only when the page says so; each stops by itself, never loops on a control that stays, and says why it stopped:
scrollToEnd(selectorto scroll an element instead of the page,itemSelectorto count items,maxScrolls1 to 200, default 50,waitMs100 to 10,000, default 1,000): scrolls to the end, waits, and again, until two rounds in a row add neither height nor items (end).loadMore(selectorof the "load more" control,itemSelector,maxClicks1 to 200, default 50,waitMs): clicks it, waits, and again, until it is gone, hidden or disabled (the attribute,aria-disabled, or adisabledclass on it or around it) (end), or two clicks in a row add nothing (no_growth); a control that was never there isselector_not_found. A control hidden or disabled before the first click, or while the items it asked for load, is waited for up to 10 s before it ends the list: many pages show theirs only once their script runs.paginate(nextSelector,itemSelector,maxPages1 to 100, default 10,waitMs): reads the page, keeps its HTML inactions.scrapes, follows Next and reads that page, until Next is gone, hidden or disabled (end) or the page at a URL shows what it showed before (repeat: Next led back or did nothing). A page is read twice 300 ms apart and only what reads the same counts, so a clock or a ticking price does not make it look new. WithitemSelector, the items are the records (each one its links and image sources and its steady words): a page whose records were already read under another URL (a first page at both/listand/list?page=1) is skipped, and two such pages in a row (a site that answers every page past the last with the last one) end the list. WithoutitemSelectorno page is skipped, so such a site is read up tomaxPages, with the warning below. Each page it reaches goes through the navigation checks before it is read; one Octocrawl does not fetch fails the step asnavigation_refused, keeping none of its pages. A click that never lands fails the step asaction_error, and the pages read before it keep their records. In a batch, each page the step reads is kept in the task's checkpoint the moment it is read: a batch cut at page N (shutdown, a crash) resumes there on the next start, passing over the kept pages along the site's own Next links and reading the rest once;actions.lists[].resumedsays how many pages came from the checkpoint. A kept page is known again by its address (when the pages have addresses of their own), by its items, or by its items' links alone when their text changed meanwhile and those links differ from page to page; a kept page known by none of those (rows without links whose text changed) is read again, its records repeat, andmaxPagescounts it twice. A check the site puts up where the next page should be stops the step there aschallenge: an interstitial (Cloudflare, a PerimeterX press-and-hold, a verification form) by its own marks before the page is read, a page of nothing but a CAPTCHA widget by the page's verdict after the steps (a page with records on it, or with other content, is a page whatever widget it carries). The pages before it keep their records, the result isblockedwith the check's reason,actions.lists[].challengenames the page, and in a batch the pages before it stay in the task's checkpoint; the batch handoff then lets you get through the check and page on in your own Chrome (see the handoff section).
actions.lists holds, per list step, { index, type, stoppedBy, rounds, items } (items the elements itemSelector matches at the end, null without one; itemsRead for paginate, over every page). A list step that stopped at its limit (max), because the deadline came (deadline) or at a check the site put up (challenge, with challenge: { page, url, reason, signals }) is not the whole list, and the result carries a list_not_exhausted warning that says so. The Markdown is the page as the steps left it (for paginate, the last page); extracting records from each page is a later step.
A scrape also takes robotsOverride ({ reason, recordedBy? }; reason of 1 to 500 characters, recordedBy of 1 to 200) on REST, the SDK and MCP: your own recorded reason for fetching this one URL although its host's robots.txt disallows it or could not be read, for example a report the publisher links from its own pages on a file host whose rules address crawlers. A local server fetches a URL a scrape names anyway (see robots.txt above); the override puts your reason and name on the record in place of user_named_url. robots.txt is still read and its verdict recorded. When it disallows the URL, the fetch goes ahead and says so everywhere: the trace carries robots_disallowed and then robots_overridden (url, appliedRules, reason, recordedBy), the result's warnings start with { code: "robots_overridden", message } naming the robots.txt URL, the rule, the reason and who recorded it (full and compact scrape responses, batch items), the browser lane's compliance record keeps the disallow with skippedFetch: false and an override that its hash covers, and the Evidence Record's robotsDecision stays disallowed with userOverride: true and overrideBasis: "robots_override". Every result of such a scrape carries the warning and these events, whichever rung or deadline produced it: when the deadline passes while the request is out (failed/timeout), or when a later rung that fetched nothing answers (in authed mode, the authed_session rung's skip when there is no session), Octocrawl adds them to that result with the lane's robots_checked event, each event naming the lane that set the rule aside, and that result's robotsDecision is read from them. The exception is an error of Octocrawl's own after the rule was set aside: the scrape answers HTTP 500 internal_error, and a batch item that fails with internal_error records only that error. The override is for fetches from your own machine: the HTTP and local browser rungs apply it, the provider lane takes none, and a scrape that set a rule aside does not go on to a vendor rung (ladder_channel_skipped in the ladder audit), so no vendor session is opened for that URL. When robots.txt allows the URL, the override does nothing and leaves no trace. An unreachable robots.txt is set aside too, its reason kept in the trace and the warning. A batch takes robotsOverrides, a list of { url, reason, recordedBy? } in which each url is one of the batch's urls, named once; only that URL is fetched past its rule, the batch's other URLs are unaffected, and the list is stored with the task, so a resumed batch keeps it. A crawl and a map take neither (they take ignoreRobotsTxt), and neither do /fc, the hosted MCP endpoint and the public preview, which accept only their own options. A hosted server (npm run api -- --hosted) takes neither as well, nor ignoreRobotsTxt: whoever holds one of its tokens, these fields are refused with HTTP 400 unsupported_parameter naming the field before anything is fetched, and a stored batch or crawl it resumes runs without its overrides. On a scrape or batch ignoreRobotsTxt is refused with HTTP 400 unsupported_parameter that names it, as is any other key inside an override.
Scrape, batch and crawl also take two labels that go into Octocrawl's own records and nowhere else: integration, yours ("nightly-prices"), and origin, the client's. The SDK sends origin: "js-sdk@<version>" (its SDK_VERSION, pinned to its package version) with every scrape, crawl and batch unless the call sets one; the MCP server records mcp-<client name>@<client version> from the client's initialize (characters outside printable ASCII written as _, cut to 100; [email protected] for a client that declared none) and refuses an origin a tool call names, so the tools expose integration alone; /fc maps the origin Firecrawl's SDKs send and takes integration too. Each is 1 to 100 printable ASCII characters without spaces, else HTTP 400 integration must be a string of 1 to 100 printable characters without spaces (the same for origin). Nothing sent to the target changes, and no response echoes them: a scrape's record carries both (below), and a batch or crawl stores them with its task, which GET /v1/batches/:id and GET /v1/crawl/:id report as attribution: { origin?, integration? } (absent when the request named neither).
A server may compress its answer although Octocrawl asked for no compression: Octocrawl's identity sends no Accept-Encoding, and some sites (www.python.org) answer gzip anyway. The HTTP lane decodes the body by its Content-Encoding, asked for or not: gzip (and x-gzip), deflate (zlib-wrapped or raw) and br, a list of codings in reverse order, after the 10 MiB wire cap (the file cap for a file) and under the 50 MiB decompressed cap (failed with decompressed_too_large past it). usage.bytesWire is the size received and bytesDecompressed the size decoded; rawSha256 and a file's saved bytes are of the decoded body; evidence.contentEncoding and the Evidence Record's contentEncoding name the codings (identity when there were none; null in the browser and provider lanes, which do not report it). Any other coding is failed with unsupported_content_encoding, and bytes that do not decode as their coding are failed with parse_error (trace content_decoding_failed): neither is read as the page. The crawl's sitemap reader decodes a sitemap file the same way before it inflates a .gz file.
A result also says when the page it read looks like a shell for data its scripts fill in. The extractor reads the page as received for a <table> with no cells beside scripts, an application root (#root, #app, #__next and their kin) with next to no text, a page of scripts with little text, an element the page hides once scripts run (a hide-if-js-enabled class and its kin), a hydration blob (__NEXT_DATA__, window.__STATE__ = and the like) on a thin page, aria-busy, or, on a page routed as a listing or collection, a list in its hydration JSON of at least 10 distinct named records (name or title, 12 characters or more) of which the visible text shows at least 3 and at most half (a Walmart category page draws 9 of the 49 products its __NEXT_DATA__ lists); every rule pairs a structural gap with script presence, so a static page with an empty table never trips one, and a noscript "enable JavaScript" notice alone counts only on a thin page (under 1,500 visible characters) or beside hydration state, since static sites carry one beside their analytics scripts. When the HTTP rung's page reads that way, its status stays what the content earned, a success or the failed/empty_unverified result that keeps the whole page as evidence when no main region was found (a page of site chrome around the script that writes its content reads that way), and the result carries { code: "client_rendered_suspected", message } in warnings (after a robots_overridden warning when there is one) naming the rule: empty_table_with_scripts, empty_app_root, script_shell, js_fallback, hydration_shell, aria_busy or hydration_list_partial. Its trace carries quality_client_rendered with the rule, the markers found, the number of empty tables and the visible-text and script sizes (and, for hydration_list_partial, listRecords: the records the list declares and how many are shown), and the ladder offers the page to the browser rung as it does a thin result (quality_low_yield): the rendered page is the answer when it holds more, otherwise the HTTP page stays, warning included, with that hop marked improved: false (trigger: "quality_client_rendered"). A failed/empty_unverified shell goes to the browser rung on its own extract_low_confidence ask, and the evidence kept when that rung then fails without a page (ladder_evidence_kept) carries the warning. The browser rung never raises it: its capture is the rendered page. The extractor's own reading is render on its output (@w2l/extract-tf): clientRendered, reason, markers, emptyTables, textChars, scriptChars and, with hydration_list_partial, listRecords. When such a thin or shell-like HTTP answer stays the answer, it also carries { code: "low_content_yield", message }, after its other warnings: The http lane extracted N tokens at confidence C; the browser lane did not improve it. when the browser rung answered with no more, or failed, and the HTTP page was kept (ladder_best_kept, the quality hop improved: false), and …; the browser lane was not available to this request. when no further rung was permitted (fastMode, a server bound to the HTTP rung, no browser rung configured); N and C are the HTTP rung's own quality_low_yield figures (else its extract confidence), and on a failed/empty_unverified shell the sentence reads The http lane found no main content at confidence C; …. A rendered page that answered carries no such warning. Every response that carries warnings also carries warning, their messages joined with a space, Firecrawl's name for it: full and compact scrape responses, batch items, crawl pages, and data.warning on /fc, which so passes the native warnings through.
A response also says what to change about the request next time, in agentHints (full and compact scrape responses, batch items, crawl pages and the scrape record; agent_hints on /fc pages), one sentence each, present only when something applies and never quoting page text, only a host, a status, a rule or a time: for a policy_denied with a robots_disallowed event, robots.txt of <host> disallows this URL for Octocrawl's identity (rule <patterns>). A local Octocrawl server fetches a URL a scrape or batch names whatever robots.txt says, and a crawl or map started there with ignoreRobotsTxt fetches the links it disallows, each on the record; a hosted server obeys robots.txt for every URL (an unreachable robots.txt says it counts as a complete disallow and that Octocrawl asks for it again after five minutes, then names the same routes); for a login_wall, the page asks for a login; Octocrawl does not create accounts; use mode authed with your own session; for cloudflare_challenge, captcha and bot_detected_generic, <host> gates automated access on the lanes tried (<lanes>); Octocrawl does not solve challenges or change its identity; a proxy or session you own is the supported route, or, on your own machine, getting through the check yourself in your own Chrome: handoff: true on a scrape (octocrawl scrape --handoff), or a batch handoff (octocrawl batch --handoff, POST /v1/batches/:id/handoff); for a retryAt, wait until <ISO time> before asking <host> again (a rate_limit block without a Retry-After says so); for a cut, the content was cut at character <n>; ask for rawHtml or a narrower includeTags; for a client_rendered_suspected warning, the page fills its data with JavaScript; the browser lane was tried or ... was not tried; for a low_content_yield warning, the http lane's content was thin and the browser lane did not improve it (or was not available) ; pass waitFor (up to 60000 ms) or a longer timeout with the browser lane available, or actions (a click, a scroll, a wait for a selector) when the data appears after an interaction; under fastMode the one sentence of that option stands alone when the HTTP rung asked for the browser rung; for an http_error that kept its page, the server answered <status>; the markdown is that error page, not the requested page (a 404 adds ; check the link, and a 404 without a page says the server answered 404; check the link); for a policy_denied that was the egress policy's rather than robots.txt's (an ssrf_denied or governance_refusal trace event), the egress policy refused <host> (<reason>) and nothing was fetched; Octocrawl reaches public addresses, and a local server the addresses its policy allowlists; for a page the http lane got blocked or an HTTP error for and the local browser lane then served, the http lane got <status>/<reason> (HTTP <n>) from <host> and the local browser lane served the page; expect other pages of <host> to need the browser lane too (read from the ladder summary's http attempt; the ladder's ordinary hop after a thin or empty http answer carries none); for a tls_error, the certificate of <host> did not verify and Octocrawl keeps verification on; a local server takes skipTlsVerification for one request, recorded in the trace and a tls_unverified warning, and a hosted server refuses it; for a timeout, no lane answered within the request's deadline; raise timeout (up to 300000 ms); for a partial result, the result is partial: the deadline passed with this much of the page read; raise timeout (up to 300000 ms) for the rest; for empty_unverified without a shell or thin-content caveat, Octocrawl found no main content on the page; onlyMainContent: false returns the whole page's Markdown as content, and includeTags names the elements to read instead (a PDF without a text layer: the PDF has no text layer, and Octocrawl runs no OCR); for an incomplete json, json is incomplete: the required field(s) <paths> ... not found on the page; modelFallback fills what the page does not state when the server has W2L_EXTRACT_BASE_URL and W2L_EXTRACT_MODEL, and for a model fallback that did not run, the json model fallback did not run: <reason>; for a file, the response was a <kind> file kept at <path>; markdown is its text layer (a CSV, JSON or text file: markdown is its text as received; an XLSX, XLS or ZIP file: it has no markdown; not saved in place of the path when the server keeps no files). A tls_unverified warning adds none, and no hint suggests stealth or a change of identity. The hints come from a fixed table keyed by warning code, status, trace event, json issue and file kind, never from page text, and at most five stay, in the table's order. A refused option Octocrawl does not offer gets a hint in the error body too (agentHints; agent_hints on /fc): stealth and proxy: "stealth" or "enhanced" (Octocrawl does not offer a stealth mode or stealth proxies; a proxy or session you own (mode authed) is the supported route), ignoreRobotsTxt on a scrape or batch (robots.txt is always read and recorded; on a local server a URL a scrape or batch names is fetched whatever it says, and ignoreRobotsTxt on a crawl or map fetches the links it disallows, on the record) and, on a hosted server, skipTlsVerification (a hosted server verifies every certificate; run Octocrawl locally to use skipTlsVerification, which is recorded in the trace and a tls_unverified warning). The SDK's W2LError.agentHints carries them (empty when there are none).
Markdown leaves out data: URIs, as Firecrawl's removeBase64Images does by default for images: an image keeps its alt text and a link its text; removeBase64Images: false (above) keeps an image's data: URI as its target, a link's never. Table cells and captions keep their links and images with absolute targets, as a paragraph does, on one line (a | in a cell, target included, is written \|); emphasis and code in a cell are plain text. A table whose nested tables hold at least half of its text, or that has a single row, lays out the page rather than holding data: its cells become paragraphs, and only the data tables inside it become GFM tables. A data table with a small table in one cell stays one GFM table, with the small table as that cell's text. A page Octocrawl extracts by its tables (document.strategy: "table", for pages whose tables hold at least half of their text) keeps as main content the lowest element that holds all its data tables, so the headings, captions and text between them stay and what lies outside that element, usually the page's own header, footer and side column, does not; a menu laid out as a table (mostly links, and no figures outside them) counts only when it is the page's largest table, and a lone <h1> that shares a container with the tables below <body> is kept with them. When the HTTP rung finds no main content and the browser rung then fails without a page, the answer is the HTTP rung's failed/empty_unverified result with its page (the ladder audit records ladder_evidence_kept); when the deadline ends the browser rung instead, it is failed/timeout with that page, and the deadline_exceeded trace event names the rung it came from (evidence).
A result whose page was extracted carries metadata, what the page's HTML says about itself, in scrape responses (full, compact and MCP), batch items and crawl pages: title (its <title>), description (<meta name="description">), language (<html lang>, or <meta http-equiv="content-language"> when <html> has no lang), keywords and robots (those <meta> values as written), favicon (the first <link rel="icon">, as an absolute http(s) URL resolved against the document base) and canonicalUrl (<link rel="canonical">, likewise). A value the page does not declare is null; Octocrawl does not substitute og:description, a heading or /favicon.ico. Beside those seven, metadata carries the Open Graph, Dublin Core and article tags the page states, under Firecrawl's names and only when stated (absent otherwise, never null): ogTitle, ogDescription, ogUrl, ogImage (og:image, else og:image:secure_url, else og:image:url), ogAudio, ogVideo (likewise), ogDeterminer, ogLocale, ogLocaleAlternate (every og:locale:alternate, as a list) and ogSiteName, read from <meta property="og:…"> then <meta name="og:…">, the first non-empty tag winning, the URL fields resolved against the document base when they parse; dcTermsCreated, dcDateCreated, dcDate, dcTermsType, dcType, dcTermsAudience, dcTermsSubject, dcSubject, dcDescription and dcTermsKeywords from <meta name="dcterms.…"> and <meta name="dc.…"> (names matched case-insensitively); and publishedTime (article:published_time), modifiedTime (article:modified_time), articleSection and articleTag (every article:tag, as a list). Each value is the page's own words with whitespace collapsed: no date normalisation, no time-zone math, and no fallback from twitter:*, govuk:*, citation_* tags, JSON-LD or <time> elements, so a page that states none gets none. A failed or blocked batch item or crawl page has no metadata. A scrape response always has one, because it also carries the facts of the call (next paragraph); on a page that was not read as content (a block page, an error page, a file) its page fields are all null: Octocrawl declares nothing from a page that is evidence rather than the page asked for. metadata.title is the page's <title>, often with the site name added; document.title is unchanged: the content's title, usually the first heading of the main content, with <title> only as its fallback. /fc puts title, description, language, keywords, robots and favicon in data.metadata when the page declares them, and the Open Graph, Dublin Core and article fields under the same names, articleTag joined with , as Firecrawl writes it and ogLocaleAlternate kept as a list.
Every scrape response (full and compact; /fc puts the same keys in data.metadata, with creditsUsed: null) carries the facts of the call under Firecrawl's names, beside the page fields: scrapeId (a UUID minted per POST /v1/scrape or /fc/v1/scrape call; the full response repeats it at the top level), sourceURL (the requested URL), url (the final URL, evidence.finalUrl), statusCode and contentType (of the response that answered url, as in snapshot), proxyUsed ("operator" when the request went through the server's environment proxy, read from evidence.envProxy and the egress_proxy trace events; "user" when the browser lane's compliance record names your own egress; null otherwise, never a guess), timezone (the IANA zone the browser lane declares, America/Los_Angeles for the desktop and the mobile identity alike; null for an HTTP-lane result, where no time zone goes on the wire), and concurrencyLimited with concurrencyQueueDurationMs (below). cacheState and cachedAt are not there: Octocrawl has no cache yet and reports no invented miss. The record of the call, without any page body, is written to <task root>/scrapes/<scrapeId>.json before the response is sent and served by GET /v1/scrapes/:id (SDK getScrape(id), MCP get_scrape): scrapeId, requestedAt, request (the parsed options, each custom header's value replaced by its name), origin, integration, status, failureReason, blockReason, budgetExceeded, lane, channelsTried, metadata, snapshot, usage (wallMs, totalMs, requestCount, attemptCount, browserMs), warnings and agentHints. An unknown or malformed id is HTTP 404 not_found. When the record cannot be written the response still carries its id, the server logs scrape_record_unwritten and the full response's trace records that event. Records are kept without retention: a hosted operator prunes the directory. A Monitor's capture mints an id for its response but writes no record (a preview persists nothing; a run's observation is the Monitor's own).
concurrencyLimited is true when the per-origin concurrency ceiling (W2L_PER_HOST_CONCURRENCY, 1 to 4, per target origin) held an attempt of the scrape back because every slot was taken, and concurrencyQueueDurationMs is how long its attempts waited for a slot in all, a concurrent cooldown and the minimum interval between requests excluded (each lane tried acquires its own permit, so a scrape that escalated adds both lanes' waits). Each lane result carries its own share as usage.timings.concurrencyWaitMs, present only when its permit was held, on batch items and crawl pages too; usage.timings.queueMs keeps its meaning, every scheduler wait but a cooldown, and includes that time. Unlike Firecrawl's concurrencyLimited, which reports a per-account limit, Octocrawl's is the per-origin politeness gate, and a crawl's own frontier scheduling (its per-host spacing) is not counted.
Request deterministic structured data with a JSON Schema alongside, or instead of, Markdown:
Octocrawl maps supported product fields directly from subject-bound HTML, JSON-LD, metadata and DOM evidence. A top-level key that no such fact covers is matched to the page's own labels, two-cell th/td table rows and dt/dd pairs in the main content, compared without case, spaces or punctuation (Number of reviews fills numberOfReviews; Price (excl. tax) fills price when no label is exactly Price). title (or pageTitle) takes the content title, url (or finalUrl) and requestUrl the fetch's final and requested URL, and pageType the page type Octocrawl classified. A number is read only from text that is one amount, such as £51.77 or 1.299,00 €, never from text like HL-1, 4.7 out of 5 or a URL (see below); labels that state different values leave the field out with a field_ambiguous issue. A missing nullable field is null with a field_unavailable issue; a missing required field is left out with a missing_required issue and the result is incomplete. Every value has an evidence entry (the same map is evidenceRecord.fieldEvidence): the label's table or list, row and label; dom with h1[0] (the main content's first heading) or title (the page's <title>) for the title; fetch with finalUrl or requestedUrl for a URL; inferred with document.pageType for the page type, and with document.product.images (prices, variants, specifications) for a list or map the product extractor reported empty, since nothing on the page locates an absence; model for a value the model fallback wrote. json.evidence also quotes the text each number was read from (text, whitespace collapsed); the Evidence Record keeps source and locator.
Numbers are read as the page writes them. An amount is an optional sign, a currency symbol or code before or after the number (€, EUR, US$, kr, 円, or a symbol before and a code after, as in $12.99 USD) and one number: its decimal separator is . or ,, its thousands are grouped by ., ,, a space (no-break and narrow no-break spaces too) or an apostrophe in groups of three, or in India's lakh groups, and ,- or .– after it ends a whole amount. So 12,99 € is 12.99; 1.299,00 €, 1 299,00 €, CHF 1'299.– and $1,299.00 are 1299; ₹1,29,999 is 129999. A single . or , before exactly three digits (1.299 €, $1,299) is 1299 in one notation and 1.299 in the other, so it is read only when the value settles it: a review count is whole, so is an amount in a currency without minor units (JPY, KRW, ISK, VND, CLP, ₩, 円, in the text or as the page's priceCurrency), and a JSON-LD or product:price:amount price writes . as its decimal point. Octocrawl does not guess from the page's language, currency or domain: an English page of a German shop can write 1.299 €, Irish shops write €1,299, and a German page can quote $1,299. Such a number, and a product price that is not a number at all (Call for price), is left out, or null when the field is nullable, with a field_unavailable issue quoting the text and where it is (a required one also gets missing_required); asked for as a string, the field is the text. An amount in the Amazon adapter's prices list that cannot be read stays its text, with a field_unavailable issue at /prices/<i>/amount.
The schema may use the JSON Schema subset Octocrawl can honour, which covers what Pydantic's model_json_schema() and zod-to-json-schema usually write:
- Structure:
type(one or a list),properties,required,items(one schema),additionalProperties,enum,const, local$ref(#,#/$defs/…,#/definitions/…or another pointer into the schema) with$defsordefinitions, andanyOf/oneOfof a schema and{ "type": "null" }(Pydantic'sOptional) or of primitive types only. - Checked on the result, never used to fill a value in:
minimum,maximum,exclusiveMinimum,exclusiveMaximum,multipleOf,minLength,maxLength,pattern,minItems,maxItems,uniqueItems, andenum/const. A page value that breaks one stays indataand the result isincompletewith afield_unavailableissue quoting the check (the model fallback, when on, may replace it). Apatternhas at most 2,000 characters, and one that can backtrack catastrophically, such as^(a+)+$, is refused withinvalid_request. The rest run on V8's linear-time engine, except a pattern with a lookaround, a backreference, a counted repetition above 16 (such as[0-9a-f]{32}), a\p{…}or\u{…}escape or a character outside the Basic Multilingual Plane, and any text holding such a character (an emoji): those are matched only against text of up to 2,048 UTF-16 code units and within 100 ms, and text they cannot decide counts as breaking the pattern. - Accepted and not acted on:
title,description,$comment,examples,deprecated,readOnly,writeOnly,format(not checked),default(never filled in: a field the page does not give stays missing), and at the root only$schema(draft-07, 2019-09 or 2020-12) and$id. - Bounds: 64 KiB, 8 levels of nesting and 100 properties.
Anything else is refused with HTTP 400 unsupported_parameter, whose details.parameters names the keyword where it was sent (for example formats[0].schema.properties.author.allOf): allOf, not, if, patternProperties, prefixItems, OpenAPI's nullable, a union of objects, arrays or references, a keyword other than an annotation beside $ref, $schema or $id below the root. A malformed value, such as an invalid pattern or a $ref that does not resolve, is invalid_request. A json format needs a schema: { "type": "json", "prompt": "…" } alone is refused with invalid_request, because Octocrawl does not extract JSON without one (that would need a model for every page); prompt only instructs the model fallback.
Model fallback is opt-in with modelFallback: true and runs only when a required field is missing or a value breaks the schema; configure an OpenAI-compatible endpoint through W2L_EXTRACT_BASE_URL, W2L_EXTRACT_MODEL and optional W2L_EXTRACT_API_KEY. Without those variables, page content is never sent to a model and the JSON result reports model_unavailable. The model receives the main-content Markdown and the values already read. It fills only what is missing, or replaces a page value that breaks the schema; every value read from the page keeps its value and evidence whatever the model answers, and every value the model wrote has model evidence. The request uses strict structured outputs (json_schema with strict: true) with a strict-safe copy of the schema: every object closed, every property required and the optional ones nullable, assertions and annotations left out. The answer is still checked against your schema, with one repair round, and a null your schema does not allow is dropped as "not found". A schema strict mode cannot express, such as an object without properties, is sent as given without strict mode; json.modelUsage.strict says which was used and strictReason why not.
Run the fixed 10-product, three-round Amazon MCP baseline with:
The setup uses an anonymous Singapore public delivery preference for this benchmark only. Round 1 pins the observed context; later unobserved or mismatched region/currency records remain in the report and do not count as comparable. Reports and raw HTML stay under ignored .w2l/amazon-baseline/; the URL manifest and schema are versioned. The signed ten-product result passed at limited concurrency, but Amazon remains beta pending the 100/1000 promotion gates. The older baseline is historical.
The concurrency-1 command can exit nonzero because its ten-page median exceeds 20 seconds; inspect its report for comparability and blocking before continuing to 2. The signed run had 37.93 seconds at 1, 19.92 at 2, and 12.39 at 4.
This signed Amazon slice was merged into main by PR #52, after the v0.4.0-rc.1 source prerelease, so that prerelease does not contain it.
Firecrawl v1 clients (partial compatibility): set the base URL to http://127.0.0.1:8787/fc so /v1/scrape and /v1/crawl hit the shim. The scrape shim maps url, the markdown, links, html, rawHtml, images and screenshot formats (screenshot@fullPage and an { type: "screenshot", fullPage, quality, viewport } entry too, returned as the data.screenshot data URI) and an { type: "attributes", selectors } entry, onlyMainContent, includeTags, excludeTags, waitFor, timeout, headers, mobile, skipTlsVerification, fastMode, blockAds, removeBase64Images the cache options maxAge, minAge, storeInCache and lockdown (an omitted maxAge reuses nothing; a lockdown scrape with no stored result is HTTP 404 SCRAPE_LOCKDOWN_CACHE_MISS), and proxy as the access choice (basic is standard; stealth and auto are enhanced, which a server without an access grant of tier enhanced refuses with HTTP 400); the crawl shim maps url, limit, maxDepth, includePaths, excludePaths, regexOnFullURL, ignoreQueryParameters, deduplicateSimilarURLs, crawlEntireDomain (and v1's allowBackwardLinks), allowSubdomains, allowExternalLinks, sitemap (and v1's ignoreSitemap: true is skip, false is include; sitemapOnly: true is only), maxConcurrency, webhook (onto the native job webhook, the receiver getting Firecrawl's payload shape) and the same scrapeOptions. The five execution options follow Octocrawl's rules, not Firecrawl's: a User-Agent or Cookie in headers is HTTP 400, skipTlsVerification is refused on a hosted server, and fastMode returns the HTTP rung's verdict instead of a render. removeBase64Images (default true) keeps an image's alt text where Firecrawl writes a placeholder, and false keeps the data: URI. Any other parameter or format (location, json, a crawl's scrapeOptions.actions, ...) is rejected with HTTP 400 and success: false, naming it. Snapshot 2026-09-18; known diffs in docs/firecrawl-shim.md. Firecrawl Search / Interact / Agent / Monitor compatibility is not implemented. Octocrawl's native Monitor and Delivery APIs use their own contracts.
Continuous Monitors and event delivery
The native SDK includes Crawl pagination/cancellation, Monitor creation/revisions/runs/control, and Delivery destinations/status/retry. The runnable example uses a controlled price source, explicit captureMode, validated baselines, conditional HTTP requests, and persisted events. The webhook receiver stores event receipts and applies a versioned product projection transactionally.
The API, Monitor scheduler and delivery worker share a persistent control database.
npm run local:mcp:install manages all three for local MCP users. w2l-api
(npm run api) runs a delivery worker of its own since job webhooks arrived,
so Monitor and job deliveries leave the API process too; the standalone worker
below is still there for a deployment that wants it separate, and running both
is safe (a delivery is leased and fenced). For a standalone API deployment, run
the Monitor worker in a separate terminal with the same W2L_TASK_ROOT as the API:
The Monitor worker defaults to public-source network policy. For the controlled local source in the onboarding example, explicitly set W2L_MONITOR_NETWORK_MODE=local in that worker terminal. A locally running worker does not inherit broader network access from the API or database.
See onboarding for the HTTPS receiver, authentication, worker configuration, and pending-delivery restart exercise. Gate 2–4 source freeze 99894bd636ecafd254a7c7bc79d26e9a97fa9199 is on main through PR #50 and is published as source prerelease v0.4.0-rc.1. Clone main or check out that tag. Workspace packages remain private and independent human installation remains pending.
The Gate 2–4 acceptance record links the process-crash, concurrent-claim, public HTTPS and agent clean-install evidence. Gate 2/3 engineering acceptance passed; Gate 4 awaits a non-author human, and Gate 5 external two-week/repeat-use validation has not started. npm run package:handoff captures review source with per-file hashes. The existing tested archive is a preserved pre-commit snapshot, not a package of subsequent roadmap edits.
C2 Monitor/Delivery MCP and its local HTTPS first-use workflow are implemented. C3 has a unified process and authenticated Streamable HTTP implementation, experimental and not deployed (archived setup). B1/B2 and C1 remain in_progress for their broader operational/adoption gates. See the first-use walkthrough and dated local evidence.
Files: PDF, CSV, XLSX, ZIP, JSON
Scrape, batch items, crawl pages, MCP and /fc take a URL that answers with a file the same way as a web page. A 2xx response is a file when its Content-Type says PDF, CSV (including +csv types such as Eurostat's SDMX-CSV), JSON (and +json), plain text, XLSX, XLS or ZIP; when it says nothing useful (application/octet-stream and its kin, or none), the bytes decide: a %PDF- header in the first 1024 bytes, a ZIP header (an XLSX when the file is named .xlsx), an OLE header named .xls, or text named .csv, .json or .txt. A PDF header at the start overrides text/html or text/plain, and an HTML document sent as text/plain is still a page. The name is the Content-Disposition filename, else the URL path.
- Saved as received, never sent to the browser. The HTTP lane saves every byte to
<W2L_TASK_ROOT>/files/<sha256>.<ext>(.w2l/api/files/by default;npm run scrapeuses the same place,npm run crawlits task directory'sfiles/), so the same bytes are stored once however many URLs serve them. A file never escalates to the browser, whatever its outcome. The result'sfileblock gives the kind, how it was detected, theContent-Typeas received, the declared and received size, the SHA-256, the path, and for a PDF its pages;evidence.rawBodySha256and the Evidence Record'srawSha256are the SHA-256 of the bytes. A request refused before anything is fetched (robots.txt, policy, DNS) saves nothing, and so does a non-2xx answer. - The browser lane catches the download. When the browser is the first rung (
waitFor) or otherwise reaches a file, it takes the file from the download the navigation starts, or from the response it displays (JSON, text), instead of failing withDownload is starting, and saves the same bytes the same way. - Size cap.
W2L_MAX_FILE_BYTESsets the largest file in bytes (default 52 428 800, 50 MiB; at most 524 288 000, 500 MiB; anything else stops the service at start). A request, batch or crawl can lower it withmaxFileBytes, never raise it (a larger value is HTTP 400invalid_request). A file over the cap isfailedwithbody_too_large, with its declared size infile.declaredByteswhen the server sent one; it is not read further and nothing is saved or truncated. Web pages keep the 10 MiB body cap. - Other binary types (images, audio, video, fonts, Word and PowerPoint documents, other archives) are
failedwithunsupported_content_type: not downloaded, not saved, not sent to the browser. - A body that stalls or breaks off after the headers is
failedwithtimeoutorconnection_error, with nothing saved (it was an internal error before).
What each kind returns:
onlyMainContent, includeTags and excludeTags do not apply to files, and waitFor is not waited for once the file arrives. A file result has no document or metadata and no links, and its html and rawHtml, when asked for, are null. The Evidence Record lists the file in artifacts as { kind: "file", path, sha256, bytes, contentType } and names the extractor pdf-text (PDF_TEXT_VERSION) for a PDF or file-text (FILE_TEXT_VERSION) for another file; outputSha256.markdown covers the Markdown delivered.
JSON extraction reads a PDF deterministically: a schema key is matched, as on a web page, to the labels of the PDF's Label: value lines (for example KPI 2: Reduction of carbon intensity fills kpi2), and fieldEvidence gives each such field { source: "pdf", locator: "page N \"label\"" }. Labels that state different values leave the field out with field_ambiguous; prose and table cells are not read, the PDF's metadata is not used, and modelFallback is not applied to PDF text (a model_unavailable issue says so), so a field not found is reported missing, never guessed. PDF text runs on the API process's thread: the 304-page IEA report takes under a second.
PDF text
pdfToMarkdown(bytes, options?) in packages/extract-tf turns the bytes of a PDF into Markdown with page numbers, so a figure quoted from a report can be traced to its page. Scrape, batch, crawl, MCP and /fc use it for every PDF they fetch (see Files).
What it does:
- Reads the PDF's own text layer with Mozilla pdf.js (
pdfjs-dist6.3.289, Apache-2.0), in Node, without rendering. - Starts each page with a line
<!-- page N -->, N being the page's position in the file, and returnspages[]: each page'stext, its printedlabelwhen the PDF declares one, and thestart/endoffsets of that text in the Markdown.pdfPagesForSpan(pages, start, end)names the pages any span of the Markdown came from. - Rebuilds lines, spaces and paragraphs from text positions, reads multi-column pages column by column and keeps table rows as lines. A word hyphenated at a line end is joined; the hyphen is removed only where the document spells the word without it elsewhere.
- Reports
infoas the PDF declares it (title, author, producer, dates, language), null where it declares nothing.
What it does not do:
- No OCR: a page without a text layer (a scan) comes back empty with a
no_text_layerwarning. - No table reconstruction: cells become lines of text, and every result with text carries
tables_unverified. - Running headers and footers stay in the text unless
repeatedLines: 'remove', which lists the removed lines per page. maxPages(default 1000) andtimeBudgetMs(default 60 000, checked before each page) stop with apage_caportime_budgetwarning and the pages read so far. Encrypted, malformed and non-PDF input returns{ ok: false, error: { code, message } }instead of throwing.
It is checked on 10 public reports, six of them the seed user's PDFs: manifest, node research/pdf-corpus/run.mjs, runs in research/pdf-corpus/runs/.
Scrape, batch and crawl take Firecrawl's parsers to choose how a PDF is read (REST, SDK, MCP and /fc); other files are unaffected:
- Absent: every PDF's text layer, with the defaults above and page markers.
[]: no PDF text. The file is saved as received and the result issuccesswithmarkdown: nulland apdf_not_parsedfile warning, as for a spreadsheet.- One
pdfentry, the string"pdf"or{ "type": "pdf", "mode", "maxPages", "pages", "pageMarkers" }:modeisfastorauto, both the text-layer reader;ocrand theimageparser are refused with HTTP 400 by name, since Octocrawl runs no OCR.maxPages(1 to 10 000) reads the first pages. A document cut by the request's ownmaxPagesstayssuccess, withfile.pdf.pagesRead,pageCountand apage_capwarning saying how much was read; only the default cap's cut ispartial.pages: trueaddspages: [{ pageNumber, markdown }], each page's text as the Markdown has it, without its marker (full and compact scrape responses, batch items, crawl pages,/fcdata.pages).pageMarkers: falseleaves the<!-- page N -->lines out;file.pdf.pagesoffsets still locate each page in the Markdown. Natively they are on by default; on/fcthey are off unless asked, as on Firecrawl.
A second entry, an unknown key or a value out of range is HTTP 400 naming it.
Benchmark
Run the full fixture suite against the bare HTTP baseline:
Expected output:
The bare HTTP baseline intentionally has a high false-success rate (no content extraction, no challenge detection, no redirect handling). A production subject should beat these numbers.
Throughput and crash recovery
npm run bench:throughput measures pages per minute and per-page time through the API process on a loopback site of 20 hosts: the HTTP lane with 32 workers over 1,000 URLs, and the browser lane with 8 over 200. The first results, with what they leave out, are in docs/benchmarks/2026-10-03-throughput.md. npm run verify:batch-crash-1000 kills the API with kill -9 partway through a batch of 1,000 URLs over 20 hosts, starts it again on the same task root and checks that the resumed batch loses no URL and records none twice. npm run verify:serve-smoke starts octocrawl serve, scrapes a page and a PDF, stops the server partway through a batch, starts it again and checks that the batch completes. CI runs the crash test on Linux and the serve test on Windows. All three need the compiled packages (npx tsc -b).
The API's engine runs W2L_WORKER_COUNT pages at once, an integer from 1 to 64 (default 4). That bound sits above the per-host limits: W2L_PER_HOST_CONCURRENCY and W2L_PER_HOST_MIN_DELAY_MS still apply to each host, so more workers help only across hosts. A crawl's maxConcurrency may go up to the worker count.
Repository Structure
Roadmap
- Contracts and ground-truth schema
- Fixture server with 56 ground-truth cases
- robots.txt ReDoS fix (token-based glob matcher)
- Benchmark pipeline with bare HTTP baseline
- extract-tf + HTML→Markdown after extract
- Browser lane (Playwright) and HTTP → browser → vendor ladder
- Honest identity bundle (UA / hints / locale / viewport must agree)
-
octocrawlproduct CLI (@octocrawl/cli,octocrawl):scrape,crawl,batch,mapandserve, every API option as a flag -
octocrawl crawl+ SQLite checkpoint resume - REST API + TypeScript SDK (
POST /v1/scrape,POST /v1/crawl,GET /v1/crawl/:id, paginated crawl pages/errors, cancel) - MCP server (
scrape,crawl,get_crawl, paginated pages/errors, and cancel over REST) - Compact MCP scrape responses, direct structured JSON/JSON Schema extraction, and Amazon subject adapter/baseline
- Firecrawl
/scrape/crawlmigration shim (snapshot 2026-09-18; not a compatibility layer) - Task-level ladder accounting, preserved per-channel attempts, and honest unknown cost/evidence fields
- Bounded multi-page workers, shared host scheduling, conditional browser settling, and runtime resource reuse
- Phase 1 Local Reliability Gate: Chromium-backed full test suite and GitHub Actions
- Phase 2 L0-L2 quality benchmark: Octocrawl ladder, verified completion, false-success, P95, escalation, and tiered reports
- Phase A4 real-task harness: AI knowledge and product-info manifests, field assertions, repeat consistency, holdout and cost/evidence records
- Phase A4 real-task gate: 100-200 permitted pages, human correction time, repeated task evidence, and complete failure taxonomy
- Phase A4 diagnostic expansion: 20 real tasks, 11 domains, 40 repeated runs, and holdout results
- A6 scale slice: 100 pages, 10 domains, two runs; labeled holdout is not independent
- A6 recovery/install evidence recorded: interrupt-resume lost 0 URLs; same-machine clean-clone first task; historical 18-minute correction record lacks human confirmation; billed USD unknown
- A6 deferred exceptions recorded: second-developer install is deferred, not passed; billed USD is unknown, not zero
- A6 unconditional pass still needs a second human install
- A5/A6 gate report: conditional alpha; billed USD remains unknown
- Gate 2 execution contract, actual process recovery, controlled changes/cache and Monitor isolation
- Gate 3 durable HTTPS delivery, same-event retry, deduplication and restart recovery
- Gate 4 SDK, docs, examples and agent clean installation
- Gate 4 independent non-author human installation and full workflow
- C2 Monitor/Delivery MCP and local conversational first-use flow
- C2 n8n and narrow task UI
- C3 unified single-instance process and authenticated Streamable HTTP implementation
- C3 hosted MCP: deployment, sign-in acceptance and hosted restart drill (experimental code; setup archived in docs/archive/hosted-mcp-pilot.md; hosting is a P5 item)
- Gate 5 two external trial users, two weeks, repeat use and real downstream consumption
- Phase 3 Benchmark Gate harness: fixed Octocrawl run, comparator evidence, and blocked-until-real-comparators decision
- Hosted Egress Gate: browser subresource policy enforcement and DNS-to-connection binding
See ROADMAP.md for the current phase plan; the Section A/B/C roadmap is archived in docs/roadmap/sections-abc-roadmap-2026-09-28.md. PRODUCT_PLAN_V2.md remains the historical detailed plan.
Contributing
We use the Developer Certificate of Origin (DCO) instead of a CLA. Every commit needs a Signed-off-by line:
See CONTRIBUTING.md for details.
License
Server-side code, the CLI (@octocrawl/cli, octocrawl) and the MCP server (@octocrawl/mcp): AGPL-3.0
Client libraries: MIT: the TypeScript SDK (@octocrawl/sdk, which includes the workspace's @w2l/contracts, also MIT) and the Python client (octocrawl-client)
See PHASE1_ENGINEERING_NOTES.md §1.3 for the rationale.
Why AGPL?
AGPL requires network-deployed modifications to remain open. Anyone can fork, modify, and host Octocrawl — as long as they share those modifications. The real differentiator is the name (trademark) and the hosted service, not the license lock.
來源:README.md,提交 a85bc50
工具
0版本歷史
1- v0.3.1最新Oct 9, 2026
