
Spicrawl
io.github.OfficialSpicrawlv0.1.4Updated Oct 8, 2026
Web scraping to Markdown and structured data, screenshots, batch jobs and browser sessions.
Overview
Lets an assistant scrape web pages into Markdown, text, HTML, JSON or PDF, extract structured data, take screenshots, and run large batch jobs.
- What it does
- Spicrawl exposes 25 tools for web scraping. spicrawl_scrape retrieves one URL as Markdown (default), text, HTML, JSON or PDF, optionally rendering JavaScript, extracting fields with CSS/XPath selectors or autoparse, capturing screenshots and PDFs, and running browser actions such as click, fill, scroll and wait. Batch tools submit up to 10,000 URLs as one async job and poll, page, retry or cancel it. Session tools keep a logged-in browser identity, and usage, request-log and docs tools report credits and errors.
- When to use it
- Use it when an assistant needs live page content or structured data from websites, including JavaScript-rendered pages, pages behind a login, or many URLs at once. It fits research, price and listing extraction, and screenshot capture inside an agent workflow.
- Requirements
- A Spicrawl API key in SPICRAWL_API_KEY (stdio) or an Authorization bearer header (hosted endpoint). The hosted endpoint at mcp.spicrawl.com needs no install; the local option needs Node.js 20+ and runs via npx @spicrawl/mcp. Network access to the Spicrawl API is required, and usage is billed in credits.
Installation
In SourceWeft
- Open Spicrawl in the dashboard and add it to a workspace.
- Enable the server for the chats that should use its tools.
Web executable via Streamable HTTP. Remote servers run from the web runtime once configured in a workspace.
Other MCP clients
Add this to your client's mcpServers config.
{
"mcpServers": {
"mcp": {
"type": "http",
"url": "https://mcp.spicrawl.com/mcp"
}
}
}README
Spicrawl MCP Server
The Spicrawl MCP server is a Model Context Protocol server that gives AI agents in Claude Code, Claude Desktop, Cursor, VS Code, Windsurf, Gemini CLI and Codex web scraping to Markdown, structured data extraction, screenshots, browser sessions and batch scraping as 25 native tools.
[npm version] [npm downloads] [license] [CI]
Docs · Get an API key · TypeScript SDK · CLI · Agent plugins · Changelog
[Add Spicrawl MCP server to Cursor] [Install Spicrawl MCP server in VS Code]
Quickstart
-
Create an API key at app.spicrawl.com. Keys start with
spicrawl_live_orspicrawl_test_. -
Connect the hosted endpoint,
https://mcp.spicrawl.com/mcp(MCP Streamable HTTP, nothing to install):Codex:
codex mcp add spicrawl --url https://mcp.spicrawl.com/mcp --bearer-token-env-var SPICRAWL_API_KEY -
Or run it locally over stdio (Node.js 20+), for clients that only launch local servers:
-
Ask your agent: "Use Spicrawl to read https://example.com/pricing as markdown and list each plan with its price." It calls
spicrawl_scrapeand gets the page back as Markdown.
Both transports serve the same 25 spicrawl_* tools, and every call runs under your key with the same scopes, limits and credits as a direct API call.
What it does
- Scrape any URL to Markdown, text, HTML, a JSON envelope or a printed PDF with one tool,
spicrawl_scrape. Markdown is the default and strips the page to its main content. - Render JavaScript with
render: true(theobscurabrowser engine, 3 credits) or pin real Chromium withengine: "chromium"(8 credits). A plain fetch costs 1 credit, andmode: "auto"escalates fetch to browser and bills only the rung that worked (price table). - Extract structured data with a CSS/XPath selector map (
extract) or from the JSON-LD, OpenGraph, Twitter Card and microdata a page publishes (autoparse). - Drive the page before capture with 8 typed browser actions:
click,fill,wait_for,wait_for_navigation,scroll(including infinite scroll),select,evaluateandscreenshot. - Return screenshots and PDFs the model can see: screenshots arrive as MCP image blocks (up to 5 MB of base64 each) and printed pages as
application/pdfresource blocks (up to 10 MB). - Batch scrape thousands of URLs as one async job, up to 10,000 items per call, with 9 tools to submit, poll, page through results, retry, cancel and keep a job open for crawling.
- Keep a logged-in browser with persistent sessions: cookies, storage and a pinned engine reused across scrapes.
- Answer "what did that cost?" from usage and request-log tools, and read the Spicrawl docs from inside the agent.
Which tool to use
- If you need one page's content or data now, use
spicrawl_scrape. - If the page is blank or a skeleton without JavaScript, call
spicrawl_scrapeagain withrender: true. - If you know the page's structure, use
extractwith CSS selectors; if you don't, tryautoparse. - If you have more than about 20 URLs, use
spicrawl_batch_submit, thenspicrawl_batch_statusandspicrawl_batch_results. - If you discover URLs as you go (crawling), submit with
open: true, append withspicrawl_batch_add_itemsand finish withspicrawl_batch_close. - If you must log in first, use
spicrawl_session_createand pass itsidassession_idtospicrawl_scrape. - If a call failed and you need to know why, use
spicrawl_request_getwith the request id, orspicrawl_requests_listwithonly_errors. - If you are unsure of a parameter or an error code, use
spicrawl_docs_searchbefore guessing.
Tools
25 tools, grouped by task. Tools reject unknown arguments, so a misspelt field fails with an error naming it instead of being silently dropped. Every tool sets a title and explicit readOnlyHint, destructiveHint and openWorldHint annotations, so a client knows which calls to confirm: spicrawl_scrape is not read-only (its method and actions can submit forms), and spicrawl_batch_cancel, spicrawl_session_release and spicrawl_session_delete are destructive. Full argument reference: docs.spicrawl.com/agents/mcp.
Scrape a website to Markdown
Batch scrape many URLs
Browser sessions
Usage and request history
Docs
Coming soon
spicrawl_browser_connect_url will mint a single-use CDP WebSocket URL for driving a Spicrawl-hosted browser from Puppeteer or Playwright; remote browsers are not available yet. The schemas also accept ai_extract, stealth, extract_preset, premium_proxy, proxy_country and sticky_key ahead of launch. Agents should not send them yet. To list only what works today, set SPICRAWL_MCP_HIDE_UNAVAILABLE=1 (see Environment variables).
Example prompts
- "Scrape https://example.com/pricing as markdown and summarise the plans."
- "Get the product name, price and stock status from this URL." (
spicrawl_scrapewithextractorautoparse) - "Take a full-page screenshot of https://example.com after the cookie banner is gone." (
render,actions,screenshot_fullpage) - "Scrape these 500 URLs and give me each page title." (
spicrawl_batch_submit, thenspicrawl_batch_results) - "Why did my last scrape fail?" (
spicrawl_requests_listwithonly_errors, thenspicrawl_request_get) - "How many credits have I used this month?" (
spicrawl_usage_summary)
Set up your MCP client
Put your key in SPICRAWL_API_KEY and reference it from the config where the client allows, so the key never lands in a committed file.
Claude Code
Hosted endpoint, stored for you only (add --scope user for every project):
To share it with a team, commit .mcp.json at the project root. Claude Code expands ${VAR}, so each teammate's own key is used:
Or install the plugin, which bundles the MCP server and the Spicrawl agent skill:
Check with claude mcp list, or /mcp in a session. More: docs.spicrawl.com/agents/claude-code.
Claude Desktop
Claude Desktop's claude_desktop_config.json starts local (stdio) servers only, and custom connectors cannot send an Authorization header, so run the package locally. The file is at ~/Library/Application Support/Claude/claude_desktop_config.json on macOS and %APPDATA%\Claude\claude_desktop_config.json on Windows:
Restart Claude Desktop after saving.
Cursor
Use the Add to Cursor button above, or add this to .cursor/mcp.json in the project (or ~/.cursor/mcp.json for every project). Cursor resolves ${env:NAME} in url and headers, so start Cursor from an environment where SPICRAWL_API_KEY is set:
VS Code (GitHub Copilot)
Use the Install in VS Code button above, or add this to .vscode/mcp.json. VS Code uses the servers key, and the input variable makes it ask for the key once and store it securely, so the file is safe to commit:
VS Code does not expand ${env:...} inside headers, so use the input variable. More: docs.spicrawl.com/agents/github-copilot.
Windsurf (Devin Desktop)
Windsurf is now named Devin Desktop. For the default Devin Local agent, add the server in local scope (saved to the gitignored .devin/mcp_config.local.json):
For the legacy Cascade agent, edit ~/.config/devin/mcp_config.json (%APPDATA%\devin\mcp_config.json on Windows). Cascade reads the URL from serverUrl and replaces ${env:...}:
Gemini CLI
Gemini CLI strips variables whose names contain KEY, TOKEN or AUTH before expanding headers, so export the key under another name:
Then add to ~/.gemini/settings.json (or .gemini/settings.json for one project). Streamable HTTP uses the httpUrl key:
Or: gemini mcp add --scope user --transport http --header 'Authorization: Bearer $SPICRAWL_MCP_BEARER' spicrawl https://mcp.spicrawl.com/mcp. Run /mcp to check. More: docs.spicrawl.com/agents/gemini-cli.
Codex
Codex reads the key from the environment each time it connects, so the key never lands in a file:
The equivalent ~/.codex/config.toml (or .codex/config.toml in a trusted project):
Local stdio instead:
Any other MCP client
A stdio-only client runs npx -y @spicrawl/mcp with SPICRAWL_API_KEY in its environment. The server accepts API keys only, not OAuth. Setup pages for more clients (OpenCode, Zed, JetBrains, Cline, Goose and others): docs.spicrawl.com/agents/mcp.
Environment variables
With no docs variable set, a server pointed at a self-hosted API reads that API's own docs at <SPICRAWL_PUBLIC_BASE_URL or SPICRAWL_BASE_URL>/docs.
Self-host the MCP server over HTTP
The hosted endpoint runs this package's Streamable HTTP server. To run your own:
SPICRAWL_BASE_URLdefaults tohttp://127.0.0.1:8080in HTTP mode (a Spicrawl API on the same host), not to the public API. Set it tohttps://api.spicrawl.comto front the public API.- MCP is served at
POST/GET/DELETE /mcp; liveness atGET /healthz(no auth). - Each caller sends their own key as
Authorization: Bearer spicrawl_live_.... The key is checked against the API before a session opens (401if refused,503withRetry-Afterif the API cannot be reached), and a session is pinned to the key that opened it. - The server speaks plain HTTP. Put it behind TLS before sending keys across a network.
ChatGPT endpoint (OAuth)
The same process can serve a second, restricted MCP endpoint for ChatGPT apps. It is off unless SPICRAWL_CHATGPT_ENABLED is on, and /mcp is unchanged either way.
Authentication takes OAuth access tokens (spicrawl_oat_…) only:
- No token:
401withWWW-Authenticate: Bearer resource_metadata="https://mcp.spicrawl.com/.well-known/oauth-protected-resource/chatgpt/mcp", scope="scrape batch". - A bad, inactive, expired or wrong-audience token, or an API key (
spicrawl_live_…/spicrawl_test_…, never accepted here): the same401pluserror="invalid_token", error_description="…". - A tool call the token's scopes do not cover:
403witherror="insufficient_scope".spicrawl_scrapeneedsscrape, the batch tools needbatch, the docs tools need only a valid token. - The authorization server unreachable or refusing this server's secret:
503withRetry-After.
Each token is checked with POST <issuer>/api/oauth/introspect (Authorization: Bearer <SPICRAWL_OAUTH_INTROSPECT_SECRET>, form body token=…). Only active: true with aud equal to the resource is accepted. A positive answer is cached by the token's SHA-256 for at most 30 seconds and never past its exp. Tools run upstream under the api_key the introspection returns; neither it nor the token is logged. If the API refuses that key mid-call, the tool returns an error whose _meta["mcp/www_authenticate"] carries a fresh invalid_token challenge, and the token is dropped from the cache.
The endpoint lists six tools, each with explicit readOnlyHint/destructiveHint/openWorldHint, a title, and securitySchemes: [{type: "oauth2", scopes: [...]}] (also in _meta.securitySchemes):
spicrawl_scrape is read-only here, unlike on /mcp: this endpoint only sends a GET, with no method, body, headers or actions. spicrawl_batch_submit is not read-only, because it creates a job.
spicrawl_batch_results fills in the content of a succeeded item that the results page lists without it, which happens while the job is still finishing, from GET /v1/batch/{id}/tasks/{seq}/content. It fetches at most 25 items and 2 MiB per page. An item that can't be read yet comes back unavailable, and one past the size cap omitted.
Every other argument is refused, not dropped. Results keep the content, the site's status and the credits charged; request ids, timestamps, engine and proxy details, storage references and project ids are left out.
Installed globally (npm install -g @spicrawl/mcp), the command is spicrawl-mcp (alias spicrawl-mcp-server), with --http, --version and --help.
FAQ
Is there a hosted Spicrawl MCP server?
Yes: https://mcp.spicrawl.com/mcp, MCP Streamable HTTP, authenticated with Authorization: Bearer <API key>, so there is nothing to install.
Is Spicrawl free?
Spicrawl bills credits per successful request, and during the beta every organization gets one plan of 1,000 credits per month. Failed requests cost 0 credits; a cache hit is billed at the same price as the fetch that stored it. See Credits.
Does it render JavaScript?
Yes. Set render: true on spicrawl_scrape to run the page in a browser, or mode: "auto" to let Spicrawl escalate from a plain fetch only when needed.
Does it work with Claude Code, Cursor and VS Code?
Yes, and with Claude Desktop, Windsurf, Gemini CLI, Codex and any client that speaks MCP Streamable HTTP with custom headers or launches stdio servers.
Can it crawl a whole website?
There is no single crawl tool. Scrape a page with links: true to get its links, then feed them to an open batch job (open: true and spicrawl_batch_add_items).
Does it support OAuth?
Not on /mcp: it accepts Spicrawl API keys only, so turn OAuth off for it in clients that try it on a 401. The separate ChatGPT endpoint at /chatgpt/mcp takes OAuth access tokens only, when the operator turns it on.
Should I use the MCP server, the SDK or the CLI?
Use the MCP server when an AI agent should call Spicrawl as tools. Use @spicrawl/sdk in your own TypeScript or JavaScript code, @spicrawl/cli in a terminal or shell script, and the REST API from any other language.
Troubleshooting
curl -i https://mcp.spicrawl.com/mcp without a key answers 401: the endpoint is reachable.
Related packages
@spicrawl/sdk: the official TypeScript SDK for the Spicrawl API.@spicrawl/cli: the Spicrawl command-line interface.- OfficialSpicrawl/agent-plugins: plugins that bundle this server and the Spicrawl skill for Claude Code, Codex, Cursor, Gemini CLI and more.
Links
- Documentation: docs.spicrawl.com/agents/mcp
- Docs for LLMs: docs.spicrawl.com/llms.txt
- Dashboard and API keys: app.spicrawl.com
- Issues: github.com/OfficialSpicrawl/mcp/issues
- Security: SECURITY.md. Never commit an API key; revoke a leaked one at app.spicrawl.com.
Publishing to the MCP Registry
server.json lists both the npm/stdio package and the hosted
Streamable HTTP endpoint under io.github.OfficialSpicrawl/mcp. API keys are
requested from the person installing the server; no key belongs in the manifest.
The existing release workflow publishes to the MCP Registry after its npm
job succeeds, on pushes to main or manual runs from main. Registry authentication
uses GitHub OIDC (id-token: write), so it needs no new secret, device login or DNS
record. The npm release keeps using the existing NPM_TOKEN secret.
For a release, run npm version --no-git-tag-version X.Y.Z, update the changelog,
and submit the changes to main. The npm version hook synchronizes both versions
in server.json. If you edit package.json by hand, run npm run registry:sync.
CI checks the manifest against its official $schema with Ajv and rejects name
or version drift before npm publication. To run that check locally:
The registry job checks that the published npm package has the matching mcpName
before publishing. An existing registry version is skipped; lookup failures stop
the job. To retry a registry failure or list the already-published current npm
version, run the release workflow on main again. Use a new version for metadata
changes to an existing listing. The publisher binary is pinned and its download
is checked against the release checksums; update its version in release.yml
when upgrading the CLI.
Verify publication at
https://registry.modelcontextprotocol.io/v0.1/servers?search=io.github.OfficialSpicrawl/mcp.
The workflow must be merged into OfficialSpicrawl/mcp before OIDC can publish
that namespace. Fork pull requests validate and test, but never run the registry job.
Official references: publishing quickstart, GitHub Actions, and remote servers.
License
Source: README.md at commit d4dbbb4
Tools
0Version history
1- v0.1.4LatestOct 8, 2026

