ScrapingBot
io.scrapingbotv1.0.0更新于 Oct 4, 2026
Public web data for AI agents: scrape any page, plus Google, TikTok, Instagram and Amazon as JSON.
概览
让助手抓取任意网页,并以 JSON 获取 Google、TikTok、Instagram 和 Amazon 的公开数据,通过托管端点接入。
- 功能
- 这是一个托管型 MCP 服务器,通过 Streamable HTTP 提供 23 个工具。它能把任意网页抓取为 markdown 或 HTML,可选 JavaScript 渲染、截图和浏览器操作,并可用 AI 提取结构化字段。其他工具覆盖 Google 搜索、图片、视频、新闻、购物、地点、地图和评论;Instagram 个人资料、媒体、粉丝和搜索;TikTok 视频、用户、评论和搜索;以及 Amazon 搜索、商品详情、榜单和搜索建议。另有工具用于列出能力和轮询抓取任务。
- 适用场景
- 当助手需要实时公开网页数据时使用:抓取并解析页面、执行浏览器操作、用 AI 提取字段,或查询公开的 Google、TikTok、Instagram 和 Amazon 数据,而无需自建爬虫或代理。它是按额度计费的付费服务,适合偶尔或中等规模的数据采集,不适合大规模持续爬取。若只处理本地文件或私有数据则无需使用。
- 运行要求
- 远程托管的 MCP 端点 Streamable HTTP,无需安装。需要 ScrapingBot 账号和 API 密钥,通过 x-api-key 请求头(或 Authorization: Bearer,或仅支持 URL 的客户端使用 api_key 查询参数)发送。需要能访问 scrapingbot.io。注册免费赠送 100 额度,无需信用卡。
安装
在 SourceWeft 中
- 打开 控制台中的 ScrapingBot,将其添加到工作区。
- 为需要使用其工具的对话启用该服务。
Web executable,通过 Streamable HTTP。 远程服务在工作区中配置后即可从网页运行时运行。
其他 MCP 客户端
把它添加到你客户端的 mcpServers 配置中。
{
"mcpServers": {
"scrapingbot": {
"type": "http",
"url": "https://scrapingbot.io/api/mcp"
}
}
}README
ScrapingBot
One API key for public web data: scrape any page (plain or JavaScript-rendered), pull fields out with AI, and get public TikTok, Instagram, Google and Amazon data as JSON. A hosted MCP server gives AI agents the same data as tools.
100 free credits when you sign up. No credit card. Get your API key
Docs · Pricing · MCP server · Examples · Agent Skills
Quickstart
That call cost 1 credit. Failed requests are refunded automatically.
Base URL: https://scrapingbot.io/api/v1
Auth: x-api-key: YOUR_API_KEY header (or Authorization: Bearer YOUR_API_KEY). Keep the key on the server.
Endpoints
The data APIs (TikTok, Instagram, Google, Amazon) are all POST with a JSON body naming the endpoint and its params:
Full parameters and response shapes: docs.
How charging works
- Data APIs charge only for a successful
200. Any error is refunded. - Website API charges when the page answers (2xx/3xx, or a real 400/404 from the site, reported as
fault: "user"). Timeouts, other 4xx such as 403 and 429, 5xx and errors on our side are refunded (credits_used: 0). - Requests rejected before they run (missing parameter, bad key, concurrency limit) cost nothing.
Limits
Limits are on requests in flight at once, not per minute. The free plan allows 1 concurrent request; paid plans allow 10 to 200. Going over returns 429 right away (free); retry with a short, growing delay. Timeouts: 45 s for scraping and Instagram, 30 s for TikTok and Amazon, 15 s for Google.
MCP server for AI agents
Hosted at https://scrapingbot.io/api/mcp (Streamable HTTP, stateless). Nothing to install. Authenticate with your API key in the x-api-key header, or Authorization: Bearer.
Claude Code
Add --scope user to use it in every project. Type /mcp in a session to see the tools.
Claude Desktop (and Claude on the web)
Open Settings → Connectors → Add custom connector, name it ScrapingBot, and paste:
Then turn it on from the tools menu in a chat. Connectors take a URL only, so the key goes in the URL: keep that URL private, like a password.
Cursor
~/.cursor/mcp.json (or .cursor/mcp.json in a project):
VS Code
.vscode/mcp.json in your project, then pick the ScrapingBot tools in Copilot's agent mode:
Any other MCP client
URL https://scrapingbot.io/api/mcp, header x-api-key: YOUR_API_KEY (or Authorization: Bearer YOUR_API_KEY). If the client only takes a URL, append ?api_key=YOUR_API_KEY. When calling it by hand, send Accept: application/json, text/event-stream.
Tools
Tool calls are ordinary API requests: same credits, same concurrency slots, same refunds.
Agent Skills
skills/ holds Agent Skills that teach an agent to call the REST API directly with SCRAPINGBOT_API_KEY: which endpoint for which task, the parameters, the response fields worth reading, error handling and costs.
Claude Code: copy a folder into ~/.claude/skills/ (or .claude/skills/ in a project). Set SCRAPINGBOT_API_KEY in the environment the agent runs in.
Examples
Small, runnable scripts in examples/, each in curl, Python (requests) and Node (built-in fetch, Node 18+):
Python needs pip install requests; the curl scripts that build JSON use jq. Each language folder has a tiny shared client (scrapingbot.py, scrapingbot.mjs) that retries refunded failures (408, 429, 5xx) with backoff.
Pricing
Free: 100 credits on sign-up, 1 concurrent request. Paid plans start at $49.99/month for 275,000 credits and 10 concurrent requests. See pricing.
Responsible use
ScrapingBot returns publicly available data. Respect each site's terms and applicable law, and handle any personal data you collect accordingly.
License
The examples and skills in this repository are MIT licensed. See LICENSE. Use of the ScrapingBot API is subject to the terms at scrapingbot.io.
来源:README.md,提交 349f2bf
工具
0版本历史
1- v1.0.0最新Oct 4, 2026
