ScrapingBot

io.scrapingbotv1.0.0更新于 Oct 4, 2026

Public web data for AI agents: scrape any page, plus Google, TikTok, Instagram and Amazon as JSON.

已验证Streamable HTTP可网页运行Web Search & ScrapingData & AnalyticsBrowser Automation

概览

AI 生成的概览

让助手抓取任意网页,并以 JSON 获取 Google、TikTok、Instagram 和 Amazon 的公开数据,通过托管端点接入。

功能
这是一个托管型 MCP 服务器,通过 Streamable HTTP 提供 23 个工具。它能把任意网页抓取为 markdown 或 HTML,可选 JavaScript 渲染、截图和浏览器操作,并可用 AI 提取结构化字段。其他工具覆盖 Google 搜索、图片、视频、新闻、购物、地点、地图和评论;Instagram 个人资料、媒体、粉丝和搜索;TikTok 视频、用户、评论和搜索;以及 Amazon 搜索、商品详情、榜单和搜索建议。另有工具用于列出能力和轮询抓取任务。
适用场景
当助手需要实时公开网页数据时使用:抓取并解析页面、执行浏览器操作、用 AI 提取字段,或查询公开的 Google、TikTok、Instagram 和 Amazon 数据,而无需自建爬虫或代理。它是按额度计费的付费服务,适合偶尔或中等规模的数据采集,不适合大规模持续爬取。若只处理本地文件或私有数据则无需使用。
运行要求
远程托管的 MCP 端点 Streamable HTTP,无需安装。需要 ScrapingBot 账号和 API 密钥,通过 x-api-key 请求头(或 Authorization: Bearer,或仅支持 URL 的客户端使用 api_key 查询参数)发送。需要能访问 scrapingbot.io。注册免费赠送 100 额度,无需信用卡。
安装前请注意
需要 ScrapingBot API 密钥(x-api-key 请求头,或 Authorization Bearer;部分客户端把密钥放在 URL 的 api_key 参数中,README 要求像密码一样保密)。调用会消耗付费额度:普通抓取 1、渲染 5、premium 代理 10、stealth 代理 75,AI 提取另加 5,Instagram 每次 5,Google 或 Amazon 每次 10;免费方案 100 额度、1 个并发请求,付费方案每月 49.99 美元起。工具调用与 API 请求同样计费。抓取的是公开数据,仍需遵守目标网站条款与适用法律,并妥善处理个人信息。

安装

在 SourceWeft 中

  1. 打开 控制台中的 ScrapingBot,将其添加到工作区。
  2. 为需要使用其工具的对话启用该服务。

Web executable,通过 Streamable HTTP。 远程服务在工作区中配置后即可从网页运行时运行。

其他 MCP 客户端

把它添加到你客户端的 mcpServers 配置中。

{
  "mcpServers": {
    "scrapingbot": {
      "type": "http",
      "url": "https://scrapingbot.io/api/mcp"
    }
  }
}

README

ScrapingBot

One API key for public web data: scrape any page (plain or JavaScript-rendered), pull fields out with AI, and get public TikTok, Instagram, Google and Amazon data as JSON. A hosted MCP server gives AI agents the same data as tools.

100 free credits when you sign up. No credit card. Get your API key

Docs · Pricing · MCP server · Examples · Agent Skills


Quickstart

bash
export SCRAPINGBOT_API_KEY="YOUR_API_KEY"
curl "https://scrapingbot.io/api/v1/scrape?url=https://example.com" \  -H "x-api-key: $SCRAPINGBOT_API_KEY"
json
{  "success": true,  "url": "https://example.com",  "html": "<!doctype html><html lang=en><head>…",  "status": 200,  "duration": "0.79",  "credits_used": 1,  "job_id": "…"}

That call cost 1 credit. Failed requests are refunded automatically.

Base URL: https://scrapingbot.io/api/v1 Auth: x-api-key: YOUR_API_KEY header (or Authorization: Bearer YOUR_API_KEY). Keep the key on the server.

Endpoints

The data APIs (TikTok, Instagram, Google, Amazon) are all POST with a JSON body naming the endpoint and its params:

bash
curl -X POST "https://scrapingbot.io/api/v1/tiktok" \  -H "x-api-key: $SCRAPINGBOT_API_KEY" \  -H "Content-Type: application/json" \  -d '{"endpoint": "/user/info", "params": {"unique_id": "tiktok"}}'
APIRouteendpoint valuesCredits
WebsiteGET or POST /api/v1/scrapeone route; options: render_js, screenshot, js_scenario, wait_for, premium_proxy, stealth_proxy, …1 plain · 5 rendered · 10 premium proxy · 75 stealth proxy
AI extraction/api/v1/scrape + ai_query or ai_schemareturns ai_result JSON+5 on top of the page
Job lookupGET /api/v1/job/:job_idre-fetch a scrape result (kept 24 hours)free
TikTokPOST /api/v1/tiktok/ (video), /user/info, /user/posts, /user/followers, /user/following, /user/search, /feed/search, /music/info, /music/posts, /comment/list, /comment/reply1
InstagramPOST /api/v1/instagram/user/by_username, /user/by_id, /medias/by_user_id, /reels/by_user_id, /medias/tagged_by_user_id, /stories/by_username, /followers/by_user_id, /following/by_user_id, /media/by_shortcode, /media/by_url, /comments/media_comments_by_id, /comments/replies, /search/users_by_keyword, /search/hashtags_by_keyword, /search/places_by_keyword, /search/global, /search/posts, /media/shortcode_to_id, /media/id_to_shortcode5
GooglePOST /api/v1/google/search, /images, /videos, /news, /shopping, /places, /maps, /reviews10
AmazonPOST /api/v1/amazon/search, /product-details, /products (up to 20 ASINs), /autocomplete, /best-sellers, /new-releases, /product-category-list, /deals-v2, /seller-profile, /seller-products10 (/products: 10 per product returned)
ChatGPTPOST /api/v1/chatgptbody {"prompt": "…"}10
MCPPOST /api/mcp23 tools over Streamable HTTPsame as the API behind each tool

Full parameters and response shapes: docs.

How charging works

  • Data APIs charge only for a successful 200. Any error is refunded.
  • Website API charges when the page answers (2xx/3xx, or a real 400/404 from the site, reported as fault: "user"). Timeouts, other 4xx such as 403 and 429, 5xx and errors on our side are refunded (credits_used: 0).
  • Requests rejected before they run (missing parameter, bad key, concurrency limit) cost nothing.

Limits

Limits are on requests in flight at once, not per minute. The free plan allows 1 concurrent request; paid plans allow 10 to 200. Going over returns 429 right away (free); retry with a short, growing delay. Timeouts: 45 s for scraping and Instagram, 30 s for TikTok and Amazon, 15 s for Google.

StatusMeaningCharged
400Missing or invalid parameter, unsupported endpointNo (except a target page's own 400 on the Website API)
401Missing or invalid API keyNo
402Not enough credits; the message says how many are neededNo
404Website API: the page doesn't exist. Data APIs: profile, post or product not foundWebsite API only
408Timed outNo
429Concurrency limit reachedNo
5xxThe site or a data source failedNo

MCP server for AI agents

Hosted at https://scrapingbot.io/api/mcp (Streamable HTTP, stateless). Nothing to install. Authenticate with your API key in the x-api-key header, or Authorization: Bearer.

Claude Code

bash
claude mcp add --transport http scrapingbot https://scrapingbot.io/api/mcp \  --header "x-api-key: YOUR_API_KEY"

Add --scope user to use it in every project. Type /mcp in a session to see the tools.

Claude Desktop (and Claude on the web)

Open Settings → Connectors → Add custom connector, name it ScrapingBot, and paste:

https://scrapingbot.io/api/mcp?api_key=YOUR_API_KEY

Then turn it on from the tools menu in a chat. Connectors take a URL only, so the key goes in the URL: keep that URL private, like a password.

Cursor

~/.cursor/mcp.json (or .cursor/mcp.json in a project):

json
{  "mcpServers": {    "scrapingbot": {      "url": "https://scrapingbot.io/api/mcp",      "headers": { "x-api-key": "YOUR_API_KEY" }    }  }}

VS Code

.vscode/mcp.json in your project, then pick the ScrapingBot tools in Copilot's agent mode:

json
{  "servers": {    "scrapingbot": {      "type": "http",      "url": "https://scrapingbot.io/api/mcp",      "headers": { "x-api-key": "YOUR_API_KEY" }    }  }}

Any other MCP client

URL https://scrapingbot.io/api/mcp, header x-api-key: YOUR_API_KEY (or Authorization: Bearer YOUR_API_KEY). If the client only takes a URL, append ?api_key=YOUR_API_KEY. When calling it by hand, send Accept: application/json, text/event-stream.

Tools

AreaToolsCredits
Any websitescrapeWebsite (markdown or HTML, optional screenshot), extractStructuredData (AI fields as JSON), runBrowserScenario (click, fill, scroll, wait)1 / 5 rendered; +5 for AI
GooglegoogleSearch (web, images, videos, news, shopping, places, maps), googleReviews10
InstagraminstagramUser, instagramSearch, instagramMedia, instagramFollowers5
TikToktiktokVideo, tiktokUser, tiktokSearch, tiktokComments, tiktokFollowers1
AmazonamazonSearch, amazonProduct, amazonProducts, amazonSuggestions, amazonRankings10
UtilitylistCapabilities, getScrapeJob, pollJobUntilDone (free), providerRequest (any endpoint above)

Tool calls are ordinary API requests: same credits, same concurrency slots, same refunds.

Agent Skills

skills/ holds Agent Skills that teach an agent to call the REST API directly with SCRAPINGBOT_API_KEY: which endpoint for which task, the parameters, the response fields worth reading, error handling and costs.

SkillUse it for
web-scrapingFetching any page, rendering JavaScript, browser actions, screenshots, AI extraction
google-searchGoogle web, images, videos, news, shopping, places, maps and reviews
tiktok-dataPublic TikTok videos, profiles, posts, sounds, comments and search
instagram-dataPublic Instagram profiles, posts, reels, comments and search
amazon-productsAmazon search, product details, batches, rankings, suggestions and deals

Claude Code: copy a folder into ~/.claude/skills/ (or .claude/skills/ in a project). Set SCRAPINGBOT_API_KEY in the environment the agent runs in.

Examples

Small, runnable scripts in examples/, each in curl, Python (requests) and Node (built-in fetch, Node 18+):

Python needs pip install requests; the curl scripts that build JSON use jq. Each language folder has a tiny shared client (scrapingbot.py, scrapingbot.mjs) that retries refunded failures (408, 429, 5xx) with backoff.

bash
export SCRAPINGBOT_API_KEY="YOUR_API_KEY"sh examples/curl/scrape.sh https://example.compython examples/python/google_search.py "best espresso machine"node examples/node/tiktok_user.mjs tiktok

Pricing

Free: 100 credits on sign-up, 1 concurrent request. Paid plans start at $49.99/month for 275,000 credits and 10 concurrent requests. See pricing.

Responsible use

ScrapingBot returns publicly available data. Respect each site's terms and applicable law, and handle any personal data you collect accordingly.

License

The examples and skills in this repository are MIT licensed. See LICENSE. Use of the ScrapingBot API is subject to the terms at scrapingbot.io.

来源:README.md,提交 349f2bf

工具

0
工具元数据尚未被收录。

版本历史

1
  1. v1.0.0最新Oct 4, 2026