
MeshArc
dev.mesharcv0.3.1更新于 Oct 6, 2026
Crawl and scrape websites to clean markdown, and track what changed between crawls.
概览
让助手把网站抓取、爬取为干净的 Markdown,映射站点地图,并跟踪页面随时间的变化。
- 功能
- MeshArc 提供 scrape_urls、extract_url、map_site、crawl_site、keep_crawl_as_project 等工具,以及项目与运行相关工具(list_projects、create_project、start_run、list_pages、get_page、get_changes、search_pages、recrawl_pages、get_job)。页面内容可按需返回 Markdown、文本、HTML、链接、结构化字段或截图,并附带状态、判定结果和错误码。爬取结果可保留为定时项目,记录新增、删除、修改的页面以及逐词差异。结果较大时返回的是索引和摘录,而不是完整页面集合。
- 适用场景
- 当助手需要把网页或整个站点读成干净文本、查看站点地图声明的内容,或按周期监控站点变化时适用。适合文档、定价和内容监测类任务,尤其是需要可重复爬取和变更记录的场景。
- 运行要求
- 可使用远程端点 Python 3.10+ 上安装 mesharc[mcp] 并以 stdio 本地运行。每次调用都需要 MeshArc API 密钥,可直接传入或通过 MESHARC_API_KEY 环境变量提供;托管端点通过浏览器登录并签发自己的密钥。需要能访问 MeshArc API 的网络。
安装
在 SourceWeft 中
- 打开 控制台中的 MeshArc,将其添加到工作区。
- 为需要使用其工具的对话启用该服务。
Web executable,通过 Streamable HTTP。 远程服务在工作区中配置后即可从网页运行时运行。
其他 MCP 客户端
把它添加到你客户端的 mcpServers 配置中。
{
"mcpServers": {
"mesharc": {
"type": "http",
"url": "https://mcp.mesharc.dev/mcp"
}
}
}README
mesharc
The Python client for the MeshArc API: a URL in, clean content out, and a record of what changed.
- Scrape one page or a batch — markdown, text, HTML, links, structured fields, a screenshot.
- Crawl a whole site with no project to set up first, and keep it as one if it turns out to be worth watching.
- Map what a site declares in its sitemaps before fetching any of it.
- Watch a site over time: projects, scheduled runs, and a change record — pages added, removed, modified, field by field.
Python 3.10 or newer. One dependency (httpx). Fully typed.
Install
Authentication
Every call needs an API key. Create one in the app under Settings → API keys — it is shown once — and give it to the client, or put it in MESHARC_API_KEY and construct the client with nothing:
The key is only ever sent as a bearer header to api.mesharc.dev.
A key carries the scopes it was made with (read, write, admin), optionally a set of projects it may see, an expiry and a rate limit. A route the key may not use answers 403; a project it may not see answers 404.
Quick start
scrape holds the request open until the page comes back (60 s by default), so there is nothing to poll for an ordinary page.
Reading pages
One page
config is any setting a project takes, by its API name (render_js, only_main_content, formats, max_tier, wait_for_selector, actions, …). The full list, with defaults, is at mesharc.dev/docs/configuration.
Many pages
A list of URLs is a batch: grouped by host, fetched in parallel where the config allows, and returned as one row per URL.
arc.scrape(urls, wait=False) returns the batch id at once; arc.batch(id, wait=True) finishes it later. Pass webhook_url= to be told instead of polling (batch.finished, signed with a secret returned once).
What a page looks like
Every page row carries the same fields, whether it came from a scrape, a crawl or a project:
A page the site refused costs 0, and so does a 404.
Crawling a site
crawl returns a handle immediately; job.pages() yields pages as they land and ends when the crawl does. wait=True blocks until it finishes; job.wait(), job.refresh(), job.cancel() do what they say; arc.get_crawl(id) reattaches to a crawl started elsewhere.
A one-shot crawl expires after 30 days. If the site is worth watching:
A webhook= in the options ({"url", "events", "metadata"}) is told about crawl.started, crawl.page (fifty pages a message) and crawl.completed; its signing secret comes back once as job.webhook_secret.
Mapping a site
A map costs one credit per sitemap file read — most sites are one file.
Watching a site: projects and runs
The workspace
Errors
Every failure raises MeshArcError:
request_id is the id the API put on the response and in its own logs, so a support conversation starts from one string.
Two more cases: a network failure or a request that hits timeout raises MeshArcError with status == 0 and code network or timeout; a job the client stopped waiting for raises MeshArcTimeoutError — both a MeshArcError and a TimeoutError — which carries job_id so you can poll it later (arc.get_crawl(id), arc.batch(id)).
Idempotency and timeouts
scrape,scrape_oneandcrawltakeidempotency_key=: send the same key again within 24 hours and you get the first answer back rather than a second job.- Waiting calls take
wait=,poll=(seconds between polls) andtimeout=(seconds beforeTimeoutError).wait=Falsereturns the envelope at once; the default polls every 3 s for up to an hour. timeout_son a single scrape is how long the API itself holds the request open (60 s by default, 120 at most); a slower page comes back as an id and is polled.MeshArc(..., timeout=150.0)is the HTTP timeout per request. A request is retried on 429, 502, 503, 504 and network failures when it is safe to repeat — a GET, a DELETE, or a POST with an idempotency key — up tomax_retriestimes (2), honouringRetry-After.
Credits
Every response says what it cost: credits on a page, creditsUsed on a job envelope, X-MeshArc-Credits on the HTTP response. A page costs the engine that read it — a plain fetch 1, a render 4 — and a refused page or a 404 costs nothing. The schedule and the plans are at mesharc.dev/docs/billing.
The MCP server
The package also ships MeshArc as an MCP server, so Claude Desktop, Claude Code, Cursor and any MCP client can scrape, crawl, map and read change records as tools. Python 3.10+.
There is a hosted one, so most people need install nothing:
It opens a browser once to sign in to mesharc.dev and approve, then works. Approving issues the app an API key of its own, which you can see and revoke under Settings → API keys; read only unless you allow changes. Any client that takes a remote MCP URL — the Claude.ai and ChatGPT connectors, Claude mobile, Cursor — takes it the same way.
Or run it yourself, over stdio, with your own key:
Tools: scrape_urls, extract_url, map_site, crawl_site, keep_crawl_as_project, list_projects, describe_project_config, get_project, create_project, update_project, start_run, list_pages, get_page, get_changes, search_pages, recrawl_pages, get_job. Every tool is a call through this client.
A result with many pages in it is a map, not the territory: a crawl answers with an index of the pages read (up to 500, fewer if the budget needs the room), an excerpt of as many as a 60,000-character budget pays for (fifty at most), and a count of each page's links rather than the links. Ask for the one page you want on its own with get_job(kind="crawl", id=…, url=…), or carry on through the crawl with cursor=. A multi-URL scrape past its budget counts the rest by status and names any that did not come back ok. One page asked for on its own -- extract_url, get_page, scrape_urls with a single url -- comes back whole, capped at 12,000 characters.
Hosted, a connection can be approved read-only, and read-only cuts in a place worth knowing: it reads the whole workspace, but it cannot reach the site. Anything that fetches is a write, because it leaves the workspace and usually spends its credits -- map_site is the one that fetches and costs nothing, and it is gated with the rest. list_projects, get_project, describe_project_config, list_pages, get_page, get_changes, search_pages and get_job work; the nine that fetch or write -- scrape_urls, extract_url, map_site, crawl_site, keep_crawl_as_project, create_project, update_project, start_run, recrawl_pages -- answer {"code": "read_only"} with what to do about it, until the app is reconnected with write access.
Hosted, a tool that would hold a connection open for minutes — a crawl, a run you asked to wait for, a multi-URL scrape — hands back a job after MESHARC_MCP_WAIT seconds (25 by default, because many MCP hosts time a tool call out sooner). The work carries on server-side and get_job picks it up. Run locally, those tools block as they always have.
Hosting it yourself
The server holds no key. Each request carries the caller's OAuth token, verified against MESHARC_OAUTH_ISSUER and then used to make the call, so one process serves many workspaces. MESHARC_API_KEY is never read in this mode; if it is set, startup says it is being ignored.
Behind a reverse proxy, note that binding 127.0.0.1 makes the MCP SDK enable DNS-rebinding protection with a localhost-only Host allow-list. The server therefore passes its own, built from MESHARC_MCP_PUBLIC_URL — without which every proxied request would be refused. Add browser-based clients to the Origin allow-list with MESHARC_MCP_ALLOWED_ORIGINS (comma-separated).
Privacy and security
The client talks to one host — the API base URL, https://api.mesharc.dev unless MESHARC_API_URL or base_url= says otherwise — and to nothing else. The key travels only as a bearer header, only over HTTPS. Nothing is written to disk and no telemetry is sent. The client reads only MESHARC_API_KEY and MESHARC_API_URL; the MCP server in HTTP mode reads MESHARC_MCP_PUBLIC_URL, MESHARC_OAUTH_ISSUER, MESHARC_INTROSPECT_SECRET, MESHARC_MCP_WAIT and MESHARC_MCP_ALLOWED_ORIGINS, and no key of its own.
What MeshArc keeps about you and about the pages you crawl, and for how long, is in the privacy policy. How the service is secured is on the security page. To report a vulnerability in this client or in the service, write to [email protected] rather than opening a public issue — see SECURITY.md.
Anything else
The client is a thin wrapper: every method is one API call and returns the API's JSON as a dict. The full reference is at mesharc.dev/docs/api. Call arc.close() when you are done, or use the client as a context manager.
- Documentation: mesharc.dev/docs
- Node client:
npm install mesharc— mesharc-node - Issues and pull requests: mesharc-python
- Questions: [email protected]
MIT.
来源:README.md,提交 68e8327
工具
0版本历史
1- v0.3.1最新Oct 6, 2026
