MeshArc

dev.mesharcv0.3.1更新於 Oct 6, 2026

Crawl and scrape websites to clean markdown, and track what changed between crawls.

概覽

AI 產生的概覽

讓助理把網站抓取、爬取成乾淨的 Markdown,對應站台地圖,並追蹤頁面隨時間的變化。

功能
MeshArc 提供 scrape_urls、extract_url、map_site、crawl_site、keep_crawl_as_project 等工具,以及專案與執行相關工具(list_projects、create_project、start_run、list_pages、get_page、get_changes、search_pages、recrawl_pages、get_job)。頁面內容可依需求回傳 Markdown、文字、HTML、連結、結構化欄位或截圖,並附上狀態、判定結果與錯誤碼。爬取結果可保留為排程專案,記錄新增、移除、修改的頁面以及逐字差異。結果較大時回傳的是索引與摘錄,而不是完整頁面集合。
適用情境
當助理需要把網頁或整個站台讀成乾淨文字、查看站台地圖宣告的內容,或定期監控站台變化時適用。適合文件、定價與內容監測類任務,尤其是需要可重複爬取與變更記錄的情境。
執行需求
可使用遠端端點 Python 3.10+ 上安裝 mesharc[mcp] 並以 stdio 在本機執行。每次呼叫都需要 MeshArc API 金鑰,可直接傳入或透過 MESHARC_API_KEY 環境變數提供;託管端點會以瀏覽器登入並簽發自己的金鑰。需要能連線到 MeshArc API 的網路。
安裝前請注意
API 金鑰屬於憑證:它以 bearer 標頭送往 MeshArc API,並帶有作用域(read、write、admin)、可選的專案範圍、有效期限與速率限制。抓取頁面被視為寫入操作,因為它會離開工作區並消耗額度,因此唯讀連線無法抓取、爬取或對應站台。爬取、執行與多 URL 抓取會消耗額度,託管模式下工具呼叫會在 MESHARC_MCP_WAIT 秒後回傳任務。抓取到的頁面資料會送往 MeshArc 服務。

安裝

在 SourceWeft 中

  1. 開啟 儀表板中的 MeshArc,將其新增到工作區。
  2. 為需要使用其工具的對話啟用該服務。

Web executable,透過 Streamable HTTP。 遠端服務在工作區中設定後即可從網頁執行環境執行。

其他 MCP 客戶端

把它新增到你客戶端的 mcpServers 設定中。

{
  "mcpServers": {
    "mesharc": {
      "type": "http",
      "url": "https://mcp.mesharc.dev/mcp"
    }
  }
}

README

mesharc

The Python client for the MeshArc API: a URL in, clean content out, and a record of what changed.

  • Scrape one page or a batch — markdown, text, HTML, links, structured fields, a screenshot.
  • Crawl a whole site with no project to set up first, and keep it as one if it turns out to be worth watching.
  • Map what a site declares in its sitemaps before fetching any of it.
  • Watch a site over time: projects, scheduled runs, and a change record — pages added, removed, modified, field by field.

Python 3.10 or newer. One dependency (httpx). Fully typed.

Install

bash
pip install mesharc

Authentication

Every call needs an API key. Create one in the app under Settings → API keys — it is shown once — and give it to the client, or put it in MESHARC_API_KEY and construct the client with nothing:

python
from mesharc import MeshArc
arc = MeshArc("mesharc_...")# or, with MESHARC_API_KEY in the environment:arc = MeshArc()# options: MeshArc(api_key, timeout=150.0, max_retries=2)

The key is only ever sent as a bearer header to api.mesharc.dev.

A key carries the scopes it was made with (read, write, admin), optionally a set of projects it may see, an expiry and a rate limit. A route the key may not use answers 403; a project it may not see answers 404.

Quick start

python
from mesharc import MeshArc
arc = MeshArc("mesharc_...")
page = arc.scrape("https://example.com/pricing")print(page["markdown"])print(page["verdict"], page["method"], page["credits"])   # ok crawler 1

scrape holds the request open until the page comes back (60 s by default), so there is nothing to poll for an ordinary page.

Reading pages

One page

python
page = arc.scrape("https://quotes.toscrape.com/js/", config={"render_js": "always"})

config is any setting a project takes, by its API name (render_js, only_main_content, formats, max_tier, wait_for_selector, actions, …). The full list, with defaults, is at mesharc.dev/docs/configuration.

python
page = arc.scrape_one(    "https://example.com/",    formats="markdown,text,cleanHtml",     # which bodies to return    timeout_s=120,                         # how long the API holds the request (120 max)    idempotency_key="pricing-2026-09-18",  # the same key returns the first answer for 24 h)

Many pages

A list of URLs is a batch: grouped by host, fetched in parallel where the config allows, and returned as one row per URL.

python
batch = arc.scrape(["https://a.com/", "https://b.com/x"], config={"concurrency": 4})for row in batch["pages"]:    print(row["url"], row["httpStatus"], row["verdict"], row["credits"])

arc.scrape(urls, wait=False) returns the batch id at once; arc.batch(id, wait=True) finishes it later. Pass webhook_url= to be told instead of polling (batch.finished, signed with a secret returned once).

What a page looks like

Every page row carries the same fields, whether it came from a scrape, a crawl or a project:

FieldMeaning
markdown, text, cleanHtml, html, links, fields, screenshotThe bodies you asked for
httpStatusThe status the site answered with
verdictok, thin (short, but a page), blocked (refused, or a 404), skipped
errorCodeOK, or what went wrong: BLOCKED, NOT_FOUND, TIER_LIMIT, CAPTCHA, LOGIN_REQUIRED, RATE_LIMITED, TIMEOUT …
shape, warnings, signalslisting / table / form for a short page whose markup says what it is; short; why the judge decided as it did
method, tier, climbedToThe engine that read it (crawler, tls, minted, browser, browser-residential, …), its tier, and how far a refused page climbed
credits, billedAsWhat it cost; the rung it is priced at when not the one that fetched it
words, language, head, reason, crawledAtSize, language, the head fields, the judge's sentence, when

A page the site refused costs 0, and so does a 404.

Crawling a site

python
job = arc.crawl(    "https://docs.example.com",    limit=200,                          # page budget    maxDepth=3,                         # link hops from the seed    includePaths=["/docs/*"],    crawlMode="sitemap_first",          # what the sitemap declares first, then links    maxTier="browser",                  # how far a refused page may climb    scrapeOptions={"formats": ["markdown", "links"]},    config={"crawl_delay_ms": 500},     # any project setting, directly)
for page in job.pages():                # follows the cursor while the crawl runs    print(page["url"], page["words"])print(job.status, job.envelope["counts"], job.envelope["creditsUsed"])

crawl returns a handle immediately; job.pages() yields pages as they land and ends when the crawl does. wait=True blocks until it finishes; job.wait(), job.refresh(), job.cancel() do what they say; arc.get_crawl(id) reattaches to a crawl started elsewhere.

A one-shot crawl expires after 30 days. If the site is worth watching:

python
project = job.keep(name="Docs", schedule="weekly")

A webhook= in the options ({"url", "events", "metadata"}) is told about crawl.started, crawl.page (fifty pages a message) and crawl.completed; its signing secret comes back once as job.webhook_secret.

Mapping a site

python
for u in arc.map("https://docs.example.com"):    print(u["url"], u["lastmod"])
details = arc.map_details("https://www.gov.uk/", search="visa", limit=500)print(details["totals"], details["creditsUsed"])   # {'files': 29, 'urls': 508431, …} 29

A map costs one credit per sitemap file read — most sites are one file.

Watching a site: projects and runs

python
project = arc.projects.create(    "https://docs.example.com",    name="Docs",    schedule="weekly",                                        # manual | hourly | daily | weekly    config={"max_pages": 300, "include_paths": ["/docs/*"]},)
run = arc.runs.start(project["id"], wait=True)                # the first run# ...a week later, or arc.runs.start again: the second run produces the change record
record = arc.changes(project["id"])print(record["change"]["counts"])   # {'added': …, 'removed': …, 'modified': …, 'withheld': …}
diff = arc.page_diff(project["id"], "https://docs.example.com/pricing")
MethodWhat it does
projects.list() · projects.get(id) · projects.update(id, name=, schedule=, retention=, config=) · projects.delete(id)The projects
runs.list(project_id) · runs.start(project_id, wait=) · runs.wait(project_id, run_id) · runs.get(project_id, run_id) · runs.cancel(project_id, run_id)Runs
pages(project_id, run_id=None) · page(project_id, url, run_id=None)The pages of a run; one page in full
changes(project_id, run_id=None) · page_diff(project_id, url, run_id=None)The change record; one page's word-level diff
search(project_id, q, mode="content" | "selector", run_id=None)Which pages say this (words, "phrases") or contain this (CSS / XPath)
recrawl(project_id, urls)Fetch these pages again, now
sources(project_id)The seed, sitemap, URL list, feeds and patterns with what the last run found through each
export(project_id, path, dataset="pages", fmt="jsonl", run_id=None, urls=None)Stream a dataset (pages, markdown, changes, fields, sitemap) as jsonl or csv to a file
python
arc.export(project["id"], "pages.csv", dataset="pages", fmt="csv")

The workspace

python
me = arc.me()               # the workspace, its plan and limits, credits used and remaining, what this key may dousage = arc.usage()         # pages per day, this month by enginemonitor = arc.monitor()     # what is queued and runningmeta = arc.meta()           # verdict meanings, engine costs, the config defaultskeys = arc.keys()key = arc.create_key("ci", scopes=["read", "write"], projects=[project["id"]], expires_in_days=90)   # key["key"], oncearc.revoke_key(key["id"])

Errors

Every failure raises MeshArcError:

python
from mesharc import MeshArc, MeshArcError
try:    arc.crawl("https://example.com", limit=1_000_000)except MeshArcError as exc:    print(exc.status, exc.code, exc.detail, exc.request_id)
codeStatusMeaning
validation400 / 422Something in the request is wrong; detail says what
unauthorized401No key, or a revoked or expired one
plan_limit402The plan does not include this, or the credits are spent
forbidden403The key's scopes do not allow it
not_found404No such thing — or not one this key may see
conflict409The request contradicts current state
rate_limited429Over the key's rate limit; X-RateLimit-Reset says when
internal500Quote request_id to support

request_id is the id the API put on the response and in its own logs, so a support conversation starts from one string.

Two more cases: a network failure or a request that hits timeout raises MeshArcError with status == 0 and code network or timeout; a job the client stopped waiting for raises MeshArcTimeoutError — both a MeshArcError and a TimeoutError — which carries job_id so you can poll it later (arc.get_crawl(id), arc.batch(id)).

Idempotency and timeouts

  • scrape, scrape_one and crawl take idempotency_key=: send the same key again within 24 hours and you get the first answer back rather than a second job.
  • Waiting calls take wait=, poll= (seconds between polls) and timeout= (seconds before TimeoutError). wait=False returns the envelope at once; the default polls every 3 s for up to an hour.
  • timeout_s on a single scrape is how long the API itself holds the request open (60 s by default, 120 at most); a slower page comes back as an id and is polled.
  • MeshArc(..., timeout=150.0) is the HTTP timeout per request. A request is retried on 429, 502, 503, 504 and network failures when it is safe to repeat — a GET, a DELETE, or a POST with an idempotency key — up to max_retries times (2), honouring Retry-After.

Credits

Every response says what it cost: credits on a page, creditsUsed on a job envelope, X-MeshArc-Credits on the HTTP response. A page costs the engine that read it — a plain fetch 1, a render 4 — and a refused page or a 404 costs nothing. The schedule and the plans are at mesharc.dev/docs/billing.

The MCP server

The package also ships MeshArc as an MCP server, so Claude Desktop, Claude Code, Cursor and any MCP client can scrape, crawl, map and read change records as tools. Python 3.10+.

There is a hosted one, so most people need install nothing:

bash
# Claude Codeclaude mcp add --transport http mesharc https://mcp.mesharc.dev/mcp

It opens a browser once to sign in to mesharc.dev and approve, then works. Approving issues the app an API key of its own, which you can see and revoke under Settings → API keys; read only unless you allow changes. Any client that takes a remote MCP URL — the Claude.ai and ChatGPT connectors, Claude mobile, Cursor — takes it the same way.

Or run it yourself, over stdio, with your own key:

bash
pip install "mesharc[mcp]"MESHARC_API_KEY=mesharc_... mesharc-mcp          # serves over stdio
# Claude Codeclaude mcp add mesharc -e MESHARC_API_KEY=mesharc_... -- mesharc-mcp

Tools: scrape_urls, extract_url, map_site, crawl_site, keep_crawl_as_project, list_projects, describe_project_config, get_project, create_project, update_project, start_run, list_pages, get_page, get_changes, search_pages, recrawl_pages, get_job. Every tool is a call through this client.

A result with many pages in it is a map, not the territory: a crawl answers with an index of the pages read (up to 500, fewer if the budget needs the room), an excerpt of as many as a 60,000-character budget pays for (fifty at most), and a count of each page's links rather than the links. Ask for the one page you want on its own with get_job(kind="crawl", id=…, url=…), or carry on through the crawl with cursor=. A multi-URL scrape past its budget counts the rest by status and names any that did not come back ok. One page asked for on its own -- extract_url, get_page, scrape_urls with a single url -- comes back whole, capped at 12,000 characters.

Hosted, a connection can be approved read-only, and read-only cuts in a place worth knowing: it reads the whole workspace, but it cannot reach the site. Anything that fetches is a write, because it leaves the workspace and usually spends its credits -- map_site is the one that fetches and costs nothing, and it is gated with the rest. list_projects, get_project, describe_project_config, list_pages, get_page, get_changes, search_pages and get_job work; the nine that fetch or write -- scrape_urls, extract_url, map_site, crawl_site, keep_crawl_as_project, create_project, update_project, start_run, recrawl_pages -- answer {"code": "read_only"} with what to do about it, until the app is reconnected with write access.

Hosted, a tool that would hold a connection open for minutes — a crawl, a run you asked to wait for, a multi-URL scrape — hands back a job after MESHARC_MCP_WAIT seconds (25 by default, because many MCP hosts time a tool call out sooner). The work carries on server-side and get_job picks it up. Run locally, those tools block as they always have.

Hosting it yourself

bash
MESHARC_MCP_PUBLIC_URL=https://mcp.example.dev/mcp MESHARC_OAUTH_ISSUER=https://api.example.dev MESHARC_INTROSPECT_SECRET=...                      mesharc-mcp --http --host 127.0.0.1 --port 8040

The server holds no key. Each request carries the caller's OAuth token, verified against MESHARC_OAUTH_ISSUER and then used to make the call, so one process serves many workspaces. MESHARC_API_KEY is never read in this mode; if it is set, startup says it is being ignored.

Behind a reverse proxy, note that binding 127.0.0.1 makes the MCP SDK enable DNS-rebinding protection with a localhost-only Host allow-list. The server therefore passes its own, built from MESHARC_MCP_PUBLIC_URL — without which every proxied request would be refused. Add browser-based clients to the Origin allow-list with MESHARC_MCP_ALLOWED_ORIGINS (comma-separated).

Privacy and security

The client talks to one host — the API base URL, https://api.mesharc.dev unless MESHARC_API_URL or base_url= says otherwise — and to nothing else. The key travels only as a bearer header, only over HTTPS. Nothing is written to disk and no telemetry is sent. The client reads only MESHARC_API_KEY and MESHARC_API_URL; the MCP server in HTTP mode reads MESHARC_MCP_PUBLIC_URL, MESHARC_OAUTH_ISSUER, MESHARC_INTROSPECT_SECRET, MESHARC_MCP_WAIT and MESHARC_MCP_ALLOWED_ORIGINS, and no key of its own.

What MeshArc keeps about you and about the pages you crawl, and for how long, is in the privacy policy. How the service is secured is on the security page. To report a vulnerability in this client or in the service, write to [email protected] rather than opening a public issue — see SECURITY.md.

Anything else

The client is a thin wrapper: every method is one API call and returns the API's JSON as a dict. The full reference is at mesharc.dev/docs/api. Call arc.close() when you are done, or use the client as a context manager.

MIT.

來源:README.md,提交 68e8327

工具

0
工具後設資料尚未被收錄。

版本歷史

1
  1. v0.3.1最新Oct 6, 2026