SNAC Archives

io.github.iandersov0.1.0更新於 Oct 6, 2026

Find which archive holds the papers: SNAC's index of archival collections, plus ArchiveGrid links.

概覽

AI 產生的概覽

讓助理檢索 SNAC 合作組織的索引,查出某個人、家族或機構的文件收藏在哪個檔案館,並產生 ArchiveGrid 檢索連結。

功能
此伺服器提供八個唯讀工具。三個用來尋找收藏:search_collections 依標題詞檢索,get_collection 回傳單筆完整記錄,repository_holdings 列出 SNAC 連結到某檔案館的收藏。三個用來尋找名稱:search_names 依個人、家族或團體機構檢索,get_name 回傳標目、日期、地點與識別碼連結,collections_in_common 找出同時連結兩個名稱的收藏。兩個不發出網路請求:archivegrid_search_link 產生供人開啟的 ArchiveGrid 檢索或記錄連結,cache_status 回報本次工作階段的即時呼叫與快取命中。
適用情境
適合檔案與族譜研究,也就是問題在於文件收藏於何處,而不是文件內容寫了什麼。可用於在數千個檔案館中定位某家族的書信、某教會的登記冊或某商號的帳冊,也可與聯邦檔案或已索引記錄的伺服器搭配使用。
執行需求
需要 Python 3.11 或更新版本以及 uv;以 uvx snac-archives-mcp 透過 stdio 執行,通常由 MCP 用戶端啟動。不需要金鑰或帳號。選用變數:SNAC_API_URL(必須為 https)、SNAC_CACHE_DIR、SNAC_TIMEOUT 與 SNAC_CONTACT。需要能連線至所設定的 API 主機。
安裝前請注意
唯讀:用戶端除六個讀取命令外拒絕送出任何命令,不會寫入 SNAC。目錄摘要與傳記說明會原樣傳給模型,屬於不可信文字。回傳的是檢索工具,不是文件內容;目錄會過時,寫信給檔案館前應確認目前索書號。不會抓取 ArchiveGrid,只產生連結供人開啟。SNAC_CONTACT 為選用且出於禮貌。

安裝

在 SourceWeft 中

  1. 開啟 儀表板中的 SNAC Archives,將其新增到工作區。
  2. 為需要使用其工具的對話啟用該服務。

Desktop only,透過 STDIO。 STDIO 服務會啟動本機處理程序,因此需要 SourceWeft 桌面主機。

其他 MCP 客戶端

參照 儲存庫 中的啟動說明。

README

snac-archives-mcp

[CI] [PyPI]

An MCP server for finding which archive holds the papers: a family's letters, a church's registers, a store's ledgers, a county office's loose papers. Most of that material has never been digitised. It exists online only as a catalogue record or a finding aid in some library, and the hard part of the research is knowing which one.

The server searches the SNAC Cooperative's index of people, families and organisations and the archival collections that hold their papers, across thousands of repositories. The data is CC0, and the API needs no key and no account. It also builds ArchiveGrid searches for you to open, because ArchiveGrid often finds what SNAC cannot, and OCLC does not permit automated access to it.

It works the way a careful genealogist does. Everything it returns is a finding aid: evidence of where to look and roughly what is there, never of what a document says. The tools say so, and they point you at the holding repository's own finding aid, which is the description to cite.

Nothing here writes anywhere, and nothing here keeps a family tree. The server finds collections; what you conclude from the papers belongs in your genealogy software, or in a family-tree MCP server running alongside this one. It sits well beside nara-catalog-mcp (federal records) and familysearch-mcp (indexed records and images).

This is an independent project. It is not affiliated with, endorsed by, or supported by the SNAC Cooperative, the University of Virginia, or OCLC.

Tools

The server publishes eight tools, all read-only and annotated so for the client. Two make no network call at all.

Finding collections

ToolPurpose
search_collectionsFind collections by words in their title. Each hit names its holding repository with a postal address, its OCLC number and its link, and flags the likely duplicates (the same papers catalogued two or three times).
get_collectionOne collection in full: the whole abstract, extent, repository, links, and the names SNAC connects to it.
repository_holdingsWhat SNAC links to one repository, filtered by title and paged.

Finding people, families and organisations

ToolPurpose
search_namesFind SNAC name records by name, optionally by kind: person, family, or corporate body (churches, businesses, societies, offices).
get_nameOne name record: headings, dates, places, biographical note and identifier links (VIAF, Library of Congress and others). With detail="full", every collection linked to it, with its role, and the related names.
collections_in_commonCollections linked to both of two names: where to look for letters between intermarried families or an ancestor's associates.

Where SNAC stops

ToolPurpose
archivegrid_search_linkBuild an ArchiveGrid search, in its fielded syntax, for you to open; or a record link, given an OCLC number. Makes no request.
cache_statusThis session's live calls and cache hits. Makes no request.

Setup

You need Python 3.11 or later and uv. There is no key to request.

Without cloning. uvx fetches it from PyPI and runs it in one step:

bash
uvx snac-archives-mcp

From a clone, which is what you want if you will change it:

bash
git clone https://github.com/ianderso/snac-archives-mcpcd snac-archives-mcpuv syncuv run snac-archives-mcp   # stdio server, usually launched by the client

Either way the server speaks MCP over stdio, so you will normally let an MCP client start it rather than run it by hand.

Claude Desktop

json
{  "mcpServers": {    "snac": {      "command": "uvx",      "args": ["snac-archives-mcp"]    }  }}

A desktop app does not always inherit your shell's PATH. If the server fails to start because uvx cannot be found, give the full path that which uvx prints as the command.

Claude Code

bash
claude mcp add snac -- uvx snac-archives-mcp

Configuration

Nothing is required. A .env file in the directory the server starts in supplies anything the environment does not; only that directory is read.

VariableMeaning
SNAC_API_URLThe SNAC REST endpoint. Default https://api.snaccooperative.org/. Must be https; its host is the only one the server will contact.
SNAC_CACHE_DIRResponse cache directory. Default ~/.cache/snac-archives-mcp.
SNAC_TIMEOUTHTTP timeout in seconds. Default 60. Holdings lists get 120.
SNAC_CONTACTAn email address or URL added to the User-Agent, so the SNAC team can reach you if your use causes trouble. Optional, and courteous.

An unusable value is reported on the first tool call as a not_configured result naming the variable.

Being a good guest

SNAC is a free service run by a small cooperative, and it publishes no rate limit. The server sends one request at a time, at least a second apart; two identical calls in flight share one request; and every answer is cached on disk for 30 days (a collection record for good). A 429 or a 5xx is retried three times with back-off, then reported as rate_limited or upstream_error, which is never the same as "nothing found". Pass refresh=true to ask again, and refresh a cached empty result before concluding anything is absent.

How to read what comes back

  • A finding aid is not the record. A collection description says the papers exist and roughly what is in them. Do not attach a citation to a fact on its strength; record a research task (request copies, plan a visit) instead.
  • Cite the repository's own finding aid, by the repository and its collection number, for example "Wilder and Anderson Family Papers #01255, Southern Historical Collection, Wilson Library, UNC-Chapel Hill". Not SNAC, and not ArchiveGrid: they are indexes to it. When link_kind is finding_aid, the link goes there; otherwise look the collection up in the repository's own catalogue.
  • Record the identifiers. The OCLC number ties a collection together across WorldCat, ArchiveGrid and SNAC; the SNAC ARK (ark:/99166/...) is the stable id of a name record. Record the repository's collection number too: WorldCat merges records, and a number can come to redirect to another.
  • Description depth varies enormously, from a one-line catalogue record ("Papers, 1881-1958", 4 boxes) to a folder-level container list. A terse description does not mean a person is absent from the papers.
  • Title search is title search. search_collections needs every word in the collection's title. "Davenport family papers" finds ten collections; "Davenport family papers Lincoln County" finds none, because the county is only in the abstract. Use search_names, or an ArchiveGrid link with place.
  • Duplicates are normal. One collection often appears as a WorldCat record and as one or two harvested finding aids, sometimes under different forms of the repository's name. possible_duplicate_of groups them by title words and years; it is a hint, not a merge.
  • A name record is not an identification. SNAC holds hundreds of records headed "Anderson family." with nothing to tell them apart but the collections they link to. In a name's collections, creatorOf means these are the name's own papers; referencedIn means the name is an index term on the collection, not that any document concerns the person. maybe_same_count above 0 means SNAC suspects a conflation.
  • The papers of slaveholding families are a primary route to enslaved ancestors. Many descriptions name enslaved people only as a category. Search the papers of the family that held them, not only the ancestor's name.
  • The catalogue ages. SNAC's collection index was largely built from 2010s extracts. Collections get reprocessed, renumbered and transferred, and survey "repositories" such as a state historical documents inventory record what a town clerk or church held when surveyed decades ago. Confirm the current call number before writing to a repository.
  • ArchiveGrid's indexes have edges. location is where the repository is; place is a place the papers are about. Its name, place and subject indexes cover catalogue records and EAD finding aids only; HTML and PDF finding aids match keywords alone. It leaves out records held by more than one library, so microfilm of county or church records is usually missing: use WorldCat or the FamilySearch Catalog for that.

Deliberately not here

  • Writing to SNAC. Its edit commands need an account and an API key. The client refuses to send any command but its six read ones, so no argument can make it edit anything.
  • Fetching ArchiveGrid. OCLC's terms of use forbid robots and automated copying, and the site blocks automated clients. The server builds addresses; a person opens them.
  • Fetching finding aids from repositories' own sites. That would widen the server from one API to thousands of hosts, with a parser for each and a much larger surface for injected text. It may come later, behind an option.

Security

  • One host. A request hook refuses any request not for the configured API host, so a value a model passes in cannot make the server fetch another site. SNAC_API_URL must be https.
  • Identifiers are validated (numeric ids as ASCII digits, ARKs against SNAC's pattern) before they reach a request.
  • Catalogue text is untrusted. Abstracts and biographical notes are written by cataloguers and contributors and reach the model verbatim. The server's instructions tell the model to treat that text as material to weigh, never as instructions; the model still decides, so review what it proposes to do.

To report a vulnerability, see SECURITY.md.

Development

bash
uv sync --extra devuv run pytest                      # mocked with respx; never touches SNACuv run ruff check .uv run ruff format --check .uv run python -m tests.live_check  # eight paced calls to the live API

The live check asks SNAC what the recorded fixtures cannot: whether its answers still have the shape the server reads. See CONTRIBUTING.md for how the suite is organised, docs/API-NOTES.md for what was observed of the API and when, and docs/DESIGN.md for why the server is shaped this way.

Credits

The data is the SNAC Cooperative's, released under CC0, with collection records contributed by its member institutions and drawn from WorldCat and finding aids. ArchiveGrid is a project of OCLC Research.

License

MIT.

來源:README.md,提交 246cc41

工具

0
工具後設資料尚未被收錄。

版本歷史

1
  1. v0.1.0最新Oct 6, 2026