
CONTENTdm State Archives
io.github.iandersov0.1.0更新于 Oct 6, 2026
Search US state archives' record images on CONTENTdm; read items, pages and transcripts.
概览
搜索美国各州档案馆的 CONTENTdm 记录图像,并读取条目、页面、转录文本和页面图像。
- 功能
- 提供八个工具:list_instances、list_collections、get_collection、search、get_item、get_pages、get_image 和 cache_status。可同时搜索一个或多个精选州档案馆站点,读取条目的元数据、转录或 OCR 文本及引用信息,列出复合对象的页面顺序,并通过站点的 IIIF 服务下载页面图像。结果带有收藏机构的条目地址和引用核心信息。
- 适用场景
- 适合在美国州档案馆和州立图书馆做家谱或历史研究,例如死亡证明、养老金档案、选民和监狱登记册、县法院文书。当你需要跨多个机构搜索、把转录文本当作线索,再打开作为实际证据的页面图像时,它很合适。
- 运行要求
- 需要 Python 3.11 或更高版本以及 uv;通过 uvx contentdm-mcp 以 stdio 运行,通常由 MCP 客户端启动。无需 API 密钥。可选环境变量:CONTENTDM_CACHE_DIR、CONTENTDM_TIMEOUT、CONTENTDM_MIN_INTERVAL、CONTENTDM_CONTACT、CONTENTDM_EXTRA_INSTANCES 和 CONTENTDM_DOWNLOAD_DIR。需要能访问这些站点的网络。
安装
在 SourceWeft 中
- 打开 控制台中的 CONTENTdm State Archives,将其添加到工作区。
- 为需要使用其工具的对话启用该服务。
Desktop only,通过 STDIO。 STDIO 服务会启动本地进程,因此需要 SourceWeft 桌面宿主。
其他 MCP 客户端
参照 仓库 中的启动说明。
README
contentdm-mcp
An MCP server for the record images that US state archives and state libraries publish on CONTENTdm: death certificates and their annual indexes, Confederate pension files, voter registers, prison registers, county court and estate papers, letters and diaries, with whatever transcript or OCR text the institution added.
CONTENTdm is OCLC's hosted digital-collections platform. The Alabama Department of Archives and History, the Georgia Archives' Virtual Vault, the Tennessee State Library and Archives, Ohio Memory, the Illinois Digital Archives, the Indiana State Library, Missouri Digital Heritage and many more run on it, and every one of them answers the same keyless JSON API and IIIF image service. The server ships a curated list of those sites, searches one or several of them at once, reads an item and its pages, and downloads a page image to read.
It works the way a careful genealogist does. The page image is the evidence. A transcript or OCR text is a lead that says which page to read, and an index card is a finding aid to the record behind it. Every result carries the holding institution's item address and a citation core (institution, collection, title, identifier, address), because the item page is what gets cited, not this server.
Nothing here writes anywhere, and nothing here keeps a family tree. It sits
well beside dpla-catalog-mcp,
which finds an item across hundreds of collections and hands you the
institution's address (get_item takes that address directly), and
nara-catalog-mcp for federal
records.
This is an independent project. It is not affiliated with, endorsed by, or supported by OCLC or any of the institutions whose sites it reads.
Tools
The server publishes eight tools. All but get_image are read-only;
get_image writes one new file and never overwrites one.
Finding
Reading
The curated sites
list_instances gives the whole list. On 2026-10-06, 18 sites answered, and 7
entries record institutions that have left CONTENTdm, so a search can say
where they went:
Any other CONTENTdm site works too, at its cdmNNNNN.contentdm.oclc.org
address: pass https://cdmNNNNN.contentdm.oclc.org as instance. Every
classic site answers there, whatever its own domain, and the source of its
pages names the number ("cdmServerUrl": "serverNNNNN.contentdm.oclc.org").
There are more than 675 of them, and no registry. To use a site by its own
domain, add it to CONTENTDM_EXTRA_INSTANCES; a tool argument cannot, for the
reason under Security.
Setup
You need Python 3.11 or later and uv. There is no key to request.
Without cloning. uvx fetches it from PyPI and runs it in one step:
From a clone, which is what you want if you will change it:
Either way the server speaks MCP over stdio, so you will normally let an MCP client start it rather than run it by hand.
Claude Desktop
A desktop app does not always inherit your shell's PATH. If the server fails
to start because uvx cannot be found, give the full path that which uvx
prints as the command.
Claude Code
Configuration
Nothing is required. A .env file in the directory the server starts in
supplies anything the environment does not; only that directory is read.
An unusable value is reported on the first tool call as a not_configured
result naming the variable.
Being a good guest
Each site is one institution's server. The client sends one request at a time
to each site, at least a second apart; a search across several sites runs
them side by side, six at a time, so no site sees more than one caller. Two
identical calls in flight share one request. Collection lists and field
definitions are cached for 30 days, items for 7, searches for a day. A 429, a
5xx or a dropped connection gets one retry, honouring Retry-After. The
User-Agent names the package, its version and this repository.
robots.txt on these sites disallows the website's search pages, not the
API this server uses. Each institution's own terms govern what you do with
its images: list_instances and get_item pass on what the sites say, and
several ask for permission before publication.
How to read what comes back
- A search is only as wide as
answered.failed,timed_outandnot_searchedlist the sites a negative does not cover. A site that redirects, asks for a bot check or lists no collections has usually moved. - Search covers metadata and text fields, never the image. Many record images have no text at all, so a name can be on a page that no search finds. Browse by collection, or by the index volume, instead.
- A compound object's text is on its pages. Volumes, files and newspaper
issues are compound objects; their own transcript field is usually empty.
Use
page_hits=true, orget_pageswith aquery, to find the page. - Field names are per collection. "Name of Father" is
nameain one collection andfatherin another.get_collectionsays which nick holds what, and which field is the transcript: that is marked by type, because the name differs from site to site. - "Transcript" can be machine OCR. The Tennessee vital-records indexes carry OCR under that label. Read the image.
exactis a phrase search (the words together, in order), not whole-field equality.allneeds every word,anyone of them.- A PDF item can hold many pages. Its IIIF image shows the first;
get_imagewithpdf_pagereaches the others. pageis a position in the object, counted from 1, not the number printed on the scan: the entry for Sgt. Alvin C. York is on page 461 of volume 8 of Tennessee's 1960-1964 death index, a scan named54784_460and stamped 457. Cite the printed number too when it differs.- Cite the item page. Use the
citationeach result carries, and the institution's owncite_asline whereget_itemfinds one (Georgia, Oklahoma and Alaska items have them). Record the identifier (alias:pointer) so the item can be found again.
Deliberately not here
- Writing to any site. CONTENTdm's write functions need a staff login.
- Sites on other platforms. The new CONTENTdm (launched 2026-09-22) has no published public API, and Quartex's is undocumented. The tools speak to an adapter layer, so either can be added without changing them; see docs/DESIGN.md.
- Working around bot checks. A site that answers with a challenge is reported as blocked and left alone.
Security
Tool arguments are written by a model, and the model reads text this server does not control: titles, transcripts, DPLA records, web pages. The server assumes that text can steer the model, and limits what a steered model can make it do.
- Which hosts. A tool reaches the curated sites, any address under
contentdm.oclc.org(OCLC's own servers), and the sites you list inCONTENTDM_EXTRA_INSTANCES. Any other host is refused before it is even looked up, because a DNS query forsecret.attacker.examplewould already deliver the name; the refusal says how to reach the site instead. - Which addresses. Each connection is checked where it is made: a name that leads to a private, loopback, link-local, CGNAT, multicast, reserved or unspecified address, IPv4 or IPv6, is refused, and the connection goes to the address that was checked. Redirects are followed only between an instance's own hosts, and are checked again. Proxy settings in the environment are not used.
- How much. A JSON answer over 10 MB, or an image over 60 MB, is refused as it streams in.
- Which files.
get_imagecreates one new file and never overwrites one. The bytes must be an image, judged by their first bytes rather than the Content-Type, and of the format the file's suffix names:.jpgor.jpeg. Never a hidden file or folder, never under~/Library, and withCONTENTDM_DOWNLOAD_DIRset, never outside it, all judged after links are resolved. A refused download leaves nothing on disk. - Arguments are validated (collection aliases, field nicks, numeric pointers) before they reach a request, and search words are stripped of the API's own syntax characters.
- Site text is untrusted. Titles, descriptions and transcripts reach the model verbatim. The server's instructions tell the model to treat that text as material to weigh, never as instructions; the model still decides, so review what it proposes to do.
To report a vulnerability, see SECURITY.md.
Development
The live check asks the sites what the recorded fixtures cannot: whether their answers still have the shape the server reads, and whether every curated site still answers. See CONTRIBUTING.md for how the suite is organised, docs/API-NOTES.md for what was observed of the API and when, and docs/DESIGN.md for why the server is shaped this way.
Credits
The images and descriptions belong to the institutions that publish them. CONTENTdm is a product of OCLC.
License
MIT.
来源:README.md,提交 00c62f4
工具
0版本历史
1- v0.1.0最新Oct 6, 2026


