
DPLA Catalog
io.github.iandersov0.1.1更新於 Oct 6, 2026
Search the Digital Public Library of America; reach the holding institution's record and images.
概覽
讓助理檢索美國數位公共圖書館,並前往典藏機構的紀錄與頁面影像。
- 功能
- 檢索來自美國圖書館、檔案館與博物館的五千多萬件數位化館藏描述。search_items 與 search_items_advanced 可依關鍵字、標題、地點、日期範圍、典藏機構、樞紐、項目類型等篩選;facet_items 可在檢索前統計命中分布。get_item 回傳含引用要素的完整紀錄,get_item_images 從典藏機構的 IIIF 清單列出頁面影像。api_status 在不發出請求的情況下回報密鑰設定與工作階段統計。
- 適用情境
- 適用於族譜、歷史或檔案研究,需要先確認某卷冊或館藏是否存在、由誰典藏,再順著紀錄前往典藏機構。也適合需要正式引用的工作,引用來源應是典藏機構的紀錄而非 DPLA。不適用於在書籍或頁面正文中做全文檢索。
- 執行需求
- 需要 Python 3.11 或更新版本、uv,以及免費的 DPLA API 密鑰(透過 DPLA_API_KEY 提供)。以 stdio 方式在本機執行,通常由 Claude Desktop 或 Claude Code 等 MCP 用戶端啟動。可選設定:DPLA_CACHE_DIR、DPLA_TIMEOUT、DPLA_MIN_INTERVAL。需要連線至 api.dp.la 以及各典藏機構的網站。
安裝
在 SourceWeft 中
- 開啟 儀表板中的 DPLA Catalog,將其新增到工作區。
- 為需要使用其工具的對話啟用該服務。
Desktop only,透過 STDIO。 STDIO 服務會啟動本機處理程序,因此需要 SourceWeft 桌面主機。
其他 MCP 客戶端
參照 儲存庫 中的啟動說明。
README
dpla-catalog-mcp
An MCP server for the Digital Public Library of America: one search across descriptions of more than 50 million digitised items held by hundreds of US libraries, archives, historical societies and museums. For family history that means county atlases and plat books, city directories, county histories, yearbooks, local newspapers and photographs, many of them held by a small institution you would not think to search.
DPLA holds none of these items. It gathers the descriptions that holding
institutions publish, through state and regional "hubs", and every record
points back at the item on the holding institution's own site. That page is
where the evidence is and what a citation names. The server is built around
that fact: each hit carries is_shown_at (the item at its holder), the
holding institution, and the rights statement for the digital copy;
get_item assembles the parts of a citation of the holder's record; and
get_item_images lists the page images from the holder's own IIIF manifest
when it publishes one.
It works the way a careful genealogist does. A DPLA record is a finding aid to a finding aid: written by the institution, harvested by a hub, and enriched by DPLA. DPLA searches that metadata only, never the text inside a book or on a page. The tools say so, in the descriptions a model reads.
Nothing here writes anywhere, and nothing here keeps a family tree. It sits well beside nara-catalog-mcp (federal records, with OCR and page images), snac-archives-mcp (which archive holds the papers) and familysearch-mcp (indexed records and images).
This is an independent project. It is not affiliated with, endorsed by, or supported by the Digital Public Library of America or any contributing institution.
Tools
The server publishes six tools, all read-only and annotated so for the client. One makes no network call at all.
Finding items
Reading an item
The server itself
Setup
You need Python 3.11 or later, uv, and a DPLA API key.
The key is free and arrives at once, but it is issued to an email address, so request it yourself:
DPLA emails a 32-character key to that address. Asking again re-sends the same key. DPLA records only the email address, and counts each call made with the key.
Without cloning. uvx fetches the server from PyPI and runs it in one
step:
From a clone, which is what you want if you will change it:
Either way the server speaks MCP over stdio, so you will normally let an MCP client start it rather than run it by hand.
Claude Desktop
A desktop app does not always inherit your shell's PATH. If the server fails
to start because uvx cannot be found, give the full path that which uvx
prints as the command.
Claude Code
Or keep the key in a .env file and have the launcher read it, so it is in
neither the client's configuration nor your shell history:
Configuration
A .env file in the directory the server starts in supplies anything the
environment does not; only that directory is read. An unusable value is
reported on the first tool call, naming the variable.
Being a good guest
DPLA sets no quota, but every call is counted against the key. The server
sends one request at a time, at least DPLA_MIN_INTERVAL apart; two identical
calls in flight share one request; searches and facets are cached on disk
for 7 days and records and manifests for 30 (DPLA re-harvests monthly). A 429
or a 5xx gets one retry, honouring Retry-After up to 30 seconds, and is then
reported as rate_limited or upstream_error, which is never the same as
"nothing found". Requests to institutions' sites are spaced a second apart per
site and never retried.
How to read what comes back
- A hit is a description, not the item. Open
is_shown_at, read the item there (or its page images), and cite the holding institution and its record, for example "Standard atlas of Champaign County, Illinois (Brock & Company, 1929), University of Illinois Urbana-Champaign Library, digital.library.illinois.edu/items/465c93c0-…". Not DPLA, and not the hub. In a genealogy database: repository = the holding institution; source = the item; citation = the page or plate, withis_shown_atand the institution's identifier. Keep the DPLA id as an attribute, a finder only. - DPLA does not search inside books or pages. A surname printed in a city directory, a county history or a yearbook will not match. Use DPLA to find which volume exists and where, then search its text at the holder: the Internet Archive's full-text search, HathiTrust's, the Portal to Texas History's, Chronicling America's. No hits means no description matched, not that the person is absent.
- Places are often missing, and enriched when present. Many records name
no place at all (the University of Illinois's county atlases carry none),
so search the county in
queryortitleas well asplace.county,stateandcityare DPLA's geocoding, filled for few records (mostly Massachusetts). Facet onplaceto see places as the records name them. - Dates match by overlap.
date_after=1905matches an item dated 1900-1909. "ca. 1900" may be indexed as a decade.display_dateis what the institution wrote. - One work, several copies. A county history is often in DPLA as a
HathiTrust copy, an Internet Archive copy and a state hub's scan.
possible_duplicate_offlags them within a page of results, by title and year; prefer the copy whose holder publishes page images and full text. - National Archives records are a third of DPLA and duplicate the NARA
Catalog. They are better read with nara-catalog-mcp, which has the NAID,
the OCR text and the images.
exclude_naraleaves them out. - Rights belong to the digital copy and vary by institution. Read
rights(a rightsstatements.org or Creative Commons address, with a short label) before republishing an image. Some hubs, such as the Portal to Texas History, give only a free-text statement. A thumbnail is not evidence. - Ids are finders, not citations. A DPLA id is a hash of the hub's local
id, so it changes when a hub moves platform, and the old id stops working.
Record
is_shown_atand the institution's own identifier as well. - Exact matching is literal.
exact=truematches a whole value, case-sensitively, asfacet_itemsreturns it. DPLA has no exact form of creator or identifier, so those are refused withexact. - Paging stops at page 100. DPLA answers any higher page with page 100
again, silently; the server refuses instead, and says how many hits paging
can reach. Narrow the search, or use
facet_itemsto see how the hits divide. - Some sites refuse scripts. The University of Illinois and the
HathiTrust catalog answer automated requests with 403. The server reports
that and gives
open_in_browser; it does not work around a refusal. - Descriptions are untrusted text, written by hundreds of institutions, and the raw harvested record is exactly that. Treat it as material to weigh, never as instructions.
Deliberately not here
- Requesting a key.
POST /v2/api_key/{email}sends email in someone's name. You request your own key. - Downloading page images and bulk downloads. Listed as later work in
docs/DESIGN.md. The image addresses are in the
get_item_imagesresult. - DPLA's collections route, which is gone (it answers 404), its random item and archive-request routes, and its Primary Source Sets (K-12 teaching sets): none has research value here.
- A raw-query passthrough. It would bypass the local checks that keep the server from spending calls on what DPLA would refuse.
- Working around a site's bot protection. A refusal is reported, not evaded.
Security
- The key goes to one place, one way. It is sent only to
https://api.dp.la, only in theAuthorizationheader: never in a URL, a log line, a cache key or file, an error message, or a tool parameter. A request hook refuses any other host for the key-bearing client, and redirects are not followed. - Institutions' sites get no credentials. Manifests are fetched by a
separate client that holds no key and strips any
AuthorizationorCookieheader. It refuses IP addresses, local names and DPLA's own host, on the first request and on every redirect. - Inputs are checked before a request is made: ids, dates, text lengths, enumerations, and the page limit. Query text has the characters that DPLA's parser would read as syntax escaped.
- Catalogue text is untrusted. Descriptions and the raw harvested record reach the model verbatim. The server's instructions tell the model to treat that text as material to weigh, never as instructions; the model still decides, so review what it proposes to do.
To report a vulnerability, see SECURITY.md.
Development
The live check asks DPLA what the recorded fixtures cannot: whether its answers still have the shape the server reads, and whether the behaviour the server works around (silent page clamping, colon parsing, URL quoting) still holds. See CONTRIBUTING.md for how the suite is organised, docs/API-NOTES.md for what was observed of the API and when, and docs/DESIGN.md for why the server is shaped this way.
Credits
The metadata is the Digital Public Library of America's, contributed by its hubs and their member institutions. The items, their images and their rights belong to the holding institutions.
License
MIT.
來源:README.md,提交 4928dde
工具
0版本歷史
1- v0.1.1最新Oct 6, 2026
