Polish Archives Indexes

io.github.iandersov0.1.1更新於 Oct 11, 2026

Search Geneteka, the Poznań Project and BaSIA for Polish vital records; fetch archive scans.

概覽

AI 產生的概覽

搜尋波蘭家譜索引(Geneteka、Poznań Project、BaSIA)中的民事登記紀錄,並從 Szukaj w Archiwach 下載檔案掃描件。

功能
提供七個工具:regions 與 geneteka_coverage 用來查看某教區已索引的登記簿和年份;geneteka_search、poznan_search 和 basia_search 依姓名、地點和年份查找洗禮、婚姻和死亡條目;swa_scan 憑 64 字元名稱從 Szukaj w Archiwach 圖片主機下載一張全尺寸掃描件;cache_status 顯示本次工作階段的請求、快取和限速。搜尋結果只是查找線索,指向登記簿和條目,並附保管機構、備註和掃描連結。除 swa_scan 外均為唯讀,swa_scan 只寫入一個新 JPEG 檔案且從不覆寫。
適用情境
適用於波蘭家族史研究:在相信零結果前先核對索引涵蓋範圍,在三個志工索引中依姓氏或教區檢索,並取得對應的檔案掃描件。適合把索引列當作線索、把掃描件當作證據的嚴謹流程。
執行需求
透過 stdio 執行的本機程序;需要 Python 3.11 或更新版本和 uv,用 uvx polish-archives-mcp 執行。無需帳號或 API 金鑰。選用環境變數:POLISH_ARCHIVES_CACHE_DIR、POLISH_ARCHIVES_CONTACT、POLISH_ARCHIVES_DOWNLOAD_DIR、POLISH_ARCHIVES_MIN_INTERVAL、POLISH_ARCHIVES_TIMEOUT。需要存取四個固定主機的網路。僅限桌面端。
安裝前請注意
swa_scan 會寫入新的 .jpg 檔案到磁碟;設定 POLISH_ARCHIVES_DOWNLOAD_DIR 後只會儲存到該資料夾內。請求會送往志工協會和檔案館的伺服器,並強制限速(至少 4 秒,BaSIA 和 Geneteka 為 10 秒)。索引備註和評論會原樣進入模型,屬於不可信內容。Szukaj w Archiwach 的單位頁面有機器人檢查,伺服器從不抓取,需人工在瀏覽器中開啟。發布掃描件前請查看檔案館的使用條款。

安裝

在 SourceWeft 中

  1. 開啟 儀表板中的 Polish Archives Indexes,將其新增到工作區。
  2. 為需要使用其工具的對話啟用該服務。

Desktop only,透過 STDIO。 STDIO 服務會啟動本機處理程序,因此需要 SourceWeft 桌面主機。

其他 MCP 客戶端

參照 儲存庫 中的啟動說明。

README

polish-archives-mcp

[CI] [PyPI]

An MCP server for the Polish volunteer indexes of vital records, and for the archive scans they point to:

  • Geneteka, the Polskie Towarzystwo Genealogiczne's index of parish and civil registers: tens of millions of baptisms, marriages and burials from thousands of parishes across Poland, and some beyond. Its coverage list says, parish by parish, which registers and years have been indexed.
  • The Poznań Project: marriages of 1800-1899 in Wielkopolska and Kujawy, the former Prussian Provinz Posen and its neighbours, with ages and parents.
  • BaSIA, the Wielkopolskie Towarzystwo Genealogiczne "Gniazdo"'s archival index, whose entries name the archive, the call number and the scan.
  • Szukaj w Archiwach, the Polish state archives' portal: this server downloads a full-size scan from its open photo host, given the scan's name.

It works the way a careful genealogist does. An index row is a finding aid, typed by a volunteer from the register: it says which register and entry to read, and the scan is the evidence. A zero covers only what was indexed, so Geneteka's coverage list comes first. And a scan is cited by the archive's own call number, the Sygnatura on the Szukaj w Archiwach unit page, never by an index's rendering of it.

Nothing here writes anywhere, and nothing here keeps a family tree. It grew from a script used in one family-history project, and sits well beside familysearch-mcp, whose catalogue holds the filmed copies of many of the same registers.

This is an independent project. It is not affiliated with, endorsed by, or supported by the Polskie Towarzystwo Genealogiczne, the Poznań Project, the Wielkopolskie Towarzystwo Genealogiczne "Gniazdo", the Naczelna Dyrekcja Archiwów Państwowych, or any archive.

Tools

The server publishes seven tools. All but swa_scan are read-only; swa_scan writes one new file and never overwrites one.

Before searching

ToolPurpose
regionsGeneteka's region codes and names (15wp is Wielkopolska). Makes no request.
geneteka_coverageThe registers Geneteka has indexed for a parish: births, marriages or deaths, the years exactly as Geneteka lists them (gaps included), and each register's rid. With year, whether each register covers it. The list is cached for a week.

Searching the indexes

ToolPurpose
geneteka_searchBirths, marriages or deaths in one region, or in one register by rid. Rows carry the custodian holding the book, the indexer's notes (often a date or a village) and any scan link. Fifty rows a page.
poznan_searchMarriages 1800-1899 by groom's and bride's surnames, or either spouse's. Given names are matched to the project's name groups; a name in no group returns the groups. Each row says who holds the original register.
basia_searchBaSIA entries by surname, with more people in the same entry, a place and radius, years, office and record type. Every row carries BaSIA's call number and scan, the unit's address on szukajwarchiwach.gov.pl, and a warning about BaSIA's series numbers.

Reading the scan

ToolPurpose
swa_scanDownload one full-size scan (about 3500 px) from Szukaj w Archiwach's photo host, given its 64-character name. Saves a new .jpg, never a hidden file or one under ~/Library.
cache_statusThis session's requests, by site, the cache, and the pacing. Makes no request.

The workflow they serve

  1. geneteka_coverage for the parish: which registers and years exist in the index. A year not listed was never searched.
  2. geneteka_search, poznan_search and basia_search. Records are in Latin, Polish or German: search Wojciech as Adalbertus and Adalbert too, Jan as Joannes and Johann, Marianna as Maria; -ski and -ska, and genitive forms ("z Kowalskich"), are one surname.
  3. Open the unit page (a BaSIA row's unit_url, or a Geneteka row's scan link) in a browser. Szukaj w Archiwach's pages sit behind a bot check and this server never fetches them. Read the Sygnatura there, and each scan's name from its image address (photos.szukajwarchiwach.gov.pl/<name>).
  4. swa_scan with that name. Read the scan before saying anything about it.
  5. Cite the archive, the Sygnatura and the scan number, with the unit page's address. Keep the index's id (Geneteka, Poznań Project or BaSIA) as the finder.

Setup

You need Python 3.11 or later and uv. There is no key to request and no account to make.

Without cloning. uvx fetches it from PyPI and runs it in one step:

bash
uvx polish-archives-mcp

From a clone, which is what you want if you will change it:

bash
git clone https://github.com/ianderso/polish-archives-mcpcd polish-archives-mcpuv syncuv run polish-archives-mcp   # stdio server, usually launched by the client

Either way the server speaks MCP over stdio, so you will normally let an MCP client start it rather than run it by hand.

Claude Desktop

json
{  "mcpServers": {    "polish-archives": {      "command": "uvx",      "args": ["polish-archives-mcp"]    }  }}

A desktop app does not always inherit your shell's PATH. If the server fails to start because uvx cannot be found, give the full path that which uvx prints as the command.

Claude Code

bash
claude mcp add polish-archives -- uvx polish-archives-mcp

Configuration

Nothing is required. A .env file in the directory the server starts in supplies anything the environment does not; only that directory is read.

VariableMeaning
POLISH_ARCHIVES_CACHE_DIRResponse cache directory. Default ~/.cache/polish-archives-mcp.
POLISH_ARCHIVES_TIMEOUTHTTP timeout in seconds for one request. Default 60. Scan downloads get 180 to read.
POLISH_ARCHIVES_MIN_INTERVALLeast seconds between two requests to one site. Default 4, and never below 4. BaSIA always gets at least 10.
POLISH_ARCHIVES_CONTACTAn email address or URL added to the User-Agent, so a site can reach you if your use causes trouble. Optional, and courteous.
POLISH_ARCHIVES_DOWNLOAD_DIRAn existing folder. When set, swa_scan saves only inside it. Set it to save into an iCloud Drive folder, which lives under ~/Library.

An unusable value is reported on the first tool call as a not_configured result naming the variable.

Being a good guest

Each site is a volunteer society's server or an archive's. The client sends one request at a time to each site, at least four seconds apart, and ten to BaSIA and Geneteka; different sites do not wait for one another. A Geneteka search fetches one page of fifty rows per call, never the whole result in a burst. Two identical calls in flight share one request. The coverage list is cached for a week, searches for a day, the Poznań Project's name groups for a month; a failure is never cached. A 429, a 5xx or a dropped connection gets one retry, honouring Retry-After. The User-Agent names the package, its version and this repository.

What each source says about automated use, as checked on 2026-10-11:

  • Geneteka. No terms of use are linked from geneteka.genealodzy.pl. Its robots.txt disallows nothing and sets Crawl-delay: 120. This server makes one request per tool call there, at least ten seconds apart.
  • The Poznań Project. Its about page says the database is "dostępna dla wszystkich użytkowników sieci poprzez bezpłatną wyszukiwarkę" (open to every internet user through a free search engine). It serves no robots.txt.
  • BaSIA. No terms of use were found on www.basia.famula.pl. Its robots.txt sets Crawl-delay: 10, disallows /search.php and other scripts, and says in a comment that robots may index information pages, "ale nie wyszukiwarkę ani skrypty pobierające dane" (but not the search engine or data-fetching scripts), because each search is a costly database query. This server posts the site's own search form at /en/, one search per tool call, at least ten seconds apart, and caches each answer for a day.
  • Szukaj w Archiwach. The portal's unit and fonds pages sit behind Imperva's bot check, so this server never requests them; a person opens them in a browser. The photo host serves no robots.txt (it answers 400). The portal's own terms could not be read without a browser; the archives set the terms for reusing their scans, so check them before publishing one.

If a site asks you to stop, stop: tell the server's operator, and open an issue so the project can change.

How to read what comes back

  • An index row is a finding aid. Every search answer says so. The row tells you which register, year and entry to read; the scan is the record.
  • Check coverage before trusting a zero. geneteka_search with a rid and a span of years names the years Geneteka has not indexed (years_not_indexed). The Poznań Project and BaSIA publish no comparable list: their zeros say less.
  • "Too many" is not a zero. The Poznań Project shows no entries when a search matches too many; poznan_search returns too_many_results with the counts, so narrow it. BaSIA lists only the first 250 entries of a search and says so (truncated).
  • Cite the archive's call number, not BaSIA's. BaSIA writes call numbers its own way, and its middle (series) number can differ from the archive's. Every BaSIA row carries that warning. Take the Sygnatura from the unit page. Some BaSIA rows still link the retired szukajwarchiwach.pl domain, which no longer opens; those rows have no unit_url, only the reference written in the old link, and a note on finding the unit.
  • Who holds the book. Geneteka rows carry the custodian Geneteka names (an archdiocesan archive, a state archive, a parish). Poznań Project rows say who holds the original register, from the project's own archive pages.
  • Names come in three languages. A Polish family can appear as Wojciech, Adalbertus and Adalbert in three records. The Poznań Project's name groups bridge this for given names; for the others, search each form.
  • Comments are other researchers' words. Poznań Project comments are returned without their authors' names or addresses.

Deliberately not here

  • Szukaj w Archiwach's own pages. They need a browser session past a bot check; a server driving a browser to get past it would be impersonating a person. The unit page stays with the person; the scan host, which is open, is what swa_scan uses.
  • Other indexes (metryki.genealodzy.pl's scans, Kartenmeister, the Pomeranian Greif index). Proposals are welcome as issues.
  • Writing to any site, including corrections and comments.
  • Working around bot checks. A site that answers with a challenge is reported as blocked and left alone.

Security

Tool arguments are written by a model, and the model reads text this server does not control: index notes, comments, web pages. The server assumes that text can steer the model, and limits what a steered model can make it do.

  • Which hosts. Four, fixed in the code: geneteka.genealodzy.pl, poznan-project.psnc.pl, www.basia.famula.pl and photos.szukajwarchiwach.gov.pl. No argument names a host: arguments only fill in a search, and a scan's name must be 64 hexadecimal characters. A request hook refuses anything else, including an address taken from a response. Redirects are followed only within the same host.
  • How much. A page over 10 MB, or a scan over 60 MB, is refused as it streams in.
  • Which files. swa_scan creates one new file and never overwrites one. The bytes must be a JPEG, judged by their first bytes, and not the photo host's "file unavailable" picture. The file must end in .jpg or .jpeg; never a hidden file or folder, never under ~/Library, and with POLISH_ARCHIVES_DOWNLOAD_DIR set, never outside it, all judged after links are resolved. A refused download leaves nothing on disk.
  • Site text is untrusted. Names, notes and comments reach the model verbatim. The server's instructions tell the model to treat that text as material to weigh, never as instructions; the model still decides, so review what it proposes to do.

To report a vulnerability, see SECURITY.md.

Development

bash
uv sync --extra devuv run pytest                      # mocked with respx; never touches a siteuv run ruff check .uv run ruff format --check .uv run python -m tests.live_check  # paced calls to the live sites

The live check asks the sites what the recorded fixtures cannot: whether their answers still have the shape the server reads. It takes a few minutes, because it keeps to the same pacing. See CONTRIBUTING.md for how the suite is organised, docs/API-NOTES.md for what was observed of each site and when, and docs/DESIGN.md for why the server is shaped this way.

Credits

The indexes are the work of thousands of volunteers of the Polskie Towarzystwo Genealogiczne, the Poznań Project and the Wielkopolskie Towarzystwo Genealogiczne "Gniazdo". The registers and their scans belong to the archives that hold them.

License

MIT.

來源:README.md,提交 c6a4dd0

工具

0
工具後設資料尚未被收錄。

版本歷史

1
  1. v0.1.1最新Oct 11, 2026