AMIRA — Africa Multiple Research Data
io.github.AM-Digital-Research-Environmentv1.20.0更新于 Oct 7, 2026
Read-only access to AMIRA, the Africa Multiple research-data platform on Omeka S.
概览
以只读方式访问非洲多元研究数据平台 AMIRA,让助手检索项目、研究条目、人物、出版物、播客和视频。
- 功能
- 通过只读工具访问基于 Omeka S 的 AMIRA 研究数据平台:集合概览,对约 4,000 条研究条目、项目、研究板块、人物、机构、研究组、集合、主题、地点和年份的关键词与分面检索。还覆盖集群文献目录及其可检索的开放获取全文、期刊、播客节目和带可检索转录文本的 YouTube 视频,并支持 BibTeX、RIS、CSL-JSON 引文导出、实体解析、有界的一跳实体图谱和文本段落检索。五个提示词封装了文献综述、项目档案等多步工作流。
- 适用场景
- 当助手需要回答有关非洲多元集群所藏非洲研究数据的问题时适用:按主题、地点、国家、语言或年份查找条目,追踪人物与项目,引用出版物或转录文本中的段落,或生成引文与表格导出。
- 运行要求
- 可作为远程 Streamable HTTP 端点运行,也可作为本地进程运行;打包的 .mcpb 可直接安装到 Claude Desktop,无需额外配置。未声明账号、API 密钥或请求头。自托管需要 Node.js 或 Docker,且服务器默认会抓取公开的 AMIRA 站点,因此除非使用内置快照,否则需要访问该站点的网络连接。
安装
在 SourceWeft 中
- 打开 控制台中的 AMIRA — Africa Multiple Research Data,将其添加到工作区。
- 为需要使用其工具的对话启用该服务。
Web executable,通过 Streamable HTTP。 远程服务在工作区中配置后即可从网页运行时运行。
其他 MCP 客户端
把它添加到你客户端的 mcpServers 配置中。
{
"mcpServers": {
"amira-mcp-server": {
"type": "http",
"url": "https://data.africamultiple.uni-bayreuth.de/mcp"
}
}
}README
AMIRA MCP Server
[CI] [Release] [Data refresh] [License: MIT] [MCP 2026-07-28]
Read-only Model Context Protocol server for the Africa Multiple Interactive Research Atlas (AMIRA), the research-data platform of the Africa Multiple Cluster of Excellence at the University of Bayreuth. AMIRA is published on the cluster's public Omeka S site at data.africamultiple.uni-bayreuth.de.
AMIRA is built and maintained by the Cluster's Digital Research Environment (DRE), its digital infrastructure unit. The DRE designs, builds, and maintains the data systems that connect researchers across the Africa Multiple Research Centres (AMRCs) and partner institutions worldwide. Curation and description are joint efforts with partners at the AMRCs in Université Joseph Ki-Zerbo, Rhodes University, the University of Lagos, and Moi University. Bayreuth hosts the coordinating DRE infrastructure and metadata layer. Federal University of Bahia is a privileged partner, not an AMRC.
Storage stays distributed by default: data remains in its local repository, while Bayreuth holds the metadata layer that points to it. Research data becomes findable without being relocated. High-quality metadata makes researchers' data more discoverable and their work more visible to a wider community of peers.
This server exposes AMIRA's projects, thematic research sections, ~4,000 digitised research items, people, institutions, groups, collections, the cluster bibliography with searchable full text (extracted from the open-access PDFs) and its journals, podcast episodes with transcripts, and the cluster's YouTube videos with searchable transcripts — as 33 core tools an LLM can query, five research-workflow prompts, and resources for citable records and bulk exports. From one MCP interface, clients can move across records and the places, languages, and subjects that connect them.
Every record carries an amira_url — its public page on the Omeka S site
(…/s/amira/item/<id>) — so findings can be cited as links back to the source.
AMIRA focuses on the Cluster's research data. For news, events, and general
information about the Africa Multiple Cluster of Excellence, visit
africamultiple.uni-bayreuth.de.
Coverage checked 6 October 2026: 3,975 research items (1,419 with digitised
media), 93 projects, 555 publications (60 with extracted full text), 87 journals,
43 podcast episodes, 140 videos and 3,085 subject authorities (646 Library of
Congress headings). Counts are a dated snapshot; use get_collection_overview for
what your running server actually holds. See the publication guide,
the changelog for what each release changed, and the
roadmap for further improvements.
How it gets its data — and why nothing else is needed
The server is self-contained and offline-first. It reads from two sources:
- A complete data snapshot bundled inside the
.mcpb— crawled from the public Omeka S REST API at build time and transformed into compact typed records, so the server works fully offline with zero setup. - (optional, on by default) the Omeka S API over HTTPS. At startup and daily, probes compare item and item-set modification signatures and totals. A changed signal triggers a complete crawl; a weekly forced crawl also catches vocabulary-only edits. A writer lock, immutable generations and an atomic pointer preserve the previous usable snapshot if refresh fails. The API origin and installation path partition the cache and must match the manifest before its records can be served under that site's citation URLs.
End users need no API key, no credentials, and no VPN — it's the same openly published data that powers the public site.
Tools
Call get_collection_overview first to scope the data, then drill in.
Prompts and resources
Five prompts package multi-step research workflows. Hosts show them as slash
commands; their arguments autocomplete from the snapshot's own authority names, and
each restates the citation rules. They are listed through prompts/list, so they
add nothing to the per-turn tool payload.
Resources expose data outside the tool calls:
Profiles and research workflows
AMIRA_TOOL_PROFILE=full is the default: 33 core tools, or 35 over HTTP.
research, discovery and visualization expose smaller documented subsets
(22/12/14 core tools respectively; HTTP adds search and fetch). Profiles are
fixed at startup; they reduce discovery cost, not access permissions. Tools
outside a profile are removed, and so are the apps that render them. See
src/toolProfiles.ts for the exact lists.
Typed ids work everywhere in either vocabulary (item: / research_item:,
pub: / publication:, section:, person:, organisation:, project:,
video:, podcast:), including the id of every get_* tool. Paged results
stop early, with response_limited: true and a correct next_offset, when a page
would exceed 40,000 characters — Claude Code moves larger tool results out of
the conversation. Error results carry the error as JSON text only.
Use resolve_entity before graph traversal; pass its typed id unchanged to
get_entity_graph. Graph edges count distinct source records, separate each
corpus, and distinguish catalogue links from co-occurrence. Neither co-occurrence
nor a shared subject establishes collaboration. Follow an edge with edge_id
and the returned snapshot_id for stable evidence paging. Graphs cap at 100 nodes,
200 edges and 42,000 JSON bytes, so check truncated.
get_text_passages takes 1–10 publication:ID (or pub:ID), video:ID or
podcast:ID identifiers. It returns at most 20 passages per page, original UTF-16
offsets and source citations; every match is reachable through next_offset
(scanned_matches_capped reports a document with more than 10,000 matches). These are extracted-text offsets, not
PDF page numbers or audio timestamps. get_snapshot_changes requires two local
same-source generations; it reports available history instead of inventing a diff.
Example questions it can answer
Companion skill
A research-workflow skill ships in .claude/skills/amira-mcp/:
cluster context, a query workflow, a tool-by-task map, the citation discipline, and coverage caveats.
Install it either way:
- Download
amira-mcp-skill.zipfrom the releases page and unzip it into your Claude skills directory (e.g.~/.claude/skills/) — it expands to~/.claude/skills/amira-mcp/. - Or copy the
.claude/skills/amira-mcp/folder from this repo there.
…or let the server hand it over (Skills over MCP)
Official extension; host support varies. SEP-2640 is now Final. The server follows the published Skills extension, including SHA-256 digests and byte sizes for every file. The zip above remains available for clients without extension support. The research tools and citation contract work either way.
The server also serves that same skill over the connection, following the
SEP-2640 extension
(io.modelcontextprotocol/skills). Nothing to download and nothing to keep in sync: a host that
implements the extension discovers the skill on connect, and the copy it gets is the one built into
the running server rather than a zip that may predate the tool surface it describes. It is the only
route that reaches the remote HTTP surface — ChatGPT, Claude.ai connectors and the APIs have no
local skills directory.
Cost when unused: zero. Skill text is never injected into the server instructions or into any
tool description, so the tools/list payload — the part re-sent every turn — is byte-identical
whether or not a host supports the extension. Disclosure timing stays a host decision; the three
reference files load only when something actually reads them.
Set AMIRA_SKILLS=0 to drop the capability and the three methods entirely.
Install (end users)
Download amira-mcp-server.mcpb from the
releases page
and double-click it. Claude Desktop shows an install dialog; click Install.
No further configuration is required. The latest tagged release carries the
current data; the rolling data-latest pre-release always tracks the freshest
site snapshot.
Use it from ChatGPT, the API, or any remote client
The .mcpb is the local, offline option for Claude Desktop. The same server can
also run as a remote Streamable HTTP endpoint — one HTTPS URL that ChatGPT,
Claude (web + desktop remote connectors), the OpenAI and Anthropic APIs, Cursor,
VS Code and other clients connect to by pasting a URL (no download, always-fresh
data). The remote surface serves the same 33 tools plus the
OpenAI-compatible search / fetch tools for ChatGPT research integrations (35
total). Access is unauthenticated — the data is public and read-only.
search takes plain keywords (matched term-by-term, not as an exact phrase),
with optional limit and types to keep the result set tight, and returns
ranked hits across research items, the bibliography (reaching into publication
full text), podcasts, videos, projects and research sections. The url on
search results and fetch documents is
always the AMIRA/Omeka public record page; DOI, YouTube/watch, and podcast/listen
URLs are kept as secondary metadata/text links. fetch returns one record's full
text by the id search hands back — for videos and podcasts the transcript is omitted by default
(metadata + description only, since a full one can run to tens of thousands of
characters), and include_transcript=true pulls it in — paged with
transcript_offset / transcript_max_chars (the same names get_video /
get_podcast use). Publication full text works the same way
(include_fulltext=true, paged with fulltext_offset / fulltext_max_chars,
matching get_publication), and max_chars caps the whole text body. The
appended window is sized against what max_chars leaves after the metadata
header, so *_returned_chars is exactly what landed in text and the next page
starts at offset + returned_chars with no gap.
Endpoints: POST /mcp (the MCP endpoint) and GET /healthz. Bind with PORT /
HOST.
-
ChatGPT → enable Developer mode under Settings → Security and login, then add your server URL from ChatGPT Plugins. Research integrations use
search+fetch; see the current OpenAI setup guide. -
Claude (web or desktop) → Settings → Connectors → Add custom connector →
https://<your-host>/mcp. -
OpenAI API (Responses) — point the
mcptool at the endpoint:
Deploy (self-hosted, e.g. alongside the amira site)
The multi-stage Dockerfile crawls a fresh snapshot at build time
and runs the self-contained bundle (no node_modules at runtime). On a Linux
host you can equally run it under systemd behind a reverse proxy:
proxy_buffering off keeps the Streamable-HTTP/SSE responses flowing. The refresh uses AMIRA_SITE_BASE (the public HTTPS API by default), even when
co-located with Omeka. It performs a full crawl when changes are detected;
co-location does not eliminate API load. The server caps request bodies at 64 KiB
and rate-limit state at 10,000 clients. Configure connection, request-rate and
concurrency limits at the reverse proxy for public deployments; the in-process
limiter is a courtesy control. Shutdown cancels refresh work and gives HTTP
connections a ten-second drain window.
For a build that consumes a previously reviewed data/ snapshot without calling
Omeka during the build:
The default SNAPSHOT_STAGE=fetch still crawls fresh data. Both stages validate the
snapshot. The image pins its base digest; npm ci uses the committed lockfile.
Develop / rebuild
Token budgets
Context is the scarcest resource a client has, so both halves of what this server
spends are measured and gated (scripts/weigh.mjs, estimating tokens as
bytes/4 — comparable to itself, not to a billing statement):
No tool description is derived from the snapshot, so the surface is a pure
function of the code: test/unit/budget.test.mjs gates it offline against no
data at all, and asserts that the empty-store measurement still matches the
baseline recorded against the full snapshot. Response weight needs real data, so
npm run weigh -- --check runs in the smoke job instead. Drift there is only a
warning — a routine npm run fetch-data must not red the build — while the
absolute cap (Claude Code truncates tool results at 25,000 tokens) is fatal, as
is any probe whose call failed, since a structured refusal is ~60 tokens and
would otherwise sail under every ceiling.
Version 1.18 deliberately raises the full-profile ceiling to 14,000 estimated
tokens to accommodate six tools and useful output schemas. Smaller profile gates
remain 10,000 (research), 6,500 (discovery) and 9,500 (visualization).
The committed baseline records the exact surface size,
text bytes, serialized wire bytes (text and structured content both count) and the
snapshot it was measured on. Every response must also stay under 45,000 characters:
Claude Code stores any tool result over 50,000 characters in a file instead of the
conversation. Paged results and the graph are byte-bounded to fit.
When a change legitimately grows either number:
Commit the diff. That is the point of the file: the cost of a new tool or a richer result summary lands in code review as a number instead of arriving unnoticed.
Pack the extension:
scripts/census.mjs re-runs the property census behind the field mapping
(scripts/census-report.json) — rerun and diff it if the instance's templates
change. The design decisions behind it are in ROADMAP.md; CHANGELOG.md
records each release.
Continuous delivery
Three GitHub Actions cover pull requests and distribution (the two distribution workflows crawl the public API — no credentials):
-
CI (
.github/workflows/ci.yml) — on pull requests and pushes tomain: type-checks once, then runs the offline unit suite (including the tool-surface token budget) on Node.js 22, 24 and 26 on Linux and Node.js 24 on Windows, then exercises both MCP transports, weighs every tool's response at its maximum limit, and validates the MCPB manifest. The complete dependency audit fails on high or critical advisories. A separate container job builds an offline fixture image and checks health. Windows CI also packs a fixture extension and checks its file list. A non-blocking job runs the official MCP conformance suite against the HTTP server. Actions are pinned to commit SHAs; every job has a timeout. -
Release (
.github/workflows/release.yml) — on a pushedv*tag: checks the tag againstpackage.json, crawls a fresh snapshot, runs unit, live and smoke tests, packs the.mcpband zips the companion skill (amira-mcp-skill.zip) in a read-only job; a separate publish job attaches both to the GitHub Release with build-provenance attestations. A non-blocking job publishesserver.jsonto the MCP Registry. -
Refresh data snapshot (
.github/workflows/refresh-data.yml) — weekly and on demand; rebuilds only when the transformed snapshot content hash changes, including item-set and vocabulary changes, updating the rollingdata-latestpre-release.
Publishing from Actions requires a writable token. If the organization disables write permissions for the default workflow token, either (a) enable Read and write permissions under Organization → Settings → Actions → General → Workflow permissions, or (b) add a repo/org secret
RELEASE_TOKEN(a PAT or GitHub App token withcontents: write). The workflows usesecrets.RELEASE_TOKENwhen present and fall back to the default token.
Configuration (environment / extension settings)
Snapshot layout and recovery
The cache lives under AMIRA_CACHE_DIR/<API identity hash>/. active.json names
the current immutable generations/<uuid>/ directory and two previous generations.
Unreferenced generations have a 24-hour cleanup grace period. Existing flat
snapshots and a same-source legacy cache/current remain readable; the first new
publication migrates a flat destination into generation history. npm run fetch-data
also writes the generation layout to data/; scripts must use loadSnapshot
rather than assuming data/manifest.json is active.
A crashed writer can leave writer.lock. Stop all processes using that cache,
confirm no writer is running, then remove that one lock file and restart. The
server deliberately never guesses whether an old lock is safe to steal. Failed
refreshes leave the last usable data available. /healthz, overview and data
quality report refresh attempt/success times, in-flight state and failure class.
Retries apply only to transient upstream failures, honor bounded Retry-After,
and abort after a 15-minute refresh deadline or shutdown. Probes cannot provide a
transactional Omeka snapshot; periodic complete crawls remain necessary.
Metadata-exposure levels (benchmark experiments)
AMIRA_EXPOSURE lets an evaluation run the same tasks under graded metadata
visibility — the "metadata mediation" condition of LLM-access studies. It is an
experiment flag, not an end-user setting; the default (full) is the normal
server. Existence flags (has_transcript / has_fulltext) stay visible at
every level; only content and relations are gated, and refusals are structured
errors (exposure_restricted / text_access_disabled) so a model can say why
it cannot answer rather than hallucinating.
Architecture
-
Pure in-memory typed records. The Omeka JSON-LD is transformed once at build time (
src/transform.ts, evidence inscripts/census-report.json); the runtime loads compact records and indexes them at startup. No native bindings; the esbuild bundles are self-contained. -
Offline-first with atomic refresh. Bundled snapshot + locked generation publication; the freshest manifest wins at startup (an old cache can never shadow a newer bundled snapshot).
-
Citations are uniform: every entity is an Omeka item, so every record carries
amira_url = <site>/s/amira/item/<o:id>. -
Accent-insensitive matching everywhere (
src/text.ts). The collection is francophone-Africa-heavy and its authority records are not consistently accented against the free text — "Côte d'Ivoire" is the subject heading while item titles carry "Cote d'Ivoire". Every keyword comparison folds both sides (NFD, drop combining marks, lowercase, and map curly quotes, dashes and the ligatures œ/æ/ß to plain forms), so the answer no longer depends on which spelling the caller guessed. Folds of large texts are memoised and dropped when a refresh replaces the snapshot. -
Interactive results (MCP Apps). Seven tools carry
_meta.ui.resourceUripointing at atext/html;profile=mcp-appresource, so hosts implementing theio.modelcontextprotocol/uiextension (Claude, Claude Desktop) render the result inline. Every other client ignores the_metaand gets exactly the same JSON.The modules live in
src/ui/:shell.tsholds the design tokens and shared bar-chart primitive;bridge.tsbundles the official Apps SDK. Each app supplies its own CSS and render function. The templates load nothing from the network — no scripts, styles, fonts or tiles — so they need no_meta.ui.cspgrants and run in the strictest sandbox. App controls call an allowlisted set of read-only tools through the host. Citation links and file downloads also go through host APIs. Every visual has a table or text alternative. See app development and provenance.Colours come from the DREVisualizations Omeka module's theme (
--primary/ the Africa Multiple brand palette), so a chart in the chat and the same chart on the AMIRA site read as one system, and were validated against those surfaces rather than eyeballed — light#007a50on#fdfcfa, dark#35a87don#1b211e(the theme's own#3fb488sits just outside the dark lightness band, so the app steps one down). Every chart is single-series, with the category on the axis label, so one accent is correct and no legend is needed. The co-occurrence hub would nominally want six hues for its six relation types, but no six-hue set from the brand palette survives all-pairs CVD separation (the best candidates bottomed out at ΔE 2.4 under deutan), so identity there is carried spatially — one labelled sector per relation type — which also survives greyscale, print and forced-colors.
Protocol posture
Verified against the current specification and TypeScript SDK documentation on 6 October 2026. The server speaks MCP 2026-07-28 on both transports, and still answers 2025-era clients unchanged.
It runs on the v2 TypeScript SDK — the scoped packages
(@modelcontextprotocol/server, /node, /client), not the frozen
@modelcontextprotocol/sdk v1 monolith. Two things are worth knowing if you
work on the entry points:
- The modern era is opt-in.
SUPPORTED_PROTOCOL_VERSIONSis the legacyinitializeladder and stops at 2025-11-25; the SDK keeps the 2026 string internal on purpose.createAmiraServernames the revision itself insupportedProtocolVersions, which is what registersserver/discover— without it that method answers-32601. - The era is owned by the entry, not the transport.
serveStdio(factory)andcreateMcpHandler(factory)decide the era per connection/request and pin an instance from the factory. Connecting a bareStdioServerTransportorNodeStreamableHTTPServerTransportby hand serves the 2025 era only.
Concretely on the wire: server/discover advertises ["2026-07-28"], results
carry resultType, and the cacheable results carry real ttlMs / cacheScope
(1 h public on tools/list, resources/list and server/discover; 24 h on
resources/read, whose ui:// app HTML is immutable per build) rather than the
SDK's conservative { ttlMs: 0, cacheScope: "private" } default. Back-compat is
the handler's legacy: 'stateless' default, which answers 2025-era traffic with
a fresh instance per request — the same shape this server always had.
The deprecations of that revision cost nothing here: Roots, Sampling, Logging
and Elicitation are all unused, logging goes to stderr, and tool order is
deterministic (registration order). CORS accepts both generations of headers
(Mcp-Session-Id/MCP-Protocol-Version and Mcp-Method/Mcp-Name/X-Mcp-Header).
Citing this software
If this server is part of how you found or analysed AMIRA material, please cite
it. The repository carries a CITATION.cff, so GitHub's Cite
this repository button will render BibTeX and APA for you.
Cite the data separately from the software: individual records are citable
by their amira_url (https://data.africamultiple.uni-bayreuth.de/s/amira/item/<id>),
which is what every tool result returns for exactly this reason.
Credits and funding
Built and maintained by the Digital Research Environment (DRE) of the Africa Multiple Cluster of Excellence, University of Bayreuth — the Cluster's digital infrastructure unit, which designs and runs the data systems behind AMIRA. Curation and description of the underlying records are joint efforts with partners at the Africa Multiple Research Centres (Université Joseph Ki-Zerbo, Rhodes University, University of Lagos, Moi University) and Federal University of Bahia.
Funded by the Deutsche Forschungsgemeinschaft (DFG, German Research Foundation) under Germany's Excellence Strategy — EXC 2052/1 — 390713894 (Africa Multiple: Reconfiguring African Studies).
License
MIT — see LICENSE. See also CONTRIBUTING.md and SECURITY.md.
来源:README.md,提交 4eeb06c
工具
0版本历史
1- v1.20.0最新Oct 7, 2026


