cern-inspire-mcp-server

io.github.cyanheadsv0.1.1更新于 Oct 1, 2026

Search INSPIRE-HEP papers, authors, experiments, HEPData records; get citation metrics and BibTeX.

概览

AI 生成的概览

让助手检索 INSPIRE-HEP 高能物理论文、作者、实验与 HEPData 记录,并计算引用指标或导出 BibTeX。

功能
八个工具查询公开的 INSPIRE-HEP API:用 INSPIRE 语法或自由文本检索文献,按 recid、arXiv ID 或 DOI 获取论文完整记录,查询作者档案、实验与协作组,以及检索 HEPData 测量记录。还可计算某位作者或某个查询的 h 指数与引用分布,导出 BibTeX 或 LaTeX 引用条目,并解释查询语法与标识符形式。结果带有可串联的标识符,如 recid 和现成的 literatureQuery。
适用场景
适合需要高能物理文献、作者或协作组资料、引用数与 h 指数,或某批论文 BibTeX 条目的场景。可用于文献综述、引用分析,以及定位论文数值表格对应的 HEPData 记录。
运行要求
以本地 stdio 进程运行,通过 npm 包 @cyanheads/cern-inspire-mcp-server 使用,需要 Node.js v24+ 或 Bun v1.4.0+,也可用 Docker;另支持 HTTP 传输。无需 API 密钥或账号。可选环境变量涉及传输方式、HTTP 主机、端口与端点路径、日志级别和认证模式(none、jwt、oauth)。需要访问 INSPIRE-HEP 的网络连接。
安装前请注意
对公开数据只读,不涉及写入、付款或凭据。INSPIRE 限制每个地址每 5 秒 15 次请求,且同一进程内所有调用者共用一个队列,因此在共享或托管部署中,单个客户端的突发请求可能延迟或拒绝其他调用。每个查询只能访问前 10,000 条结果,且格式错误的 INSPIRE 语法会扩大或清空结果而不会报错。不返回 HEPData 表格数值。

安装

在 SourceWeft 中

  1. 打开 控制台中的 cern-inspire-mcp-server,将其添加到工作区。
  2. 为需要使用其工具的对话启用该服务。

Desktop only,通过 STDIO。 STDIO 服务会启动本地进程,因此需要 SourceWeft 桌面宿主。

其他 MCP 客户端

参照 仓库 中的启动说明。

README

@cyanheads/cern-inspire-mcp-server

Search INSPIRE-HEP papers, authors, experiments, HEPData records; get citation metrics and BibTeX via MCP. STDIO or Streamable HTTP.

8 Tools • 1 Resource


Overview

High-energy-physics literature from INSPIRE-HEP, including its index of HEPData measurement records. Search papers, authors, and experiments, read a paper's full record, compute citation summaries and h-indices, export BibTeX or LaTeX entries, and find the HEPData record that holds a paper's numerical tables. Runs as a stdio process or a local Streamable HTTP server.

Tools

ToolDescription
cern_inspire_search_literatureSearch papers with INSPIRE query syntax or free text, filtered by document type, subject, and year
cern_inspire_get_paperFetch one paper's full record by recid, arXiv ID, or DOI, with its HEPData availability
cern_inspire_export_citationsExport INSPIRE's BibTeX or LaTeX \bibitem entries for the papers a query selects
cern_inspire_search_authorsFind physicist profiles by name, BAI, ORCID, INSPIRE ID, or author recid
cern_inspire_get_citation_summaryh-index, citation totals, and citation buckets for one author or any literature query
cern_inspire_search_experimentsFind experiments, collaborations, and facilities, with a query that selects their papers
cern_inspire_search_hepdataFind HEPData measurement records by process, observable, energy, or collaboration
cern_inspire_list_referenceDecode query syntax, identifier forms, filter values, citation buckets, and HEPData DOIs

Resources

ResourceDescription
inspire://literature/{recid}One literature record as the cern_inspire_get_paper dossier in JSON

The same record is reachable through cern_inspire_get_paper for clients that don't surface resources.

Capability reference

cern_inspire_search_literature tool

  • INSPIRE query syntax or free text, with sort (relevance, mostrecent, mostcited), document_types and subjects (up to 4 values each, all of which must hold), and year_from / year_to; size 1–100 (default 10), paged by page
  • Only the first 10,000 results of a query are reachable: page × size beyond that fails as beyond_result_window, and a reversed year range as invalid_year_range
  • Hits carry recid, title, first author, date, citation counts, arXiv ID, DOI, publication, and a 300-character abstract snippet; totalCount, nextPage, and appliedFilters come back with the page, and a notice flags any query matching over 100,000 records

cern_inspire_get_paper tool

  • paper takes a recid, arXiv ID, DOI, inspirehep.net literature URL, or HEPData ins<recid> / hepdata.net record URL; resolvedAs names the form that matched, and a miss fails as paper_not_found
  • max_authors 0–500 (default 25) caps the author list; authorCount always gives the full number
  • hepdata.status is available, none, or lookup_failed, with recordDoi, latestVersion, tableCount, and hepdataUrl when available; citingQuery and referencesQuery feed cern_inspire_search_literature

cern_inspire_export_citations tool

  • Any literature query (recid:451647 or arxiv:1207.7214 for named papers); format is bibtex (default), latex-eu, or latex-us; size 1–50 (default 10)
  • Entries arrive verbatim from INSPIRE, each with its texkey; truncated is set when more papers matched than size

cern_inspire_search_authors tool

  • A name or one identifier (BAI, ORCID, INSPIRE ID, author recid); limit 1–25 (default 5)
  • matchedAs reports the route: orcid, inspire_id, bai, and recid match exactly, while name runs a free-text search whose ranked candidates are returned for the caller to choose from
  • Profiles carry recid, bai, ORCID, positions, advisors, arXiv categories, awards, and a literatureQuery selecting the person's papers

cern_inspire_get_citation_summary tool

  • Exactly one of author (BAI, ORCID, INSPIRE ID, or author recid) or query (any literature query); otherwise missing_target, and a name passed as author fails as author_not_identifier
  • document_types, subjects, and year_from / year_to narrow every figure; exclude_self_citations recounts without self-citations
  • h-index, citation totals and averages, and paper counts in the buckets 0, 1–9, 10–49, 50–99, 100–249, 250–499, 500+, each for all citeable and for published papers

cern_inspire_search_experiments tool

  • An experiment, collaboration, accelerator, or facility name, an INSPIRE legacy name (CERN-LHC-CMS), or an experiment recid (digits only); limit 1–25 (default 5)
  • Records carry the accelerator, host institutions, collaboration, lifecycle dates, ongoing (omitted when INSPIRE records neither state), INSPIRE's paper count, and a literatureQuery for cern_inspire_search_literature or cern_inspire_get_citation_summary

cern_inspire_search_hepdata tool

  • Free text or INSPIRE syntax over HEPData submissions (collaborations.value:LHCb, literature.control_number:<recid>); sort is relevance or mostrecent; size 1–50 (default 10), within the same 10,000-result window
  • Records carry paperRecids, collaborations, keywords (reactions, observables, centre-of-mass energies), recordDoi, latestVersion, tableCount, and hepdataUrl; table values are not returned

cern_inspire_list_reference tool

  • topic: search_syntax, identifiers, document_types, subjects, citation_buckets, or hepdata
  • Static term / meaning / example entries with no upstream call

inspire://literature/{recid} resource

  • The cern_inspire_get_paper dossier for one recid as application/json, listing the first 25 authors
  • Takes a recid only; use the tool for arXiv IDs, DOIs, or a higher author cap

Features

Built on @cyanheads/mcp-ts-core: stdio and Streamable HTTP transports, pluggable auth (none / jwt / oauth), swappable storage (in-memory, filesystem, Supabase, Cloudflare KV/R2/D1), structured logging with optional OpenTelemetry tracing.

INSPIRE-specific:

  • One process-wide pacer under INSPIRE's published 15 requests per 5 s: 12 request starts per 5 s, at most 4 in flight, and a shared cooldown after a 429 that starts at 5 s and doubles on each consecutive 429, up to 30 s
  • One 55 s budget per tool call covers queue wait, up to 2 retries, and every request the call makes; a request that can't start in time fails at once as pacer_shed with a retryAfter
  • JSON requests select only the fields a tool returns, under an 8 MiB response ceiling and a strict query-parameter allowlist, since INSPIRE silently ignores parameters it doesn't know; author email addresses are never requested
  • Forgiving inputs: paper identifiers accept an arXiv: or doi: prefix, a version suffix, an arxiv.org, doi.org, or inspirehep.net URL, and HEPData's ins<recid>; author identifiers accept an orcid.org URL; document_types and subjects take an array or a comma-joined string in any case

Agent-friendly output:

  • Chainable identifiers: hits carry recid, author profiles and experiments carry a ready literatureQuery, and cern_inspire_get_paper returns citingQuery and referencesQuery, so the next call needs no query building
  • Query echo and paging state: totalCount, truncated / shown / cap, nextPage, appliedFilters, and effectiveQuery, plus a notice with next-step text on empty, capped, or suspiciously broad results
  • Discriminated fields: hepdata.status, resolvedAs, matchedAs, and target.kind let callers branch on data, and typed failure reasons (paper_not_found, author_not_found, beyond_result_window, inspire_rate_limited) each carry a recovery hint
  • No fabrication: a field INSPIRE leaves out stays absent and prints as "Not available" or "not recorded"; upstream strings are escaped in the text output and kept verbatim in structuredContent

Data and licensing

  • INSPIRE-HEP metadata is mostly CC0 under INSPIRE's terms of use; credit INSPIRE when you reuse it.
  • HEPData records are CC0; cite the HEPData record DOI (recordDoi) when you reuse the data.
  • INSPIRE allows 15 requests per 5 seconds per address, and the server paces its own requests under that limit.
  • This is an independent project, not affiliated with or endorsed by INSPIRE-HEP, HEPData, or CERN.

Known limitations

  • No HEPData table values. hepdata.net's bot challenge refuses the server's User-Agent, so tools that read hepdata.net directly are deferred. cern_inspire_get_paper and cern_inspire_search_hepdata return the record DOI and the hepdata.net page where the values are read.
  • Malformed INSPIRE syntax doesn't fail. An unparsed operator widens or empties the match instead; zero hits or a very large totalCount usually means a syntax slip (cern_inspire_list_reference topic search_syntax).
  • 10,000-result window. Only the first 10,000 results of a query are reachable; narrow the query to reach the rest.
  • One request queue per process, one rate limit per address. Every caller of a server process shares one queue under INSPIRE's 15 requests per 5 s, so on a shared deployment one client's burst can delay or shed everyone else's calls with a retryable rate-limit error. A hosted deployment needs a per-client rate limit in front of /mcp, and should run one replica per egress IP, since INSPIRE counts requests per address; a per-caller share inside the server waits on the framework (cyanheads/mcp-ts-core#618).

Getting started

Add the following to your MCP client configuration file. No API key is needed.

json
{  "mcpServers": {    "cern-inspire-mcp-server": {      "type": "stdio",      "command": "bunx",      "args": ["@cyanheads/cern-inspire-mcp-server@latest"],      "env": {        "MCP_TRANSPORT_TYPE": "stdio",        "MCP_LOG_LEVEL": "info"      }    }  }}

Or with npx (no Bun required):

json
{  "mcpServers": {    "cern-inspire-mcp-server": {      "type": "stdio",      "command": "npx",      "args": ["-y", "@cyanheads/cern-inspire-mcp-server@latest"],      "env": {        "MCP_TRANSPORT_TYPE": "stdio",        "MCP_LOG_LEVEL": "info"      }    }  }}

Or with Docker:

json
{  "mcpServers": {    "cern-inspire-mcp-server": {      "type": "stdio",      "command": "docker",      "args": ["run", "-i", "--rm", "-e", "MCP_TRANSPORT_TYPE=stdio", "ghcr.io/cyanheads/cern-inspire-mcp-server:latest"]    }  }}

For Streamable HTTP, set the transport and start the server:

sh
MCP_TRANSPORT_TYPE=http MCP_HTTP_PORT=3010 bun run start:http# Server listens at http://localhost:3010/mcp

Prerequisites

  • Bun v1.4.0 or higher (or Node.js v24+).
  • Nothing else: INSPIRE-HEP's API is public and keyless.

Installation

  1. Clone the repository:
sh
git clone https://github.com/cyanheads/cern-inspire-mcp-server.git
  1. Navigate into the directory:
sh
cd cern-inspire-mcp-server
  1. Install dependencies:
sh
bun install
  1. Configure environment (optional):
sh
cp .env.example .env# edit .env to change the transport, logging, or session settings

Configuration

The server has no settings of its own: INSPIRE needs no key, and the request pacing is fixed in code. These framework variables cover most deployments.

VariableDescriptionDefault
MCP_TRANSPORT_TYPETransport: stdio or http.stdio
MCP_HTTP_PORTHTTP server port.3010
MCP_SESSION_MODEHTTP session mode: stateless, stateful, or auto. .env.example and the Docker image set stateless.auto
MCP_AUTH_MODEAuthentication: none, jwt, or oauth.none
MCP_LOG_LEVELLog level (debug, info, notice, warning, error, etc.). The Docker image sets info.debug
LOGS_DIRDirectory for log files (Node.js only).<app-root>/logs
STORAGE_PROVIDER_TYPEStorage backend: in-memory, filesystem, supabase, cloudflare-kv/r2/d1.in-memory
OTEL_ENABLEDEnable OpenTelemetry.false

See .env.example for the full list of framework overrides.

Running the server

Local development

  • Build and run the production version:

    sh
    # One-time buildbun run rebuild
    # Run the built serverbun run start:http# orbun run start:stdio
  • Run checks and tests:

    sh
    bun run devcheck  # Lints, formats, type-checks, and morebun run test      # Runs the test suite

Project structure

DirectoryPurpose
src/index.tscreateApp() entry point: registers the tools and resource, sets the server instructions, starts and disposes the INSPIRE service.
src/mcp-server/toolsTool definitions (*.tool.ts), eight tools, plus shared input schemas (inputs.ts).
src/mcp-server/resourcesResource definitions. The literature record resource.
src/services/inspireINSPIRE service: request pacer, per-call budget, retries, identifier routing, normalization, and the controlled vocabularies.
src/services/httpBounded fetch: per-attempt timeout and response byte ceiling.
src/utilsEscaping for upstream text in tool output and error messages (render.ts).
tests/Vitest suites for the tools, resource, services, and shared helpers, with INSPIRE response fixtures.
docs/design.mdTool surface, verified INSPIRE behavior, design decisions, and the deferred HEPData-direct tools.

Development guide

See CLAUDE.md for development guidelines and architectural rules. The short version:

  • Handlers throw, framework catches — no try/catch in tool logic
  • Every INSPIRE request goes through InspireService, with the call opened by beginCall(ctx); handlers never fetch directly
  • Register new tools and resources in the barrels at src/mcp-server/tools/definitions/index.ts and src/mcp-server/resources/definitions/index.ts
  • Wrap external API calls: validate raw → normalize to domain type → return output schema; never fabricate missing fields

Contributing

Issues are welcome. Run checks and tests before submitting:

sh
bun run devcheckbun run test

License

This project is licensed under the Apache 2.0 License. See the LICENSE file for details.

来源:README.md,提交 9740298

工具

0
工具元数据尚未被收录。

版本历史

1
  1. v0.1.1最新Oct 1, 2026