Cern Opendata Mcp Server

io.github.cyanheadsv0.1.1更新于 Oct 1, 2026

Search CERN Open Data, fetch records, files, analysis environments, CMS good-run lists, HLT paths.

概览

AI 生成的概览

让助手搜索并读取 CERN 开放数据门户的记录、文件、分析环境、CMS 有效运行列表和触发路径。

功能
它通过七个只读工具访问 CERN 开放数据门户:用精确词表过滤器和实时分面计数搜索记录;按 recid、DOI、CMS 数据集路径或文档 slug 获取最多 20 条记录的完整元数据;分页浏览记录的文件索引,含 XRootD URI、HTTPS 链接、大小和校验和。它还能组装记录的分析环境(容器镜像、CMSSW 版本、global tag、指南)、返回带亮度区间的 CMS 有效运行列表、查询 CMS 高级触发路径,并解释门户词表。cern-opendata://record/{recid} 资源返回单条记录的元数据、许可和引用信息。
适用场景
适合处理粒子物理开放数据的场景:查找 ALICE、ATLAS、CMS 或 LHCb 的数据集,核对许可与引用,定位数据文件,或准备 CMS 分析环境和运行选择。适用于需要门户元数据而非本地数据处理的研究、教学和可复现分析工作流。
运行要求
以本地 stdio 进程(或本地 Streamable HTTP 服务器)运行,npm 包为 @cyanheads/cern-opendata-mcp-server,需要 Bun v1.4.0+ 或 Node.js v24+,也可用 Docker。无需 API 密钥或账号,门户是公开的。需要能访问 CERN 开放数据门户的网络。可选框架变量包括 MCP_TRANSPORT_TYPE、MCP_HTTP_PORT、MCP_HTTP_HOST、MCP_AUTH_MODE 和 MCP_LOG_LEVEL。
安装前请注意
只读:不会暂存磁带文件,也不写入任何内容。门户元数据和数据集为 CC0,但软件、容器镜像和文档按每条记录单独授权,CERN 要求复用者引用每个数据集的 DOI。门户限制每个客户端 IP 每分钟 60 次请求;服务器自我限速为 50 次,托管部署会让同一出口 IP 后的所有用户共享该额度。标记为 on demand 的文件存放在磁带上,须先在门户页面申请。该项目独立,与 CERN 无隶属关系。

安装

在 SourceWeft 中

  1. 打开 控制台中的 Cern Opendata Mcp Server,将其添加到工作区。
  2. 为需要使用其工具的对话启用该服务。

Desktop only,通过 STDIO。 STDIO 服务会启动本地进程,因此需要 SourceWeft 桌面宿主。

其他 MCP 客户端

参照 仓库 中的启动说明。

README

@cyanheads/cern-opendata-mcp-server

Search CERN Open Data, fetch records, files, analysis environments, CMS good-run lists, HLT paths via MCP. STDIO or Streamable HTTP.

7 Tools • 1 Resource


Overview

Particle-physics data from the CERN Open Data Portal: collision, simulated and derived datasets, analysis software, environments and documentation from ALICE, ATLAS, CMS, LHCb and other experiments. Search it with exact-vocabulary filters and live facet counts, open records with their license and citation, list the files that hold the data, assemble a record's analysis environment, and look up CMS good-run lists and trigger paths. Runs as a stdio process or a local Streamable HTTP server.

Tools

ToolDescription
cern_opendata_search_recordsSearch datasets, software, environments, documentation and supplementary records with exact-vocabulary filters and live facet counts
cern_opendata_get_recordsFetch full metadata for 1–20 records by recid, DOI, CMS dataset path or documentation slug, with license and citation
cern_opendata_list_filesPage through a record's file indexes and files: XRootD URIs, HTTPS URLs, sizes, checksums, tape availability
cern_opendata_get_analysis_envAssemble a record's analysis environment: container images, CMSSW release, global tag, linked environment and software records, guide sections
cern_opendata_get_validated_runsGet a CMS validated-run (good-run) list for a dataset, a list or a run period, with luminosity-section ranges
cern_opendata_search_trigger_pathsLook up CMS High-Level Trigger paths by name or prefix, parsed into run ranges, versions and L1 seeds
cern_opendata_list_referenceDecode the vocabulary the other tools accept: experiments, record types, energies, formats, identifiers, query syntax, licensing, run periods

Resources

ResourceDescription
cern-opendata://record/{recid}One record's metadata, license and citation, in the cern_opendata_get_records record shape

Tool-only clients get the same data from cern_opendata_get_records.

Capability reference

cern_opendata_search_records tool

  • Optional query (an OpenSearch query_string, up to 500 characters) plus OR-list filters type, experiment, collision_energy, collision_type, file_type, availability and collection, each an array or a comma-separated string; year_from/year_to and min_events/max_events bound the data-taking year and the event count
  • sort (bestmatch, mostrecent, title, title_desc), limit 1–50 (default 10) and page from 1; paging reaches the first 10,000 matches, and page × limit past that fails as page_window_exceeded
  • Compact hits with recids, plus eight live facets that each ignore their own filter; applied_filters echoes what ran, with values outside the verified vocabulary listed under unrecognized_values

cern_opendata_get_records tool

  • ids: 1–20 recids, DOIs, CMS dataset paths (/Primary/Era/TIER) or documentation slugs, mixed in one array or comma-separated string
  • Each record carries a license with its basis (record, cern_terms_default, not_stated) and, when it has a DOI, a ready citation; documentation and news bodies are cut at 30,000 characters
  • Identifiers that resolve to nothing land in missing with interpreted_as and guidance instead of failing the call; file lists come from cern_opendata_list_files

cern_opendata_list_files tool

  • recid required; without index, returns the record's file indexes and regular files, and with an index key, that index's files
  • limit 1–500 (default 50), continued with next_cursor; each file carries xrootd_uri, https_url, size_in_bytes, checksum and availability, and each index a uri_list_url listing every XRootD URI in it
  • Files marked on demand sit on tape and must be requested on the record's portal page first; an umbrella record with no files of its own returns its children recids

cern_opendata_get_analysis_env tool

  • recid required; software carries the record's own container images, CMSSW release, global tag and environment recid
  • environment_records (condition, VM, validation) for the record's run periods and example_software that declares it works with the record, up to 50 between them; guides quotes the linked section of the first two portal guides, each capped at 12,000 characters
  • Always separately_licensed: true; linked records or guides that can't be read leave a notice instead of failing the call

cern_opendata_get_validated_runs tool

  • Exactly one of recid (a CMS collision dataset or a validated-run list) or run_period (Run2012B; 2012B also matches); variant full or muons_only; run_min/run_max; limit 1–2000 (default 200)
  • A dataset recid bounds the runs to the first and last run the dataset lists, echoed in run_bounds; when several lists match, matched_lists names them and no runs are read
  • Each run carries lumi_sections and lumi_ranges, and list.https_url downloads the whole list file; CMS only, so other records fail as no_validated_runs

cern_opendata_search_trigger_paths tool

  • path: an exact name (HLT_IsoMu24) or a prefix with one trailing * (HLT_IsoMu*); HLT_ is added when missing and a _v<n> version suffix dropped; optional year, limit 1–50 (default 10) and page
  • Each per-year record is parsed into first_seen, last_seen, per-version run ranges with their l1_seed, and HLT menu record links; parsed: false marks a record to read from its abstract_html
  • CMS open data from 2010–2016; prescale tables are not published

cern_opendata_list_reference tool

  • Optional topic: experiments, record_types, collision_energies, collision_types, file_types, availability, identifiers, query_syntax, licensing or run_periods; omit it for every table
  • Static and offline, with no portal requests; run_periods is a dated snapshot, while cern_opendata_get_validated_runs reads the live list collection

cern-opendata://record/{recid} resource

  • One record by recid (leading zeros ignored) as application/json, in the cern_opendata_get_records record shape: metadata, license and citation, without file lists
  • recid comes from cern_opendata_search_records; reads carry a 15-minute public cache hint

Features

Built on @cyanheads/mcp-ts-core: stdio and Streamable HTTP transports, pluggable auth (none / jwt / oauth), swappable storage (in-memory, filesystem, Supabase, Cloudflare KV/R2/D1), structured logging with optional OpenTelemetry tracing.

CERN Open Data-specific:

  • Keyless, read-only client for the portal's record, documentation and file routes; it never stages tape files or writes anything
  • One shared pacer at 50 requests a minute, under the portal's published 60 per client IP, and one 50-second deadline per call across queue wait and retries
  • Filter values canonicalized against the portal's verified vocabulary (13 tev → 13TeV, lhcb → LHCb, Pb-Pb → PbPb, dataset/collision → Dataset::Collision); unknown values are sent as given and flagged
  • Tape-resident (ondemand) records included in every search and lookup, where the portal otherwise drops them silently; every hit and file states its availability
  • File manifests read once and cached for 15 minutes, so paging through a record's files costs one portal request

Agent-friendly output:

  • Provenance on every response: portal_url on each hit and record, a license with its basis, a DOI citation, and applied_filters or effectiveQuery echoing what ran
  • Graceful partial results: cern_opendata_get_records returns unresolved ids under missing with guidance, and cern_opendata_get_analysis_env reports unreadable linked records or guides in a notice rather than failing
  • Discriminated outputs: kind, license.basis, interpreted_as, scope, variant, run_bounds.source and parsed let callers branch on data, not string parsing
  • Portal text kept as data: titles, descriptions, guide sections and file names are fenced or escaped in content[] and relayed as received (HTML in _html fields) in structuredContent

Data and licensing

Portal metadata and datasets are CC0 under the CERN Open Data Terms of Use. Software, container images, documentation and guide code are licensed separately, per record (software is commonly GPL). cern_opendata_get_records reports each record's license and its basis: record when the record states one, cern_terms_default for a dataset that states none (CC0 under the Terms of Use), and not_stated otherwise. cern_opendata_get_analysis_env marks container images, software and guide code as separately licensed.

CERN asks reusers to cite each dataset's DOI in applications and publications. cern_opendata_get_records returns a ready citation for every record with a DOI.

This server is an independent project and is not affiliated with or endorsed by CERN.

Known limitations

  • 60 requests a minute per client IP. The portal publishes this limit. The server paces itself to 50 a minute, and a call that cannot start within its deadline fails as rate_limited with retryAfter. A hosted deployment shares that one budget across every user behind its egress IP, and the server has no per-user quota, so a hosted deployment needs a per-client rate limit at its edge. cern_opendata_get_analysis_env and cern_opendata_get_validated_runs cost 2–4 requests each.
  • 10,000-result window. Search and trigger-path paging reach only the first 10,000 matches; deeper result sets must be narrowed with filters.
  • Facet lists are partial. Terms facets return the first 10 values alphabetically (file_type up to 100), with the rest counted in other_count. A filter does not narrow its own facet, only the hits and the other facets.
  • Tape-resident files. Files with availability on demand must be requested on the record's portal page before download; staging them is a write and out of scope. A record whose availability is ondemand lists none of its files through the API, so cern_opendata_list_files reports only the count and size its metadata states.
  • Run lists are CMS-only, and trigger records cover CMS 2010–2016 only. Muons-only lists do not exist for every period: Commissioning2010, Run2010B and the 2011 ReReco list have none.
  • No prescale tables. Trigger detail is limited to what each record's abstract states, and fields the abstract omits are absent.
  • Glossary entries are not served. The portal's glossary links answer 404, so search excludes them.
  • Umbrella records hold no files themselves. Their files sit in child records, which cern_opendata_list_files returns under children.

Getting started

Add the following to your MCP client configuration file.

json
{  "mcpServers": {    "cern-opendata-mcp-server": {      "type": "stdio",      "command": "bunx",      "args": ["@cyanheads/cern-opendata-mcp-server@latest"],      "env": {        "MCP_TRANSPORT_TYPE": "stdio",        "MCP_LOG_LEVEL": "info"      }    }  }}

Or with npx (no Bun required):

json
{  "mcpServers": {    "cern-opendata-mcp-server": {      "type": "stdio",      "command": "npx",      "args": ["-y", "@cyanheads/cern-opendata-mcp-server@latest"],      "env": {        "MCP_TRANSPORT_TYPE": "stdio",        "MCP_LOG_LEVEL": "info"      }    }  }}

Or with Docker:

json
{  "mcpServers": {    "cern-opendata-mcp-server": {      "type": "stdio",      "command": "docker",      "args": ["run", "-i", "--rm", "-e", "MCP_TRANSPORT_TYPE=stdio", "ghcr.io/cyanheads/cern-opendata-mcp-server:latest"]    }  }}

For Streamable HTTP, set the transport and start the server:

sh
MCP_TRANSPORT_TYPE=http MCP_HTTP_PORT=3010 bun run start:http# Server listens at http://localhost:3010/mcp

Prerequisites

  • Bun v1.4.0 or higher (or Node.js v24+).
  • No API key or account: the CERN Open Data Portal is public.

Installation

  1. Clone the repository:
sh
git clone https://github.com/cyanheads/cern-opendata-mcp-server.git
  1. Navigate into the directory:
sh
cd cern-opendata-mcp-server
  1. Install dependencies:
sh
bun install
  1. Configure environment:
sh
cp .env.example .env# optional: adjust transport, logging, or telemetry settings

Configuration

The server has no settings of its own; these framework variables apply.

VariableDescriptionDefault
MCP_TRANSPORT_TYPETransport: stdio or http.stdio
MCP_HTTP_PORTHTTP server port.3010
MCP_HTTP_HOSTHTTP server host.127.0.0.1
MCP_SESSION_MODEHTTP session mode: stateless, stateful, or auto.stateless
MCP_AUTH_MODEAuthentication: none, jwt, or oauth.none
MCP_LOG_LEVELLog level (debug, info, warning, error, etc.).info
LOGS_DIRDirectory for log files (Node.js only).<app-root>/logs
OTEL_ENABLEDEnable OpenTelemetry.false

See .env.example for the common framework overrides.

Running the server

Local development

  • Build and run the production version:

    sh
    # One-time buildbun run rebuild
    # Run the built serverbun run start:http# orbun run start:stdio
  • Run checks and tests:

    sh
    bun run devcheck  # Lints, formats, type-checks, and morebun run test      # Runs the test suite

Project structure

DirectoryPurpose
src/index.tscreateApp() entry point: registers the tools and resource, sets the server instructions, starts the portal client.
src/mcp-server/toolsTool definitions (*.tool.ts), plus the shared input helpers and list enrichment.
src/mcp-server/resourcesResource definitions. The cern-opendata://record/{recid} resource.
src/mcp-server/record-schema.tsThe record output schema shared by cern_opendata_get_records and the resource.
src/services/cern-opendataPortal client (pacing, retries, per-call deadline, byte ceilings, caches), normalization, vocabulary tables, text rendering, trigger parsing.
tests/Unit and tool tests, mirroring the src/ structure, run against fixture portal responses.
docs/design.mdTool surface, verified portal behavior, and design decisions.

Development guide

See CLAUDE.md for development guidelines and architectural rules. The short version:

  • Handlers throw, framework catches — no try/catch in tool logic
  • Use ctx.log for logging and ctx.enrich for notices and paging context
  • Register new tools and resources in the barrels in src/mcp-server/*/definitions/index.ts
  • Wrap external API calls: validate raw → normalize to domain type → return output schema; never fabricate missing fields

Contributing

Issues are welcome. Run checks and tests before submitting:

sh
bun run devcheckbun run test

License

This project is licensed under the Apache 2.0 License. See the LICENSE file for details.

来源:README.md,提交 ae6cf2b

工具

0
工具元数据尚未被收录。

版本历史

1
  1. v0.1.1最新Oct 1, 2026