
Cern Opendata Mcp Server
io.github.cyanheadsv0.1.1Updated Oct 1, 2026
Search CERN Open Data, fetch records, files, analysis environments, CMS good-run lists, HLT paths.
Overview
Lets an assistant search and read CERN Open Data Portal records, files, analysis environments, CMS good-run lists and trigger paths.
- What it does
- It exposes seven read-only tools over the CERN Open Data Portal: search records with exact-vocabulary filters and live facet counts, fetch full metadata for up to 20 records by recid, DOI, CMS dataset path or documentation slug, and page through a record's file indexes with XRootD URIs, HTTPS URLs, sizes and checksums. It also assembles a record's analysis environment (container images, CMSSW release, global tag, guides), returns CMS validated-run lists with luminosity sections, looks up CMS High-Level Trigger paths, and decodes the portal vocabulary. A cern-opendata://record/{recid} resource returns one record's metadata, license and citation.
- When to use it
- Useful when working with particle-physics open data: finding datasets from ALICE, ATLAS, CMS or LHCb, checking licensing and citations, locating data files, or preparing a CMS analysis environment and run selection. It fits research, teaching and reproducible-analysis workflows that need portal metadata rather than local data processing.
- Requirements
- Runs locally as a stdio process (or a local Streamable HTTP server) via npm package @cyanheads/cern-opendata-mcp-server, using Bun v1.4.0+ or Node.js v24+, or Docker. No API key or account is needed; the portal is public. Network access to the CERN Open Data Portal is required. Optional framework variables include MCP_TRANSPORT_TYPE, MCP_HTTP_PORT, MCP_HTTP_HOST, MCP_AUTH_MODE and MCP_LOG_LEVEL.
Installation
In SourceWeft
- Open Cern Opendata Mcp Server in the dashboard and add it to a workspace.
- Enable the server for the chats that should use its tools.
Desktop only via STDIO. STDIO servers start a local process, so they need the SourceWeft desktop host.
Other MCP clients
Follow the launch instructions in the repository.
README
@cyanheads/cern-opendata-mcp-server
Search CERN Open Data, fetch records, files, analysis environments, CMS good-run lists, HLT paths via MCP. STDIO or Streamable HTTP.
Overview
Particle-physics data from the CERN Open Data Portal: collision, simulated and derived datasets, analysis software, environments and documentation from ALICE, ATLAS, CMS, LHCb and other experiments. Search it with exact-vocabulary filters and live facet counts, open records with their license and citation, list the files that hold the data, assemble a record's analysis environment, and look up CMS good-run lists and trigger paths. Runs as a stdio process or a local Streamable HTTP server.
Tools
Resources
Tool-only clients get the same data from cern_opendata_get_records.
Capability reference
cern_opendata_search_records tool
- Optional
query(an OpenSearchquery_string, up to 500 characters) plus OR-list filterstype,experiment,collision_energy,collision_type,file_type,availabilityandcollection, each an array or a comma-separated string;year_from/year_toandmin_events/max_eventsbound the data-taking year and the event count sort(bestmatch,mostrecent,title,title_desc),limit1–50 (default 10) andpagefrom 1; paging reaches the first 10,000 matches, andpage × limitpast that fails aspage_window_exceeded- Compact
hitswith recids, plus eight livefacetsthat each ignore their own filter;applied_filtersechoes what ran, with values outside the verified vocabulary listed underunrecognized_values
cern_opendata_get_records tool
ids: 1–20 recids, DOIs, CMS dataset paths (/Primary/Era/TIER) or documentation slugs, mixed in one array or comma-separated string- Each record carries a
licensewith itsbasis(record,cern_terms_default,not_stated) and, when it has a DOI, a readycitation; documentation and news bodies are cut at 30,000 characters - Identifiers that resolve to nothing land in
missingwithinterpreted_asand guidance instead of failing the call; file lists come fromcern_opendata_list_files
cern_opendata_list_files tool
recidrequired; withoutindex, returns the record's file indexes and regular files, and with an index key, that index's fileslimit1–500 (default 50), continued withnext_cursor; each file carriesxrootd_uri,https_url,size_in_bytes,checksumandavailability, and each index auri_list_urllisting every XRootD URI in it- Files marked
on demandsit on tape and must be requested on the record's portal page first; an umbrella record with no files of its own returns itschildrenrecids
cern_opendata_get_analysis_env tool
recidrequired;softwarecarries the record's own container images, CMSSW release, global tag and environment recidenvironment_records(condition, VM, validation) for the record's run periods andexample_softwarethat declares it works with the record, up to 50 between them;guidesquotes the linked section of the first two portal guides, each capped at 12,000 characters- Always
separately_licensed: true; linked records or guides that can't be read leave anoticeinstead of failing the call
cern_opendata_get_validated_runs tool
- Exactly one of
recid(a CMS collision dataset or a validated-run list) orrun_period(Run2012B;2012Balso matches);variantfullormuons_only;run_min/run_max;limit1–2000 (default 200) - A dataset
recidbounds the runs to the first and last run the dataset lists, echoed inrun_bounds; when several lists match,matched_listsnames them and no runs are read - Each run carries
lumi_sectionsandlumi_ranges, andlist.https_urldownloads the whole list file; CMS only, so other records fail asno_validated_runs
cern_opendata_search_trigger_paths tool
path: an exact name (HLT_IsoMu24) or a prefix with one trailing*(HLT_IsoMu*);HLT_is added when missing and a_v<n>version suffix dropped; optionalyear,limit1–50 (default 10) andpage- Each per-year record is parsed into
first_seen,last_seen, per-version run ranges with theirl1_seed, and HLT menu record links;parsed: falsemarks a record to read from itsabstract_html - CMS open data from 2010–2016; prescale tables are not published
cern_opendata_list_reference tool
- Optional
topic:experiments,record_types,collision_energies,collision_types,file_types,availability,identifiers,query_syntax,licensingorrun_periods; omit it for every table - Static and offline, with no portal requests;
run_periodsis a dated snapshot, whilecern_opendata_get_validated_runsreads the live list collection
cern-opendata://record/{recid} resource
- One record by
recid(leading zeros ignored) asapplication/json, in thecern_opendata_get_recordsrecord shape: metadata, license and citation, without file lists recidcomes fromcern_opendata_search_records; reads carry a 15-minute public cache hint
Features
Built on @cyanheads/mcp-ts-core: stdio and Streamable HTTP transports, pluggable auth (none / jwt / oauth), swappable storage (in-memory, filesystem, Supabase, Cloudflare KV/R2/D1), structured logging with optional OpenTelemetry tracing.
CERN Open Data-specific:
- Keyless, read-only client for the portal's record, documentation and file routes; it never stages tape files or writes anything
- One shared pacer at 50 requests a minute, under the portal's published 60 per client IP, and one 50-second deadline per call across queue wait and retries
- Filter values canonicalized against the portal's verified vocabulary (
13 tev→13TeV,lhcb→LHCb,Pb-Pb→PbPb,dataset/collision→Dataset::Collision); unknown values are sent as given and flagged - Tape-resident (
ondemand) records included in every search and lookup, where the portal otherwise drops them silently; every hit and file states its availability - File manifests read once and cached for 15 minutes, so paging through a record's files costs one portal request
Agent-friendly output:
- Provenance on every response:
portal_urlon each hit and record, alicensewith itsbasis, a DOIcitation, andapplied_filtersoreffectiveQueryechoing what ran - Graceful partial results:
cern_opendata_get_recordsreturns unresolved ids undermissingwith guidance, andcern_opendata_get_analysis_envreports unreadable linked records or guides in anoticerather than failing - Discriminated outputs:
kind,license.basis,interpreted_as,scope,variant,run_bounds.sourceandparsedlet callers branch on data, not string parsing - Portal text kept as data: titles, descriptions, guide sections and file names are fenced or escaped in
content[]and relayed as received (HTML in_htmlfields) instructuredContent
Data and licensing
Portal metadata and datasets are CC0 under the CERN Open Data Terms of Use. Software, container images, documentation and guide code are licensed separately, per record (software is commonly GPL). cern_opendata_get_records reports each record's license and its basis: record when the record states one, cern_terms_default for a dataset that states none (CC0 under the Terms of Use), and not_stated otherwise. cern_opendata_get_analysis_env marks container images, software and guide code as separately licensed.
CERN asks reusers to cite each dataset's DOI in applications and publications. cern_opendata_get_records returns a ready citation for every record with a DOI.
This server is an independent project and is not affiliated with or endorsed by CERN.
Known limitations
- 60 requests a minute per client IP. The portal publishes this limit. The server paces itself to 50 a minute, and a call that cannot start within its deadline fails as
rate_limitedwithretryAfter. A hosted deployment shares that one budget across every user behind its egress IP, and the server has no per-user quota, so a hosted deployment needs a per-client rate limit at its edge.cern_opendata_get_analysis_envandcern_opendata_get_validated_runscost 2–4 requests each. - 10,000-result window. Search and trigger-path paging reach only the first 10,000 matches; deeper result sets must be narrowed with filters.
- Facet lists are partial. Terms facets return the first 10 values alphabetically (
file_typeup to 100), with the rest counted inother_count. A filter does not narrow its own facet, only the hits and the other facets. - Tape-resident files. Files with availability
on demandmust be requested on the record's portal page before download; staging them is a write and out of scope. A record whose availability isondemandlists none of its files through the API, socern_opendata_list_filesreports only the count and size its metadata states. - Run lists are CMS-only, and trigger records cover CMS 2010–2016 only. Muons-only lists do not exist for every period: Commissioning2010, Run2010B and the 2011 ReReco list have none.
- No prescale tables. Trigger detail is limited to what each record's abstract states, and fields the abstract omits are absent.
- Glossary entries are not served. The portal's glossary links answer 404, so search excludes them.
- Umbrella records hold no files themselves. Their files sit in child records, which
cern_opendata_list_filesreturns underchildren.
Getting started
Add the following to your MCP client configuration file.
Or with npx (no Bun required):
Or with Docker:
For Streamable HTTP, set the transport and start the server:
Prerequisites
- Bun v1.4.0 or higher (or Node.js v24+).
- No API key or account: the CERN Open Data Portal is public.
Installation
- Clone the repository:
- Navigate into the directory:
- Install dependencies:
- Configure environment:
Configuration
The server has no settings of its own; these framework variables apply.
See .env.example for the common framework overrides.
Running the server
Local development
-
Build and run the production version:
-
Run checks and tests:
Project structure
Development guide
See CLAUDE.md for development guidelines and architectural rules. The short version:
- Handlers throw, framework catches — no
try/catchin tool logic - Use
ctx.logfor logging andctx.enrichfor notices and paging context - Register new tools and resources in the barrels in
src/mcp-server/*/definitions/index.ts - Wrap external API calls: validate raw → normalize to domain type → return output schema; never fabricate missing fields
Contributing
Issues are welcome. Run checks and tests before submitting:
License
This project is licensed under the Apache 2.0 License. See the LICENSE file for details.
Source: README.md at commit ae6cf2b
Tools
0Version history
1- v0.1.1LatestOct 1, 2026

