
Gnomad Genetics Mcp Server
io.github.cyanheadsv0.4.0更新於 Oct 8, 2026
Look up allele frequencies by ancestry, gene constraint, variants, and coverage over gnomAD.
概覽
讓助理查詢 gnomAD 的依祖源等位基因頻率、基因約束、基因變異清單、定序覆蓋度以及 ClinVar 臨床意義。
- 功能
- 提供七個以 gnomAD(Broad Institute)為基礎並合併 ClinVar 臨床意義的工具:單一或多個變異的完整族群記錄(含依祖源的等位基因頻率)、基因功能喪失約束指標、基因或區域變異目錄、定序覆蓋深度,以及基因層級的 ClinVar 記錄。過大的結果集可暫存到 DuckDB 畫布並以唯讀 SQL 查詢。兩個資源對應變異與基因約束查詢,另有一個變異分診提示依頻率、約束、可檢出性順序串接。
- 適用情境
- 適合遺傳學或罕見疾病變異解讀情境,助理需要族群頻率、基因約束、覆蓋度或 ClinVar 背景資訊時使用。也適合以 SQL 對大型變異清單做探索性分析。若只需一般網頁或文獻檢索則不必安裝。
- 執行需求
- 可作為遠端 Streamable HTTP 端點使用,也可在本機透過 npm 搭配 Node.js v24+(或 Bun v1.4.0+)、npx、bunx 或 Docker 執行。gnomAD 無需金鑰,因此不需要憑證;選用的 NCBI_API_KEY 可提高 ClinVar 速率限制。SQL 畫布路徑需設定 CANVAS_PROVIDER_TYPE=duckdb。需要連線至 gnomAD 與 NCBI 的網路。
安裝
在 SourceWeft 中
- 開啟 儀表板中的 Gnomad Genetics Mcp Server,將其新增到工作區。
- 為需要使用其工具的對話啟用該服務。
Web executable,透過 Streamable HTTP。 遠端服務在工作區中設定後即可從網頁執行環境執行。
其他 MCP 客戶端
把它新增到你客戶端的 mcpServers 設定中。
{
"mcpServers": {
"gnomad-genetics-mcp-server": {
"type": "http",
"url": "https://gnomad-genetics.caseyjhand.com/mcp"
}
}
}README
@cyanheads/gnomad-genetics-mcp-server
Look up variant allele frequencies by ancestry, gene loss-of-function constraint, gene variant lists, and sequencing coverage over gnomAD — with ClinVar significance joined in — via MCP. STDIO or Streamable HTTP.
Public Hosted Server: https://gnomad-genetics.caseyjhand.com/mcp
Overview
Population genetics over gnomAD (Broad Institute), with ClinVar clinical significance joined in from NCBI. Look up per-ancestry allele frequencies, gene loss-of-function constraint, gene variant catalogs, and sequencing coverage, then query large result sets with SQL from any MCP client. Runs as a stdio process, a local Streamable HTTP server, or the public hosted endpoint above.
Tools
Resources
All resource data is also reachable via tools. The list tools (gnomad_list_gene_variants, gnomad_get_coverage, gnomad_search_clinvar) return analytical row sets rather than stable single-URI documents, so they are not exposed as resources — call the tools instead.
Prompts
Capability reference
gnomad_get_variant tool
- Batch up to 25 IDs per call (default; raise via
GNOMAD_MAX_VARIANT_BATCH), each achrom-pos-ref-altvariantId on chromosome1–22,X, orYwith an optionalchrprefix (e.g.1-55051215-G-GA) or an rsID (e.g.rs11591147) - Per-item partial success — a malformed or absent ID lands in
failed[]without failing the others - Each
failed[]item is{ variant, error, reason, recovery }:reasonis typed (invalid_variant_id,variant_not_found,upstream_unavailable, …) andrecoveryis the hint the tool declares for it; an ambiguous rsID addscandidates - Mitochondrial IDs (
M,MT,chrM) are refused per item asmitochondrial_unsupportedwith no fetch — gnomAD models mitochondrial variants separately, and they are outside this server - Per-ancestry frequency vector is returned in full, never collapsed to a single global AF
- Reports which callset(s) (
exome/genome) carry the variant, quality flags, transcript consequence, in-silico predictor scores, and the ClinVar significance gnomAD joins per variant - An empty
found[]for a well-formed ID means the variant is not in the chosen dataset — pair withgnomad_get_coverageto confirm the position is callable before concluding true absence
gnomad_get_gene_constraint tool
-
Accepts an HGNC symbol (
PCSK9) or an Ensembl gene ID (ENSG00000169174) -
Returns pLI (>0.9 intolerant), LOEUF /
oe_lof_upperwith its lower bound, observed/expected ratios for LoF / missense / synonymous, and the three Z-scores -
constraint_releasenames the release the metrics come from: -
ExAC publishes pLI, the Z-scores, and observed/expected counts only, so on
exacthe ratios and LOEUF are null andconstraint_flagsis empty -
Many genes have null constraint (sparse upstream) — null fields are reported as such, never fabricated
-
constraint_flagscarries the caveat flags gnomAD attaches to a gene's constraint (e.g.no_exp_lof,syn_outlier)
gnomad_list_gene_variants tool
- Supply exactly one of
gene,transcript_id, orregion(chrom-start-stop, 1-based inclusive, chromosome1–22,X, orYwith an optionalchrprefix) - A region must span less than 2,500,000 bp (stop − start) and hold at most ~30,000 variants; an unserved chromosome or out-of-range coordinate fails as
invalid_regionand an over-wide span asregion_too_large, both before any fetch, while a region over the variant ceiling fails asregion_too_largeafter one request, carrying gnomAD's own message - Mitochondrial targets (an
M/MTregion, or a gene or transcript gnomAD places on chromosome M) fail asmitochondrial_unsupportedinstead of returning an empty list - Optional filters: one
consequence_class(lof/missense/synonymous/other) and/or a maximum allele frequency - A result too large to inline (the preview holds about 14,000 characters of rows, keeping a response near 24 KB) is staged on a DataCanvas table named
gene_variants, returned ascanvas_idandtable_namebeside the preview — inspect it withgnomad_dataframe_describe, then query it withgnomad_dataframe_queryto rank by AF, count by consequence, or group across every row. The response notice names the table and both tools - A result that fits inline stages no table and uses no canvas (
canvas_idis empty) unless you pass acanvas_id - Passing a
canvas_idalways writes the result togene_variantson that canvas, REPLACING the previous table (it never appends), even when the result fits inline; a result with no variants removes the table - When the canvas is disabled (
CANVAS_PROVIDER_TYPE!=duckdb) the tool returns the same capped inline preview (as many rows as fit about 14,000 characters) and the SQL path is unavailable - A blank
geneortranscript_idcounts as omitted
gnomad_get_coverage tool
- Supply exactly one of
gene,transcript_id, orregion; a blankgeneortranscript_idcounts as omitted regiontakes the samechrom-start-stopform asgnomad_list_gene_variants— chromosome1–22,X, orY, optionalchrprefix, a span under 2,500,000 bp — and fails withinvalid_regionorregion_too_largebefore any fetch otherwise- Mitochondrial targets fail as
mitochondrial_unsupportedinstead of returning empty coverage - Returns mean and median read depth plus the mean fraction of samples covered at each threshold (1× through 100×), summarized per callset track
coverage_sourcenarrows to one track (exome/genome); omit to return every available track- A variant missing from a well-covered region is informative; one missing from a poorly-covered region is not
gnomad_search_clinvar tool
- Returns a gene's classified ClinVar variants — clinical significance, review status with a 0–4 star rating, associated conditions, molecular consequences, and submission counts
- Each row carries gnomAD-compatible identifiers:
canonical_spdi,rsids, andgrch38_variant_id(chrom-pos-ref-alt, set for SNVs, MNVs, and delins), whichgnomad_get_variantresolves in the GRCh38 datasets - Optional filters:
clinical_significance(e.g.pathogenic; blank means no filter) and a minimum star rating (min_review_stars, 0–4) - Returns one window of up to 500 ClinVar records per call:
total_foundis ClinVar's candidate count for the gene and filter terms (taken before the significance and star filters narrow each window),truncatedandnext_offsetsay whether more remain, andoffset/limit(1–500) page through them.limitcounts records before the filters, so a window can return fewer rows - VariationIDs ClinVar returns no summary for are listed in
unavailable_idsrather than returned as blank rows - Accepts an HGNC symbol only — ClinVar's gene index doesn't resolve Ensembl gene IDs, unlike the other gnomAD tools. An Ensembl gene ID returns guidance to resolve its symbol and searches nothing, leaving any
canvas_idyou pass untouched - A window too large to inline (the preview holds about 11,000 characters of rows, keeping a response near 24 KB) is staged on the
clinvar_variantscanvas table — inspect it withgnomad_dataframe_describe, then query it withgnomad_dataframe_query. A window that fits inline stages no table unless you pass acanvas_id; passing one always writes the window toclinvar_variants, REPLACING the previous table, and a window with no rows removes it. With the canvas disabled, the preview is the same capped preview (as many rows as fit about 11,000 characters) - Keyless, but honors
NCBI_API_KEYfor a higher rate limit (10 vs 3 req/s)
gnomad_dataframe_query tool
- Runs single-statement, read-only SQL
SELECTs against a canvas table staged bygnomad_list_gene_variantsorgnomad_search_clinvar— writes, DDL, and file/HTTP table functions are rejected by the canvas gate - Reference tables by the name the staging tool returned (
gene_variantsorclinvar_variants) - Returns one page of the result:
offset(default 0) andlimit(default 100, max 500) select it, and a page also ends before its rows pass 10,000 characters of JSON, which keeps a full page under about 24 KB - Output:
rows(dynamic columns per the SQL projection),columns,offset,returned,total(exact row count;nullwhen the result exceeds the canvas row cap),truncated(rows exist after this page), andnext_offset(nullon the last page) — follownext_offsetuntil it isnull - Each page re-runs the SQL: stable paging needs an
ORDER BYover a unique key (such asvariant_id) and an unchanged table. Paging stops at the canvas row cap (CANVAS_DEFAULT_ROW_LIMIT, 10,000 by default); filter or aggregate in SQL to reach rows past it - A row over the 10,000-character budget fails with
row_too_large— select fewer or narrower columns - Requires
CANVAS_PROVIDER_TYPE=duckdb— otherwise fails with acanvas_disablederror
gnomad_dataframe_describe tool
- Lists every table staged on a canvas with its row count and column schema (name and DuckDB type)
- Call it before writing SQL for
gnomad_dataframe_query - Requires
CANVAS_PROVIDER_TYPE=duckdb— otherwise fails with acanvas_disablederror
gnomad_dataframe_drop tool
- Drops a named table from a canvas to reclaim memory — a deliberate mutation (
readOnlyHint: false,destructiveHint: true) on an otherwise read-only surface - Opt-in via
GNOMAD_DATAFRAME_DROP_ENABLED=true; absent fromtools/listwhen off, since per-table TTL already reclaims memory automatically - Requires
CANVAS_PROVIDER_TYPE=duckdb— otherwise fails with acanvas_disablederror
gnomad://variant/{dataset}/{variantId} resource
- Population record for one variant as
application/json— mirrorsgnomad_get_variant; thedatasetsegment keeps the URI self-describing variantIdaccepts a chrom-pos-ref-alt ID or an rsID, same grammar as the tool- Typed errors:
invalid_variant_id(outside the coordinate/rsID grammar),mitochondrial_unsupported(anM/MT/chrMID),variant_not_found,ambiguous_rsid(withcandidates), and the gnomAD failuresgraphql_error,upstream_build_mismatch,upstream_unavailable,upstream_timeout,upstream_access, andinvalid_upstream_response— each with its declared recovery hint
gnomad://gene/{dataset}/{gene}/constraint resource
- Gene loss-of-function constraint as
application/json— mirrorsgnomad_get_gene_constraint, including the sameconstraint_releasefor eachdatasetsegment (exacserves ExAC r0.3 constraint;gnomad_r3serves the GRCh38 gnomAD v4.1.2 table) geneaccepts an HGNC symbol or Ensembl gene ID- Typed errors:
gene_not_foundwhen no gene matches in the requested build,invalid_constraint_datawhen gnomAD's metrics fall outside their valid ranges, and the gnomAD failuresgraphql_error,upstream_unavailable,upstream_timeout,upstream_access, andinvalid_upstream_response— each with its declared recovery hint
gnomad_variant_triage prompt
- Arguments:
variantrequired (chrom-pos-ref-alt or rsID);geneanddatasetoptional.datasetis one ofgnomad_r4,gnomad_r3,gnomad_r2_1,exac— any other value is rejected; omitted or blank, the emitted calls carry nodatasetand use the server default. A blankgenecounts as omitted - Emits a three-step chain as one user message: population frequency (
gnomad_get_variant) → gene constraint (gnomad_get_gene_constraint) → callability check (gnomad_get_coverageon the exact position, not gene-level) - Rejects a malformed
variantwith a validation error before generating the chain
Features
Built on @cyanheads/mcp-ts-core: stdio and Streamable HTTP transports, pluggable auth (none / jwt / oauth), swappable storage (in-memory, filesystem, Supabase, Cloudflare KV/R2/D1), structured logging with optional OpenTelemetry tracing.
gnomAD-specific:
- Single keyless GraphQL source for the entire core surface — ClinVar significance is joined per variant inside gnomAD's own response
datasetandreference_genomeare distinct, coherence-validated parameters (v4/v3 ⇒ GRCh38, v2.1/ExAC ⇒ GRCh37); both are echoed in every tool's output so a wrong-build coordinate mismatch is visible- Polite client — conservative concurrency cap (
GNOMAD_MAX_CONCURRENCY, default 2) and exponential backoff against a community-funded, rate-limited API - In-conversation SQL analytics:
gnomad_list_gene_variantsandgnomad_search_clinvarstage results too large to inline on a DuckDB-backed canvas table —gnomad_dataframe_describelists its columns,gnomad_dataframe_queryruns SQL over every row
Agent-friendly output:
- Per-ancestry allele-frequency vector returned in full, never collapsed to a single global AF — the cross-ancestry contrast is the signal clinical interpretation needs
- Graceful partial failure —
gnomad_get_variantreturns per-itemfailed[]rows, each with a typed reason and its recovery hint, instead of failing the whole batch - Provenance on every response — effective
datasetandreference_genomeechoed back; null upstream fields preserved as null, never fabricated - Every error carries a typed
reasonand the recovery hint its tool or resource declares for that reason, so callers know the next move: input problems (incoherent_build,invalid_target,invalid_variant_id,invalid_region,region_too_large,mitochondrial_unsupported,ambiguous_rsid), absences (gene_not_found,variant_not_found), gnomAD refusals (graphql_error,upstream_build_mismatch,invalid_constraint_data), upstream faults (upstream_unavailable,upstream_timeout,upstream_access,invalid_upstream_response), and the canvas (canvas_disabled,row_too_large).gnomad_get_variantputs the same reason and hint on eachfailed[]item
Getting started
Public Hosted Instance
A public instance is available at https://gnomad-genetics.caseyjhand.com/mcp — no installation required. Point any MCP client at it via Streamable HTTP:
Self-Hosted / Local
Add the following to your MCP client configuration file. gnomAD is a free, keyless API — no credentials required.
Or with npx (no Bun required):
Or with Docker:
For Streamable HTTP, set the transport and start the server:
To enable the SQL analytics path, also set CANVAS_PROVIDER_TYPE=duckdb — @duckdb/node-api ships as a dependency, so nothing extra to install.
Prerequisites
- Bun v1.4.0 or higher (or Node.js v24+).
- No API key — gnomAD's GraphQL endpoint is keyless. An optional
NCBI_API_KEYraises thegnomad_search_clinvarrate limit.
Installation
- Clone the repository:
- Navigate into the directory:
- Install dependencies:
- Configure environment:
Configuration
All variables are optional; the server runs keyless with the defaults below.
See .env.example for the full list of optional overrides.
Running the server
Local development
-
Build and run:
-
Run checks and tests:
Docker
The Dockerfile defaults to HTTP transport, stateless session mode, and logs to /var/log/gnomad-genetics-mcp-server. OpenTelemetry peer dependencies are installed by default — build with --build-arg OTEL_ENABLED=false to omit them.
Project structure
Development guide
See CLAUDE.md/AGENTS.md for development guidelines and architectural rules. The short version:
- Handlers throw, framework catches — no
try/catchin tool logic - Use
ctx.logfor request-scoped logging,ctx.statefor tenant-scoped storage - Register new tools and resources in the
createApp()arrays insrc/index.ts - Wrap external API calls: validate raw → normalize to domain type → return output schema; never fabricate missing fields
Data attribution
gnomAD data is provided by the Genome Aggregation Database (Broad Institute). ClinVar data is provided by NCBI.
Contributing
Issues are welcome. Run checks and tests before submitting:
License
Apache-2.0 — see LICENSE for details.
來源:README.md,提交 14557fc
工具
0版本歷史
4- v0.4.0最新Sep 30, 2026
- v0.3.0Sep 24, 2026
- v0.2.2Sep 20, 2026
- v0.2.1Sep 16, 2026

