
Chembl Mcp Server
io.github.cyanheadsv0.3.2更新於 Oct 8, 2026
Link compounds to protein targets, rank bioactivity, and look up drug mechanisms and indications.
概覽
查詢 ChEMBL 藥物研發資料:檢索化合物與標靶、排序生物活性,並查詢藥物作用機制與適應症。
- 功能
- 把 ChEMBL(EBI)REST API 包裝成七個工具與兩個資源。工具可依名稱、ChEMBL ID、InChIKey 或 SMILES 結構(精確、相似、子結構)檢索分子,用基因符號或 UniProt 登錄號解析蛋白質標靶,並回傳依 pchembl_value 排序的生物活性測量值(IC50/Ki/EC50),可針對化合物、標靶或化合物與標靶配對。其他工具回傳藥物作用機制、分子標靶、首次核准年份與適應症、附 ChEMBL 信心分數的實驗來源,以及對溢出到 DuckDB 畫布的大量活性資料執行唯讀 SQL。
- 適用情境
- 適合需要藥物研發或化學資訊學事實的情境:把化合物與其蛋白質標靶關聯、比較效價測量值、查詢藥物機制或適應症,或判斷兩項實驗是否可比較。SQL 畫布路徑適合需要排序或聚合的大量活性資料集。
- 執行需求
- 可作為公開遠端 Streamable HTTP 端點使用,無需安裝;也可透過 bunx/npx/Docker 在本機執行。本機執行需要 Bun v1.4.0+ 或 Node.js v24+。ChEMBL 無需金鑰,不需要帳號或 API key。選用的 SQL 畫布工具需要 CANVAS_PROVIDER_TYPE=duckdb;其他變數(CHEMBL_API_BASE_URL、CHEMBL_DEFAULT_LIMIT、CHEMBL_MAX_SPILL_ROWS、MCP_HTTP_PORT、MCP_AUTH_MODE、MCP_LOG_LEVEL)皆為選用覆寫項。
安裝
在 SourceWeft 中
- 開啟 儀表板中的 Chembl Mcp Server,將其新增到工作區。
- 為需要使用其工具的對話啟用該服務。
Web executable,透過 Streamable HTTP。 遠端服務在工作區中設定後即可從網頁執行環境執行。
其他 MCP 客戶端
把它新增到你客戶端的 mcpServers 設定中。
{
"mcpServers": {
"chembl-mcp-server": {
"type": "http",
"url": "https://chembl.caseyjhand.com/mcp"
}
}
}README
@cyanheads/chembl-mcp-server
Link compounds to protein targets, rank bioactivity (IC50/Ki/EC50), and look up drug mechanisms and indications over ChEMBL via MCP. STDIO or Streamable HTTP.
Public Hosted Server: https://chembl.caseyjhand.com/mcp
Overview
Drug-discovery data over ChEMBL (EBI) — the curated link between compounds, protein targets, and measured bioactivity (IC50/Ki/EC50), plus drug mechanisms and indications. Search compounds by name, ID, or structure, resolve protein targets, rank bioactivity measurements, and look up drug mechanisms and indications from any MCP client. Runs as a stdio process, a local Streamable HTTP server, or the public hosted endpoint above.
Tools
Resources
All resource data is also reachable via the tools, so tool-only MCP clients lose nothing. There are no prompts — the canonical workflows are short tool chains an agent composes directly, and the cross-server chain guidance ships as server-level instructions instead.
Capability reference
chembl_search_molecules tool
- Default
search_type=namematches drug names, synonyms, ChEMBL IDs, and InChIKeys in one query; a query that is exactly a ChEMBL ID or InChIKey routes to ChEMBL's single-record lookup instead of the fuzzy text index (totalCount: 1) - Structure search via
search_type:exact,similarity(Tanimoto ≥similarity_threshold, integer 40–100, default 70), orsubstructure— supplystructureas a SMILES;max_phase_min(name search only) restricts to compounds at or above a max clinical phase - Every row carries
max_phase, MW, AlogP, Lipinski rule-of-five violations, and QED; onlysearch_type=similarityresults carry a Tanimotosimilaritypercent — absent, not null, on other modes - Paginated via
nextCursor/cursor, omitted (not null) on the last page — redeem a cursor with the same filters that minted it - Chain
molecule_chembl_idintochembl_get_bioactivitiesorchembl_get_drug_info
chembl_get_bioactivities tool
- Supply at least one of
molecule_chembl_idortarget_chembl_id; supplying both narrows to that compound–target pair — neither is amissing_filtererror - Filter by
standard_type(IC50/Ki/EC50/…),pchembl_value_min,assay_type,organism; ranked onpchembl_value— comparable only within onestandard_type potency_viewselectspotency_ranked(default, measurements with a derivablepchembl_value) ornull_potency(the excluded rows);pchembl_value_minwithnull_potencyis acontradictory_potency_filtererror, andtotalCountspans both views- Numerics are coerced to
number | nullat the service boundary — a missing potency reads asnull, never0 - Large sets spill to a DataCanvas table per view (
bioactivities/bioactivities_null_potency), capped atCHEMBL_MAX_SPILL_ROWS(default 50,000) and reportedtruncated: true+staged_row_countwhen hit; requiresCANVAS_PROVIDER_TYPE=duckdb - The inline preview is always capped at
limit(default 25) regardless of spill status; the optionalcanvas_idreuses a canvas, but re-querying the same view replaces its prior rows
chembl_search_targets tool
- Supply at least one of
accession(UniProt, e.g.P00533),gene_symbol, orquery(free-text); narrow withorganismandtarget_type— none supplied is amissing_inputerror - A UniProt accession is the most precise input — chain it from a
uniprot/proteinserver - Each row carries target type, organism, and component UniProt accessions + gene symbols, flattened from ChEMBL's nested component synonyms
- Paginated via
nextCursor/cursor, the same contract aschembl_search_molecules - Chain
target_chembl_idintochembl_get_bioactivities
chembl_get_drug_info tool
- Supply
molecule_chembl_id; returns mechanism(s) of action, molecular target(s), action type, first-approval year, and clinical indications with the max phase reached for each - Mechanisms and indications are fetched with
Promise.allSettled, so a rejected list degrades to a disclosed partial result rather than failing the call - Each list carries its own
mechanisms_status/indications_status(complete/truncated/failed) next to a*_total_count— an empty array is authoritative only when the status iscomplete - A mechanism's
target_chembl_idchains intochembl_get_bioactivities
chembl_get_assay tool
- Supply
assay_chembl_idfrom achembl_get_bioactivitiesrow - Returns description, assay type (binding / functional / ADMET / toxicity), the target measured, organism, and ChEMBL's 1–9 confidence score (9 = direct assay on the protein target, lower = homologous or indirect)
- Call it to judge whether two measurements are comparable before ranking them together
chembl_dataframe_query tool
- Accepts a single read-only
SELECTagainst acanvas_idfrom a spilledchembl_get_bioactivitiescall; writes, DDL, and non-SELECT statements are rejected by the framework SQL gate - Reference each staged table by the name
chembl_get_bioactivitiesreturned —bioactivities(potency_ranked) orbioactivities_null_potency(null_potency); discover columns withchembl_dataframe_describefirst - Two independent bounds, each disclosed:
truncatedis the canvas engine's own query-result cap;rendered_rowsis how many rows thecontent[]markdown table holds under its character budget — either can trip without the other; page past both with SQLLIMIT/OFFSET structuredContent.rowsalways carries the full materialized result regardless of the render bound- Requires
CANVAS_PROVIDER_TYPE=duckdb, else acanvas_disablederror
chembl_dataframe_describe tool
- Supply a
canvas_idfrom a spilledchembl_get_bioactivitiescall - Returns each staged table/view with its row count, kind, and column names + types
- Requires
CANVAS_PROVIDER_TYPE=duckdb, else acanvas_disablederror
chembl_dataframe_drop tool
- Opt-in — registered only when
CHEMBL_DATAFRAME_DROP_ENABLED=true; absent fromtools/listwhen off, though it still appears in the server manifest carrying the enable hint - Drops a named staged table by
canvas_id+table_name; returnsdropped: trueif it existed,falseif already gone - Rarely needed — per-table and per-canvas TTL already reclaim staged tables; reach for it only to free a large table early in a long session
- Requires
CANVAS_PROVIDER_TYPE=duckdb
chembl://molecule/{chemblId} resource
- Molecule record as
application/json— the same shape achembl_search_moleculesrow carries (ID, names, structures, properties, max clinical phase) chemblIdis validated against theCHEMBL\d+pattern- Fully covered by the tool surface — a convenience injectable-context mirror of the per-record fetch
chembl://target/{chemblId} resource
- Target record as
application/json— preferred name, type, organism, and component UniProt accessions + gene symbols chemblIdis validated against theCHEMBL\d+pattern- Fully covered by the tool surface — a convenience injectable-context mirror of the per-record fetch
Features
Built on @cyanheads/mcp-ts-core: stdio and Streamable HTTP transports, pluggable auth (none / jwt / oauth), swappable storage (in-memory, filesystem, Supabase, Cloudflare KV/R2/D1), structured logging with optional OpenTelemetry tracing.
ChEMBL-specific:
- Bidirectional bioactivity bridge — one tool serves both compound→target and target→compound, ranked on
pchembl_value - Structure search (exact / similarity / substructure) consolidated under one discovery tool via a
search_typeenum - String →
number | nullnumeric coercion at the service boundary — a missing potency becomesnull, never0 - DataCanvas spill on the flagship tool — tens-of-thousands-of-row activity sets stream to a DuckDB table you inspect with
chembl_dataframe_describeand query viachembl_dataframe_query - Server-level
instructionscarry the cross-server chain guidance and the ChEMBL CC BY-SA 3.0 attribution requirement
Agent-friendly output:
- Provenance on every response — total-found counts, applied-filter echo, and a spill notice so agents know whether the preview is the full set or a slice of a canvas table
- Truncation disclosure — capped searches report
shown/cap/totalCount, and spilled tables reporttruncated+staged_row_count, so a page or slice is never mistaken for the complete set - Typed, recoverable errors —
missing_filter/missing_input/contradictory_potency_filter/canvas_disabledcarry recovery hints, so callers correct the call without parsing prose - Never fabricates — normalization and
format()preservenullpotency / units; a missing measurement renders as "not reported", never0
Getting started
Public Hosted Instance
A public instance is available at https://chembl.caseyjhand.com/mcp — no installation required. Point any MCP client at it via Streamable HTTP:
Self-Hosted / Local
ChEMBL is keyless — no API key or account is required.
Add the following to your MCP client configuration file:
Or with npx (no Bun required):
Or with Docker:
For Streamable HTTP, set the transport and start the server:
To unlock the analytical SQL path (the bioactivities spill and the chembl_dataframe_* tools), add "CANVAS_PROVIDER_TYPE": "duckdb" to the env.
Prerequisites
- Bun v1.4.0 or higher (or Node.js v24+).
- Optional: set
CANVAS_PROVIDER_TYPE=duckdbto enable the DataCanvas SQL path for large bioactivity sets.
Installation
- Clone the repository:
- Navigate into the directory:
- Install dependencies:
- Configure environment:
Configuration
All configuration is validated at startup via Zod schemas in src/config/server-config.ts. ChEMBL is keyless, so every variable is optional.
See .env.example for the full list of optional overrides.
Running the server
Local development
-
Build and run:
-
Run checks and tests:
Docker
The Dockerfile defaults to HTTP transport, stateless session mode, and logs to /var/log/chembl-mcp-server. OpenTelemetry peer dependencies are installed by default — build with --build-arg OTEL_ENABLED=false to omit them. Dependency installation and security scanning run on the builder's native platform, with Bun's --cpu and --os selecting DuckDB bindings for the target linux/amd64 or linux/arm64 image. The slim runtime receives that production dependency tree, so CANVAS_PROVIDER_TYPE=duckdb works on either image architecture.
Project structure
Development guide
See CLAUDE.md/AGENTS.md for development guidelines and architectural rules. The short version:
- Handlers throw, framework catches — no
try/catchin tool logic - Use
ctx.logfor request-scoped logging,ctx.statefor tenant-scoped storage - Register new tools and resources in the
createApp()arrays - Wrap the ChEMBL API: validate raw → normalize to the flat domain type → return the output schema; never fabricate missing fields (absent numerics become
null, never0)
Contributing
Issues are welcome. Run checks and tests before submitting:
License
Apache-2.0 — see LICENSE for details.
來源:README.md,提交 220c08d
工具
0版本歷史
3- v0.3.2最新Sep 27, 2026
- v0.3.0Sep 20, 2026
- v0.2.4Sep 16, 2026

