BioHarbor

io.github.danielLuo2v0.1.4更新于 Oct 7, 2026

Run real bioinformatics from your AI agent. Reliable, reproducible, on your own GPUs.

已验证STDIO仅桌面AI & MLData & Analytics

概览

AI 生成的概览

让 AI 助手在本机 GPU 上真正运行生物信息学计算:序列统计、翻译、ORF 查找、MMseqs2 同源搜索和 ESMFold 结构预测。

功能
BioHarbor 提供的生物信息学工具是让智能体执行分析,而不只是查询。内联工具包括序列校验与统计、DNA/RNA 单框或六框翻译,以及双链最长 ORF 查找。耗时任务以作业形式排队:针对本地数据库的 MMseqs2 同源搜索,以及 ESMFold 结构预测,返回 pLDDT 区间、低置信区域和 pTM。get_job、list_jobs、cancel_job、describe_tool、list_databases、gpu_status、read_file 等运行时工具用于管理作业,每次运行都会连同溯源信息记录到 SQLite。
适用场景
当助手需要真正执行序列分析、同源搜索或蛋白质结构预测,而不仅是查询生物数据库记录时适用。适合拥有自有 GPU 硬件的实验室,包括通过 SSH 隧道访问共享 GPU 服务器的场景,也适合需要可复现、带溯源记录运行的工作流。
运行要求
需要本地 Python 环境,并通过 pip 安装 bioharbor 这个 PyPI 包;清单声明不需要认证、环境变量或请求头。同源搜索需要 MMseqs2 以及用 setup-db 配置的参考数据库。GPU 结构预测需要 esmfold 附加组件,RTX 50xx 显卡还需 CUDA 12.8+ 的 PyTorch。仅支持桌面客户端;HTTP 模式为可选项且无认证。
安装前请注意
项目处于 alpha(v0.1)阶段,scrna_pipeline 等部分工具仍在计划中。HTTP 模式没有认证,应保持绑定 127.0.0.1 并通过 SSH 隧道访问。作业会向磁盘写入文件与溯源记录,GPU 作业会占用共享显存和机时,在多用户机器上需注意调度与取消行为。

安装

在 SourceWeft 中

  1. 打开 控制台中的 BioHarbor,将其添加到工作区。
  2. 为需要使用其工具的对话启用该服务。

Desktop only,通过 STDIO。 STDIO 服务会启动本地进程,因此需要 SourceWeft 桌面宿主。

其他 MCP 客户端

参照 仓库 中的启动说明。

README

BioHarbor

Run real bioinformatics from your AI agent. Reliable, reproducible, on your own GPUs.

[CI] [PyPI] [License]

🧪 Alpha (v0.1). Sequence tools, homology search (MMseqs2) and structure prediction (ESMFold, validated on RTX 5090) work. Feedback welcome — see the roadmap.

BioHarbor is an MCP server that lets AI agents such as Claude, Cursor and Codex execute bioinformatics tools — not just look things up. Agents ask for an analysis; BioHarbor validates the input, schedules it on a GPU with room, records exactly how it ran, and hands back a compact, agent-readable summary.

[BioHarbor demo]

Why another bio MCP server?

Most bio MCP servers wrap databases (UniProt, PDB, PubMed…). Use them — BioHarbor complements them by running the compute:

Database MCP serversBioHarbor
Runs real analyses (search, fold, cluster)❌✅
Validates inputs before burning GPU time❌✅
GPU-aware queue, polite on shared GPUs❌✅
Long jobs return a job_id instead of timing out❌✅
Compact summaries + files on disk (saves tokens)❌✅
Provenance for every run, export to a pipeline❌✅ (export: planned)

Quick start

bash
pip install bioharborbioharbor doctor             # checks Python, GPUs, workspace, toolsbioharbor setup-db swissprot # reference database for search_homologs (needs MMseqs2)bioharbor install            # shows how to connect Claude, Cursor or Codex

For structure prediction on a GPU: pip install "bioharbor[esmfold]" — see docs/gpu-setup.md (RTX 50xx needs a CUDA 12.8+ PyTorch).

Connect your agent

BioHarbor is a standard MCP server, so it works with any MCP client. One command sets up the popular ones (it writes an absolute path, so GUI apps find it even outside your venv):

ClientSet up
Claude Codeclaude mcp add bioharbor -- bioharbor serve
Claude Desktopbioharbor install claude-desktop --write, then restart the app
Cursorbioharbor install cursor --write, or [Add to Cursor]
Codex (CLI, IDE extension, app)codex mcp add bioharbor -- bioharbor serve, or bioharbor install codex --write
Anything elserun bioharbor serve (stdio) or bioharbor serve --http (Streamable HTTP)

Long-running tools return a job_id within ~20 s instead of blocking, so they stay within every client's tool-call timeout.

[Codex chaining BioHarbor's find_orfs and seq_stats tools on a DNA sequence]

Step-by-step setup (local or on a GPU server, with troubleshooting): docs/connect-clients.md.

Then ask your agent something like:

Find the longest ORF in this contig, translate it, search Swiss-Prot for homologs and predict its structure. Which regions are low confidence?

Use it without an agent

Every tool is also a CLI command, with identical behaviour:

bash
bioharbor tools listbioharbor run find_orfs [email protected] min_aa=100 --brief   # human-readablebioharbor run seq_stats sequence=MKTAYIAKQRQISFVKSHFSRQbioharbor jobs

Shared GPU server

GPUs on a lab server, agent on your laptop? Run BioHarbor on the server and reach it through an SSH tunnel; no extra port is opened on the server:

bash
# on the GPU serverbioharbor serve --http --host 127.0.0.1 --port 8765# on your laptop, then point Cursor / Codex / Claude Code at http://127.0.0.1:8765/mcpssh -N -L 8765:127.0.0.1:8765 you@gpu-server

⚠️ HTTP mode has no authentication yet (on the roadmap), so keep it on 127.0.0.1 and use the tunnel. Details: docs/connect-clients.md.

Tools

ToolWhat it doesRuns
seq_statsValidate sequences; type, length, GC%, molecular weightinline
translate_sequenceDNA/RNA → protein, one or all six framesinline
find_orfsLongest ORFs on both strands, with coordinatesinline
search_homologsMMseqs2 search (protein, or translated DNA) vs local DBsjob
predict_structureESMFold structure, pLDDT bands, low-confidence regions, pTMjob (GPU)
scrna_pipelinescanpy QC → clustering → markersplanned

Runtime tools: get_job, list_jobs, cancel_job, describe_tool, list_databases, gpu_status, read_file.

How it works

Agent ──MCP──▶ validate input ─▶ inline? ──yes──▶ run ─┐                                   │ no                  ├─▶ provenance + summary ─▶ Agent                                   ▼                     │                     job queue (SQLite) ─▶ GPU placement ┘                     (waits politely for a GPU with free memory)
  • Every call is a job recorded in SQLite with params, versions, timings and GPU used, plus a provenance.json next to its outputs.
  • GPU placement reads live free memory and utilisation (NVML or nvidia-smi), keeps headroom, and reserves memory for jobs it has started so two jobs never grab the same space. Other users' processes are respected.
  • Fail fast: input, binaries and databases are checked before a job is queued, so a bad request never waits behind a busy GPU.
  • Results are agent-shaped: summary, message, files, suggestions. Errors carry a hint and a retryable flag.

Details: docs/design.md.

Writing a tool

python
from pydantic import BaseModel, Fieldfrom bioharbor.registry import Resources, RunContext, toolfrom bioharbor.results import ToolResult

class FoldParams(BaseModel):    sequence: str = Field(..., description="Protein sequence")

@tool(    version="1",    slow=True,    resources=Resources(gpu=True, gpu_mem_gb=lambda p: 4 + len(p.sequence) / 100),)def predict_structure(params: FoldParams, ctx: RunContext) -> ToolResult:    """Predict a protein structure with ESMFold."""    ...    return ToolResult(summary={"mean_plddt": 87.1}, files=["model.pdb"])

Plugins can ship tools in their own package via the bioharbor.tools entry-point group. See CONTRIBUTING.md.

License

Apache-2.0

来源:README.md,提交 f0a0884

工具

0
工具元数据尚未被收录。

版本历史

1
  1. v0.1.4最新Oct 7, 2026