BioHarbor

io.github.danielLuo2v0.1.4更新於 Oct 7, 2026

Run real bioinformatics from your AI agent. Reliable, reproducible, on your own GPUs.

已驗證STDIO僅桌面AI & MLData & Analytics

概覽

AI 產生的概覽

讓 AI 助理在自己的 GPU 上實際執行生物資訊運算:序列統計、轉譯、ORF 尋找、MMseqs2 同源搜尋與 ESMFold 結構預測。

功能
BioHarbor 提供的生物資訊工具讓代理真正執行分析,而不只是查詢。內嵌工具涵蓋序列驗證與統計、DNA/RNA 單一框架或六框架轉譯,以及雙股最長 ORF 尋找。耗時工作以佇列作業執行:對本機資料庫的 MMseqs2 同源搜尋,以及 ESMFold 結構預測,回傳 pLDDT 區段、低信心區域與 pTM。get_job、list_jobs、cancel_job、describe_tool、list_databases、gpu_status、read_file 等執行期工具負責管理這些作業,每次執行都會連同溯源資訊記錄到 SQLite。
適用情境
當助理需要真正執行序列分析、同源搜尋或蛋白質結構預測,而不只是查詢生物資料庫紀錄時適用。適合擁有自有 GPU 硬體的實驗室,包括透過 SSH 通道存取共用 GPU 伺服器的情境,也適合需要可重現、具溯源紀錄之執行流程的工作。
執行需求
需要本機 Python 環境,並以 pip 安裝 bioharbor 這個 PyPI 套件;清單宣告不需要驗證、環境變數或標頭。同源搜尋需要 MMseqs2 以及用 setup-db 設定的參考資料庫。GPU 結構預測需要 esmfold 附加元件,RTX 50xx 顯示卡還需 CUDA 12.8+ 的 PyTorch。僅支援桌面用戶端;HTTP 模式為選用且無驗證。
安裝前請注意
專案處於 alpha(v0.1)階段,scrna_pipeline 等部分工具仍在規劃中。HTTP 模式沒有驗證,應維持綁定 127.0.0.1 並透過 SSH 通道存取。作業會將檔案與溯源紀錄寫入磁碟,GPU 作業會佔用共用顯示記憶體與機時,在多使用者機器上需留意排程與取消行為。

安裝

在 SourceWeft 中

  1. 開啟 儀表板中的 BioHarbor,將其新增到工作區。
  2. 為需要使用其工具的對話啟用該服務。

Desktop only,透過 STDIO。 STDIO 服務會啟動本機處理程序,因此需要 SourceWeft 桌面主機。

其他 MCP 客戶端

參照 儲存庫 中的啟動說明。

README

BioHarbor

Run real bioinformatics from your AI agent. Reliable, reproducible, on your own GPUs.

[CI] [PyPI] [License]

🧪 Alpha (v0.1). Sequence tools, homology search (MMseqs2) and structure prediction (ESMFold, validated on RTX 5090) work. Feedback welcome — see the roadmap.

BioHarbor is an MCP server that lets AI agents such as Claude, Cursor and Codex execute bioinformatics tools — not just look things up. Agents ask for an analysis; BioHarbor validates the input, schedules it on a GPU with room, records exactly how it ran, and hands back a compact, agent-readable summary.

[BioHarbor demo]

Why another bio MCP server?

Most bio MCP servers wrap databases (UniProt, PDB, PubMed…). Use them — BioHarbor complements them by running the compute:

Database MCP serversBioHarbor
Runs real analyses (search, fold, cluster)❌✅
Validates inputs before burning GPU time❌✅
GPU-aware queue, polite on shared GPUs❌✅
Long jobs return a job_id instead of timing out❌✅
Compact summaries + files on disk (saves tokens)❌✅
Provenance for every run, export to a pipeline❌✅ (export: planned)

Quick start

bash
pip install bioharborbioharbor doctor             # checks Python, GPUs, workspace, toolsbioharbor setup-db swissprot # reference database for search_homologs (needs MMseqs2)bioharbor install            # shows how to connect Claude, Cursor or Codex

For structure prediction on a GPU: pip install "bioharbor[esmfold]" — see docs/gpu-setup.md (RTX 50xx needs a CUDA 12.8+ PyTorch).

Connect your agent

BioHarbor is a standard MCP server, so it works with any MCP client. One command sets up the popular ones (it writes an absolute path, so GUI apps find it even outside your venv):

ClientSet up
Claude Codeclaude mcp add bioharbor -- bioharbor serve
Claude Desktopbioharbor install claude-desktop --write, then restart the app
Cursorbioharbor install cursor --write, or [Add to Cursor]
Codex (CLI, IDE extension, app)codex mcp add bioharbor -- bioharbor serve, or bioharbor install codex --write
Anything elserun bioharbor serve (stdio) or bioharbor serve --http (Streamable HTTP)

Long-running tools return a job_id within ~20 s instead of blocking, so they stay within every client's tool-call timeout.

[Codex chaining BioHarbor's find_orfs and seq_stats tools on a DNA sequence]

Step-by-step setup (local or on a GPU server, with troubleshooting): docs/connect-clients.md.

Then ask your agent something like:

Find the longest ORF in this contig, translate it, search Swiss-Prot for homologs and predict its structure. Which regions are low confidence?

Use it without an agent

Every tool is also a CLI command, with identical behaviour:

bash
bioharbor tools listbioharbor run find_orfs [email protected] min_aa=100 --brief   # human-readablebioharbor run seq_stats sequence=MKTAYIAKQRQISFVKSHFSRQbioharbor jobs

Shared GPU server

GPUs on a lab server, agent on your laptop? Run BioHarbor on the server and reach it through an SSH tunnel; no extra port is opened on the server:

bash
# on the GPU serverbioharbor serve --http --host 127.0.0.1 --port 8765# on your laptop, then point Cursor / Codex / Claude Code at http://127.0.0.1:8765/mcpssh -N -L 8765:127.0.0.1:8765 you@gpu-server

⚠️ HTTP mode has no authentication yet (on the roadmap), so keep it on 127.0.0.1 and use the tunnel. Details: docs/connect-clients.md.

Tools

ToolWhat it doesRuns
seq_statsValidate sequences; type, length, GC%, molecular weightinline
translate_sequenceDNA/RNA → protein, one or all six framesinline
find_orfsLongest ORFs on both strands, with coordinatesinline
search_homologsMMseqs2 search (protein, or translated DNA) vs local DBsjob
predict_structureESMFold structure, pLDDT bands, low-confidence regions, pTMjob (GPU)
scrna_pipelinescanpy QC → clustering → markersplanned

Runtime tools: get_job, list_jobs, cancel_job, describe_tool, list_databases, gpu_status, read_file.

How it works

Agent ──MCP──▶ validate input ─▶ inline? ──yes──▶ run ─┐                                   │ no                  ├─▶ provenance + summary ─▶ Agent                                   ▼                     │                     job queue (SQLite) ─▶ GPU placement ┘                     (waits politely for a GPU with free memory)
  • Every call is a job recorded in SQLite with params, versions, timings and GPU used, plus a provenance.json next to its outputs.
  • GPU placement reads live free memory and utilisation (NVML or nvidia-smi), keeps headroom, and reserves memory for jobs it has started so two jobs never grab the same space. Other users' processes are respected.
  • Fail fast: input, binaries and databases are checked before a job is queued, so a bad request never waits behind a busy GPU.
  • Results are agent-shaped: summary, message, files, suggestions. Errors carry a hint and a retryable flag.

Details: docs/design.md.

Writing a tool

python
from pydantic import BaseModel, Fieldfrom bioharbor.registry import Resources, RunContext, toolfrom bioharbor.results import ToolResult

class FoldParams(BaseModel):    sequence: str = Field(..., description="Protein sequence")

@tool(    version="1",    slow=True,    resources=Resources(gpu=True, gpu_mem_gb=lambda p: 4 + len(p.sequence) / 100),)def predict_structure(params: FoldParams, ctx: RunContext) -> ToolResult:    """Predict a protein structure with ESMFold."""    ...    return ToolResult(summary={"mean_plddt": 87.1}, files=["model.pdb"])

Plugins can ship tools in their own package via the bioharbor.tools entry-point group. See CONTRIBUTING.md.

License

Apache-2.0

來源:README.md,提交 f0a0884

工具

0
工具後設資料尚未被收錄。

版本歷史

1
  1. v0.1.4最新Oct 7, 2026