
BioHarbor
io.github.danielLuo2v0.1.4更新於 Oct 7, 2026
Run real bioinformatics from your AI agent. Reliable, reproducible, on your own GPUs.
概覽
讓 AI 助理在自己的 GPU 上實際執行生物資訊運算:序列統計、轉譯、ORF 尋找、MMseqs2 同源搜尋與 ESMFold 結構預測。
- 功能
- BioHarbor 提供的生物資訊工具讓代理真正執行分析,而不只是查詢。內嵌工具涵蓋序列驗證與統計、DNA/RNA 單一框架或六框架轉譯,以及雙股最長 ORF 尋找。耗時工作以佇列作業執行:對本機資料庫的 MMseqs2 同源搜尋,以及 ESMFold 結構預測,回傳 pLDDT 區段、低信心區域與 pTM。get_job、list_jobs、cancel_job、describe_tool、list_databases、gpu_status、read_file 等執行期工具負責管理這些作業,每次執行都會連同溯源資訊記錄到 SQLite。
- 適用情境
- 當助理需要真正執行序列分析、同源搜尋或蛋白質結構預測,而不只是查詢生物資料庫紀錄時適用。適合擁有自有 GPU 硬體的實驗室,包括透過 SSH 通道存取共用 GPU 伺服器的情境,也適合需要可重現、具溯源紀錄之執行流程的工作。
- 執行需求
- 需要本機 Python 環境,並以 pip 安裝 bioharbor 這個 PyPI 套件;清單宣告不需要驗證、環境變數或標頭。同源搜尋需要 MMseqs2 以及用 setup-db 設定的參考資料庫。GPU 結構預測需要 esmfold 附加元件,RTX 50xx 顯示卡還需 CUDA 12.8+ 的 PyTorch。僅支援桌面用戶端;HTTP 模式為選用且無驗證。
安裝
在 SourceWeft 中
- 開啟 儀表板中的 BioHarbor,將其新增到工作區。
- 為需要使用其工具的對話啟用該服務。
Desktop only,透過 STDIO。 STDIO 服務會啟動本機處理程序,因此需要 SourceWeft 桌面主機。
其他 MCP 客戶端
參照 儲存庫 中的啟動說明。
README
BioHarbor
Run real bioinformatics from your AI agent. Reliable, reproducible, on your own GPUs.
🧪 Alpha (v0.1). Sequence tools, homology search (MMseqs2) and structure prediction (ESMFold, validated on RTX 5090) work. Feedback welcome — see the roadmap.
BioHarbor is an MCP server that lets AI agents such as Claude, Cursor and Codex execute bioinformatics tools — not just look things up. Agents ask for an analysis; BioHarbor validates the input, schedules it on a GPU with room, records exactly how it ran, and hands back a compact, agent-readable summary.
Why another bio MCP server?
Most bio MCP servers wrap databases (UniProt, PDB, PubMed…). Use them — BioHarbor complements them by running the compute:
Quick start
For structure prediction on a GPU: pip install "bioharbor[esmfold]" — see
docs/gpu-setup.md (RTX 50xx needs a CUDA 12.8+ PyTorch).
Connect your agent
BioHarbor is a standard MCP server, so it works with any MCP client. One command sets up the popular ones (it writes an absolute path, so GUI apps find it even outside your venv):
Long-running tools return a job_id within ~20 s instead of blocking, so they stay
within every client's tool-call timeout.
Step-by-step setup (local or on a GPU server, with troubleshooting): docs/connect-clients.md.
Then ask your agent something like:
Find the longest ORF in this contig, translate it, search Swiss-Prot for homologs and predict its structure. Which regions are low confidence?
Use it without an agent
Every tool is also a CLI command, with identical behaviour:
Shared GPU server
GPUs on a lab server, agent on your laptop? Run BioHarbor on the server and reach it through an SSH tunnel; no extra port is opened on the server:
⚠️ HTTP mode has no authentication yet (on the roadmap), so keep it on
127.0.0.1and use the tunnel. Details: docs/connect-clients.md.
Tools
Runtime tools: get_job, list_jobs, cancel_job, describe_tool, list_databases,
gpu_status, read_file.
How it works
- Every call is a job recorded in SQLite with params, versions, timings and GPU used,
plus a
provenance.jsonnext to its outputs. - GPU placement reads live free memory and utilisation (NVML or
nvidia-smi), keeps headroom, and reserves memory for jobs it has started so two jobs never grab the same space. Other users' processes are respected. - Fail fast: input, binaries and databases are checked before a job is queued, so a bad request never waits behind a busy GPU.
- Results are agent-shaped:
summary,message,files,suggestions. Errors carry ahintand aretryableflag.
Details: docs/design.md.
Writing a tool
Plugins can ship tools in their own package via the bioharbor.tools entry-point group.
See CONTRIBUTING.md.
License
來源:README.md,提交 f0a0884
工具
0版本歷史
1- v0.1.4最新Oct 7, 2026


