
BioHarbor
io.github.danielLuo2v0.1.4Updated Oct 7, 2026
Run real bioinformatics from your AI agent. Reliable, reproducible, on your own GPUs.
Overview
Lets an AI assistant run real bioinformatics compute — sequence stats, translation, ORF finding, MMseqs2 homology search and ESMFold structure prediction — on…
- What it does
- BioHarbor exposes bioinformatics tools that an agent can execute rather than just query. Inline tools cover sequence validation and stats, DNA/RNA translation in one or all six frames, and longest-ORF finding on both strands. Longer work runs as queued jobs: MMseqs2 homology search against local databases and ESMFold structure prediction with pLDDT bands, low-confidence regions and pTM. Runtime tools such as get_job, list_jobs, cancel_job, describe_tool, list_databases, gpu_status and read_file manage those jobs, and every run is recorded in SQLite with provenance.
- When to use it
- Use it when an assistant should actually perform sequence analysis, homology search or protein structure prediction instead of only looking up records in biological databases. It suits labs with their own GPU hardware, including a shared GPU server reached over an SSH tunnel, and workflows where reproducible, provenance-tracked runs matter.
- Requirements
- A local Python environment with the bioharbor PyPI package installed via pip; the manifest declares no authentication, environment variables or headers. Homology search needs MMseqs2 and a reference database set up with setup-db. GPU structure prediction needs the esmfold extra and a CUDA 12.8+ PyTorch build on RTX 50xx cards. Desktop clients only; HTTP mode is optional and unauthenticated.
Installation
In SourceWeft
- Open BioHarbor in the dashboard and add it to a workspace.
- Enable the server for the chats that should use its tools.
Desktop only via STDIO. STDIO servers start a local process, so they need the SourceWeft desktop host.
Other MCP clients
Follow the launch instructions in the repository.
README
BioHarbor
Run real bioinformatics from your AI agent. Reliable, reproducible, on your own GPUs.
🧪 Alpha (v0.1). Sequence tools, homology search (MMseqs2) and structure prediction (ESMFold, validated on RTX 5090) work. Feedback welcome — see the roadmap.
BioHarbor is an MCP server that lets AI agents such as Claude, Cursor and Codex execute bioinformatics tools — not just look things up. Agents ask for an analysis; BioHarbor validates the input, schedules it on a GPU with room, records exactly how it ran, and hands back a compact, agent-readable summary.
Why another bio MCP server?
Most bio MCP servers wrap databases (UniProt, PDB, PubMed…). Use them — BioHarbor complements them by running the compute:
Quick start
For structure prediction on a GPU: pip install "bioharbor[esmfold]" — see
docs/gpu-setup.md (RTX 50xx needs a CUDA 12.8+ PyTorch).
Connect your agent
BioHarbor is a standard MCP server, so it works with any MCP client. One command sets up the popular ones (it writes an absolute path, so GUI apps find it even outside your venv):
Long-running tools return a job_id within ~20 s instead of blocking, so they stay
within every client's tool-call timeout.
Step-by-step setup (local or on a GPU server, with troubleshooting): docs/connect-clients.md.
Then ask your agent something like:
Find the longest ORF in this contig, translate it, search Swiss-Prot for homologs and predict its structure. Which regions are low confidence?
Use it without an agent
Every tool is also a CLI command, with identical behaviour:
Shared GPU server
GPUs on a lab server, agent on your laptop? Run BioHarbor on the server and reach it through an SSH tunnel; no extra port is opened on the server:
⚠️ HTTP mode has no authentication yet (on the roadmap), so keep it on
127.0.0.1and use the tunnel. Details: docs/connect-clients.md.
Tools
Runtime tools: get_job, list_jobs, cancel_job, describe_tool, list_databases,
gpu_status, read_file.
How it works
- Every call is a job recorded in SQLite with params, versions, timings and GPU used,
plus a
provenance.jsonnext to its outputs. - GPU placement reads live free memory and utilisation (NVML or
nvidia-smi), keeps headroom, and reserves memory for jobs it has started so two jobs never grab the same space. Other users' processes are respected. - Fail fast: input, binaries and databases are checked before a job is queued, so a bad request never waits behind a busy GPU.
- Results are agent-shaped:
summary,message,files,suggestions. Errors carry ahintand aretryableflag.
Details: docs/design.md.
Writing a tool
Plugins can ship tools in their own package via the bioharbor.tools entry-point group.
See CONTRIBUTING.md.
License
Source: README.md at commit f0a0884
Tools
0Version history
1- v0.1.4LatestOct 7, 2026


