BioHarbor

io.github.danielLuo2v0.1.4Updated Oct 7, 2026

Run real bioinformatics from your AI agent. Reliable, reproducible, on your own GPUs.

VerifiedSTDIODesktop onlyAI & MLData & Analytics

Overview

AI-generated overview

Lets an AI assistant run real bioinformatics compute — sequence stats, translation, ORF finding, MMseqs2 homology search and ESMFold structure prediction — on…

What it does
BioHarbor exposes bioinformatics tools that an agent can execute rather than just query. Inline tools cover sequence validation and stats, DNA/RNA translation in one or all six frames, and longest-ORF finding on both strands. Longer work runs as queued jobs: MMseqs2 homology search against local databases and ESMFold structure prediction with pLDDT bands, low-confidence regions and pTM. Runtime tools such as get_job, list_jobs, cancel_job, describe_tool, list_databases, gpu_status and read_file manage those jobs, and every run is recorded in SQLite with provenance.
When to use it
Use it when an assistant should actually perform sequence analysis, homology search or protein structure prediction instead of only looking up records in biological databases. It suits labs with their own GPU hardware, including a shared GPU server reached over an SSH tunnel, and workflows where reproducible, provenance-tracked runs matter.
Requirements
A local Python environment with the bioharbor PyPI package installed via pip; the manifest declares no authentication, environment variables or headers. Homology search needs MMseqs2 and a reference database set up with setup-db. GPU structure prediction needs the esmfold extra and a CUDA 12.8+ PyTorch build on RTX 50xx cards. Desktop clients only; HTTP mode is optional and unauthenticated.
Before you install
The project is alpha (v0.1) and some tools, such as scrna_pipeline, are still planned. HTTP mode has no authentication, so it should stay bound to 127.0.0.1 and be reached through an SSH tunnel. Jobs write files and provenance records to disk, and GPU jobs consume shared GPU memory and time, so placement and cancellation behaviour matter on multi-user machines.

Installation

In SourceWeft

  1. Open BioHarbor in the dashboard and add it to a workspace.
  2. Enable the server for the chats that should use its tools.

Desktop only via STDIO. STDIO servers start a local process, so they need the SourceWeft desktop host.

Other MCP clients

Follow the launch instructions in the repository.

README

BioHarbor

Run real bioinformatics from your AI agent. Reliable, reproducible, on your own GPUs.

[CI] [PyPI] [License]

🧪 Alpha (v0.1). Sequence tools, homology search (MMseqs2) and structure prediction (ESMFold, validated on RTX 5090) work. Feedback welcome — see the roadmap.

BioHarbor is an MCP server that lets AI agents such as Claude, Cursor and Codex execute bioinformatics tools — not just look things up. Agents ask for an analysis; BioHarbor validates the input, schedules it on a GPU with room, records exactly how it ran, and hands back a compact, agent-readable summary.

[BioHarbor demo]

Why another bio MCP server?

Most bio MCP servers wrap databases (UniProt, PDB, PubMed…). Use them — BioHarbor complements them by running the compute:

Database MCP serversBioHarbor
Runs real analyses (search, fold, cluster)❌✅
Validates inputs before burning GPU time❌✅
GPU-aware queue, polite on shared GPUs❌✅
Long jobs return a job_id instead of timing out❌✅
Compact summaries + files on disk (saves tokens)❌✅
Provenance for every run, export to a pipeline❌✅ (export: planned)

Quick start

bash
pip install bioharborbioharbor doctor             # checks Python, GPUs, workspace, toolsbioharbor setup-db swissprot # reference database for search_homologs (needs MMseqs2)bioharbor install            # shows how to connect Claude, Cursor or Codex

For structure prediction on a GPU: pip install "bioharbor[esmfold]" — see docs/gpu-setup.md (RTX 50xx needs a CUDA 12.8+ PyTorch).

Connect your agent

BioHarbor is a standard MCP server, so it works with any MCP client. One command sets up the popular ones (it writes an absolute path, so GUI apps find it even outside your venv):

ClientSet up
Claude Codeclaude mcp add bioharbor -- bioharbor serve
Claude Desktopbioharbor install claude-desktop --write, then restart the app
Cursorbioharbor install cursor --write, or [Add to Cursor]
Codex (CLI, IDE extension, app)codex mcp add bioharbor -- bioharbor serve, or bioharbor install codex --write
Anything elserun bioharbor serve (stdio) or bioharbor serve --http (Streamable HTTP)

Long-running tools return a job_id within ~20 s instead of blocking, so they stay within every client's tool-call timeout.

[Codex chaining BioHarbor's find_orfs and seq_stats tools on a DNA sequence]

Step-by-step setup (local or on a GPU server, with troubleshooting): docs/connect-clients.md.

Then ask your agent something like:

Find the longest ORF in this contig, translate it, search Swiss-Prot for homologs and predict its structure. Which regions are low confidence?

Use it without an agent

Every tool is also a CLI command, with identical behaviour:

bash
bioharbor tools listbioharbor run find_orfs [email protected] min_aa=100 --brief   # human-readablebioharbor run seq_stats sequence=MKTAYIAKQRQISFVKSHFSRQbioharbor jobs

Shared GPU server

GPUs on a lab server, agent on your laptop? Run BioHarbor on the server and reach it through an SSH tunnel; no extra port is opened on the server:

bash
# on the GPU serverbioharbor serve --http --host 127.0.0.1 --port 8765# on your laptop, then point Cursor / Codex / Claude Code at http://127.0.0.1:8765/mcpssh -N -L 8765:127.0.0.1:8765 you@gpu-server

⚠️ HTTP mode has no authentication yet (on the roadmap), so keep it on 127.0.0.1 and use the tunnel. Details: docs/connect-clients.md.

Tools

ToolWhat it doesRuns
seq_statsValidate sequences; type, length, GC%, molecular weightinline
translate_sequenceDNA/RNA → protein, one or all six framesinline
find_orfsLongest ORFs on both strands, with coordinatesinline
search_homologsMMseqs2 search (protein, or translated DNA) vs local DBsjob
predict_structureESMFold structure, pLDDT bands, low-confidence regions, pTMjob (GPU)
scrna_pipelinescanpy QC → clustering → markersplanned

Runtime tools: get_job, list_jobs, cancel_job, describe_tool, list_databases, gpu_status, read_file.

How it works

Agent ──MCP──▶ validate input ─▶ inline? ──yes──▶ run ─┐                                   │ no                  ├─▶ provenance + summary ─▶ Agent                                   ▼                     │                     job queue (SQLite) ─▶ GPU placement ┘                     (waits politely for a GPU with free memory)
  • Every call is a job recorded in SQLite with params, versions, timings and GPU used, plus a provenance.json next to its outputs.
  • GPU placement reads live free memory and utilisation (NVML or nvidia-smi), keeps headroom, and reserves memory for jobs it has started so two jobs never grab the same space. Other users' processes are respected.
  • Fail fast: input, binaries and databases are checked before a job is queued, so a bad request never waits behind a busy GPU.
  • Results are agent-shaped: summary, message, files, suggestions. Errors carry a hint and a retryable flag.

Details: docs/design.md.

Writing a tool

python
from pydantic import BaseModel, Fieldfrom bioharbor.registry import Resources, RunContext, toolfrom bioharbor.results import ToolResult

class FoldParams(BaseModel):    sequence: str = Field(..., description="Protein sequence")

@tool(    version="1",    slow=True,    resources=Resources(gpu=True, gpu_mem_gb=lambda p: 4 + len(p.sequence) / 100),)def predict_structure(params: FoldParams, ctx: RunContext) -> ToolResult:    """Predict a protein structure with ESMFold."""    ...    return ToolResult(summary={"mean_plddt": 87.1}, files=["model.pdb"])

Plugins can ship tools in their own package via the bioharbor.tools entry-point group. See CONTRIBUTING.md.

License

Apache-2.0

Source: README.md at commit f0a0884

Tools

0
Tool metadata has not been indexed yet.

Version history

1
  1. v0.1.4LatestOct 7, 2026