
Benzi
net.varianttechv0.1.1Updated Oct 2, 2026
Compiler-backed code intelligence: calls, data flow, references and class hierarchy as tools.
Overview
Benzi compiles a local code index and exposes it over MCP so an assistant can query symbols, callers, data flow, and class hierarchies.
- What it does
- Benzi parses a repository with a tree-sitter-based compiler into a queryable index of symbols, call edges, references, inheritance, and data flow, then exposes that index as MCP tools. Sample tools include get_callers, call_tree, trace_path, forwardflow, backflow, profile, get_definition, search_symbols, get_hierarchy, skim_source, and execute_from. It also offers gated writes with auto-revert on broken parses, snapshot-based rollback, and a Python-only runtime tracer.
- When to use it
- Use it when an assistant needs whole-codebase structural answers rather than grep or embedding search, such as tracing where a value came from, finding every caller of a function, or mapping a class hierarchy. It is aimed at local development workflows in VS Code, headless CLI, or an existing MCP client.
- Requirements
- Runs locally as a stdio MCP server. Install via pip or uvx (benzi-mcp). Requires a model provider API key set up through benzi_login; the README says BYOK and that the CLI and MCP are free but you pay your own model provider. No environment variables or headers are declared in the manifest.
Installation
In SourceWeft
- Open Benzi in the dashboard and add it to a workspace.
- Enable the server for the chats that should use its tools.
Desktop only via STDIO. STDIO servers start a local process, so they need the SourceWeft desktop host.
Other MCP clients
Follow the launch instructions in the repository.
README
Language Agnostic · Model Agnostic · Compiles Locally · BYOK · MCP Compatible
[PyPI version] [Python versions] [MCP compatible] [License]
[Benzi -- compiler-backed code intelligence]
[Try now]
[Download]
[FAQ]
[Benchmarks]
Contents
What is Benzi — how it works in one paragraph
Live demos — StallionSwipe, VS Code's own source, or any repo you paste
What people say — what people wrote about it
SWE-bench Verified — 391/500 (78.2%) for $37.33
How it works — compile, query, edit, verify
Tools — 16 of the 35+ the index makes possible
What the index actually changes — lines read vs three other harnesses
Features · Language support · Getting started
FAQ — privacy, API keys, pricing, limits
What is Benzi
Most AI coding agents dump a repo into a context window and hope the model finds what matters. Benzi parses every file first — a real compiler, built on tree-sitter — into a precise, queryable map. Every symbol, every call edge, every reference, every class in its inheritance chain. One pass, done.
Every file parsed, imports resolved, class ancestry built, every identifier traced to its definition — before a single question is answered. Call flow and data flow join at every call site, so a bad value traces to its origin in one tool call. Claude Code greps; Cursor embeds; Aider maps signatures; Benzi resolves — in O(1). Ten languages so far, plus a markup engine for HTML, CSS, and DOM-JS — see Language support.
You can try pasting this repo's link to Benzi in the live demo too!
[Benzi on SQLite]
Benzi on SQLite
Live demos
StallionSwipe · Python, HTML, CSS, JS — a dating app for horses, greenfielded by Benzi from scratch in a single chat session. No image is a file: every horse portrait is procedural SVG, generated in code. Match with one and it flirts back through a real model, live. Frontend, backend, and the prompts — all written by Benzi. Try it live.
[StallionSwipe swipe deck] [StallionSwipe profile detail] [StallionSwipe live AI chat] [StallionSwipe profile creation]
VS Code's own source, resolved · TypeScript — the real microsoft/vscode repo is 1.8M lines; this indexes 923k of them: the editor core (src/vs/editor + src/vs/base), the platform services layer, and workbench's shell/API/browser plumbing — deliberately excluding the 747k-line grab-bag of individual features in workbench/contrib. Built once, in just over two minutes, then cached. Try it live (chat panel, near the bottom of the page).
Or, try any repo of your choice at all here — point Benzi at any public GitHub repo and it builds the index live. varianttech.net/demo.
[Image]What people say
"78.2% for $37 is a slap in the face to the 'brute force wins' school." — Alex Xiang, zicode (translated)
"Benzi is proving that the core competency of coding tools is shifting from simple 'reading comprehension' to 'structural grasping ability.'" — Gi Pyeong Lee, Tech Blog
"Fewer tokens, no context drift. Wild idea, honestly." — prompt 🤖 AI News
[Image]"It analyzes changes before writing them — and beats Claude Code on benchmarks." — Ponte al dIA (translated)
SWE-bench Verified
The full SWE-bench Verified set — 500 real GitHub issues from twelve Python repositories — run end to end on DeepSeek v4-flash, one attempt per instance, graded by the official swebench.harness.run_evaluation inside its own per-instance Docker images, with network access to GitHub and PyPI blocked inside every container.
Full technical report: swebench/SWE_BENCH_REPORT.md (web version). Every instance's cost, tokens, turns, and lines read: varianttech.net/benchmark_swebench. The cross-harness efficiency comparison below (and the full 24-bug chart): varianttech.net/benchmark.
[Image]How it works
- Compile. Tree-sitter parses every file, resolves imports, builds class ancestry, traces every identifier to its definition. Output is an index, not text.
- Query. The agent reads and plans through structured tools over that index —
profile,get_callers,backflow,trace_path,skim_source, ~30 more. - Edit, gated. Every write is checked against the real parser; a broken parse auto-reverts. Blast radius — the changed symbol, its callers, its holders, the relevant tests — is checked going in and again once the write lands.
- Verify by running it. A focused repro runs under a runtime tracer alongside the tests blast-radius flagged. Real values, real dispatch: this proves the change and settles the map — ambiguous edges collapse onto whatever target actually fired.
- Reindex, incrementally. Every turn re-parses only what changed on disk — your edits and the agent's, treated the same.
- Revert, from a snapshot. Every write snapshots the index first, so undo reloads that snapshot instead of re-deriving it.
- Back to step 2. The next question always answers against the code as it is now.
Tools
A sample of 16 of Benzi's 35+ tools — what falls out of actually resolving the code, from the index itself to the gates on every write.
What the index actually changes
Same 24 bugs, one run each, four harness/model combinations. Lines read counts only what came back from file-read calls — grep and shell output are search, not reading, so this is the one figure that means the same thing in every harness.
Every harness opens more source as bugs get harder — the question is the slope. Benzi's stays flatter because it answers most of what a bug needs from the map instead of by reading.
[Source lines read per bug, all four harnesses]
Benzi reads the least source on every bug and the gap widens as bugs get harder — the index answers most of what a fix needs before a file is ever opened.
[Wall-clock time per bug, all four harnesses]
Wall-clock time tracks close across all four — reading less doesn't make Benzi slower to think, just cheaper to look.
[Cost per fix, all four harnesses]
Benzi on DeepSeek costs about a cent a bug; Claude Code climbs to $0.18 a step as bugs get harder — roughly 18x.
More detail, per-bug breakdowns, and full methodology: varianttech.net/benchmark.
[Image]Features
- Six states, never a guess — every call site and every file carries one: resolved (proven in-repo edge), external (into a library, with the import evidence), candidate (ambiguous — the bounded set of possible targets, kept in full), unresolved (seen but not settled, carrying why), observed (confirmed by an actual run), unindexed (never parsed, with the reason). One rule throughout: whatever static analysis can't settle is flagged as unsettled rather than guessed — and running the program is what settles it.
- Runtime tracer — hooks every call during execution and overlays the observations back onto the static map.
- Reasoning you can click — the same map that drives the tools drives a live call graph beside the chat; when the agent names a function, that node lights up.
- Persistent memory — durable per-repo facts survive restarts; conventions learned once aren't re-derived every session.
- Dual-engine: code + markup — a separate index for HTML/CSS/DOM-JS with cascade resolution and selector specificity, including frontend embedded inside Python strings.
- Model-agnostic — Anthropic, OpenAI, or any compatible API; the agent can escalate itself to a larger model mid-task when a problem outgrows the one running it.
Language support
Python · JavaScript · TypeScript · Java · C# · C++ · C · Go · Rust · Ruby
One compiler, ten languages — each is a tree-sitter grammar plugin, so the core of the map (symbols, call edges, references, inheritance, data flow) is built the same way everywhere.
Depth is uneven, and we'd rather say so than let you find out. Python is deepest, and the only one with the runtime tracer. Every language reaches the core of the map, but each has its own constructs, not all modelled yet — a question specific to your language may come back thinner than the same question in Python.
Incremental reindexing is also less optimized for C, C++, Rust, and Ruby — it works, just not as fast on a large edit loop.
Wrong or thin answer in your language? Open an issue with the repo, the question, and what it got wrong — that's how the uneven parts get found.
[Image]Getting started
Benzi is completely free to use.
- In the browser — paste any public GitHub repo at varianttech.net/demo; no install, no signup. Read-only: ask it questions, explore the map, nothing writes to the repo. This is the demo — click here to see what it can do.
- In VS Code — the same compiler, but with edit access: chat, graph, and Benzi actually writing code in your own project. VS Code Marketplace. This is the real tool — click here to use it.
- MCP — the same compiled index, exposed as tools over MCP for whatever agent you already run: Claude Code, Cursor, or your own harness.
pip install benzi, then point your MCP client atbenzi-mcp. This is the benzi index without the agentic loop — output quality will depend on your agent/harness. - Headless — the same agent as VS Code, from your own terminal:
pip install benzi, thenbenzi <repo> "your question". This is Benzi for scripts and CI — no editor needed. (https://pypi.org/project/benzi/)
Run benzi_login once to authenticate before using the VS Code extension, MCP, or headless — the same command lets you update your model or key again later too.
FAQ
Does my code leave my machine? No. In VS Code, the CLI, or MCP, the compiler and index run locally — nothing uploaded, no copy kept. Only the snippets the agent actually reads go to your model provider, same as any AI assistant, and less of them: 9,125 lines read vs Claude Code's 20,704 on the same 24 bugs. The browser demo differs — it fetches a public repo server-side, read-only, deletes it after your session.
Do I need an API key?
No, for the browser demo. Yes for VS Code, MCP, and the CLI — run benzi_login once with your own key.
Is it actually free? Yes, Benzi doesn't charge. The CLI and MCP are BYOK, so you pay your own model provider. Early and in development — that's the trade, not a paywall.
Can I point it at a private repo? Not the web demo (public GitHub API only). Everywhere else, yes — the compiler runs locally on whatever path you give it.
How large a repo can it handle?
VS Code handles real codebases — microsoft/vscode, 923k lines, indexes in ~2 minutes, then caches. The browser demo caps at 2,000 files, 2 MB each.
How is this different from Cursor, Copilot, or Claude Code? They search — grep or embeddings. Benzi resolves first: a real index of symbols, calls, inheritance, data flow, queried instead of guessed. Same 24 bugs, 2.3× less source read than Claude Code. Details: what the index actually changes.
How is this different from an LSP-backed MCP server? An LSP answers at a cursor, in one open file: go-to-definition or find-references, one position and one hop at a time. Benzi compiles the whole repo up front into one index, so the questions are whole-codebase ones: transitive call trees, the path between two functions, and data flow, meaning where a bad value came from or where a return value lands. Every answer also carries a confidence tier: resolved, candidate, unresolved (with the reason), or observed. An LSP gives an answer or nothing. The runtime tracer then settles what static analysis can't by watching what actually fires. And the compiler itself is language agnostic: ten languages run through one pipeline into one index format. An LSP setup needs a separate server per language, each installed, configured and kept running.
How is this different from CodeQL or Sourcegraph's SCIP indexers? They need a working build and index in batch: dependencies installed, the project compiling, a CI job measured in minutes. Benzi needs no build, works on half-finished code, and re-parses only what changed on every turn. That's what makes gated writes possible: each edit is checked against a fresh index before the next step, not after a rebuild. One pipeline covers all ten languages instead of one indexer per language. Where that costs precision, Benzi flags the call site as candidate or unresolved instead of guessing.
How is this different from CodeGraph? Both index instead of search, but CodeGraph retrieves — ranked candidates from a queried database. Benzi resolves — settles what a name binds to before answering, and refuses rather than guesses when a call site is ambiguous. It also models code the way an engineer reads it — file → scopes → call flow → data/control flow — not a flat symbol graph.
On CodeGraph's own benchmark (their repos, their questions, their methodology), four models blind-judged Benzi's MCP answers first. Full results: varianttech.net/benchmark_codegraph.
My language isn't Python — how much do I lose? The structural index — symbols, calls, references, inheritance, data flow — is the same across all ten languages. Only the runtime tracer is Python-only, and depth varies by language — see Language support.
[Image]Source: README.md at commit 1264f60
Tools
0Version history
1- v0.1.1LatestOct 2, 2026

