Benzi

net.varianttechv0.1.1更新於 Oct 2, 2026

Compiler-backed code intelligence: calls, data flow, references and class hierarchy as tools.

已驗證STDIO僅桌面Developer ToolsAI & ML

概覽

AI 產生的概覽

Benzi 在本機編譯程式碼索引,並透過 MCP 提供給助理,用來查詢符號、呼叫者、資料流與類別階層。

功能
Benzi 使用以 tree-sitter 為基礎的編譯器,將程式碼庫解析成可查詢的索引,包含符號、呼叫邊、參照、繼承與資料流,再透過 MCP 工具提供該索引。範例工具包括 get_callers、call_tree、trace_path、forwardflow、backflow、profile、get_definition、search_symbols、get_hierarchy、skim_source 與 execute_from。它也提供受閘控的寫入(解析失敗自動回復)、以快照為基礎的回復,以及僅限 Python 的執行時期追蹤器。
適用情境
當助理需要整個程式碼庫的結構性答案,而非 grep 或嵌入搜尋時使用,例如追蹤某個值的來源、尋找函式的所有呼叫者,或對應類別階層。它鎖定 VS Code、無介面 CLI 或現有 MCP 用戶端中的本機開發工作流程。
執行需求
以本機 stdio MCP 伺服器執行。透過 pip 或 uvx 安裝(benzi-mcp)。需要透過 benzi_login 設定模型供應商的 API 金鑰;README 說明採用 BYOK,CLI 與 MCP 免費,但須自行支付模型供應商費用。清單中未宣告環境變數或標頭。
安裝前請注意
MCP 伺服器需要透過 benzi_login 使用您自己的模型供應商金鑰進行驗證,因此涉及憑證。它可以寫入與回復目標儲存庫中的檔案,執行時期追蹤器會執行程式碼。代理讀取的程式碼片段會傳送到您的模型供應商;編譯器與索引保留在本機。各語言支援深度不一,執行時期追蹤器僅支援 Python。

安裝

在 SourceWeft 中

  1. 開啟 儀表板中的 Benzi,將其新增到工作區。
  2. 為需要使用其工具的對話啟用該服務。

Desktop only,透過 STDIO。 STDIO 服務會啟動本機處理程序,因此需要 SourceWeft 桌面主機。

其他 MCP 客戶端

參照 儲存庫 中的啟動說明。

README

[圖片]

[Benzi]

Language Agnostic · Model Agnostic · Compiles Locally · BYOK · MCP Compatible

[PyPI version] [Python versions] [MCP compatible] [License]

[Benzi -- compiler-backed code intelligence]

[Try now]

[Download]      [FAQ]      [Benchmarks]

[圖片]

Contents

What is Benzi — how it works in one paragraph
Live demos — StallionSwipe, VS Code's own source, or any repo you paste
What people say — what people wrote about it
SWE-bench Verified — 391/500 (78.2%) for $37.33
How it works — compile, query, edit, verify
Tools — 16 of the 35+ the index makes possible
What the index actually changes — lines read vs three other harnesses
Features · Language support · Getting started
FAQ — privacy, API keys, pricing, limits

[圖片]

What is Benzi

Most AI coding agents dump a repo into a context window and hope the model finds what matters. Benzi parses every file first — a real compiler, built on tree-sitter — into a precise, queryable map. Every symbol, every call edge, every reference, every class in its inheritance chain. One pass, done.

Every file parsed, imports resolved, class ancestry built, every identifier traced to its definition — before a single question is answered. Call flow and data flow join at every call site, so a bad value traces to its origin in one tool call. Claude Code greps; Cursor embeds; Aider maps signatures; Benzi resolves — in O(1). Ten languages so far, plus a markup engine for HTML, CSS, and DOM-JS — see Language support.

You can try pasting this repo's link to Benzi in the live demo too!

[Benzi on SQLite]
Benzi on SQLite

[圖片]

Live demos

StallionSwipe · Python, HTML, CSS, JS — a dating app for horses, greenfielded by Benzi from scratch in a single chat session. No image is a file: every horse portrait is procedural SVG, generated in code. Match with one and it flirts back through a real model, live. Frontend, backend, and the prompts — all written by Benzi. Try it live.

[StallionSwipe swipe deck] [StallionSwipe profile detail] [StallionSwipe live AI chat] [StallionSwipe profile creation]

VS Code's own source, resolved · TypeScript — the real microsoft/vscode repo is 1.8M lines; this indexes 923k of them: the editor core (src/vs/editor + src/vs/base), the platform services layer, and workbench's shell/API/browser plumbing — deliberately excluding the 747k-line grab-bag of individual features in workbench/contrib. Built once, in just over two minutes, then cached. Try it live (chat panel, near the bottom of the page).

Or, try any repo of your choice at all here — point Benzi at any public GitHub repo and it builds the index live. varianttech.net/demo.

[圖片]

What people say

"78.2% for $37 is a slap in the face to the 'brute force wins' school." — Alex Xiang, zicode (translated)

"Benzi is proving that the core competency of coding tools is shifting from simple 'reading comprehension' to 'structural grasping ability.'" — Gi Pyeong Lee, Tech Blog

"Fewer tokens, no context drift. Wild idea, honestly." — prompt 🤖 AI News

"It analyzes changes before writing them — and beats Claude Code on benchmarks." — Ponte al dIA (translated)

[圖片]

SWE-bench Verified

The full SWE-bench Verified set — 500 real GitHub issues from twelve Python repositories — run end to end on DeepSeek v4-flash, one attempt per instance, graded by the official swebench.harness.run_evaluation inside its own per-instance Docker images, with network access to GitHub and PyPI blocked inside every container.

Resolved391 / 500 — 78.2%
Total cost, all 500 instances$37.33
Cost per instance resolved$0.095
Source lines read (total / median)231,574 / 379
Model turns (total / median)16,091 / 27
Input tokens served from cache97%
Output tokens22.0M

Full technical report: swebench/SWE_BENCH_REPORT.md (web version). Every instance's cost, tokens, turns, and lines read: varianttech.net/benchmark_swebench. The cross-harness efficiency comparison below (and the full 24-bug chart): varianttech.net/benchmark.

[圖片]

How it works

  1. Compile. Tree-sitter parses every file, resolves imports, builds class ancestry, traces every identifier to its definition. Output is an index, not text.
  2. Query. The agent reads and plans through structured tools over that index — profile, get_callers, backflow, trace_path, skim_source, ~30 more.
  3. Edit, gated. Every write is checked against the real parser; a broken parse auto-reverts. Blast radius — the changed symbol, its callers, its holders, the relevant tests — is checked going in and again once the write lands.
  4. Verify by running it. A focused repro runs under a runtime tracer alongside the tests blast-radius flagged. Real values, real dispatch: this proves the change and settles the map — ambiguous edges collapse onto whatever target actually fired.
  5. Reindex, incrementally. Every turn re-parses only what changed on disk — your edits and the agent's, treated the same.
  6. Revert, from a snapshot. Every write snapshots the index first, so undo reloads that snapshot instead of re-deriving it.
  7. Back to step 2. The next question always answers against the code as it is now.
[圖片]

Tools

A sample of 16 of Benzi's 35+ tools — what falls out of actually resolving the code, from the index itself to the gates on every write.

ToolWhat it answers
get_callersEvery call site that reaches a function — the code that will feel a change.
call_treeThe transitive call closure from one function, forward or in reverse.
trace_pathThe call chain connecting two functions, and the data carried along it.
external_callsWhich libraries a scope leans on, and where it calls into them.
forwardflowWhere a function's return value ends up, everywhere it has to match.
backflowWhere a wrong value came from, without opening every caller.
profileThe full 360 on one symbol in a single call.
get_definitionThe declaration card — signature, docs and location.
search_symbolsCase-insensitive substring search across every symbol in the repo.
get_hierarchyA type's resolved bases and its direct subclasses.
skim_sourceA body's one-level outline, so you know which lines are worth reading.
execute_fromRuns a file under the call tracer and records what actually happened.
check_last_executionReads back the last recorded run's facts, no re-run needed.
execute_generated_testcaseWrites a self-contained repro and runs it to debug its own change.
rollback_editUndoes the last writes by snapshot reload, not by re-editing.
upgrade_to_proEscalates itself to a larger reasoning budget mid-task.
[圖片]

What the index actually changes

Same 24 bugs, one run each, four harness/model combinations. Lines read counts only what came back from file-read calls — grep and shell output are search, not reading, so this is the one figure that means the same thing in every harness.

Harness · modelLines readvs Benzi
Benzi · Sonnet9,125—
Benzi · DeepSeek16,4071.8×
Claude Code · Sonnet20,7042.3×
DeepSeek Harness · DeepSeek43,5984.8×

Every harness opens more source as bugs get harder — the question is the slope. Benzi's stays flatter because it answers most of what a bug needs from the map instead of by reading.

[Source lines read per bug, all four harnesses]

Benzi reads the least source on every bug and the gap widens as bugs get harder — the index answers most of what a fix needs before a file is ever opened.

[Wall-clock time per bug, all four harnesses]

Wall-clock time tracks close across all four — reading less doesn't make Benzi slower to think, just cheaper to look.

[Cost per fix, all four harnesses]

Benzi on DeepSeek costs about a cent a bug; Claude Code climbs to $0.18 a step as bugs get harder — roughly 18x.

More detail, per-bug breakdowns, and full methodology: varianttech.net/benchmark.

[圖片]

Features

  • Six states, never a guess — every call site and every file carries one: resolved (proven in-repo edge), external (into a library, with the import evidence), candidate (ambiguous — the bounded set of possible targets, kept in full), unresolved (seen but not settled, carrying why), observed (confirmed by an actual run), unindexed (never parsed, with the reason). One rule throughout: whatever static analysis can't settle is flagged as unsettled rather than guessed — and running the program is what settles it.
  • Runtime tracer — hooks every call during execution and overlays the observations back onto the static map.
  • Reasoning you can click — the same map that drives the tools drives a live call graph beside the chat; when the agent names a function, that node lights up.
  • Persistent memory — durable per-repo facts survive restarts; conventions learned once aren't re-derived every session.
  • Dual-engine: code + markup — a separate index for HTML/CSS/DOM-JS with cascade resolution and selector specificity, including frontend embedded inside Python strings.
  • Model-agnostic — Anthropic, OpenAI, or any compatible API; the agent can escalate itself to a larger model mid-task when a problem outgrows the one running it.
[圖片]

Language support

Python · JavaScript · TypeScript · Java · C# · C++ · C · Go · Rust · Ruby

One compiler, ten languages — each is a tree-sitter grammar plugin, so the core of the map (symbols, call edges, references, inheritance, data flow) is built the same way everywhere.

Depth is uneven, and we'd rather say so than let you find out. Python is deepest, and the only one with the runtime tracer. Every language reaches the core of the map, but each has its own constructs, not all modelled yet — a question specific to your language may come back thinner than the same question in Python.

Incremental reindexing is also less optimized for C, C++, Rust, and Ruby — it works, just not as fast on a large edit loop.

Wrong or thin answer in your language? Open an issue with the repo, the question, and what it got wrong — that's how the uneven parts get found.

[圖片]

Getting started

Benzi is completely free to use.

  • In the browser — paste any public GitHub repo at varianttech.net/demo; no install, no signup. Read-only: ask it questions, explore the map, nothing writes to the repo. This is the demo — click here to see what it can do.
  • In VS Code — the same compiler, but with edit access: chat, graph, and Benzi actually writing code in your own project. VS Code Marketplace. This is the real tool — click here to use it.
  • MCP — the same compiled index, exposed as tools over MCP for whatever agent you already run: Claude Code, Cursor, or your own harness. pip install benzi, then point your MCP client at benzi-mcp. This is the benzi index without the agentic loop — output quality will depend on your agent/harness.
  • Headless — the same agent as VS Code, from your own terminal: pip install benzi, then benzi <repo> "your question". This is Benzi for scripts and CI — no editor needed. (https://pypi.org/project/benzi/)

Run benzi_login once to authenticate before using the VS Code extension, MCP, or headless — the same command lets you update your model or key again later too.

[圖片]

FAQ

Does my code leave my machine? No. In VS Code, the CLI, or MCP, the compiler and index run locally — nothing uploaded, no copy kept. Only the snippets the agent actually reads go to your model provider, same as any AI assistant, and less of them: 9,125 lines read vs Claude Code's 20,704 on the same 24 bugs. The browser demo differs — it fetches a public repo server-side, read-only, deletes it after your session.

Do I need an API key? No, for the browser demo. Yes for VS Code, MCP, and the CLI — run benzi_login once with your own key.

Is it actually free? Yes, Benzi doesn't charge. The CLI and MCP are BYOK, so you pay your own model provider. Early and in development — that's the trade, not a paywall.

Can I point it at a private repo? Not the web demo (public GitHub API only). Everywhere else, yes — the compiler runs locally on whatever path you give it.

How large a repo can it handle? VS Code handles real codebases — microsoft/vscode, 923k lines, indexes in ~2 minutes, then caches. The browser demo caps at 2,000 files, 2 MB each.

How is this different from Cursor, Copilot, or Claude Code? They search — grep or embeddings. Benzi resolves first: a real index of symbols, calls, inheritance, data flow, queried instead of guessed. Same 24 bugs, 2.3× less source read than Claude Code. Details: what the index actually changes.

How is this different from an LSP-backed MCP server? An LSP answers at a cursor, in one open file: go-to-definition or find-references, one position and one hop at a time. Benzi compiles the whole repo up front into one index, so the questions are whole-codebase ones: transitive call trees, the path between two functions, and data flow, meaning where a bad value came from or where a return value lands. Every answer also carries a confidence tier: resolved, candidate, unresolved (with the reason), or observed. An LSP gives an answer or nothing. The runtime tracer then settles what static analysis can't by watching what actually fires. And the compiler itself is language agnostic: ten languages run through one pipeline into one index format. An LSP setup needs a separate server per language, each installed, configured and kept running.

How is this different from CodeQL or Sourcegraph's SCIP indexers? They need a working build and index in batch: dependencies installed, the project compiling, a CI job measured in minutes. Benzi needs no build, works on half-finished code, and re-parses only what changed on every turn. That's what makes gated writes possible: each edit is checked against a fresh index before the next step, not after a rebuild. One pipeline covers all ten languages instead of one indexer per language. Where that costs precision, Benzi flags the call site as candidate or unresolved instead of guessing.

How is this different from CodeGraph? Both index instead of search, but CodeGraph retrieves — ranked candidates from a queried database. Benzi resolves — settles what a name binds to before answering, and refuses rather than guesses when a call site is ambiguous. It also models code the way an engineer reads it — file → scopes → call flow → data/control flow — not a flat symbol graph.

On CodeGraph's own benchmark (their repos, their questions, their methodology), four models blind-judged Benzi's MCP answers first. Full results: varianttech.net/benchmark_codegraph.

My language isn't Python — how much do I lose? The structural index — symbols, calls, references, inheritance, data flow — is the same across all ten languages. Only the runtime tracer is Python-only, and depth varies by language — see Language support.

[圖片]

來源:README.md,提交 1264f60

工具

0
工具後設資料尚未被收錄。

版本歷史

1
  1. v0.1.1最新Oct 2, 2026