
Crg
io.github.n24q02mv3.29.8更新於 Oct 5, 2026
Token-efficient code review knowledge graph: semantic search and call-graph resolution.
概覽
為程式碼庫建立本機 Tree-sitter 知識圖譜,讓助理能執行語意搜尋、呼叫圖查詢、影響分析與審查脈絡產生。
- 功能
- 以 Tree-sitter 將程式碼倉庫解析成函式、類別與匯入的結構化圖譜,並存於本機 SQLite 資料庫。六個分組工具涵蓋圖譜生命週期(建立、更新、統計、嵌入、匯出、摘要)、查詢(呼叫者、被呼叫者、匯入、影響範圍、差異、語意搜尋)、審查脈絡、設定、安全掃描與文件。語意搜尋預設使用本機 ONNX 嵌入模型,也可選用雲端嵌入與摘要鏈。
- 適用情境
- 當助理需要大型程式碼庫的精確脈絡、而非讀完整個檔案時適用:查找呼叫者或被呼叫者、評估變更的影響範圍、產生節省 token 的審查脈絡,或進行程式碼語意搜尋。也適合本機安全掃描與程式碼庫入門了解。
- 執行需求
- 透過 uvx 或 pip 以本機程序執行(Python 3.13);CLI 是主要使用方式,MCP stdio 為次要介面。圖譜狀態存於倉庫內的 .crg/graph.db。預設本機嵌入模型會在首次使用時下載(約 570 MB)。雲端嵌入或摘要需要對應供應商的金鑰,例如 COHERE_API_KEY 或 OPENROUTER_API_KEY,以及選用的 EMBEDDING_API_BASE 或 LLM_API_BASE。Semgrep 安全引擎需要額外安裝。
安裝
README
Better Code Review Graph
Renamed (2026-09-13): repo is now
crg— CLI-first (crgcommand). PyPI package stayscrg; MCP server is a secondary surface.
mcp-name: io.github.n24q02m/crg
Knowledge graph for token-efficient code reviews -- semantic search and call-graph resolution across your codebase.
[Mode] [CI] [codecov] [PyPI] [License: Apache-2.0]
[Python] [MCP] [semantic-release] [Renovate]
Sister projects from n24q02m (click to expand)
An MCP server that parses your codebase with Tree-sitter, builds a structural graph of functions/classes/imports, and gives Claude (or any MCP client) precise context so it reads only what matters instead of the whole tree. Semantic search runs through the local ONNX model registry from fastretrieval by default (zero config, no API key), with an optional cloud embedding chain. Fork of code-review-graph with fixed multi-word search, qualified call resolution, dual-mode embeddings, output pagination, and production CI/CD.
v2.0 migration (BREAKING)
v2.0 adds temporal columns (valid_from_sha / valid_to_sha on every node + edge) and an opt-in security scanner. The schema migration is auto-applied on first GraphStore open, and a backup of the pre-2.0 DB is saved to <graph_db>.pre-2.0.bak so you can roll back. See BREAKING_CHANGES.md for the full schema-change list, behavior changes, environment requirements, and the downgrade procedure (CRG_DOWNGRADE_TO_1_X=1 uv run crg).
Table of contents
- v2.0 migration (BREAKING)
- Install
- Usage
- Smithery
- Configuration
- Tools
- CLI
- Features
- Comparison
- Security
- Build from source
- Trust model
- Migration & changelog
- Documentation
- License
Install
For OMP and other local coding harnesses, the primary surface is the package CLI
plus the bundled skills/ workflows. The skills invoke the CLI directly and do
not require an MCP server mapping.
The optional Semgrep engine for deeper security scans is a separate extra:
MCP stdio remains a secondary protocol adapter for clients that require it:
Install matrix (stdio unless noted; the CLI-first usage above stays the primary surface):
Install with an AI agent -- paste this to your AI coding agent:
Install MCP server
crgfollowing the steps at https://raw.githubusercontent.com/n24q02m/claude-plugins/main/plugins/better-code-review-graph/setup-with-agent.md
Full CLI usage is in CLI. Optional per-client MCP setup is at mcp.n24q02m.com/servers/better-code-review-graph/setup/.
Usage
Two ways to run the server, plus the surfaces to consume it.
Dev: uv, no auth (loopback only)
Always-on: docker (token auth)
The compose file builds the Dockerfile's http target (image crg:http),
publishes the container on host loopback only
(127.0.0.1:${CRG_PORT:-8772}:8080), and mounts
./docker-config/config.toml read-only at the container's config dir
(CRG_CONFIG_DIR=/data/config); graph state persists in the crg-data
named volume. All state stays on your machine.
Consuming: CLI or MCP
CLI (local graphs, no server needed):
MCP over HTTP (remote/always-on): point any MCP client at
http://127.0.0.1:8772/mcp with Authorization: Bearer <token>:
Per-task model configuration
Each task (embed / rerank / chat / jev_score) resolves its own OpenAI-spec
provider cell from [models.<task>] in the instance config: an independent
base_url + api_key + model. Cloud or local is purely a config choice —
point base_url at OpenRouter, a vendor, or your own local gateway
(e.g. http://host.docker.internal:11434/v1). API keys are host-only
material: keep them in config.toml (gitignored under docker-config/) or
supply via HULL_<TASK>_API_KEY env; they are never end-user supplied.
Local-first boundary
CRG is local-first for coding workflows:
- CLI and bundled Skills are the primary surfaces for graph build/query, impact analysis, review context, security scans, and repository onboarding.
- MCP stdio is the secondary protocol adapter over the same local domain services; it does not maintain a separate graph implementation.
- Graph state stays in
<repo>/.crg/graph.dbunless an explicit multi-user/self-host configuration selects another data directory. - PyPI, CI, security scanning, GitHub releases, and eligible stable MCP Registry publication remain active. Historical public OCI tags are retained, but new public Docker Hub/GHCR images are no longer published.
- CRG has no hosted Cloudflare runtime in the target topology.
Smithery
The repo ships a smithery.yaml so the server can be built and
run through Smithery. It deploys over stdio and needs
no startup configuration -- the config schema is empty, and any optional cloud
embedding/summary keys are supplied at runtime through the server's own config
flow (see Configuration below). The launch command is the same
uvx invocation as a local install:
Configuration
Everything works out of the box with zero configuration -- semantic search
uses the local ONNX registry from fastretrieval
(Qwen3-Embedding-0.6B is the current built-in reference entry, ~570 MB
downloaded on first graph embed). This reference entry is not a Qwen-only
boundary: any built-in registry ID or valid non-Qwen artifact manifest follows
the same resolver. All environment variables below are optional and only needed
for cloud embeddings, LLM summaries, or an explicit BYO local artifact.
Model selection
Embeddings select the first provider/model entry in EMBEDDING_MODELS; later
entries are retained as configuration but are not runtime fallbacks. Summaries
select the first SUMMARY_MODELS entry too, without runtime fallback. Providers
are inferred from model prefixes and use the matching <PROVIDER>_API_KEY.
Cohere embed-v4.0 requests and stores 1024 dimensions; other backends retain
768-dimensional storage. CRG never slices, pads, or silently accepts a different
provider width. The embedding row's model and byte width must match before reuse.
Run graph(action="embed") after changing models or upgrading an old 768-wide
Cohere index. Searches reject incompatible widths before a provider call; graph
nodes are retained and re-embedding replaces only stale vectors.
Provider API keys
Cloud models need the provider key for the selected model prefix. Keys alone never select models: an empty embedding chain stays local, and an empty summary chain stays disabled. A configured cloud error does not fall back to local or another provider. Summarizers require a chat-completion model.
Advanced
When LOCAL_RERANK_MODEL is configured, semantic vector search retrieves a
bounded candidate pool of min(max(limit * 4, limit), 100) rows, applies the
existing kind, repo, and live-row filters, then reranks that pool and returns
at most limit rows. The response uses search_mode="semantic_reranked" and
adds rerank_score while preserving similarity_score. Blank keeps the
existing limit * 2 vector path and search_mode="semantic". Configured
reranker failures return an explicit error; CRG does not silently fall back to
vector or keyword results. Keyword searches, including as_of snapshots, do
not invoke the reranker.
Example -- cloud embeddings + summaries
Cohere embedding is paid. Authorize a bounded budget before a live index/query; the Minimax-free completion choice does not make embeddings free. This example does not add a process-wide model override: missing subject credentials fail closed rather than inheriting the server environment.
CRG currently has no cloud rerank call: LOCAL_RERANK_MODEL is its only
reranking path. Setting RERANK_MODELS or RERANK_API_BASE does not enable one.
Tools
Six tools, each grouping related actions to keep the tool surface small.
graph -- Graph lifecycle
Actions: build | update | stats | embed | export | summarize
query -- Graph queries
Actions: query | search | impact | large_functions | spot_check | renamed_in_diff | diff
Most read actions accept as_of=<sha> for temporal (point-in-time) snapshots
and repo=<repo_id> to scope a federated multi-repo graph.
review -- Code review context
Actions: context (default) | delta
Token-optimized review context with structural summary, impacted nodes, source
snippets, and review guidance. context auto-detects changed files from the
git diff; delta (with from_sha/to_sha, optional show_line_shifts)
surfaces refactor moves between two commits.
config -- Server configuration and credential setup
Actions: status | set | cache_clear | setup_status | setup_start | setup_skip | setup_reset | setup_complete
security -- Security scanning
Actions: scan | report | suppress | rule_list
The semgrep engine requires the [security] extra and runs Semgrep's
p/auto registry pack plus a 3-rule curated overlay.
help -- Full documentation
Topics: graph | query | review | config | security | recipes
Returns complete documentation for each tool. Use when the compressed descriptions above are insufficient.
CLI
The package installs two console scripts: crg (primary) and
crg (legacy long name). Running either with no
arguments starts the MCP server over stdio; a leading positional argument
routes to a local CLI subcommand that calls the same domain services used by
the MCP adapter. Run them directly after pip install, or without a
persistent install via uvx --python 3.13 --from better-code-review-graph crg ....
CLI subcommands print structured JSON and exit non-zero on an error.
Features
What this fork fixes versus the upstream code-review-graph:
Comparison
How crg stacks up against direct competitors in each pillar:
Sources: Greptile · Greptile pricing · Sourcegraph MCP · CodeGraph. Cells marked ? are capabilities the competitor does not publicly document, not confirmed absences.
Security
- Explicit selection -- Cloud embedding errors are reported; the runtime does not silently switch models or fall back to local ONNX.
- Error handling -- Tools return error strings with fix suggestions, never crash.
- Read-only mount -- Docker mode mounts the repo as
:ro(read-only). - SSRF-guarded endpoints -- Custom
EMBEDDING_API_BASE/LLM_API_BASEURLs are validated before any outbound call.
To report a vulnerability, see SECURITY.md.
Build from source
Requirements: Python 3.13, uv.
Trust model
This plugin implements TC-Local (machine-bound, single trust principal). See the mcp-core trust model for full classification.
Migration & changelog
Graph, security scan cache, and suppression state now use the package-owned
.crg/ directory. Run graph(action="build", full_rebuild=true)
once after upgrading, followed by graph(action="embed") if semantic search is
needed. The old .better-code-review-graph/ state directory plus the ambiguous
.code-review-graph/ and .code-review-graph.db paths and their SQLite
sidecars are left untouched, not migrated. Review and reapply any desired
suppression rules explicitly.
The v2.0 release added temporal columns (valid_from_sha / valid_to_sha
on every node and edge) plus an opt-in security scanner. The schema migration
is auto-applied on first GraphStore open, and a backup of the pre-2.0 DB is
written to <graph_db>.pre-2.0.bak. To downgrade and restore it:
Full schema-change list, behavior changes, and rollback procedure: BREAKING_CHANGES.md. Release-by-release history: CHANGELOG.md.
Documentation
Full docs at mcp.n24q02m.com/servers/better-code-review-graph/setup/:
- Setup -- install methods for Claude Code, Codex, Gemini CLI, Cursor, Windsurf, mcp.json
- Modes overview -- stdio / local-relay / remote-relay / remote-oauth
- Multi-user setup -- per-JWT-sub credential model
Use the help tool from any MCP client for inline per-tool reference.
License
Apache-2.0 -- See LICENSE.
來源:README.md,提交 732ba3c
工具
0版本歷史
1- v3.29.8最新Oct 5, 2026


