
Crg
io.github.n24q02mv3.29.8Updated Oct 5, 2026
Token-efficient code review knowledge graph: semantic search and call-graph resolution.
Overview
Builds a local Tree-sitter knowledge graph of a codebase so an assistant can run semantic search, call-graph queries, impact analysis, and review context.
- What it does
- Parses a repository with Tree-sitter into a structural graph of functions, classes, and imports stored in a local SQLite database. Six grouped tools cover graph lifecycle (build, update, stats, embed, export, summarize), queries (callers, callees, imports, impact, diff, semantic search), review context, configuration, security scanning, and documentation. Semantic search uses a local ONNX embedding model by default, with an optional cloud embedding and summarizer chain.
- When to use it
- Useful when an assistant needs precise context about a large codebase instead of reading whole files: finding callers or callees, estimating the blast radius of a change, generating token-efficient review context, or searching code semantically. Also relevant for local security scans and repository onboarding.
- Requirements
- Runs as a local process via uvx or pip (Python 3.13); the CLI is the primary surface and MCP stdio is secondary. Graph state is stored in the repository under .crg/graph.db. The default local embedding model is downloaded on first use (about 570 MB). Cloud embeddings or summaries need the matching provider key, for example COHERE_API_KEY or OPENROUTER_API_KEY, plus optional EMBEDDING_API_BASE or LLM_API_BASE. The Semgrep security engine needs an extra install.
Installation
In SourceWeft
- Open Crg in the dashboard and add it to a workspace.
- Enable the server for the chats that should use its tools.
Desktop only via STDIO. STDIO servers start a local process, so they need the SourceWeft desktop host.
Other MCP clients
Follow the launch instructions in the repository.
README
Better Code Review Graph
Renamed (2026-09-13): repo is now
crg— CLI-first (crgcommand). PyPI package stayscrg; MCP server is a secondary surface.
mcp-name: io.github.n24q02m/crg
Knowledge graph for token-efficient code reviews -- semantic search and call-graph resolution across your codebase.
[Mode] [CI] [codecov] [PyPI] [License: Apache-2.0]
[Python] [MCP] [semantic-release] [Renovate]
Sister projects from n24q02m (click to expand)
An MCP server that parses your codebase with Tree-sitter, builds a structural graph of functions/classes/imports, and gives Claude (or any MCP client) precise context so it reads only what matters instead of the whole tree. Semantic search runs through the local ONNX model registry from fastretrieval by default (zero config, no API key), with an optional cloud embedding chain. Fork of code-review-graph with fixed multi-word search, qualified call resolution, dual-mode embeddings, output pagination, and production CI/CD.
v2.0 migration (BREAKING)
v2.0 adds temporal columns (valid_from_sha / valid_to_sha on every node + edge) and an opt-in security scanner. The schema migration is auto-applied on first GraphStore open, and a backup of the pre-2.0 DB is saved to <graph_db>.pre-2.0.bak so you can roll back. See BREAKING_CHANGES.md for the full schema-change list, behavior changes, environment requirements, and the downgrade procedure (CRG_DOWNGRADE_TO_1_X=1 uv run crg).
Table of contents
- v2.0 migration (BREAKING)
- Install
- Usage
- Smithery
- Configuration
- Tools
- CLI
- Features
- Comparison
- Security
- Build from source
- Trust model
- Migration & changelog
- Documentation
- License
Install
For OMP and other local coding harnesses, the primary surface is the package CLI
plus the bundled skills/ workflows. The skills invoke the CLI directly and do
not require an MCP server mapping.
The optional Semgrep engine for deeper security scans is a separate extra:
MCP stdio remains a secondary protocol adapter for clients that require it:
Install matrix (stdio unless noted; the CLI-first usage above stays the primary surface):
Install with an AI agent -- paste this to your AI coding agent:
Install MCP server
crgfollowing the steps at https://raw.githubusercontent.com/n24q02m/claude-plugins/main/plugins/better-code-review-graph/setup-with-agent.md
Full CLI usage is in CLI. Optional per-client MCP setup is at mcp.n24q02m.com/servers/better-code-review-graph/setup/.
Usage
Two ways to run the server, plus the surfaces to consume it.
Dev: uv, no auth (loopback only)
Always-on: docker (token auth)
The compose file builds the Dockerfile's http target (image crg:http),
publishes the container on host loopback only
(127.0.0.1:${CRG_PORT:-8772}:8080), and mounts
./docker-config/config.toml read-only at the container's config dir
(CRG_CONFIG_DIR=/data/config); graph state persists in the crg-data
named volume. All state stays on your machine.
Consuming: CLI or MCP
CLI (local graphs, no server needed):
MCP over HTTP (remote/always-on): point any MCP client at
http://127.0.0.1:8772/mcp with Authorization: Bearer <token>:
Per-task model configuration
Each task (embed / rerank / chat / jev_score) resolves its own OpenAI-spec
provider cell from [models.<task>] in the instance config: an independent
base_url + api_key + model. Cloud or local is purely a config choice —
point base_url at OpenRouter, a vendor, or your own local gateway
(e.g. http://host.docker.internal:11434/v1). API keys are host-only
material: keep them in config.toml (gitignored under docker-config/) or
supply via HULL_<TASK>_API_KEY env; they are never end-user supplied.
Local-first boundary
CRG is local-first for coding workflows:
- CLI and bundled Skills are the primary surfaces for graph build/query, impact analysis, review context, security scans, and repository onboarding.
- MCP stdio is the secondary protocol adapter over the same local domain services; it does not maintain a separate graph implementation.
- Graph state stays in
<repo>/.crg/graph.dbunless an explicit multi-user/self-host configuration selects another data directory. - PyPI, CI, security scanning, GitHub releases, and eligible stable MCP Registry publication remain active. Historical public OCI tags are retained, but new public Docker Hub/GHCR images are no longer published.
- CRG has no hosted Cloudflare runtime in the target topology.
Smithery
The repo ships a smithery.yaml so the server can be built and
run through Smithery. It deploys over stdio and needs
no startup configuration -- the config schema is empty, and any optional cloud
embedding/summary keys are supplied at runtime through the server's own config
flow (see Configuration below). The launch command is the same
uvx invocation as a local install:
Configuration
Everything works out of the box with zero configuration -- semantic search
uses the local ONNX registry from fastretrieval
(Qwen3-Embedding-0.6B is the current built-in reference entry, ~570 MB
downloaded on first graph embed). This reference entry is not a Qwen-only
boundary: any built-in registry ID or valid non-Qwen artifact manifest follows
the same resolver. All environment variables below are optional and only needed
for cloud embeddings, LLM summaries, or an explicit BYO local artifact.
Model selection
Embeddings select the first provider/model entry in EMBEDDING_MODELS; later
entries are retained as configuration but are not runtime fallbacks. Summaries
select the first SUMMARY_MODELS entry too, without runtime fallback. Providers
are inferred from model prefixes and use the matching <PROVIDER>_API_KEY.
Cohere embed-v4.0 requests and stores 1024 dimensions; other backends retain
768-dimensional storage. CRG never slices, pads, or silently accepts a different
provider width. The embedding row's model and byte width must match before reuse.
Run graph(action="embed") after changing models or upgrading an old 768-wide
Cohere index. Searches reject incompatible widths before a provider call; graph
nodes are retained and re-embedding replaces only stale vectors.
Provider API keys
Cloud models need the provider key for the selected model prefix. Keys alone never select models: an empty embedding chain stays local, and an empty summary chain stays disabled. A configured cloud error does not fall back to local or another provider. Summarizers require a chat-completion model.
Advanced
When LOCAL_RERANK_MODEL is configured, semantic vector search retrieves a
bounded candidate pool of min(max(limit * 4, limit), 100) rows, applies the
existing kind, repo, and live-row filters, then reranks that pool and returns
at most limit rows. The response uses search_mode="semantic_reranked" and
adds rerank_score while preserving similarity_score. Blank keeps the
existing limit * 2 vector path and search_mode="semantic". Configured
reranker failures return an explicit error; CRG does not silently fall back to
vector or keyword results. Keyword searches, including as_of snapshots, do
not invoke the reranker.
Example -- cloud embeddings + summaries
Cohere embedding is paid. Authorize a bounded budget before a live index/query; the Minimax-free completion choice does not make embeddings free. This example does not add a process-wide model override: missing subject credentials fail closed rather than inheriting the server environment.
CRG currently has no cloud rerank call: LOCAL_RERANK_MODEL is its only
reranking path. Setting RERANK_MODELS or RERANK_API_BASE does not enable one.
Tools
Six tools, each grouping related actions to keep the tool surface small.
graph -- Graph lifecycle
Actions: build | update | stats | embed | export | summarize
query -- Graph queries
Actions: query | search | impact | large_functions | spot_check | renamed_in_diff | diff
Most read actions accept as_of=<sha> for temporal (point-in-time) snapshots
and repo=<repo_id> to scope a federated multi-repo graph.
review -- Code review context
Actions: context (default) | delta
Token-optimized review context with structural summary, impacted nodes, source
snippets, and review guidance. context auto-detects changed files from the
git diff; delta (with from_sha/to_sha, optional show_line_shifts)
surfaces refactor moves between two commits.
config -- Server configuration and credential setup
Actions: status | set | cache_clear | setup_status | setup_start | setup_skip | setup_reset | setup_complete
security -- Security scanning
Actions: scan | report | suppress | rule_list
The semgrep engine requires the [security] extra and runs Semgrep's
p/auto registry pack plus a 3-rule curated overlay.
help -- Full documentation
Topics: graph | query | review | config | security | recipes
Returns complete documentation for each tool. Use when the compressed descriptions above are insufficient.
CLI
The package installs two console scripts: crg (primary) and
crg (legacy long name). Running either with no
arguments starts the MCP server over stdio; a leading positional argument
routes to a local CLI subcommand that calls the same domain services used by
the MCP adapter. Run them directly after pip install, or without a
persistent install via uvx --python 3.13 --from better-code-review-graph crg ....
CLI subcommands print structured JSON and exit non-zero on an error.
Features
What this fork fixes versus the upstream code-review-graph:
Comparison
How crg stacks up against direct competitors in each pillar:
Sources: Greptile · Greptile pricing · Sourcegraph MCP · CodeGraph. Cells marked ? are capabilities the competitor does not publicly document, not confirmed absences.
Security
- Explicit selection -- Cloud embedding errors are reported; the runtime does not silently switch models or fall back to local ONNX.
- Error handling -- Tools return error strings with fix suggestions, never crash.
- Read-only mount -- Docker mode mounts the repo as
:ro(read-only). - SSRF-guarded endpoints -- Custom
EMBEDDING_API_BASE/LLM_API_BASEURLs are validated before any outbound call.
To report a vulnerability, see SECURITY.md.
Build from source
Requirements: Python 3.13, uv.
Trust model
This plugin implements TC-Local (machine-bound, single trust principal). See the mcp-core trust model for full classification.
Migration & changelog
Graph, security scan cache, and suppression state now use the package-owned
.crg/ directory. Run graph(action="build", full_rebuild=true)
once after upgrading, followed by graph(action="embed") if semantic search is
needed. The old .better-code-review-graph/ state directory plus the ambiguous
.code-review-graph/ and .code-review-graph.db paths and their SQLite
sidecars are left untouched, not migrated. Review and reapply any desired
suppression rules explicitly.
The v2.0 release added temporal columns (valid_from_sha / valid_to_sha
on every node and edge) plus an opt-in security scanner. The schema migration
is auto-applied on first GraphStore open, and a backup of the pre-2.0 DB is
written to <graph_db>.pre-2.0.bak. To downgrade and restore it:
Full schema-change list, behavior changes, and rollback procedure: BREAKING_CHANGES.md. Release-by-release history: CHANGELOG.md.
Documentation
Full docs at mcp.n24q02m.com/servers/better-code-review-graph/setup/:
- Setup -- install methods for Claude Code, Codex, Gemini CLI, Cursor, Windsurf, mcp.json
- Modes overview -- stdio / local-relay / remote-relay / remote-oauth
- Multi-user setup -- per-JWT-sub credential model
Use the help tool from any MCP client for inline per-tool reference.
License
Apache-2.0 -- See LICENSE.
Source: README.md at commit 732ba3c
Tools
0Version history
1- v3.29.8LatestOct 5, 2026


