
agentRamen
io.github.palrajjpv0.2.5更新於 Oct 9, 2026
Local-first codebase intelligence for coding agents: indexes code and Git history into SQLite.
概覽
將本機程式碼倉庫的結構與 Git 歷史索引到 SQLite,讓助理能檢索與任務相關的檔案、程式碼片段與變更影響。
- 功能
- agentRamen 為倉庫的檔案、符號、匯入關係與 Git 歷史建立私有 SQLite 圖,並據此回答檢索請求。其 MCP 工具包括 repo_context(依任務排序的檔案與片段)、repo_impact 與 repo_explain(相依性與變更影響)、repo_history、repo_search、repo_architecture、repo_hotspots、repo_graph 和 repo_tests。另提供經審核的團隊記憶工具,例如 repo_stage_memory、repo_approve_memory、repo_publish_memory 與 repo_team_memory_search。
- 適用情境
- 當編碼助理需要聚焦的倉庫脈絡而非大範圍讀取檔案時,或需要依據已索引的 Git 歷史做相依性、熱點與變更影響分析時使用。適合想保留私有本機索引的個別開發者,也可選擇透過 Git 分享經審核的記憶。
- 執行需求
- 以 stdio 方式執行的本機程序,從 PyPI 安裝 agentramen,並在倉庫工作目錄中執行 agentramen mcp。需要 Python;預設安裝沒有執行期相依套件。選用擴充提供 Tree-sitter 解析、本機語意嵌入、以分詞器為基礎的計數,以及集中式快照服務。stdio 伺服器不需要 API 金鑰。
安裝
在 SourceWeft 中
- 開啟 儀表板中的 agentRamen,將其新增到工作區。
- 為需要使用其工具的對話啟用該服務。
Desktop only,透過 STDIO。 STDIO 服務會啟動本機處理程序,因此需要 SourceWeft 桌面主機。
其他 MCP 客戶端
參照 儲存庫 中的啟動說明。
README
agentRamen
[CI] [Python 3.10+] [PyPI] [Apache-2.0] [GitHub stars]
Apache-2.0 licensed. See the security policy and community code of conduct.
agentRamen indexes source structure and Git history, then retrieves task-relevant files and excerpts. It is local-first by default: each developer keeps a private SQLite index. Teams can optionally publish sanitized, commit-pinned snapshots to a centrally hosted, OIDC-protected read-only MCP service.
Quick start
Install agentRamen, then run it from the repository you want to explore:
[agentramen-demo]agentramen init creates the configuration, indexes the repository, and offers an optional GitHub Actions caller workflow. The index is stored locally in .agentramen/graph.db.
Try it in 30 seconds with examples/demo.sh. See the roadmap and contributing guide.
A minimal VS Code extension is also available.
What it can do
- Find task context: rank files using paths, symbols, imports, Git history, and optional local semantic similarity.
- Explain change impact: inspect dependencies, dependents, co-changes, hotspots, and file history.
- Map repository structure: explore architecture, export the graph, and query cached commit snapshots.
- Connect to coding agents: use local MCP stdio, the browser UI, or an optional authenticated centralized MCP service.
- Keep local indexes private: in local mode, the database and optional embeddings stay in each checkout; central mode transfers only an explicit sanitized snapshot to infrastructure you operate.
How it works
The default install uses the Python standard library and conservative analysis for other languages. Optional extras add Tree-sitter parsing, local semantic retrieval, and model-aware token counting.
Optional extras
The default install has no runtime dependencies. Tree-sitter, semantic embeddings, tokenizers, and the remote MCP service are opt-in. Semantic model setup may download weights on first use; inference runs locally.
agentramen init creates .agentramen.yml and a GitHub Actions caller workflow if they do not exist, then creates/updates .agentramen/graph.db. agentramen update and agentramen index incrementally refresh changed files. The database is ignored by Git.
In local mode, source excerpts and optional embeddings stay in .agentramen/graph.db. Central snapshot mode transfers indexed source that passes credential heuristics to your private central service; review your organization's data-handling requirements before enabling it.
Publishing releases
Configure a PyPI trusted publisher for palrajjp/agentRamen, using the
.github/workflows/publish.yml workflow and the pypi environment. To publish
a release, update the version in pyproject.toml and agentramen/__init__.py,
create a matching v-prefixed Git tag, and publish a GitHub release for that
tag. The workflow builds the distributions and publishes them to PyPI using OIDC.
Configuration
agentRamen reads the following fields from .agentramen.yml:
indexing.incremental: false reparses every eligible file on each index. git.history: false removes stored commit, rename, and co-change history; enabling it later rebuilds the configured recent history. history.commits controls retained commits (1–100,000), and git.co_changes toggles co-change edges. context.default_budget is used by the CLI, MCP, and HTTP API when no budget is passed. Configuration is a dependency-free YAML subset; unsupported/malformed values are rejected with a setting-specific error. .agentramenignore patterns are added to, rather than replacing, the built-in and YAML ignore rules.
semantic.enabled: true opts into local embedding generation and semantic context ranking; install agentramen[semantic] first. semantic.model selects a FastEmbed model. agentramen[treesitter] opts into Tree-sitter-based parsing for supported non-Python languages; without it, or if a grammar is unavailable, agentRamen falls back to regex analysis.
context.tokenizer_model optionally selects a model supported by tiktoken for tokenizer-based context budgets; install agentramen[tokenizer] to use it. The tokenizer vocabulary may be downloaded and cached on first use, then counting runs locally. Leave it empty to use the dependency-free approximate character-based fallback. Context responses identify token_count_method (tiktoken or approximate) and the tokenizer model when configured. The reported total budgets the selected files, team memories, and path/digest/commit evidence together; per-file counts are provided as estimated_tokens.
Commands
Example context response:
Context includes source excerpts only when the file still matches its indexed hash and does not contain a detected credential pattern. Token counts use the configured model tokenizer when available; otherwise they are approximate character-based estimates, not tokenizer measurements or performance claims.
GitHub Actions
The workflow created by agentramen init calls the reusable workflow in this repository:
The reusable workflow fetches Git history, restores a cache, tests the installed package, indexes the checked-out revision, summarizes pull requests, and publishes the local SQLite artifact. Use a full-depth checkout when running the CLI outside this reusable workflow to retain history.
Use from gha-cd or another workflow
Call the reusable workflow before a deployment or agent job. Supplying task also creates a bounded JSON context artifact; the workflow does not send repository content to an AI provider.
The artifact is named agentramen-context-${{ github.sha }}. A downstream job can fetch it and pass .agentramen-context/agentramen-context.json to its agent step:
The context budget defaults to 2000 if omitted. Treat the artifact like source code: it can contain excerpts and follows the repository's GitHub Actions access and retention policy.
Agent integrations
agentRamen's MCP server runs locally over stdio. Install agentRamen once on each developer machine, then add a portable .mcp.json at the repository root so VS Code Copilot and Claude Code can use the same configuration:
For Claude Code, the project-scoped command creates or updates .mcp.json:
For GitHub Copilot CLI, save the same mcpServers object in $COPILOT_HOME/mcp-config.json, or ~/.copilot/mcp-config.json when COPILOT_HOME is unset. See the VS Code MCP configuration guide and Claude Code MCP guide for client-specific setup and trust prompts.
In VS Code, trust the workspace and start the agentramen MCP server from the MCP Servers view. Each developer keeps an independent .agentramen/graph.db; commit .agentramen.yml and .mcp.json for shared settings, but do not put a live SQLite database on a network share.
Keep agent context focused
Add an instruction like this to the project's AGENTS.md, CLAUDE.md, or Copilot instructions:
For repository analysis, call
repo_contextwith the current task before broad file reads. Start with a 1200-token budget, use the returned paths as the investigation boundary, then callrepo_explainorrepo_impactfor targeted follow-up. Expand the search only when the indexed context is insufficient. Ask foragentramen updateif the index appears stale.
The budget caps retrieved context, not the model's entire conversation. Counts use an approximate character-based estimate by default; install agentramen[tokenizer] and set context.tokenizer_model when model-specific counting is important. Run agentramen update to incrementally refresh a developer's local index after changes.
ChatGPT and GPT Store
Custom GPTs cannot start a developer's local stdio process. A GPT Action or hosted ChatGPT app needs a reachable HTTPS service with authentication. agentramen serve binds to localhost and has no authentication, so do not expose it directly to the internet; a secure remote integration requires an authenticated gateway or a separately hosted MCP service. GPT creation and publishing availability depends on the current ChatGPT plan and workspace policy; see OpenAI's GPT creation guide.
MCP
Run agentramen mcp from a repository and configure your MCP-compatible agent to launch that command in the repository working directory. Tools include repo_context, repo_status, repo_history, repo_search, repo_explain, repo_impact, repo_dependencies, repo_tests, repo_changes, repo_architecture, repo_hotspots, repo_graph, repo_stage_memory, repo_review_queue, repo_approve_memory, repo_reject_memory, repo_publish_memory, repo_team_memory_search, and repo_memory_audit. The stdio server needs no API keys and does not access a network service.
Local HTTP API
agentramen serve listens only on 127.0.0.1 by default. The browser UI is available at / and /ui. The API provides GET /health, GET /api/v1/repository, /architecture, /graph, /files/{path}, /impact?path=..., /history?path=..., /hotspots, and POST /api/v1/context or /api/v1/search. Append ?at=<commit> to /api/v1/graph or /api/v1/architecture to query a cached, commit-addressable graph snapshot. POST requests accept JSON objects such as {"task":"Add OAuth","token_budget":2000}. No authentication is provided; do not bind to a public interface without placing an authenticated access-control layer in front.
Benchmarking
Run agentramen benchmark --files 10 1000 to measure synthetic initial indexing, a one-file incremental update, and context retrieval. Add --repository /path/to/repo to measure a temporary copy of a real repository's tracked working-tree files as well. Use --semantic-mode both to compare semantic retrieval off/on; the enabled run requires agentramen[semantic] and downloads its configured model if needed. If model setup is unavailable, the comparison reports that mode as an error and still returns measurements for the successful mode. Output includes lexical candidate count, database size, and retrieval/index timings.
Add --quality to report task-to-file and affected-test precision/recall on a deterministic, labeled synthetic fixture; test_evidence shows the import chain used to confirm each suggested test:
The current fixture scores 1.0 mean precision and recall for both files and tests. This is a fixture correctness check, not a real-repository accuracy claim. For adoption decisions, curate labels from your own repository and measure those tasks; all timings and scores depend on the repo, task labels, and environment.
Privacy and exclusions
The index stays in .agentramen/graph.db on the local machine or in the configured CI artifact. With semantic retrieval disabled, the database stores source metadata and hashes; when enabled, it additionally stores locally generated embeddings. agentRamen skips common generated directories, environment files, private-key files, oversized files, binary files, and files containing recognizable private-key or credential assignment patterns. Add more path patterns to .agentramenignore (one pattern per line). Review your ignore rules before publishing generated artifacts.
Retrieval and memory
Context retrieval uses an incremental SQLite inverted index for lexical candidates, then fuses lexical, optional semantic, and graph/history rankings with reciprocal-rank fusion before applying the token budget. Reviewed memory can be staged, approved, or rejected; approving a new fact supersedes conflicting active facts with the same subject.
Share approved memories with a team
AgentRamen keeps the live SQLite database private to each checkout. To share a reviewed fact, stage it, approve it, then publish it to .agentramen-shared/memories/<id>.json. Commit that file in a pull request so teammates can review changes, see history, and receive the update through Git. One file per memory keeps unrelated additions from colliding.
MCP tools: repo_context returns source path, digest, and commit evidence for its selected files. A subsequent repo_stage_memory attaches evidence from the latest context call to the private pending proposal; review shows the evidence and whether its source bytes are still current. Approval records the Git user.name as reviewer (falling back to the local account name); publication preserves that reviewer, the source commit, digests, and validity window. Approval and publication reject changed evidence. Stale approved memories appear at the top of repo_review_queue with a restage_with_fresh_context action; they cannot be reapproved as current. repo_team_memory_search omits shared memories whose evidence is stale or missing. repo_memory_audit reports unique stale-memory counts, mismatched paths, source commits, and unverified local/shared records without editing shared Git files. Superseded facts are marked in their existing file when a newer approved fact with the same subject is published. Credential-like content is blocked from publication; still review memory content for confidential or personal data before committing. Agents should never approve or publish a memory without explicit human direction.
CLI search and publish:
Shared memory is opt-in and version-controlled; rejected or merely staged notes stay local. Existing memories created before source evidence was captured are reported as unverified and are not included in task context until restaged with evidence. A changed or missing source file makes linked memories stale even before agentramen update; refresh context and propose a new memory rather than silently treating an old fact as current. Do not sync SQLite over a network share. For 100-person teams, use protected branches and normal pull-request review for shared-memory changes.
Centralized MCP for a shared graph
Local stdio remains the default. To publish an immutable graph snapshot for remote agents, enable the optional central dependencies and the reusable workflow input:
The workflow validates the archive with agentramen snapshot verify before uploading agentramen-snapshot-${{ github.sha }}. It contains the pinned graph, hash-verified indexed source that passed credential filtering, and committed team memories. Developer-private SQLite memories and local exclusion rules are stripped. Artifacts expire after 14 days, so copy accepted snapshots to private durable company storage before expiry. If agentramen init generated .github/workflows/agentramen.yml, commit that file before exporting; snapshot export requires a clean source checkout.
Deploy one read-only MCP service per repository snapshot. It validates company OIDC JWTs against JWKS, issuer, audience, expiry, scope, and optionally group claims; DNS-rebinding protection requires an explicit Host allowlist. Terminate TLS at a trusted ingress and keep the archive and mounted snapshot inside company-controlled infrastructure.
Configure MCP clients to use the HTTPS /mcp endpoint and complete the company OAuth/OIDC flow. For example, Claude Code can add it with claude mcp add --transport http agentramen-central https://agentramen.example.com/mcp; VS Code supports remote HTTP MCP servers. Set AGENTRAMEN_OIDC_GROUP_CLAIM when your IdP uses a different group claim (default groups).
Remote tools are read-only: status, context, search, architecture, graph, and approved team-memory search. Every response includes repository_id and snapshot_commit so agents can identify stale advice.
For deployment freshness and rollback, keep each accepted archive immutable and keyed by commit SHA. Verify it, start a candidate service against that archive, and call repo_status with an authorized client to confirm the expected commit. Promote by atomically switching the ingress/service pointer to the candidate; retain the previous service/archive and roll back by switching the pointer back. Do not overwrite the archive beneath a running service. agentramen mcp-http runs one process; use your platform's process manager/load balancer and a shared read-only snapshot mount for replicas. Cloud-specific storage upload and pointer switching are deployment-pipeline responsibilities, not built into this package.
Snapshots contain sanitized but real source code. Credential heuristics are not a substitute for repository ACLs or data-loss review; use private company-controlled storage with retention and access controls matching the source repository.
Known limitations
- Import resolution handles common file and module layouts, but not every workspace alias or language-specific build system.
- Architecture and pull-request summaries use indexed dependency edges; they do not apply semantic boundary rules.
- Semantic retrieval calculates exact similarity rather than using an approximate-nearest-neighbor index.
- Token counts are approximate character-based estimates unless
context.tokenizer_modelis configured with the optional tokenizer extra. - History keeps 100 commits by default, configurable from 1 to 100,000.
Development
If agentRamen is useful in your workflow, a GitHub star helps other developers find it. See CONTRIBUTING.md, SECURITY.md, and CODE_OF_CONDUCT.md. Contributions that add language analyzers should keep parser-specific behavior separate from the indexing and context interfaces.
來源:README.md,提交 507c6f8
工具
0版本歷史
1- v0.2.5最新Oct 9, 2026


