Miru

ai.takarav1.10.1更新于 Oct 6, 2026

Hybrid semantic and keyword code search for coding agents.

已验证STDIO仅桌面Developer ToolsKnowledge & Memory

概览

AI 生成的概览

Miru 为编码助手提供语义与关键词混合的代码搜索工具,让助手直接找到相关代码片段,而不必退回到 grep。

功能
Miru 为本地代码仓库建立索引,并通过 MCP 暴露 search、locate、expand、find_related 工具(基准模式下还有 read_benchmark)。search 结合向量嵌入、BM25、融合与重排序,返回带路径、行号和相关性分数的精简片段;locate 用于查找环境变量、符号或错误码等精确子串。索引在首次调用时构建,会话内缓存,并随本地文件变化增量更新。
适用场景
当助手需要按含义而非精确文本探索代码库时适用,例如查找认证中间件在哪里配置,或跨文件追踪相似模式。它面向支持 MCP 的 IDE 与命令行编码助手,在反复 grep 和整文件读取浪费上下文的大型仓库中尤其有用。
运行要求
以 npm 包 @takara-ai/miru-code 通过 stdio 在本地运行,使用 bunx 启动,因此需要 Bun 或 Node.js 环境。必须认证:使用 TAKARA_API_KEY 环境变量,或通过交互式设备码登录并在本地保存凭据。索引与查询嵌入需要访问 Takara 推理 API 的网络连接。
安装前请注意
建立索引和嵌入查询时,文件内容会通过 HTTPS 发送到 Takara 推理 API,API 用量可能按 Takara 套餐产生费用;索引专有代码前应确认符合安全与合规要求。TAKARA_API_KEY 属于机密,保存的凭据存放在明文的 credentials.json 中。该服务会读取本地仓库文件,并写入索引缓存、IDE 配置文件和基准历史;MIRU_WORKSPACE_ROOT 仅在设置时限制 MCP 的 repo 路径。

安装

在 SourceWeft 中

  1. 打开 控制台中的 Miru,将其添加到工作区。
  2. 为需要使用其工具的对话启用该服务。

Desktop only,通过 STDIO。 STDIO 服务会启动本地进程,因此需要 SourceWeft 桌面宿主。

其他 MCP 客户端

参照 仓库 中的启动说明。

README

Miru (見る)

[CI] [coverage] [license] [bun]

Hybrid code search for AI coding agents. Find code by meaning, not grep.

Your AI agent finds the code it needs with up to 50% fewer tokens.

Miru returns the best chunks (path, lines, snippet) for questions like "where is auth middleware configured?" — plugged directly into Claude Code, Cursor, Copilot, Codex, and 9+ other agents via MCP.


How it fits into your agent's workflow

[How Miru fits into an agent's workflow]

Miru replaces the grep/glob-style search agents fall back on today. Install once, and every connected agent gets search and find_related MCP tools automatically.

Install via a plugin marketplace

Prefer this when your IDE has a plugin marketplace — no bun add -g step, and updates go through the IDE's own plugin flow. Run miru setup first if you have not authenticated yet (see Set up credentials).

Claude Code:

/plugin marketplace add takara-ai/miru-code/plugin install miru

Restart Claude Code (or reload plugins) when prompted. Update: /plugin marketplace update miru then reinstall. Remove: /plugin uninstall miru.

Cursor:

  1. Dashboard → Plugins → Team Marketplaces → Add Marketplace → Import from Repo → takara-ai/miru-code
  2. Customize (sidebar) → find miru → Install → choose project or user scope

Any other IDE, or if marketplace install isn't available: use the CLI install below.

What you get from a plugin install vs the CLI

Plugin installs don't carry full CLI parity — what each one gives you depends on the IDE's own plugin capabilities, not just Miru's packaging:

Claude CodeCodexCursor
MCP tools (search, locate, expand, find_related)✅✅❌ (plugin ships skills + rules only — no MCP entry yet)
miru / Caveman / STE skills✅✅✅
Dedicated sub-agent (miru:miru-code)✅——
Benchmark mode toggle (/plugin configure)✅——
Credentialsown plugin-scoped dirown plugin-scoped dirn/a

— means the IDE's plugin schema has no equivalent mechanism to port these to (not a packaging gap we can close): Codex's and Cursor's plugin manifests have no agents or userConfig fields, so the dedicated sub-agent and the benchmark toggle are Claude-Code-only.

Credentials are plugin-scoped, not shared with your CLI install. Claude Code and Codex plugins each store auth in their own IDE-managed data directory (survives plugin updates, removed on uninstall). This means miru setup done via the CLI does not carry over to a plugin install, or vice versa — each authenticates independently on first use (interactive device-code login bootstraps automatically). If you use both the CLI and a plugin, you'll sign in twice.

Install (CLI)

bash
bun add -g @takara-ai/miru-code

Set up credentials

bash
miru setup

Interactive miru setup defaults to device-code login and saves the resulting credentials locally. Manual bearer-token entry is still available with --key. If credentials are missing, the interactive MCP/plugin path can bootstrap the same device flow automatically on first use.

bash
miru setup --device           # explicit device-code loginmiru setup --key YOUR_TOKEN   # store a bearer token directlymiru setup --clear            # remove stored credentials

Miru stores versioned credentials in credentials.json and automatically loads or refreshes them for MCP and CLI use. TAKARA_API_KEY still overrides stored credentials when set explicitly.

Add to your IDE

bash
miru install

Interactive TUI — ↑↓ move, space toggle, a all, enter confirm. Pick agents and integrations:

IntegrationWhat it does
MCP serversearch, locate, expand, and find_related tools in the agent
InstructionsSearch policy in CLAUDE.md / AGENTS.md / GEMINI.md
Sub-agentDedicated miru-code agent file
Cursor rulesAlways-on .cursor/rules/miru-code.mdc (Cursor only)
Old Miru hooksRemoves stale search hooks from older releases
Caveman (experimental)On-demand chat compression skill (/caveman)
STE writing (experimental)On-demand clear technical English for docs (/ste)

Restart the IDE when done.

bash
miru uninstall   # remove miru config

Supported: Cursor · Claude Code · Gemini CLI · Kiro · OpenCode · GitHub Copilot · Codex · VS Code · Visual Studio (Windows) · Windsurf / Devin Desktop

IDEMCPInstructions / rulesCaveman (experimental)STE (experimental)
Cursor~/.cursor/mcp.json~/.cursor/rules/miru-code.mdc~/.agents/skills/caveman/SKILL.md~/.agents/skills/ste/SKILL.md
Claude Code~/.claude.json~/.claude/CLAUDE.md~/.claude/skills/caveman/SKILL.md~/.claude/skills/ste/SKILL.md
Gemini CLI~/.gemini/settings.json~/.gemini/GEMINI.md~/.agents/skills/caveman/SKILL.md~/.agents/skills/ste/SKILL.md
Kiro~/.kiro/settings/mcp.json~/.kiro/steering/miru.md~/.kiro/skills/caveman/SKILL.md~/.kiro/skills/ste/SKILL.md
OpenCode$XDG_CONFIG_HOME/opencode/opencode.json(c) (else ~/.config/opencode/…)…/AGENTS.md~/.agents/skills/caveman/SKILL.md~/.agents/skills/ste/SKILL.md
GitHub Copilot~/.copilot/mcp-config.json—~/.agents/skills/caveman/SKILL.md~/.agents/skills/ste/SKILL.md
Codex~/.codex/config.toml~/.codex/AGENTS.md~/.agents/skills/caveman/SKILL.md~/.agents/skills/ste/SKILL.md
VS Code…/Code/User/mcp.json—~/.agents/skills/caveman/SKILL.md~/.agents/skills/ste/SKILL.md
Visual Studio%USERPROFILE%\.mcp.json—~/.agents/skills/caveman/SKILL.md~/.agents/skills/ste/SKILL.md
Windsurf——~/.agents/skills/caveman/SKILL.md~/.agents/skills/ste/SKILL.md

Plugin packaging

  • Codex: .codex-plugin/plugin.json, .agents/plugins/marketplace.json, shares the root mcp.json
  • Claude Code: .claude-plugin/plugin.json, .claude-plugin/marketplace.json, its own .claude-plugin/mcp.json (needed for the userConfig benchmark toggle — see What you get from a plugin install vs the CLI)
  • Cursor: plugin.json and .cursor/rules/miru-code-search.mdc

Current limitation:

  • these plugin manifests launch the published Miru runtime through bunx @takara-ai/miru-code@latest mcp
  • that means local source edits do not affect plugin behavior until a package version is published
  • and a fully self-contained “no Bun required” plugin install is still future work

On-demand skills

Sub-agent files are also written where supported (see miru install plan). Windsurf hooks only (experimental) — no MCP entry yet. Caveman is an on-demand Agent Skill (default off): invoke with /caveman or “talk like caveman”; stop with “normal mode”. Invocation UI varies by IDE (/caveman, $caveman, @caveman, etc.). Most IDEs share ~/.agents/skills/caveman/SKILL.md (including Copilot / VS Code / Visual Studio); Claude Code and Kiro keep vendor-native skill dirs. Ownership is tracked on the shared path so uninstalling one IDE keeps the skill while another still owns it; selecting all owners removes it once. STE is an on-demand Agent Skill (default off): invoke with /ste or “de-slop this”; keep articles and complete sentences. Most IDEs share ~/.agents/skills/ste/ the same way (ownership via miru-owners.json); Claude Code and Kiro keep vendor-native STE dirs.

STE writing (experimental)

STE helps write clear technical English for docs, runbooks, errors, and release notes (pragmatic ASD-STE100-inspired rules). Invoke with /ste or “de-slop this”. Keep articles and complete sentences — not telegraph-style omission.

Not ASD-certified. Full dictionary compliance needs the official standard at asd-ste100.org. Miru does not ship the copyrighted ASD dictionary.

Not for marketing or brand copy. Off by default; enable STE in the installer for any supported IDE. Restart the IDE (or reload skills) after install. For Codex, install also sets [features] skills = true in ~/.codex/config.toml (same as Caveman).

Caveman mode (experimental)

Caveman compresses live chat replies (less filler, max meaning). Intensities: /caveman lite|full|ultra (default full). Persisted artifacts (commits, PRs, customer docs) stay normal prose unless you ask otherwise.

Security / destructive warnings use clear normal prose (auto-clarity) — brevity never hides risk. Session token savings vary; the skill itself costs input tokens. No guaranteed %.

Off by default at install time. Enable Caveman in the installer for any supported IDE. Restart the IDE (or reload skills) after install. For Codex, the installer also sets [features] skills = true in ~/.codex/config.toml (required for Codex to load skills).

Team sub-agent in a repo (optional):

bash
miru init --agent claude --force

Try it

Meaning-based questions → search. Exact strings (env vars, symbols, error codes) → locate.

bash
miru search "auth middleware" ./srcmiru locate REDIS_HOST ./srcmiru expand src/auth.ts 42 ./srcmiru find-related src/auth.ts 42 ./src

Terminal output is human-readable; use --json for scripts. One-off without installing:

bash
bunx @takara-ai/miru-code@latest search "auth middleware" ./src

MCP tools

When wired via miru install, the MCP server exposes search, locate, expand, and find_related. read_benchmark appears only in benchmark mode. repo is optional: it defaults to the server's startup directory. Set it for another repo, or if your MCP client starts elsewhere. The index is built on the first call and cached for the session.

ToolWhen to use
searchDefault for code exploration — hybrid semantic + keyword search. One call per question.
locateExact substrings (env vars, symbols, error codes) — prefer over Grep.
expandMore context in the same file when a hit has truncated: true.
find_relatedSimilar code in other files from a file_path + anchor_line.
read_benchmarkCumulative token-savings rollup (benchmark mode only).

Workflow

  1. search with query — returns compact snippets (~±15 lines) and relevance scores.
  2. If a hit has truncated: true, call expand with file_path and anchor_line — not another search or a full-file read.
  3. To trace similar patterns elsewhere, call find_related with the same file_path and anchor_line.
  4. Use your editor's Read on absolute_path only when editing or when expand still lacks context.

Keep repo on follow-up calls when set.

Prefer these tools over Grep, Glob, or SemanticSearch when Miru MCP is connected — hooks and instructions enforce that when enabled.

Local repo hits include absolute_path for one-click navigation. Parameter reference is below under MCP parameters.

How it works

Hybrid search: Takara embeddings + BM25 + fusion + reranking. Index code, docs, config, or all with --content.

MCP watches local files and updates the index incrementally. Package upgrades invalidate stale caches via the version epoch.

OSIndex cache
macOS~/Library/Caches/miru
Linux~/.cache/miru
Windows%LOCALAPPDATA%\miru\Cache

Chunking & languages

Miru chunks source in tiers: AST (tree-sitter, default) → structural heuristics → line splits.

AST chunking — 26 languages (syntax-aware boundaries via vendored web-tree-sitter grammars):

LanguageTypical extensions
astro.astro
bash.sh, .bash, .zsh
c.c
cpp.cpp, .h, .hpp, etc.
csharp.cs
css.css
dart.dart
elixir.ex, .exs
embeddedtemplate.erb, .ejs
go.go
haskell.hs
html.html, .htm
java.java
javascript.js, .jsx, .mjs, .cjs
json.json
ocaml.ml, etc.
php.php
python.py, .pyi
ruby.rb
rust.rs
scala.scala
solidity.sol
sql.sql
svelte.svelte
typescript.ts, .tsx, .mts, .cts
vue.vue

Structural fallback (brace/indent heuristics when AST is unavailable): python, go, typescript, javascript, cpp, c.

Line fallback: everything else that gets indexed (kotlin, swift, etc.) — still searchable, coarser chunks.

Set MIRU_AST_CHUNKING=0 to disable AST and use structural → lines only.

CLI reference

Run miru in a terminal or miru -h for the command list, or miru <command> -h for details.

CommandPurpose
miru setupAuthenticate and store credentials
miru installConfigure IDE (global)
miru uninstallRemove IDE config
miru search <query> [path]Search (-k N, --content, --json)
miru locate <literal> [path]Exact substring in the index
miru expand <file> <line> [path]Adjacent chunks in the same file
miru find-related <file> <line> [path]Related chunks
miru benchmark on/off/status/clearToggle MCP benchmark mode / clear report
miru init --agent <id>Project-local sub-agent
miru clear [path]Drop index cache (use after big CLI-only refactors)
miru mcpStart MCP server (--benchmark for comparisons)

CLI uses hyphens (find-related); MCP tool names use underscores (find_related).

Library

bash
bun add @takara-ai/miru-code
ts
import { MiruIndex } from "@takara-ai/miru-code";
const index = await MiruIndex.fromPath("./src");const results = await index.search({ query: "BM25 tokenize" });

Environment

VariableNotes
TAKARA_API_KEYRequired
MIRU_OPENAI_BASE_URLDefault https://infer.takara.ai/v1
MIRU_OPENAI_EMBEDDING_MODELDefault ds1-miru-int8
MIRU_WORKSPACE_ROOTOptional: restrict MCP local repo paths to this directory
MIRU_MAX_INDEX_FILESCap files indexed per operation
MIRU_ALLOW_HTTP_GITSet 1 to allow plain http:// git clones
MIRU_MCP_WATCHSet 0 to disable MCP file watch
MIRU_AST_CHUNKINGSet 0 to disable tree-sitter AST chunking
MIRU_BENCHMARK_HISTORY_PATHOverride; see Benchmark mode
MIRU_CACHE_HOMEOverride the platform cache root (see How it works)
MIRU_QUIETSet 1 to skip the framed CLI banner (subtitle only on color terminals)
NO_COLORDisable CLI colors

See .env.example for more.

Privacy and API usage

Miru sends file contents to the Takara inference API when building an index and when embedding search queries. Chunks from your repo are transmitted over HTTPS to generate embeddings. API usage may incur cost depending on your Takara plan.

If you index proprietary code, make sure that sending snippets to Takara's endpoint fits your security and compliance requirements. MIRU_WORKSPACE_ROOT is an opt-in boundary for MCP local repo paths only, and restricts indexing to a single workspace directory when set.

Enterprise self-hosted embeddings (no Takara egress): see docs/self-hosted-sagemaker.md.

Benchmark mode

Optional measurement of how many tokens Miru saves versus a simple Grep workflow (ripgrep + reading the top matched file). Useful when evaluating Miru; leave it off day-to-day. Only local repo paths are compared — git URL repos skip the comparison and return benchmark_skipped: "local_repo_only".

bash
miru benchmark on       # add --benchmark to installed MCP configsmiru benchmark statusmiru benchmark off      # prefer off when finished measuringmiru benchmark clear    # delete the global report

Restart agents after changing mode. miru install keeps --benchmark if it was already enabled.

While on, each search / locate response includes a compact benchmark block (save_pct, miru_tok, grep_tok, saved_tok, rank1). Call read_benchmark for a cumulative rollup (or ask the agent when you want totals).

History is global (not per-repo) under Miru's state directory:

OSDefault path
macOS~/Library/Application Support/miru/benchmark-history.json
Linux~/.config/miru/benchmark-history.json ($XDG_CONFIG_HOME/miru/… when set)
Windows%APPDATA%\miru\benchmark-history.json

Override with MIRU_BENCHMARK_HISTORY_PATH. Append-only JSONL of compact token deltas (no query text); read_benchmark returns cumulative totals. Stored in plaintext — miru benchmark clear or uninstall on shared machines.

MCP parameters

**search**

ParamRequiredNotes
queryyesNatural language or code query
reponoStartup directory by default; set for another repo
includenoGitignore-style glob patterns; only matching files are searched (same as locate.include)
excludenoGitignore-style glob patterns; matching files are skipped (same as locate.exclude)
dedupe_by_filenoKeep best hit per file (default true)

**locate**

ParamRequiredNotes
literalyesExact substring to find
reponoStartup directory by default; set for another repo
includenoGitignore-style glob patterns; only matching files are searched
excludenoGitignore-style glob patterns; matching files are skipped
modenocount · locations · lines (default). Prefer count/locations when possible
limitnoCap returned hits. Omit to return all matches
ignore_casenoCase-insensitive match (default false)

**expand**

ParamRequiredNotes
file_pathyesFrom hit file_path or absolute_path (local repos)
anchor_lineyesFrom the search hit (anchor_line when truncated, else start_line)
reponoSame repo as the search, if set
before / afternoExtra chunks before/after anchor (default 1 each)

**find_related**

ParamRequiredNotes
file_pathyesFrom a search hit
anchor_lineyesFrom the search hit
reponoSame repo as the search, if set

**read_benchmark** (benchmark mode only)

ParamRequiredNotes
reponoFilter rollup to one local path or git URL. Omit for all saved queries

Manual MCP (skip miru install)

json
{  "miru": {    "command": "miru",    "args": ["mcp"]  }}

For benchmark mode, set "args": ["mcp", "--benchmark"] (or append the flag). Prefer miru benchmark on after a normal install — it updates every agent config.

Run miru setup once so the server can load credentials from credentials.json. If the MCP server starts in an interactive terminal without stored credentials, it will start device login automatically.

Use bunx + @takara-ai/miru-code@latest if miru is not global. The installer uses this command so each new MCP server launch can pick up a published version without a global update. A running server keeps its current version until restarted. Wrapper key varies by IDE (mcpServers, servers, or mcp).

Older headless MCP configs that launch miru without a subcommand, with or without MCP flags, continue to work. Run miru install again to update managed entries to the explicit mcp subcommand, then restart your coding agent. For manual MCP configs, add mcp after the executable or package name.

Developing

bash
git clone https://github.com/takara-ai/miru-code.git && cd miru-codebun install && cp .env.example .env.localbun test && bun run typecheck

See CONTRIBUTING.md for pre-commit hooks, commit message conventions, and the PR process.

Local MCP: "command": "bun", "args": ["/path/to/miru-code/src/cli.ts"]

miru -h / setup print a framed wordmark on color terminals (MIRU_QUIET=1 for subtitle only). Crane art lives in src/brand-banner.ts; regenerate with bun run scripts/render-crane-art.ts (ImageMagick). The crane is a registered mark of Takara.ai Ltd.

Credits

Miru uses work by MinishLab. Thank you to its authors.

  • semble (MIT). Miru ports parts of semble. These parts include the file walker, the index layout, the BM25 tokenizer, the hybrid search pipeline, and the ranking signals.
  • potion-code and Model2Vec (MIT). The WordPiece tokenizer in tokenizer/tokenizer.json comes from potion-code.

See NOTICE for the licence text and the Model2Vec citation.

License

MIT

来源:README.md,提交 3625c24

工具

0
工具元数据尚未被收录。

版本历史

1
  1. v1.10.1最新Oct 6, 2026