cogvault

io.github.NBibikovv0.11.2更新于 Oct 5, 2026

Local hybrid-search memory for AI agents over plain Markdown. Multi-tenant, offline, no LLM.

概览

AI 生成的概览

面向 AI 代理的本地离线混合检索记忆,将回忆内容与事实以纯 Markdown 文件保存在各代理独立的租户目录中。

功能
Cogvault 将 Markdown 记忆文件索引到本地 SQLite 索引,并暴露两个 MCP 工具:cogvault_recall 用于自然语言混合检索(向量加 BM25 关键词排序并做排名融合,可选时间衰减与结果多样性),cogvault_record 用于把一条事实保存为新的 Markdown 卡片。它支持按 frontmatter 类型过滤和 wiki 链接关联卡片,也能索引 Obsidian 库这类嵌套目录树。命令行还提供索引、搜索、评测、健康检查与修复功能。
适用场景
适合多个编码代理各自需要隔离、确定性的记忆命名空间,且记忆以人类可读的 Markdown 保存、摄取与检索都不经过 LLM、也不依赖云服务的场景。也适合希望用自己的真实查询衡量召回效果、并把记忆纳入 git 版本管理的团队。
运行要求
通过 stdio 在本地运行,通常使用 uvx 或 pip/uv 安装 cogvault 包,需要 Python 与 uv。首次运行会下载嵌入模型(约 0.2-0.5 GB)并在本机运行。可选环境变量:COGVAULT_MODEL 覆盖嵌入模型,COGVAULT_LOG 关闭或迁移回忆日志。模型下载完成后无需账号、API 密钥或网络服务。
安装前请注意
调用 cogvault_record 时服务器会向租户目录写入 Markdown 卡片,索引与修复命令也可能改动文件;修复默认是试运行,需加 --apply 才会写入。回忆查询默认记录到 JSONL 日志,可能包含敏感查询文本,可设 COGVAULT_LOG=off 关闭。记忆仅按租户目录隔离,把某个代理指向另一个代理的目录就会暴露其记忆。

安装

在 SourceWeft 中

  1. 打开 控制台中的 cogvault,将其添加到工作区。
  2. 为需要使用其工具的对话启用该服务。

Desktop only,通过 STDIO。 STDIO 服务会启动本地进程,因此需要 SourceWeft 桌面宿主。

其他 MCP 客户端

参照 仓库 中的启动说明。

README

Get started in 30 seconds

Claude Code plugin — MCP server plus a skill that tells the agent when to recall and when to record:

text
/plugin install cogvault --marketplace NBibikov/cogvault

Claude Code before 2.1.275: /plugin marketplace add NBibikov/cogvault, then /plugin install cogvault@cogvault.

Memory lives in ~/.cogvault/memory (set COGVAULT_TENANT to change it). The plugin adds /cogvault:remember.

Any MCP client, one line (needs uv):

bash
claude mcp add cogvault -- uvx cogvault mcp --tenant ~/agent/memory
Claude Desktop, Cursor, Windsurf, Cline — JSON config

Add to claude_desktop_config.json, ~/.cursor/mcp.json, or your client's MCP config:

json
{  "mcpServers": {    "cogvault": {      "command": "uvx",      "args": ["cogvault", "mcp", "--tenant", "~/agent/memory"]    }  }}

Already have Markdown memory? Point the tenant at it — nothing to import. Claude Code's auto-memory works as-is (frontmatter type, [[links]] and all):

bash
M=~/.claude/projects/<project>/memoryuvx cogvault index  --tenant $M --ignore MEMORY.md   # first pass embeds, later passes are incrementaluvx cogvault search --tenant $M "how do we deploy"

The MCP server indexes on start by itself; the CLI search reads the existing index. --ignore MEMORY.md keeps the index file from competing with the cards it points to.

[A real Claude Code session with the cogvault plugin: asked why the worker keeps restarting, the agent recalls a memory card and answers with the fix, then saves a new card when told to remember something]

Unedited answers from a real session (Sonnet, plugin installed, a demo memory of six cards); only the waiting time is cut.


Why

Most "AI agent memory" tools want to be an autonomous LLM daemon that summarizes your work into an opaque database or a graph you can't read. For a fleet of coding agents that just need to reliably recall a decision, a bug fix, or an infra detail, that's the wrong trade.

cogvault makes the opposite bet:

  • Your Markdown files are the source of truth. Open them, edit them, git diff them. The SQLite index is a derived cache — delete it and it rebuilds from the files.
  • One library, many tenants. Each agent gets an isolated memory namespace via its own directory. A process loads each embedding model once and shares it across every tenant it touches; with the stdio MCP server that means one small process per agent session, not a central daemon.
  • No LLM in the loop. Ingest and retrieval are deterministic. Your agent is the LLM — it doesn't need a second one to remember.
  • Local, private, offline. FastEmbed runs on-device. Nothing leaves your machine.

Compared to

Checked against each project's README and docs on 2026-10-05.

cogvaultbasic-memorymem0GraphitiLetta Code
Source of truthMarkdown filesMarkdown filesVector DB (Qdrant / pgvector)Graph DB (Neo4j, FalkorDB, …)Markdown in a git repo per agent
LLM needed to store or recallNoNo (optional reranker)Yes by default (add() extracts facts)Yes to ingestYes — the agent edits its memory
RetrievalVector + BM25, RRF, decay, MMRFull-text + vector, optional rerankSemantic + BM25 + entitiesSemantic + BM25 + graphFile search; hybrid optional
InfraOne process, SQLite fileOne process, SQLite (Postgres optional)Library, or Docker + Postgres serverA graph databaseLetta backend

Pick something else when: you want an LLM to distil and merge facts for you (mem0), relationships between entities are the point (Graphiti), you want the agent to manage its own memory (Letta), or you want a richer notes app around the same Markdown idea, with Obsidian sync and a hosted option (basic-memory — the closest to cogvault).

Pick cogvault when: you run several agents and want each one's memory isolated in its own directory behind one process; you want recall to be deterministic and offline; and you want to measure it — cogvault eval scores recall on your agents' real queries and cogvault analyze lists what they tried to recall and couldn't.

How it works

[Markdown files are indexed into a derived SQLite database (vectors + FTS5) and recalled through MCP, CLI or Python]

Anatomy of a recall

[A query runs through a semantic and a keyword ranker, fused with RRF, then temporal decay, MMR and one-hit-per-card]

Hybrid retrieval fuses semantic (vector) and keyword (BM25/FTS5) ranking with Reciprocal Rank Fusion, then applies optional temporal decay (recent memory outranks stale) and MMR (diverse top results, not five near-duplicates). Each card contributes only its best chunk, so one long file can't fill the whole result list.

Benchmark

[Bar chart, 65 real agent queries: e5-small with summary chunk hit@1 0.57, hit@5 0.91, MRR 0.70; MiniLM default hit@5 0.80; bge-small-en hit@5 0.77]

Measured on real recall traffic, not synthetic questions: 66 queries sampled from the query logs of 6 live agent tenants (58% Ukrainian, the rest English), each judged against the actual cards — including answers that no configuration returned. One query has no answer in memory and counts as a gap, so 65 are scored. Every configuration was re-indexed from scratch on copies of the same tenants with cogvault 0.11.0.

Configurationhit@1hit@5MRR@10
multilingual-e5-small, chunk_chars = 700, summary chunk0.570.910.70
multilingual-e5-small, chunk_chars = 700, no summary chunk0.550.860.69
paraphrase-multilingual-MiniLM-L12-v2 (built-in default)0.540.800.66
bge-small-en-v1.5 (English-only)0.570.770.65

What the numbers do and don't say:

  • hit@1 is a tie. All four land within 0.54–0.57, and the 95% bootstrap intervals overlap almost completely. Real agent queries read like card titles, so the right card usually wins on its name alone.
  • The gap is in the top 5. e5-small with the summary chunk puts the answer in the top 5 for 91% of queries vs. 80% for the default MiniLM and 77% for English-only bge on this mixed-language memory. That's what an agent reading 5 results actually feels.
  • 65 queries is still a small sample. Treat differences under ~0.1 as noise. The aggregate numbers are in assets/benchmark.json; the queries are private and stay in each tenant.

Run the same check on your own memory: put judged queries in <tenant>/.cogvault-golden.jsonl ({"query": "...", "relevant": ["file.md"]}, empty relevant = a known gap) and run cogvault eval --tenant DIR.

Choosing an embedding model

Agent memory is often not English-only. The default is multilingual so nothing is broken out of the box — but pick the model that matches your fleet's language mix (set COGVAULT_MODEL, or Config(model=...)). Switching models auto-rebuilds the index.

Model (COGVAULT_MODEL)DimSizeReal-query hit@5*Cyrillic / multilingualWhen
paraphrase-multilingual-MiniLM-L12-v2 (default)3840.22 GB0.80✅ worksMixed-language fleets; safe default
BAAI/bge-small-en-v1.53840.13 GB0.77❌ Cyrillic vectors breakEnglish-only memory
intfloat/multilingual-e5-small3840.47 GB0.91✅ best per GB (512-token window)Mixed-language fleets; use chunk_chars = 700
intfloat/multilingual-e5-large10242.24 GBnot measured✅ bestMax quality, RAM to spare

*From the benchmark above: 65 real queries over mixed EN/UK memory, cogvault 0.11.0. On English-only memory bge-small-en is a fine choice; on Ukrainian content it returns a negative relevance margin (a distractor outranks the answer), so it is unsafe for non-English memory. Run cogvault eval on your own vault to decide.

bash
COGVAULT_MODEL=BAAI/bge-small-en-v1.5 cogvault index --tenant ~/agent/memory

Pin the model per tenant so it travels with the data instead of relying on every command exporting COGVAULT_MODEL (forget it once and a model mismatch silently re-embeds the whole index). Drop a .cogvault.toml at the tenant root:

toml
# ~/agent/memory/.cogvault.tomlmodel = "BAAI/bge-small-en-v1.5"# optional: recursive = true, strip_frontmatter = true, ignore_globs = ["Templates/*"]

Now cogvault search --tenant ~/agent/memory "…" uses the right model with no env var. Precedence: explicit --model / $COGVAULT_MODEL > .cogvault.toml > built-in default.

Install

bash
uv tool install cogvault        # CLI on PATH# orpip install cogvault

Or skip installing and run it on demand with uvx cogvault …. The first run downloads the embedding model (~0.2–0.5 GB, once per machine).

Add it to Claude Code as an MCP server in one line:

bash
claude mcp add cogvault -- uvx cogvault mcp --tenant ~/agent/memory

Also listed in the official MCP Registry as io.github.NBibikov/cogvault. Wheels are attached to each GitHub release.

Quickstart

bash
# index a tenant's markdown memorycogvault index --tenant ~/agent/memory
# search (hybrid semantic + keyword)cogvault search --tenant ~/agent/memory "how do I restart the worker service"
# enable temporal decay (recent wins) and tune diversitycogvault search --tenant ~/agent/memory "deployment steps" --half-life 30 --mmr 0.5
# only cards of one frontmatter type (user / feedback / project / reference / …)cogvault search --tenant ~/agent/memory "hard rules for deploys" --type feedback

Memory cards

cogvault understands two lightweight Markdown conventions (both optional — plain files index fine):

  • Frontmatter type — either flat (type: reference) or nested (metadata: → type: reference). Parsed at index time and filterable at search time (--type, MCP type param, search(card_type=...)). Every hit carries a type field; cards without frontmatter get null.
  • [[wiki-links]] — link targets are indexed, and the top search result includes a related list of linked cards that exist in the index (ghost links are dropped; matching is by exact filename stem).

As an MCP server (Claude Code, Cursor, any MCP client)

bash
claude mcp add cogvault -- uvx cogvault mcp --tenant ~/agent/memory

Exposes two tools:

  • cogvault_recall — natural-language hybrid search over this agent's memory. Optional type param filters to one frontmatter card type; the top result includes a Related: line built from its [[wiki-links]].
  • cogvault_record — save a fact; it's written as a Markdown card and indexed

Indexing a folder tree (Obsidian vaults, knowledge bases)

By default a tenant is one flat directory of .md files. For a nested vault (e.g. Obsidian, with 01-Projects/…, frontmatter, and folders to skip), opt in:

bash
cogvault index --tenant ~/vault \  --recursive \  --strip-frontmatter \  --ignore ".obsidian/*" --ignore ".trash/*" --ignore "Templates/*"
  • --recursive walks subdirectories; files keep their path relative to the tenant, so two notes named Tasks.md in different folders never collide.
  • --strip-frontmatter drops a leading YAML --- … --- block so its keys don't pollute the embedding.
  • --ignore GLOB (repeatable) skips paths relative to the tenant root.

Same flags exist on search and mcp, and as Config(recursive=True, strip_frontmatter=True, ignore_globs=(...)) for the library. Indexing is incremental: the first pass embeds everything, later passes only re-embed changed files. (Reference: a ~3,500-note vault → ~9,500 chunks, first index ≈ 3–4 min, then warm recall in single-digit milliseconds.)

As a library

python
from cogvault import Vault, Config
vault = Vault("~/agent/memory", Config(half_life_days=30))vault.reindex()for hit in vault.search("where are credentials stored"):    print(hit["score"], hit["file"], hit["snippet"])

Keeping a tenant healthy

cogvault doctor --tenant DIR reports what silently degrades recall: cards with no frontmatter or type, legacy timestamp filenames, frontmatter wrapped inside frontmatter, duplicate name: slugs, and [[links]] that resolve to nothing. Links resolve by frontmatter name:, filename stem, either separator style, and with or without the card-type prefix; links inside code and paths to files outside the tenant are not counted.

cogvault repair --tenant DIR fixes the mechanical half (dry run by default, --apply to write): unwraps nested frontmatter, infers a missing type, renames card-<timestamp>-….md to <type>_<slug>.md and rewrites every reference to it (MEMORY.md included), and adds minimal frontmatter to <type>_*.md cards that lack it. Healthy cards are left alone, and repaired cards keep their mtime so temporal decay is not reset.

Effectiveness logging

Every recall is logged (one JSONL line) so you can measure whether the memory is actually helping. cogvault analyze turns the log into a report:

bash
cogvault analyze            # recalls, latency, per-tenant hit distance, weakest hits, card writescogvault analyze --json     # machine-readable

The section that matters is weakest hits: each tenant's worst decile by vector distance — queries where recall returned something (hybrid search almost always does) but probably not the answer. That list is what your agents tried to recall and couldn't, i.e. the memory gaps to fill. Set COGVAULT_LOG=off to disable, or COGVAULT_LOG=/path.jsonl to relocate.

Multi-tenant fleets

Point one process at many tenants — each directory is an isolated namespace, proven by the test suite (test_multi_tenant_isolation). Agent B can never recall Agent A's memory unless you point B at A's directory.

[Four agents, each pointed at its own memory directory with its own index; tenants are isolated]

Design notes

DecisionWhy
Markdown = source of truthHuman-readable, git-versionable, editable, never locked in a DB
SQLite + sqlite-vec + FTS5Zero-infra hybrid search; one portable .db file; rebuildable
FastEmbed (multilingual MiniLM default, 384-d)In-process ONNX, no server, no API key, ~220 MB
Content-hash cacheRe-indexing only embeds changed chunks
RRF + decay + MMRPrecision, recency, and diversity without a graph DB
One process, many tenantsFleet infra, not a single-user desktop sidecar
WAL + incremental reindexConcurrent agents read while one writes; only changed files re-embed

Roadmap

  • Importers (migrate existing memory from other stores)
  • Pluggable embedders (Ollama, OpenAI-compatible endpoint)
  • valid_until per-card temporal validity
  • Optional FastMCP transport

License

MIT — your memory, your files, your infrastructure. Forever.

Figures are hand-built SVG from assets/make_graphics.py — python assets/make_graphics.py --png regenerates them.

来源:README.md,提交 2a8256b

工具

0
工具元数据尚未被收录。

版本历史

1
  1. v0.11.2最新Oct 5, 2026