
Mnemo
io.github.n24q02mv2.19.1Updated Oct 5, 2026
Persistent AI memory with hybrid search and embedded sync. Open, free, unlimited.
Overview
Gives an assistant a persistent local memory store with hybrid keyword and vector search, typed capture, and optional multi-machine sync.
- What it does
- Mnemo stores memories in a local SQLite database and retrieves them with hybrid search: FTS5 full-text plus vector similarity, fused with reciprocal rank fusion and then re-ranked. It exposes granular tools for adding, searching, listing, updating, deleting, exporting, importing, archiving, restoring, and consolidating memories, plus a config tool for status, sync, and passport export/import. Memories carry context types such as conversation, fact, preference, skill, task, and decision, with deduplication, importance scoring, a knowledge graph, and bitemporal history queries. A CLI surface (mnemo) operates directly on the database without starting a server.
- When to use it
- Worth adding when you want an assistant to remember preferences, decisions, and facts across sessions and machines, and to recall them by meaning rather than exact wording. It suits users who prefer a self-hosted, single-file store over a cloud memory service, and who may want optional Google Drive or S3-compatible sync.
- Requirements
- Runs locally over stdio, typically via uvx from the PyPI package mnemo-mcp, so a Python/uv runtime is needed. Zero-config local defaults use bundled local embedding and reranking models, which may be downloaded on first use. Optional cloud providers and sync require credentials: the API_KEYS environment variable (PROVIDER_API_KEY:key format) and model selection via EMBEDDING_MODELS, RERANK_MODELS, and LLM_MODELS; Google Drive sync uses OAuth.
Installation
In SourceWeft
- Open Mnemo in the dashboard and add it to a workspace.
- Enable the server for the chats that should use its tools.
Desktop only via STDIO. STDIO servers start a local process, so they need the SourceWeft desktop host.
Other MCP clients
Follow the launch instructions in the repository.
README
Mnemo MCP Server
Renamed (2026-09-13): repo is now
mnemo— CLI-first (mnemocommand). PyPI package staysmnemo-mcp; MCP server remains a secondary surface.
mcp-name: io.github.n24q02m/mnemo
Persistent AI memory with hybrid search and embedded sync. Open, free, unlimited.
[Mode] [CI] [codecov] [PyPI] [License: Apache-2.0] [SafeSkill 91/100]
[Python] [SQLite] [MCP] [semantic-release] [Renovate]
Sister projects from n24q02m (click to expand)
Table of contents
- Features
- Quick install
- Status
- Documentation
- Smithery
- Tools
- Security
- Build from Source
- CLI
- Self-hosting (local HTTP instance)
- Remote (HTTP mode)
- Deploy to Cloudflare
- Trust Model
- License
Roadmap (current = Phase 3 / v2.x)
Features
- Hybrid retrieval -- FTS5 + vector search (sqlite-vec locally, Vectorize on Cloudflare), fused via Reciprocal Rank Fusion (k=60), then re-ranked by a configurable rerank chain (
RERANK_MODELS, order = litellm fallback; empty -> Fastretrieval's local Qwen3 reranker) with temporal decay and importance boost - Typed capture --
memory(action="capture")with 6 context_types (conversation/fact/preference/skill/task/decision), embedding-based dedup, and a configurable LLM chain (LLM_MODELS, order = litellm fallback) - Knowledge graph -- Automatic entity extraction and relation tracking; top results boosted by graph proximity
- Importance scoring + archive policy -- LLM-scored 0.0-1.0 importance; soft-archive when
recency_factor * (1 - importance) > 1.0; restore action available - Auto-archive trigger -- Background sweep every Nth capture (default 100) -- no cron required
- STM-to-LTM consolidation -- LLM summarization of related memories in a category
- Duplicate detection -- Warns before adding semantically similar memories
- Zero config -- Fastretrieval's built-in local registry resolves Qwen3 ONNX embedding + reranking, no API keys needed. Optional cloud providers (Jina AI, Gemini, OpenAI, Cohere)
- Multi-machine sync -- JSONL-based merge sync via Google Drive (bundled Desktop OAuth public client)
- Plugin trinity -- Ships
/recall-context+/memory-commitskills and SessionStart + opt-in PostToolUse hooks (see docs/ARCHITECTURE.md) - Proactive memory -- Tool descriptions and skills guide AI to save preferences, decisions, facts at the right moment
- LLM compression -- Per-turn compression via the multi-provider dispatcher targets ~3x token reduction at >=0.9 fact retention; graceful skip when no provider configured (see docs/compression.md)
- Encrypted passport sync -- AES-256-GCM bundles + Argon2id KDF, S3 (R2 / B2 / MinIO) and Google Drive backends, delta-sync with last-write-wins per row (see docs/passport.md). Bootstrap via the
passport-bootstrapskill. - Temporal knowledge graph -- Bitemporal columns (
valid_from/valid_to/superseded_by) on every memory + entity-resolution dedup (embedding KNN at default 0.85 cosine threshold) + audit trail (memory_audittable with prev/new state hashes) + new actions (entity_search/entity_graph/history) + opt-inKG_AUTO_ENABLEDauto-extract on capture. BREAKING for clients that calledmemory.getexpecting historical-inclusive results: passas_offor time-travel; default now filters to current-state (valid_to IS NULL).
Quick install
Install matrix (stdio unless noted; see the Setup page for full steps):
Example stdio config (zero-config local defaults):
Comparison vs. peers
Status
2026-05-02 -- Architecture stabilization update
Past months saw significant churn around credential handling and the daemon-bridge auto-spawn pattern. This caused multi-process races, browser tab spam, and inconsistent setup UX across plugins. The architecture is now stable: 2 clean modes (stdio + HTTP), no daemon-bridge layer, no auto-spawn from stdio.
Apologies for the instability period. If you encountered issues with prior versions, please update to the latest release and follow the current setup docs -- most prior workarounds are no longer needed.
Related plugins from the same author:
- wet-mcp -- Web search + content extraction
- imagine-mcp -- Image/video understanding + generation
- better-notion-mcp -- Notion API
- better-email-mcp -- Email management
- better-telegram-mcp -- Telegram
- better-godot-mcp -- Godot Engine
- better-code-review-graph -- Code review knowledge graph
All plugins share the same architecture -- install once, learn pattern transfers.
Documentation
Full docs at mcp.n24q02m.com/servers/mnemo-mcp/setup/:
- Setup -- install methods for Claude Code, Codex, Gemini CLI, Cursor, Windsurf, mcp.json
- Modes overview -- stdio / local-relay / remote-relay / remote-oauth
- Multi-user setup -- per-JWT-sub credential model
Install with AI agent -- paste this to your AI coding agent:
Install MCP server
mnemo-mcpfollowing the steps at https://raw.githubusercontent.com/n24q02m/claude-plugins/main/plugins/mnemo-mcp/setup-with-agent.md
Smithery
mnemo-mcp is packaged for Smithery -- install or run it straight from the registry. It starts over stdio via uvx mnemo-mcp with no configuration required to launch; credentials are configured at runtime through the server's own config flow (see Documentation). The published start command lives in smithery.yaml.
Tools
15 MCP tools, 17 memory actions. The memory surface is exposed both as 11 specialized single-purpose tools and a deprecated legacy memory dispatcher (same actions), plus config, help, and config__open_relay:
Plugin trinity (Claude Code marketplace install):
MCP Resources
MCP Prompts
Security
- Graceful fallbacks -- Cloud → Local embedding, no cross-mode fallback
- Sync token security -- OAuth tokens stored at
~/.mnemo/tokens/with 600 permissions - Input validation -- Sync provider, folder, remote validated against allowlists
- Error sanitization -- No credentials in error messages
Build from Source
CLI
The package ships two distinct console scripts:
mnemo-- CLI-first memory surface (primary for scripts/agents; it never starts a server):capture,recall,reflect,fetch, and thestanding-*family operate directly on a SQLite memory DB.mnemo-pilotis a legacy alias of the same entry point.mnemo-mcp-- the MCP server plus one-shot operator subcommands. A bare invocation (or any---prefixed flag) starts the server; a leading subcommand runs an action and exits.
CLI-first memory surface (mnemo; every subcommand takes --db <path>,
prints a JSON envelope, and exits with a taxonomy-mapped code):
Server operator CLI (mnemo-mcp):
Self-hosting (local HTTP instance)
Two ways to run the server for MCP clients on your machine.
Dev: start with uv (no-auth, loopback only)
no-auth refuses non-loopback binds, so this is localhost-only by construction —
fine for trying the server locally. The MCP endpoint is
http://127.0.0.1:8000/mcp. For a real config, bootstrap one and edit it:
Always-on: docker compose (token auth, loopback-published port)
docker-compose.http.yml is self-contained (builds the image, persists state
in the mnemo-data volume) and publishes only on loopback:
Token setup (also documented in the example config):
For auth = "multi" (per-user namespaces) also mount users.toml — see the
commented line in docker-compose.http.yml.
CLI consumer (no server needed)
The mnemo surface talks straight to the memory DB — handy for scripts and
agents:
Every subcommand prints a JSON envelope and takes --db <path>. See
CLI for the full surface (reflect, standing-*, doctor, …).
Pointing an MCP client at the instance
Register the HTTP endpoint (Streamable HTTP transport):
- Claude Code:
claude mcp add --transport http mnemo http://127.0.0.1:8771/mcp - Any OpenAI-spec MCP client: server URL
http://127.0.0.1:8771/mcp; withauth = "token"send the shared token as the Bearer credential.
Config: local vs cloud, per task
Each task cell in mnemo-config/config.toml ([models.embed], rerank,
chat, jev_score) is independent: base_url + api_key + model, OpenAI-spec
HTTP. Mix freely — e.g. cloud OpenRouter for chat while embed/rerank
point at a local OpenAI-spec server, or all cloud. Keys are host-only
(end users never see them) and may alternatively come from the
HULL_<TASK>_API_KEY env vars.
Remote (HTTP mode)
Deployed over HTTP, mnemo speaks Streamable HTTP transport and is OAuth-gated. Point any MCP client that supports remote HTTP + OAuth at https://<your-host>/mcp and authenticate on first connect; each authenticated user gets an isolated per-user credential store (see Trust Model). To stand up an instance, see Deploy to Cloudflare.
Public OCI image publication is discontinued. Existing historical registry tags remain untouched; new container deployments build from source or use the Cloudflare-managed registry.
Deploy to Cloudflare
Run your own mnemo instance serverless on Cloudflare (Containers + D1 + Vectorize + KV).
Paused 2026-09-13 (maintained instance only): the CF deploy token was removed from the account as off-manifest (process violation), so the CD
deploy-cfjob no-ops behind theCF_DEPLOY_ENABLEDrepo variable. The maintained instance freezes at its last deployed release until a token is re-established via the documented process and the variable is set totrue. Self-hosting on your own account (below) is unaffected.
Prerequisites: a Cloudflare account on the Workers Paid plan — required for Containers, D1, and Vectorize (the Cloudflare free tier does not include them) — and the wrangler CLI.
git clone https://github.com/n24q02m/mnemo && cd mnemowrangler login- Provision the storage bindings mnemo uses -- the memories database, the embedding
index, and the encrypted credential store:
Paste the returned D1 database ID and KV namespace ID into
wrangler.jsonc(the Vectorize index binds by name, so no ID is needed), then create the memories schema (tables, indexes, and the FTS5 full-text index) in the database you just made: The SQL lives inmigrations/0001_init.sql, and the D1 binding inwrangler.jsoncpoints at that folder viamigrations_dir: "migrations". Full-text search uses FTS5, which D1 ships; vector similarity is served by Vectorize rather than by an in-database extension, because D1 cannot load one. - Build the HTTP container from this checkout and push it to your Cloudflare managed registry (CF Containers cannot pull from external registries directly), then set
<YOUR_ACCOUNT_ID>inwrangler.jsonc: - Set
<YOUR_PUBLIC_URL>(e.g.https://mnemo.example.com) and<YOUR_WORKER_DOMAIN>(e.g.mnemo.example.com) inwrangler.jsonc, then set the secrets: wrangler deployand complete setup in the browser relay form at your Worker domain. Save each subject's models, endpoints and provider keys there, not in Worker environment variables. The managed route uses Minimax-free completion and paid Cohere embedding/reranking through Cloudflare AI Gateway -- obtain the required budget authorization before exercising the paid tiers (Provider Spend Gate); see the per-task configuration. Storage maps to Cloudflare viaMCP_STORAGE_BACKEND=cf-kv(credentials / tokens, encrypted),MEMORY_DB_BACKEND=cf-d1(the memories database + FTS5 full-text; unset orsqlitekeeps the local SQLite file atDB_PATH), and Vectorize (embeddings, cosine). Cloud embedding, reranking and completion resolve per authenticated subject. Remote startup does not probe shared provider credentials or download local Fastretrieval models. Missing subject configuration never selects a process-wide provider/model fallback.
Authority & Sync Boundary
On Cloudflare deployments, Cloudflare D1 + Vectorize + KV is the sole production authority:
- D1 (
MEMORY_DB_BACKEND=cf-d1): Authoritative storage for memory rows, metadata, bitemporal valid ranges, and FTS5 search. - Vectorize (
MCP_VECTORIZE_IDX): Dense vector index for semantic similarity search. - KV (
MCP_STORAGE_BACKEND=cf-kv): Encrypted per-user credential and session store. - Sync boundary:
MEMORY_DB_BACKEND=cf-d1disables Google Drive OAuth and all external sync paths even ifSYNC_ENABLEDis toggled on or stale S3/Google settings remain.SYNC_ENABLED=falseindependently disables sync on non-CF deployments. - Local & self-host bootstrap: Local stdio (
~/.mnemo/memories.db) and self-hosted instances retain optional passport sync (Google Drive Device Code OAuth or S3/R2/B2) for workstation migration.
Deployment (maintained instance)
Every tagged release deploys automatically: the CD deploy-cf job checks out
the released tag, builds the http-slim image, pushes it to the Cloudflare-managed
registry as immutable :<release-tag>, deploys the Worker, and gates on a canary
health check -- a release is live at exactly its own version. A beta dispatch
redeploys the beta; a stable dispatch is maintainer-gated. Manual wrangler deploy
against the maintained instance is not permitted: it would break the
release-tag ↔ live-image correspondence. Self-hosting on your own Cloudflare
account (the button above) is unaffected.
Trust Model
This plugin implements TC-Local (machine-bound, single trust principal). The mode/storage/encryption breakdown below is the full classification.
Workspace username (HTTP setup form)
The browser setup form has an optional workspace username field. Entering the
same username always lands you in the same per-sub bucket, so your credentials
and memories stay reachable across a re-authorization and across devices, instead
of being tied to the one-off subject minted for each /authorize round-trip.
Leaving it blank keeps the previous per-authorize behaviour.
Trust boundary: when the form is gated by a shared MCP_RELAY_PASSWORD, the
username is a partition key, not a secret -- anyone who knows that password can
type any username and reach that bucket. That is fine for a trusted group; an
untrusted multi-tenant deployment needs a per-user secret or delegated OAuth
instead.
One-time migration: existing users must re-enter their credentials once after this change. Nothing is deleted; credentials stored under the old random subject are simply no longer addressed.
License
Apache-2.0 -- See LICENSE.
Source: README.md at commit b0ebd04
Tools
0Version history
1- v2.19.1LatestOct 5, 2026

