
llmwiki
io.github.atomicstratav1.4.1Updated Oct 2, 2026
Compile documents into a cited knowledge wiki. Retrieve evidence, query, and review via local MCP.
Overview
Compiles documents into a citation-traceable markdown knowledge wiki that an assistant can ingest, query, lint, and export over MCP.
- What it does
- llmwiki builds a durable wiki from raw sources such as papers, notes, transcripts, PDFs, images, or web pages. A two-phase LLM pipeline extracts concepts and generates typed pages, then hybrid retrieval (semantic chunk search, BM25 reranking, wikilink graph expansion) answers queries and builds evidence packs. Its MCP server exposes ingest, compile, query, read, lint, status, eval, context-pack, and Open Knowledge Format exchange tools. Generated pages carry citations, freshness, confidence, and review state.
- When to use it
- Use it when a reusable corpus of documents, research, or project knowledge should compound into reviewed, citable pages rather than being re-read from raw files each time. It suits research folders, codebase docs, team handbooks, standards, and decision logs. A one-off file read or live web search does not need a wiki.
- Requirements
- Local process started with npx from the npm package llm-wiki-compiler; Node.js 24 or newer. An LLM provider is needed for compile and query: by default Anthropic via ANTHROPIC_API_KEY or ANTHROPIC_AUTH_TOKEN, or another provider selected with LLMWIKI_PROVIDER (for example OPENAI_API_KEY, OLLAMA_HOST, GITHUB_TOKEN, ATLASCLOUD_API_KEY, ORCAROUTER_API_KEY). Read-only tools work without provider credentials.
Installation
In SourceWeft
- Open llmwiki in the dashboard and add it to a workspace.
- Enable the server for the chats that should use its tools.
Desktop only via STDIO. STDIO servers start a local process, so they need the SourceWeft desktop host.
Other MCP clients
Follow the launch instructions in the repository.
README
llmwiki
New in 1.4.0 — Review answers before publishing.
- Choose what becomes part of your wiki. Review generated answers and check their citations before saving them as published pages.
- Approve pages together. Review a batch of generated pages and publish the ones you choose in one operation.
- See the pages behind an answer. Follow citations back to your wiki and see which pages were used to answer your question.
Release notes · Upgrade guide · Install llmwiki
New in 1.3 — A fresh look for your wiki.
Meet Scientific Clay, with soft surfaces and rounded typography, and Minimal, which follows your system’s light or dark setting. Switch instantly between four themes, including Nebula Light and Dark.
The same page in a demonstration wiki. Click either screenshot for a closer look.
Recursive source folders, path exclusions, project-specific compile instructions, and storage for larger embedding indexes.
New in 1.2 — Explore your records and their evidence.
Browse profile-defined categories, declared fields, connected records, provenance, and supporting source passages in the local viewer.
New in 1.1 — Publish and share domain templates.
Create signed template distributions, discover them through explicitly trusted catalogs, and install or update them with compatibility checks.
New in 1.0 — Configurable Lifecycle Profiles.
Build a knowledge system around the way you work. Define your records, relationships, review gates, and workflows in one validated profile. Start with AutoSci for research or Newsroom for editorial work, or create your own.
Explore Configurable Lifecycle Profiles →
What llmwiki does
Compile raw sources into an interlinked, citation-traceable markdown wiki that agents and humans can browse, query, lint, export, and reuse. The default profile preserves the classic concepts-and-queries layout; optional profiles add domain-specific types and workflows without adding domain branches to the compiler.
llmwiki implements the LLM Wiki pattern: instead of re-discovering knowledge from raw files at query time, compile it once into durable pages that accumulate structure, provenance, review state, and retrieval metadata over time.
When to use this repo
Use llmwiki when you need a persistent knowledge base from raw material:
- Compile papers, notes, READMEs, transcripts, PDFs, images, or web pages into typed wiki pages.
- Give agents a stable, citation-aware context pack instead of a pile of loose files.
- Keep generated knowledge auditable with source citations, review queues, freshness checks, and quality gates.
- Browse the result locally, query it from the CLI, expose it over MCP, or embed it through the SDK.
- Exchange compiled knowledge with other tools using Open Knowledge Format (OKF), JSON, JSON-LD, GraphML, Marp, and
llms.txt.
Do not use llmwiki as a general static-site generator, a heavy ontology database, or a replacement for ad-hoc search over fast-changing raw logs. It is strongest when source knowledge is worth compiling, reviewing, and reusing.
What you get
- Compiled wiki, not chunks. A two-phase LLM pipeline extracts concepts, then generates typed pages:
concept,entity,comparison, andoverview. - Configurable Lifecycle Profiles. A fail-closed
.llmwiki/profile.jsoncan declare entity schemas, typed relations, lifecycle state machines, transition requirements, workflows, artifacts, connectors, content tiers, and retrieval policy. - Installable domain templates.
llmwiki template init autoscicreates a research project with papers, ideas, experiments, manuscripts, evidence artifacts, workflows, and Crossref import.newsroomdemonstrates the same machinery for editorial work. - Runtime trust gates. Relation, evidence, artifact, and human/agent gates are enforced by the write path rather than left as prompt conventions; standing lint detects drift after the fact.
- Citation-traceable output. Paragraphs and claims cite source files and line ranges, and
llmwiki lintvalidates the links. - Hybrid retrieval. Semantic chunk search, BM25 reranking, and wikilink graph expansion build compact evidence packs for queries and agents.
- Local viewer.
llmwiki viewopens a read-only browser UI with search, page metadata, graph exploration, source-freshness badges, and citation chips. - Review policy. Generated pages can be auto-held for review when confidence, contradiction, schema, or provenance rules trip.
- Freshness repair.
llmwiki lintandllmwiki nextsurface stale/orphaned pages;llmwiki refresh --stalerepairs changed knowledge without compiling unrelated new sources. - Eval harness.
llmwiki evalreports health score, a per-page health distribution that flags the worst pages, wikilink-graph health, citation coverage/precision, corpus stats, regression deltas, and optional judge-model citation support. - MCP server.
llmwiki serveexposes ingest, compile, query, lint, read, status, eval, context-pack, and OKF exchange tools to MCP-compatible agents. - SDK.
createWiki({ root })drives ingest, compile, query, context, status, export, eval, and OKF import/export from TypeScript without shelling out. - Open Knowledge Format exchange. Export and import OKF bundles for portable, markdown-native knowledge exchange. External OKF imports are staged through the review queue by default; trusted bundles can be written live explicitly.
- Other portable exports. Export JSON, JSON-LD, GraphML, Marp slides, and
llms.txtfor downstream systems. - Provider portable. Anthropic, Claude Agent SDK local login, OpenAI Codex CLI local login, OpenAI-compatible servers, Ollama, GitHub Copilot, Atlas Cloud, OrcaRouter, and local OpenAI-compatible runtimes.
Configurable Lifecycle Profiles (CLP)
CLP turns llmwiki's knowledge compiler into a reusable substrate for domain-specific knowledge systems. A validated .llmwiki/profile.json is the single contract for:
- typed entities, fields, and directed relations;
- lifecycle states, transition evidence, and trust gates;
- multi-stage workflows and declared actions;
- hash-pinned artifacts and first-party connector bindings; and
- content tiers and retrieval behavior.
These rules are enforced by the runtime, not left as prompt conventions. The CLI, SDK, MCP server, viewer, context builder, lint, status, export, and OKF exchange surfaces all operate from the same profile contract. Invalid profiles and writes that bypass a declared gate fail closed.
CLP is backward-compatible by construction: a project without .llmwiki/profile.json uses the built-in default concepts-and-queries profile and preserves the pre-1.0 behavior. You can start three ways — scaffold your own profile, install a built-in or local template, or install a signed template from a trusted tap:
autosci is a practical research system with papers, ideas, experiments, manuscripts, evidence artifacts, workflows, and Crossref ingestion. newsroom applies the same generic machinery to articles, desks, bylines, and editorial workflows. Templates contain configuration and examples, never executable plugin code.
Templates can also be distributed securely. Publishers build signed, offline distributions with llmwiki template publish — Ed25519 signing, key rotation, and package revocation — and verify them with template publish verify. Consumers add explicitly trusted taps, discover and inspect signed catalogs, and install or update templates with continuity, revocation, and compatibility checks enforced under lock.
Read the CLP concept guide, follow the AutoSci research workflow, or explore the Newsroom editorial workflow.
Karpathy's LLM Wiki pattern
Andrej Karpathy described the LLM Wiki pattern as a way to turn raw material into compiled knowledge that future agents can reuse. llmwiki is a concrete compiler for that pattern.
The key shift is moving work from query time to compile time. Traditional RAG repeatedly retrieves raw chunks and asks the model to reconstruct relationships for each question. llmwiki first turns sources into typed, interlinked pages with citations, metadata, and review state. Queries, context packs, exports, and MCP tools then operate over that compiled artifact.
That makes llmwiki useful when knowledge should compound: concepts shared across sources become one page, saved answers become future context, stale pages can be detected and repaired, and agents can consume a stable evidence pack instead of re-reading the same raw files from scratch.
See docs/concepts/karpathy-pattern.mdx for the deeper explanation.
Agent decision guide
Use llmwiki for a reusable corpus of documents, research, or project knowledge.
Start with wiki_status and get_context_pack when you need evidence for your
own reasoning. A one-off file read or live web search does not need a wiki.
The npm package includes an Agent Skill. See the
MCP setup guide for installation.
The MCP Registry identifier is io.github.atomicstrata/llmwiki; the npm package
is llm-wiki-compiler and the executable is llmwiki.
Choose the entry point that matches the task:
Quick start
quickstart ingests one source, compiles pages, and opens the viewer. Inside an existing project, run llmwiki next when you want the safest next action.
To start with a domain model instead of the default concepts-and-queries layout:
Template installation is for a new or empty typed project. It materializes the chosen profile into .llmwiki/profile.json; normal project loading never depends on a template registry or lockfile.
Demo
Try it on any article or document:
The examples/basic/ directory includes a small pre-generated wiki you can inspect without an API key.
Core commands
Full command docs live in docs/cli/.
Open Knowledge Format
llmwiki is an Open Knowledge Format (OKF) producer and consumer. OKF is a Google Cloud initiative for sharing compiled knowledge as portable markdown files with structured frontmatter.
OKF import is intentionally review-first: untrusted bundles become review candidates, not live wiki pages. The importer preserves foreign OKF metadata, stores llmwiki provenance under x-llmwiki, and re-exports imported pages honestly after local edits, including safe original nested paths.
See docs/guides/open-knowledge-format.mdx, docs/cli/export.mdx, and docs/cli/import.mdx.
What llmwiki creates
A project has raw inputs in sources/, compiled markdown in wiki/, and compiler state under .llmwiki/:
Compiled pages are plain markdown with YAML frontmatter, plus enough metadata for agents to reason about citations, freshness, confidence, contradictions, and review state. See docs/concepts/wiki-model.mdx.
Sources are top-level Markdown files by default. Opt into nested source folders and exclusions in project config. Excluding a compiled source retires its contribution on the next ordinary compile without deleting the source file.
Agent integration
MCP
Run:
MCP clients can ingest sources, compile, query, search pages, read pages, lint, run eval, inspect status, request context packs, and exchange OKF bundles. Read-only tools work without provider credentials; LLM-backed tools validate provider credentials at call time. The run_eval tool runs its fast suite without a provider; its full suite (which LLM-judges citation support) requires one.
See docs/guides/mcp-agent-integration.mdx.
SDK
See docs/guides/sdk.mdx. llm-wiki-compiler is the
supported entry point; its scoped supporting packages are implementation
dependencies. The local workflow engine and an application's own coordinator
are two tiers, not a migration path: llmwiki workflow runs a profile-declared
workflow through local invocations with persisted state between sessions.
An external coordinator calls the same SDK without using local workflow
execution methods; the supporting engine dependency remains installed. See
SDK package selection,
SDK upgrade notes, and
When to use the local engine.
Configuration
Minimum requirement: Node.js 24 or newer.
The default provider is Anthropic:
Provider selection is environment-driven:
See docs/configuration/providers.mdx and docs/configuration/environment-variables.mdx.
Quality and safety model
llmwiki is designed for auditable generated knowledge:
- Review before write. Use
compile --reviewor.llmwiki/config.jsonreview policy to hold risky pages as candidates. - Profile floors are runtime checks. Field contracts, lifecycle transitions, relation counts, evidence, and artifact requirements are enforced across page, lifecycle, workflow, import, and approval write surfaces.
- External connector data is untrusted. First-party connectors use confined fetches and stage fenced review candidates; approval is pinned to the exact body the operator reviewed.
- Artifacts are content-addressed evidence. Artifact reads and writes are path-confined, size-capped, schema-checked, and verified against hash-pinned references.
- Fail-closed config. Invalid review-policy config aborts compile instead of silently disabling review.
- Source confinement. Source snippets and import/export paths are confined to the project.
- Freshness is explicit. Pages can be fresh, stale, orphaned, or unverified; stale pages are flagged and repairable. The JSON export is active-page-only: it carries freshness for live pages (
fresh/stale/unverified); computed-orphaned pages (all sources deleted) surface only as lint and viewer signals and are dropped from the export. - Imported compiled knowledge is staged by default. External bundles go through the review queue unless explicitly trusted.
- CI gates are supported.
llmwiki lintandllmwiki evalcan enforce quality thresholds.
See docs/configuration/review-policy.mdx, docs/troubleshooting/stale-pages.mdx, and docs/guides/ci-quality-gates.mdx.
Scale and what works
llmwiki is still early software, but it is no longer a toy pipeline for a handful of notes.
- Incremental compilation means unchanged sources do not flow back through the LLM.
- Parallel compile runs concept extraction and page generation concurrently under a configurable cap (
--concurrency/LLMWIKI_COMPILE_CONCURRENCY), cutting wall-clock on large compiles. - Chunk-level embeddings narrow large wikis before BM25 reranking and graph expansion.
- Content-hash-aware embedding updates avoid recomputing vectors for unchanged pages and chunks.
- Batch embedding sends page and chunk vectors to the provider in batches rather than one request at a time, cutting latency on cold starts and large refreshes.
- Binary embedding storage automatically handles stores above the 64 MiB JSON limit; smaller stores can opt in. Binary selection persists, and retrieval still loads the index into memory. See limits, recovery and downgrade guidance.
- Cached citation judgements make repeated
eval --suite fullruns cheaper. - Lexical fallback keeps query/context workflows usable when the active provider has no embedding endpoint.
- Prompt budgeting and ingest truncation metadata make large sources explicit instead of silently pretending they fit.
The current sweet spot is a durable project or domain wiki: research folders, codebase docs, team handbooks, standards, design notes, decision logs, or curated source packs. The less ideal fit is a high-churn firehose where raw search is enough and compiled structure would go stale faster than it can be reviewed.
Documentation
The full docs site source is in docs/:
- Start here:
docs/introduction.mdx - Quickstart:
docs/quickstart.mdx - Installation:
docs/installation.mdx - Karpathy's LLM Wiki pattern:
docs/concepts/karpathy-pattern.mdx - How the compiler works:
docs/concepts/how-it-works.mdx - Wiki model:
docs/concepts/wiki-model.mdx - Configurable Lifecycle Profiles:
docs/concepts/configurable-lifecycle-profiles.mdx - AutoSci research workflow:
docs/guides/autosci-research-workflow.mdx - Newsroom editorial workflow:
docs/guides/newsroom-editorial-workflow.mdx - Profile templates:
docs/configuration/profile-templates.mdx - CLI reference:
docs/cli/ - Open Knowledge Format:
docs/guides/open-knowledge-format.mdx - MCP integration:
docs/guides/mcp-agent-integration.mdx - SDK:
docs/guides/sdk.mdx - SDK packages and upgrade notes:
docs/guides/sdk-packages.mdx,docs/guides/sdk-upgrade.mdx - Architecture and ownership boundaries:
ARCHITECTURE.md - Atomic Memory bridge:
docs/guides/atomic-memory-bridge.mdx
Preview the docs locally with Node 24:
Current release
Release 1.4.1: Documentation update for the 1.4.0 features below.
-
Keep personal Claude coding instructions out of generated pages and answers.
-
Clarify wiki-link alias targets; newly generated pages record prompt version
v6. -
Review answers before publication with
query --save --review, and approve candidate batches with one finalization. -
Reuse extraction for unchanged sources, share compile and query prompt prefixes, and inspect stage timing with an absolute
LLMWIKI_STAGE_TIMING_FILE. -
Report only pages supplied to the answer model; embedding failures fall back to page selection with a warning by default.
-
Keep embedding retry budgets tied to page content, with an opt-out that leaves embedding and retry stores untouched.
-
Add generic domain records, preparations, retained artifacts and local workflow composition through the standard SDK. The engine-free core and local engine ship as exact matching supporting dependencies.
-
Compare source excerpts and pending proposals in the experimental, read-only source-review cockpit, explicitly enabled through the SDK on loopback.
The compiler and its supporting packages ship together at 1.4.1 through the
npm latest dist-tag. Install the standard entry point:
Existing CLI and standard createWiki entry points remain available. Experimental
artifact types and workflow statuses can require source updates; read the
SDK upgrade notes before upgrading. Standalone
AutoSci/Newsroom process packages are separate; their builtin ontology templates
remain supported. See the changelog for the release's changes
and limits.
Released 1.3.0:
- Four viewer themes, with Scientific Clay as the new default and Minimal following your system setting. Saved light/dark preferences migrate to Nebula Light/Dark.
- Recursive source discovery, literal path exclusions, and explicit project-specific compile instructions.
- Bounded binary storage for larger embedding indexes, configurable completion budgets, and protected gateway request extensions.
- OrcaRouter embedding routing, retryable rule-language changes, and a public test type-check ratchet.
See the 1.3.0 release summary and changelog for migration details, contributor credits, and limits.
Released 1.2.0:
- Profile-aware viewer with document categories, record relationships, provenance, source previews, verified local artifact access, responsive navigation and fitted graphs.
- Additional providers and OpenAI-compatible reasoning-model controls, independent embedding configuration, and SDK embedding opt-out.
- Optional Sources sections and additive SDK system policies with prompt-change invalidation and per-page provenance.
- Shared-page recovery after source removal, citation line-list fixes, and conservative abbreviated-wikilink repair.
Windows path fixes are included, but native Windows validation and CI remain outstanding. Full filesystem-confinement guarantees on Windows remain unverified. Use a supported Linux environment for untrusted projects; see #217.
Released 1.1.0:
- Template distribution ecosystem: publishers author signed, offline distributions with
template publish init | add | build | rotate | revoke(Ed25519 signing, key rotation, package revocation) and verify them withtemplate publish verify. - Consumers configure explicitly trusted template taps, discover and inspect signed catalogs, and install or update templates with continuity, revocation, and compatibility checks under lock.
llmwiki statuscommand: a readable snapshot of page and source counts, last compile, stale and orphaned pages, pending changes, the review queue, active profile, and state-file health.
Released 1.0.0:
- Configurable Lifecycle Profiles across CLI, SDK, MCP, viewer, context, lint, status, export, and profile-aware OKF exchange.
- Built-in
autosciandnewsroomtemplates, typed workflows and actions, first-class artifacts, typed relations and runtime lifecycle gates, plus a hardened first-party connector substrate with Crossref. - Typed-page semantic search and retrieval controls, batch embeddings, parallel compile, and fail-closed state recovery.
See CHANGELOG.md for release history.
Companion: Atomic Memory
llmwiki and Atomic Memory are complementary open context infrastructure:
- llmwiki compiles source material into durable, inspectable knowledge.
- Atomic Memory gives agents runtime memory that is searchable, scoped, correctable, and inspectable.
Use them independently or together. The @atomicmemory/llmwiki bridge imports llmwiki export --target json --project-id <id> as durable memory records.
Contributing
Contributions are welcome. If llmwiki is missing something you need, open an issue or PR and describe the workflow you are trying to support - need-driven improvements are often the best ones. If you want to contribute more generally, roadmap items are a good place to start. For larger changes to core compile, review, import/export, or retrieval semantics, please start with an issue or design discussion so we can align on the contract first.
Before committing code changes, run:
See CONTRIBUTING.md.
License
MIT
Disclaimer
No LLMs were harmed in the making of this repo.
Source: README.md at commit 0db26f2
Tools
0Version history
1- v1.4.1LatestOct 2, 2026


