
Clef Mcp
io.github.HighlyLoadedEgov0.2.1Updated Oct 3, 2026
Local MCP server for the Cloudflare Clef-Flash decision model: structured decisions, fully offline
Overview
Runs a local Clef-Flash decision model so an assistant can get probability distributions over options instead of prose.
- What it does
- Exposes one tool, clef_decide, that takes a state plus typed questions and returns strict JSON probabilities per option in a single forward pass. Question types are choice, score and noul (yes/no). Up to 64 questions can be batched per call, and the server also ships four decision-shaped prompts (incident-triage, next-action, ticket-routing, security-review) and three read-only resources describing capabilities and the bundled eval dataset.
- When to use it
- Useful at decision points inside agent loops — next action, routing, classification, severity, yes/no judgment — where a small, deterministic, offline probability estimate is preferable to asking a chat model. Best when the same kind of decision is asked many times per task and privacy or token cost matters.
- Requirements
- A local Node.js runtime to run the npm package over stdio. The model (~6 GB) and a llama.cpp runtime are fetched by an explicit install command; the MLX runtime is macOS/Apple Silicon only and needs uv on PATH. Optional configuration uses environment variables such as CLEF_HOME, CLEF_RUNTIME, CLEF_MODEL and CLEF_LLAMA_BIN. No accounts or API keys.
Installation
In SourceWeft
- Open Clef Mcp in the dashboard and add it to a workspace.
- Enable the server for the chats that should use its tools.
Desktop only via STDIO. STDIO servers start a local process, so they need the SourceWeft desktop host.
Other MCP clients
Follow the launch instructions in the repository.
README
clef-mcp
[npm version] [npm downloads] [CI] [Official MCP Registry] [License: Apache-2.0]
A "reflex" for AI coding agents: structured decisions with probabilities — not prose.
Local MCP server that gives agents (Claude Code, Codex, Cursor, ZCode, …) access to the Clef-Flash decision model (9B, Apache-2.0 by Cloudflare) through a single tool: clef_decide. Pass a state and typed questions, get a probability distribution over your options in one forward pass. Fully local, offline, no tokens burned.
Quick start
clef-mcp (step 3) never downloads anything. If the model is missing, tool calls return a structured MODEL_NOT_INSTALLED error with a hint. Steps 1 and 2 combine: npx clef-mcp install --setup.
Why not just ask the LLM?
Sweet spot: decision points inside agent loops — next action, routing, classification, severity, yes/no judgment — asked dozens of times per task.
The clef_decide tool
Response — strictly structured, never prose:
Batch up to 64 questions per call — they are scored in one forward pass. state is treated strictly as data: never executed, never interpreted as instructions for the server.
Prompts & resources
The server ships four MCP prompts (canned, decision-shaped asks — your client lists them via prompts/list):
And three resources (read-only, no model needed):
Measured, not marketed
Apple M4 Pro, Clef-Flash Q4_K_M (6 GB), single request through the full MCP stdio path:
Quality gate: a 30-case evaluation dataset (coding / security / classification / routing / yes-no) — 86.7% pass on the live model. Run it yourself: clef-mcp evals.
Register with your MCP client
Claude Code
Codex — ~/.codex/config.toml
Cursor — .cursor/mcp.json
ZCode — ~/.zcode/cli/config.json (user scope, auto-connect)
Ready-made snippets: examples/.
Teach your agent (skill)
The schema tells the client what clef_decide accepts; agents also need to know when to reach for it and how to frame decisions. The bundled clef-decisions skill covers decision patterns, batching, criteria writing, distribution interpretation and error recovery:
Architecture
The ClefRuntime interface (load / decide / unload / health) isolates the engine: MLX or remote runtimes plug in without changing the MCP API. The wire format is POST /v1/systemone — the same contract across llama.cpp and other Clef runtimes.
CLI
Flags: install --quant Q8_0 --yes --skip-probe, install --setup, setup --clients zcode,cursor --no-skill, doctor --deep (re-hash the model file), uninstall --runtime --yes.
Runtimes: llama.cpp and MLX
Two local runtimes behind the same ClefRuntime interface:
Switch at runtime with CLEF_RUNTIME=mlx (must be set for the MCP server process — e.g. in the client's env block). Both speak the same POST /v1/systemone contract. The MLX snapshot is fetched into CLEF_HOME via a uv-managed huggingface_hub (no global Python state) at a pinned revision.
Configuration
Runtime resolution: CLEF_LLAMA_BIN → managed binary in CLEF_HOME/runtime → llama-server on PATH.
Storage: CLEF_HOME/models/<model>/<quant>/ (model + manifest.json with repo/revision/sha256/license) and CLEF_HOME/runtime/llama.cpp/.
The model is downloaded from the pinned official GGUF conversion (ggml-org/Clef-Flash-GGUF) and sha256-verified against Hugging Face's content hash. It is never repackaged by clef-mcp. Also listed in the official MCP Registry as io.github.HighlyLoadedEgo/clef-mcp.
Error handling
Codes: MODEL_NOT_INSTALLED, MODEL_LOAD_FAILED, RUNTIME_NOT_FOUND, RUNTIME_INIT_FAILED, RUNTIME_NOT_SUPPORTED, INVALID_INPUT, CLEF_INFERENCE_FAILED, UNSUPPORTED_PLATFORM, OUT_OF_MEMORY, CHECKSUM_MISMATCH, DOWNLOAD_FAILED. Input exceeding the 16k-token model context is rejected with a hint to reduce the state.
Security & data handling
- No network servers, no telemetry, no accounts; everything runs locally.
- The model downloads only on an explicit
install, over HTTPS, checksum-verified. statecontent is passed to the model as data; the server never executes or instruction-interprets it.- Filesystem access is limited to
CLEF_HOME(plus reading standard MCP client config paths indoctor). - The managed runtime is the official llama.cpp build; pin it with
CLEF_LLAMA_RELEASE_TAG.
See SECURITY.md for the full policy.
Development
See CONTRIBUTING.md and tests/evals/dataset.jsonl.
License
- Code: Apache-2.0.
- Clef / Clef-Flash model: © Cloudflare, Apache-2.0 — see NOTICE.
- llama.cpp runtime: © its authors, MIT-licensed; downloaded as an official prebuilt binary.
Source: README.md at commit 4e83764
Tools
0Version history
1- v0.2.1LatestOct 3, 2026


