Clef Mcp

io.github.HighlyLoadedEgov0.2.1Updated Oct 3, 2026

Local MCP server for the Cloudflare Clef-Flash decision model: structured decisions, fully offline

VerifiedSTDIODesktop onlyDeveloper ToolsAI & ML

Overview

AI-generated overview

Runs a local Clef-Flash decision model so an assistant can get probability distributions over options instead of prose.

What it does
Exposes one tool, clef_decide, that takes a state plus typed questions and returns strict JSON probabilities per option in a single forward pass. Question types are choice, score and noul (yes/no). Up to 64 questions can be batched per call, and the server also ships four decision-shaped prompts (incident-triage, next-action, ticket-routing, security-review) and three read-only resources describing capabilities and the bundled eval dataset.
When to use it
Useful at decision points inside agent loops — next action, routing, classification, severity, yes/no judgment — where a small, deterministic, offline probability estimate is preferable to asking a chat model. Best when the same kind of decision is asked many times per task and privacy or token cost matters.
Requirements
A local Node.js runtime to run the npm package over stdio. The model (~6 GB) and a llama.cpp runtime are fetched by an explicit install command; the MLX runtime is macOS/Apple Silicon only and needs uv on PATH. Optional configuration uses environment variables such as CLEF_HOME, CLEF_RUNTIME, CLEF_MODEL and CLEF_LLAMA_BIN. No accounts or API keys.
Before you install
The install step downloads roughly 6 GB of model and runtime over the network and writes them under CLEF_HOME; the server itself does not download. Filesystem access is limited to CLEF_HOME plus reading MCP client config paths during doctor. State content is sent to the local model as data. The managed llama.cpp build can be pinned with CLEF_LLAMA_RELEASE_TAG.

Installation

In SourceWeft

  1. Open Clef Mcp in the dashboard and add it to a workspace.
  2. Enable the server for the chats that should use its tools.

Desktop only via STDIO. STDIO servers start a local process, so they need the SourceWeft desktop host.

Other MCP clients

Follow the launch instructions in the repository.

README

[clef_decide — probability distributions for a production incident]

clef-mcp

[npm version] [npm downloads] [CI] [Official MCP Registry] [License: Apache-2.0]

A "reflex" for AI coding agents: structured decisions with probabilities — not prose.

Local MCP server that gives agents (Claude Code, Codex, Cursor, ZCode, …) access to the Clef-Flash decision model (9B, Apache-2.0 by Cloudflare) through a single tool: clef_decide. Pass a state and typed questions, get a probability distribution over your options in one forward pass. Fully local, offline, no tokens burned.


Quick start

bash
# 1. Detect hardware, download the model (~6 GB) + llama.cpp runtime, verify checksum + inferencenpx clef-mcp install
# 2. Register the MCP server + agent skill in your clients (zcode, claude-code, codex, cursor)npx clef-mcp setup
# 3. Run the MCP server (stdio)npx clef-mcp

clef-mcp (step 3) never downloads anything. If the model is missing, tool calls return a structured MODEL_NOT_INSTALLED error with a hint. Steps 1 and 2 combine: npx clef-mcp install --setup.

Why not just ask the LLM?

Chat LLMclef_decide
Outputprose, you parse itstrict JSON: probability per option
Determinismvaries per runsingle forward pass, no sampling
Latency (1 decision)seconds of generation~0.5 s local
Context costgrows with every decisionfixed, small schema
Privacydepends on provider100% on-device, works offline
Calibrationvibessoftmax over trained option scores

Sweet spot: decision points inside agent loops — next action, routing, classification, severity, yes/no judgment — asked dozens of times per task.

The clef_decide tool

json
{  "state": {    "task": "Fix failing tests",    "error": "TypeError: Cannot read properties of undefined"  },  "questions": {    "next_action": {      "type": "choice",      "instructions": "What should the coding agent do next?",      "criteria": {        "inspect": "Inspect the code and gather more information",        "modify": "Modify the code",        "test": "Run additional tests",        "ask_user": "Ask the user for clarification"      }    },    "confidence": {      "type": "score",      "instructions": "How confident are you in this decision?",      "criteria": ["very_low", "low", "medium", "high", "very_high"]    },    "is_outage": { "type": "noul", "instructions": "Is a service down?" }  }}
TypeCriteriaAnswer
choicemap option id → description, or a plain listprobability per option
scoreordered list (index = score)probability per level
nouloptional {"true": "...", "false": "..."}{"true": p, "false": 1-p}

Response — strictly structured, never prose:

json
{  "model": "clef-flash",  "decisions": {    "next_action": { "answer": { "inspect": 0.72, "modify": 0.12, "test": 0.14, "ask_user": 0.02 } },    "confidence":  { "answer": { "very_low": 0.01, "low": 0.04, "medium": 0.18, "high": 0.61, "very_high": 0.16 } },    "is_outage":   { "answer": { "true": 0.9, "false": 0.1 } }  }}

Batch up to 64 questions per call — they are scored in one forward pass. state is treated strictly as data: never executed, never interpreted as instructions for the server.

Prompts & resources

The server ships four MCP prompts (canned, decision-shaped asks — your client lists them via prompts/list):

PromptPurpose
incident-triageaction + severity + user-impact questions for a production incident
next-actionwhat the coding agent should do next + confidence
ticket-routingclassify a message into a team + urgency
security-reviewvulnerability yes/no, risk scale, first mitigation

And three resources (read-only, no model needed):

URIContents
clef-mcp://capabilitieslive JSON: model, runtime, limits, error codes
clef-mcp://evals/schemahow to write eval cases
clef-mcp://evals/datasetthe bundled 30-case dataset

Measured, not marketed

Apple M4 Pro, Clef-Flash Q4_K_M (6 GB), single request through the full MCP stdio path:

ScenarioLatency
Cold start (incl. model load, once per session)~4.4 s
1 question~0.5 s
10 questions, one call~2.9 s
64 questions, one call~18.6 s

Quality gate: a 30-case evaluation dataset (coding / security / classification / routing / yes-no) — 86.7% pass on the live model. Run it yourself: clef-mcp evals.

Register with your MCP client

Claude Code
bash
claude mcp add clef-mcp -- clef-mcp# or, without a global install:claude mcp add clef-mcp -- npx -y clef-mcp
Codex — ~/.codex/config.toml
toml
[mcp_servers.clef-mcp]command = "clef-mcp"args = []
Cursor — .cursor/mcp.json
json
{  "mcpServers": {    "clef-mcp": { "command": "clef-mcp", "args": [] }  }}
ZCode — ~/.zcode/cli/config.json (user scope, auto-connect)
json
{  "mcp": {    "servers": {      "clef-mcp": { "command": "clef-mcp", "args": [], "type": "stdio" }    }  }}

Ready-made snippets: examples/.

Teach your agent (skill)

The schema tells the client what clef_decide accepts; agents also need to know when to reach for it and how to frame decisions. The bundled clef-decisions skill covers decision patterns, batching, criteria writing, distribution interpretation and error recovery:

bash
clef-mcp setup                                                    # automaticcp -r skills/clef-decisions ~/.agents/skills/                     # manual, from repocp -r "$(npm root -g)/clef-mcp/skills/clef-decisions" ~/.agents/skills/  # from npm package

Architecture

mermaid
flowchart LR    subgraph clients [MCP clients]        CC[Claude Code]        CX[Codex]        CU[Cursor]        ZC[ZCode]    end    clients -- MCP stdio --> S[clef-mcp<br/>validation · limits · structured errors]    S -- SystemOne adapter --> R[ClefRuntime<br/>llama.cpp subprocess<br/>127.0.0.1]    R -- single forward pass --> M[("Clef-Flash<br/>9B · GGUF · local")]    M -. probabilities .-> S -. strict JSON .-> clients

The ClefRuntime interface (load / decide / unload / health) isolates the engine: MLX or remote runtimes plug in without changing the MCP API. The wire format is POST /v1/systemone — the same contract across llama.cpp and other Clef runtimes.

CLI

bash
clef-mcp              # run the MCP server on stdio (default command)clef-mcp install      # detect hardware → download model + runtime → verify checksum → verify inferenceclef-mcp setup        # register the MCP server + agent skill in zcode / claude-code / codex / cursorclef-mcp models       # list models/quantizations and install statusclef-mcp status       # runtime, model, memory summaryclef-mcp doctor       # full diagnosis (platform, RAM, GPU, binary, model, checksum*, inference, MCP config)clef-mcp uninstall    # remove the model (and optionally the managed runtime)clef-mcp evals        # run the evaluation dataset against the installed model

Flags: install --quant Q8_0 --yes --skip-probe, install --setup, setup --clients zcode,cursor --no-skill, doctor --deep (re-hash the model file), uninstall --runtime --yes.

Runtimes: llama.cpp and MLX

Two local runtimes behind the same ClefRuntime interface:

llama-cpp (default)mlx
PlatformsmacOS, Linux, WindowsmacOS / Apple Silicon only
ModelGGUF from ggml-org/Clef-Flash-GGUFMLX 4-bit from mlx-community/clef-flash-4bit
Extrasnoneuv on PATH (managed Python env)
Installclef-mcp installclef-mcp install --runtime mlx

Switch at runtime with CLEF_RUNTIME=mlx (must be set for the MCP server process — e.g. in the client's env block). Both speak the same POST /v1/systemone contract. The MLX snapshot is fetched into CLEF_HOME via a uv-managed huggingface_hub (no global Python state) at a pinned revision.

Configuration

VariableDefaultMeaning
CLEF_MODELclef-flashModel id (per-call model also accepted)
CLEF_HOME~/.cache/clef-mcpCache/model home
CLEF_RUNTIMEllama-cppllama-cpp | mlx
CLEF_LOG_LEVELerrorerror | warn | info | debug (stderr only)
CLEF_LLAMA_BIN–Explicit llama-server binary path (llama-cpp runtime)
CLEF_LLAMA_RELEASE_TAGlatest nightlyPin the managed llama.cpp build
CLEF_LLAMA_BATCH8192llama.cpp physical batch (multi-question requests)
CLEF_MLX_UVuv on PATHExplicit uv binary (MLX runtime)
CLEF_MAX_QUESTIONS64Max questions per call
CLEF_MAX_STATE_BYTES1048576Max serialized state size
CLEF_MAX_INSTRUCTION_CHARS10000Max chars per question instructions

Runtime resolution: CLEF_LLAMA_BIN → managed binary in CLEF_HOME/runtime → llama-server on PATH.

Storage: CLEF_HOME/models/<model>/<quant>/ (model + manifest.json with repo/revision/sha256/license) and CLEF_HOME/runtime/llama.cpp/.

The model is downloaded from the pinned official GGUF conversion (ggml-org/Clef-Flash-GGUF) and sha256-verified against Hugging Face's content hash. It is never repackaged by clef-mcp. Also listed in the official MCP Registry as io.github.HighlyLoadedEgo/clef-mcp.

Error handling

json
{  "error": {    "code": "MODEL_NOT_INSTALLED",    "message": "Clef model \"clef-flash\" is not installed.",    "hint": "Run `clef-mcp install`."  }}

Codes: MODEL_NOT_INSTALLED, MODEL_LOAD_FAILED, RUNTIME_NOT_FOUND, RUNTIME_INIT_FAILED, RUNTIME_NOT_SUPPORTED, INVALID_INPUT, CLEF_INFERENCE_FAILED, UNSUPPORTED_PLATFORM, OUT_OF_MEMORY, CHECKSUM_MISMATCH, DOWNLOAD_FAILED. Input exceeding the 16k-token model context is rejected with a hint to reduce the state.

Security & data handling

  • No network servers, no telemetry, no accounts; everything runs locally.
  • The model downloads only on an explicit install, over HTTPS, checksum-verified.
  • state content is passed to the model as data; the server never executes or instruction-interprets it.
  • Filesystem access is limited to CLEF_HOME (plus reading standard MCP client config paths in doctor).
  • The managed runtime is the official llama.cpp build; pin it with CLEF_LLAMA_RELEASE_TAG.

See SECURITY.md for the full policy.

Development

bash
npm installnpm run buildnpm test          # unit + integration (fake llama-server, no model needed)npm run evals     # needs an installed model; exit code reflects pass rate

See CONTRIBUTING.md and tests/evals/dataset.jsonl.

License

  • Code: Apache-2.0.
  • Clef / Clef-Flash model: © Cloudflare, Apache-2.0 — see NOTICE.
  • llama.cpp runtime: © its authors, MIT-licensed; downloaded as an official prebuilt binary.

Source: README.md at commit 4e83764

Tools

0
Tool metadata has not been indexed yet.

Version history

1
  1. v0.2.1LatestOct 3, 2026