
Clef Mcp
io.github.HighlyLoadedEgov0.2.1更新于 Oct 3, 2026
Local MCP server for the Cloudflare Clef-Flash decision model: structured decisions, fully offline
概览
在本地运行 Clef-Flash 决策模型,让助手得到各选项的概率分布,而不是散文式回答。
- 功能
- 只暴露一个工具 clef_decide:传入 state 和带类型的问题,一次前向计算即返回每个选项概率的严格 JSON。问题类型有 choice、score 和 noul(是/否)。每次调用最多可批量提交 64 个问题;服务器还附带四个决策类提示(incident-triage、next-action、ticket-routing、security-review)以及三个只读资源,用于说明能力与内置评测数据集。
- 适用场景
- 适合智能体循环中的决策点——下一步动作、路由、分类、严重程度、是/否判断——在这些场景下,小型、确定性、离线的概率估计比询问聊天模型更合适。当同一类决策在每个任务中被反复询问,且在意隐私或 token 成本时尤为适用。
- 运行要求
- 需要本地 Node.js 运行环境,通过 stdio 运行该 npm 包。模型(约 6 GB)和 llama.cpp 运行时由显式的安装命令下载;MLX 运行时仅支持 macOS/Apple Silicon,且需要 PATH 中有 uv。可选配置通过环境变量,如 CLEF_HOME、CLEF_RUNTIME、CLEF_MODEL、CLEF_LLAMA_BIN。无需账号或 API 密钥。
安装
在 SourceWeft 中
- 打开 控制台中的 Clef Mcp,将其添加到工作区。
- 为需要使用其工具的对话启用该服务。
Desktop only,通过 STDIO。 STDIO 服务会启动本地进程,因此需要 SourceWeft 桌面宿主。
其他 MCP 客户端
参照 仓库 中的启动说明。
README
clef-mcp
[npm version] [npm downloads] [CI] [Official MCP Registry] [License: Apache-2.0]
A "reflex" for AI coding agents: structured decisions with probabilities — not prose.
Local MCP server that gives agents (Claude Code, Codex, Cursor, ZCode, …) access to the Clef-Flash decision model (9B, Apache-2.0 by Cloudflare) through a single tool: clef_decide. Pass a state and typed questions, get a probability distribution over your options in one forward pass. Fully local, offline, no tokens burned.
Quick start
clef-mcp (step 3) never downloads anything. If the model is missing, tool calls return a structured MODEL_NOT_INSTALLED error with a hint. Steps 1 and 2 combine: npx clef-mcp install --setup.
Why not just ask the LLM?
Sweet spot: decision points inside agent loops — next action, routing, classification, severity, yes/no judgment — asked dozens of times per task.
The clef_decide tool
Response — strictly structured, never prose:
Batch up to 64 questions per call — they are scored in one forward pass. state is treated strictly as data: never executed, never interpreted as instructions for the server.
Prompts & resources
The server ships four MCP prompts (canned, decision-shaped asks — your client lists them via prompts/list):
And three resources (read-only, no model needed):
Measured, not marketed
Apple M4 Pro, Clef-Flash Q4_K_M (6 GB), single request through the full MCP stdio path:
Quality gate: a 30-case evaluation dataset (coding / security / classification / routing / yes-no) — 86.7% pass on the live model. Run it yourself: clef-mcp evals.
Register with your MCP client
Claude Code
Codex — ~/.codex/config.toml
Cursor — .cursor/mcp.json
ZCode — ~/.zcode/cli/config.json (user scope, auto-connect)
Ready-made snippets: examples/.
Teach your agent (skill)
The schema tells the client what clef_decide accepts; agents also need to know when to reach for it and how to frame decisions. The bundled clef-decisions skill covers decision patterns, batching, criteria writing, distribution interpretation and error recovery:
Architecture
The ClefRuntime interface (load / decide / unload / health) isolates the engine: MLX or remote runtimes plug in without changing the MCP API. The wire format is POST /v1/systemone — the same contract across llama.cpp and other Clef runtimes.
CLI
Flags: install --quant Q8_0 --yes --skip-probe, install --setup, setup --clients zcode,cursor --no-skill, doctor --deep (re-hash the model file), uninstall --runtime --yes.
Runtimes: llama.cpp and MLX
Two local runtimes behind the same ClefRuntime interface:
Switch at runtime with CLEF_RUNTIME=mlx (must be set for the MCP server process — e.g. in the client's env block). Both speak the same POST /v1/systemone contract. The MLX snapshot is fetched into CLEF_HOME via a uv-managed huggingface_hub (no global Python state) at a pinned revision.
Configuration
Runtime resolution: CLEF_LLAMA_BIN → managed binary in CLEF_HOME/runtime → llama-server on PATH.
Storage: CLEF_HOME/models/<model>/<quant>/ (model + manifest.json with repo/revision/sha256/license) and CLEF_HOME/runtime/llama.cpp/.
The model is downloaded from the pinned official GGUF conversion (ggml-org/Clef-Flash-GGUF) and sha256-verified against Hugging Face's content hash. It is never repackaged by clef-mcp. Also listed in the official MCP Registry as io.github.HighlyLoadedEgo/clef-mcp.
Error handling
Codes: MODEL_NOT_INSTALLED, MODEL_LOAD_FAILED, RUNTIME_NOT_FOUND, RUNTIME_INIT_FAILED, RUNTIME_NOT_SUPPORTED, INVALID_INPUT, CLEF_INFERENCE_FAILED, UNSUPPORTED_PLATFORM, OUT_OF_MEMORY, CHECKSUM_MISMATCH, DOWNLOAD_FAILED. Input exceeding the 16k-token model context is rejected with a hint to reduce the state.
Security & data handling
- No network servers, no telemetry, no accounts; everything runs locally.
- The model downloads only on an explicit
install, over HTTPS, checksum-verified. statecontent is passed to the model as data; the server never executes or instruction-interprets it.- Filesystem access is limited to
CLEF_HOME(plus reading standard MCP client config paths indoctor). - The managed runtime is the official llama.cpp build; pin it with
CLEF_LLAMA_RELEASE_TAG.
See SECURITY.md for the full policy.
Development
See CONTRIBUTING.md and tests/evals/dataset.jsonl.
License
- Code: Apache-2.0.
- Clef / Clef-Flash model: © Cloudflare, Apache-2.0 — see NOTICE.
- llama.cpp runtime: © its authors, MIT-licensed; downloaded as an official prebuilt binary.
来源:README.md,提交 4e83764
工具
0版本历史
1- v0.2.1最新Oct 3, 2026


