
Clef Mcp
io.github.HighlyLoadedEgov0.2.1更新於 Oct 3, 2026
Local MCP server for the Cloudflare Clef-Flash decision model: structured decisions, fully offline
概覽
在本機執行 Clef-Flash 決策模型,讓助理取得各選項的機率分布,而不是散文式回答。
- 功能
- 只提供一個工具 clef_decide:傳入 state 與帶類型的問題,一次前向運算就回傳每個選項機率的嚴格 JSON。問題類型有 choice、score 與 noul(是/否)。每次呼叫最多可批次送出 64 個問題;伺服器另外附帶四個決策類提示(incident-triage、next-action、ticket-routing、security-review),以及三個唯讀資源,用來說明能力與內建評測資料集。
- 適用情境
- 適合代理迴圈中的決策點——下一步動作、路由、分類、嚴重程度、是/否判斷——在這些情境下,小型、確定性、離線的機率估計比詢問聊天模型更合適。當同一類決策在每個任務中被反覆詢問,且在意的隱私或 token 成本時尤其適用。
- 執行需求
- 需要本機 Node.js 執行環境,透過 stdio 執行這個 npm 套件。模型(約 6 GB)與 llama.cpp 執行環境由明確的安裝指令下載;MLX 執行環境僅支援 macOS/Apple Silicon,且需要 PATH 中有 uv。選用設定透過環境變數,例如 CLEF_HOME、CLEF_RUNTIME、CLEF_MODEL、CLEF_LLAMA_BIN。不需要帳號或 API 金鑰。
安裝
在 SourceWeft 中
- 開啟 儀表板中的 Clef Mcp,將其新增到工作區。
- 為需要使用其工具的對話啟用該服務。
Desktop only,透過 STDIO。 STDIO 服務會啟動本機處理程序,因此需要 SourceWeft 桌面主機。
其他 MCP 客戶端
參照 儲存庫 中的啟動說明。
README
clef-mcp
[npm version] [npm downloads] [CI] [Official MCP Registry] [License: Apache-2.0]
A "reflex" for AI coding agents: structured decisions with probabilities — not prose.
Local MCP server that gives agents (Claude Code, Codex, Cursor, ZCode, …) access to the Clef-Flash decision model (9B, Apache-2.0 by Cloudflare) through a single tool: clef_decide. Pass a state and typed questions, get a probability distribution over your options in one forward pass. Fully local, offline, no tokens burned.
Quick start
clef-mcp (step 3) never downloads anything. If the model is missing, tool calls return a structured MODEL_NOT_INSTALLED error with a hint. Steps 1 and 2 combine: npx clef-mcp install --setup.
Why not just ask the LLM?
Sweet spot: decision points inside agent loops — next action, routing, classification, severity, yes/no judgment — asked dozens of times per task.
The clef_decide tool
Response — strictly structured, never prose:
Batch up to 64 questions per call — they are scored in one forward pass. state is treated strictly as data: never executed, never interpreted as instructions for the server.
Prompts & resources
The server ships four MCP prompts (canned, decision-shaped asks — your client lists them via prompts/list):
And three resources (read-only, no model needed):
Measured, not marketed
Apple M4 Pro, Clef-Flash Q4_K_M (6 GB), single request through the full MCP stdio path:
Quality gate: a 30-case evaluation dataset (coding / security / classification / routing / yes-no) — 86.7% pass on the live model. Run it yourself: clef-mcp evals.
Register with your MCP client
Claude Code
Codex — ~/.codex/config.toml
Cursor — .cursor/mcp.json
ZCode — ~/.zcode/cli/config.json (user scope, auto-connect)
Ready-made snippets: examples/.
Teach your agent (skill)
The schema tells the client what clef_decide accepts; agents also need to know when to reach for it and how to frame decisions. The bundled clef-decisions skill covers decision patterns, batching, criteria writing, distribution interpretation and error recovery:
Architecture
The ClefRuntime interface (load / decide / unload / health) isolates the engine: MLX or remote runtimes plug in without changing the MCP API. The wire format is POST /v1/systemone — the same contract across llama.cpp and other Clef runtimes.
CLI
Flags: install --quant Q8_0 --yes --skip-probe, install --setup, setup --clients zcode,cursor --no-skill, doctor --deep (re-hash the model file), uninstall --runtime --yes.
Runtimes: llama.cpp and MLX
Two local runtimes behind the same ClefRuntime interface:
Switch at runtime with CLEF_RUNTIME=mlx (must be set for the MCP server process — e.g. in the client's env block). Both speak the same POST /v1/systemone contract. The MLX snapshot is fetched into CLEF_HOME via a uv-managed huggingface_hub (no global Python state) at a pinned revision.
Configuration
Runtime resolution: CLEF_LLAMA_BIN → managed binary in CLEF_HOME/runtime → llama-server on PATH.
Storage: CLEF_HOME/models/<model>/<quant>/ (model + manifest.json with repo/revision/sha256/license) and CLEF_HOME/runtime/llama.cpp/.
The model is downloaded from the pinned official GGUF conversion (ggml-org/Clef-Flash-GGUF) and sha256-verified against Hugging Face's content hash. It is never repackaged by clef-mcp. Also listed in the official MCP Registry as io.github.HighlyLoadedEgo/clef-mcp.
Error handling
Codes: MODEL_NOT_INSTALLED, MODEL_LOAD_FAILED, RUNTIME_NOT_FOUND, RUNTIME_INIT_FAILED, RUNTIME_NOT_SUPPORTED, INVALID_INPUT, CLEF_INFERENCE_FAILED, UNSUPPORTED_PLATFORM, OUT_OF_MEMORY, CHECKSUM_MISMATCH, DOWNLOAD_FAILED. Input exceeding the 16k-token model context is rejected with a hint to reduce the state.
Security & data handling
- No network servers, no telemetry, no accounts; everything runs locally.
- The model downloads only on an explicit
install, over HTTPS, checksum-verified. statecontent is passed to the model as data; the server never executes or instruction-interprets it.- Filesystem access is limited to
CLEF_HOME(plus reading standard MCP client config paths indoctor). - The managed runtime is the official llama.cpp build; pin it with
CLEF_LLAMA_RELEASE_TAG.
See SECURITY.md for the full policy.
Development
See CONTRIBUTING.md and tests/evals/dataset.jsonl.
License
- Code: Apache-2.0.
- Clef / Clef-Flash model: © Cloudflare, Apache-2.0 — see NOTICE.
- llama.cpp runtime: © its authors, MIT-licensed; downloaded as an official prebuilt binary.
來源:README.md,提交 4e83764
工具
0版本歷史
1- v0.2.1最新Oct 3, 2026


