
ember
io.github.shapeandsharev0.8.0更新于 Oct 9, 2026
Calibrated advice for coding agents: typed questions in, probabilities out. Apple Silicon.
概览
为编码代理提供本地决策模型,用类型化问题换取校准后的概率,而不是生成文本。
- 功能
- Ember 在本地运行 Cloudflare 的 Clef-Flash 决策模型,并暴露一个 MCP 工具 advise:输入一个状态和一组类型化问题,返回每个选项的概率。问题类型包括 noul(是/否)、choice(具名选项)和 score(有序选项,返回期望分数)。它还提供 ember://guide 资源和 ember-advise 操作手册技能,教代理何时咨询以及如何解读答案。
- 适用场景
- 当代理需要校准的判断而非生成的文字时使用,例如判断缺陷报告的紧急程度、在推送前评估变更风险,或把工作分派给某个团队。它面向能够本地承载模型的 Apple Silicon Mac,或指向远程 ember 端点的客户端。
- 运行要求
- 本地使用需要 macOS 上的 Apple Silicon Mac(M 系列),flash(9B)需 32 GB 以上统一内存,full(27B)需 64 GB 以上,并在 Hugging Face 缓存中占用约 18 GiB 或 55 GiB 磁盘。需要由 uv 管理的 Python 3.12,从 PyPI 安装 ember-advise。可选远程使用需要 EMBER_SERVER_URL 和 EMBER_AUTH_TOKEN;托管在 NVIDIA/CUDA 上时使用 EMBER_DEVICE=cuda。
安装
在 SourceWeft 中
- 打开 控制台中的 ember,将其添加到工作区。
- 为需要使用其工具的对话启用该服务。
Desktop only,通过 STDIO。 STDIO 服务会启动本地进程,因此需要 SourceWeft 桌面宿主。
其他 MCP 客户端
参照 仓库 中的启动说明。
README
A local gut feeling for coding agents. ember runs a decision model — Cloudflare's
Clef-Flash, from its public Hugging Face repo
— on your Apple Silicon Mac and gives agents one MCP tool, advise: describe a situation,
ask typed questions, and get back a calibrated feeling about every option. It's a little
buddy for judgment calls — it advises; the agent decides.
Website: shapeandshare.github.io/ember
How it works
A decision model is not a chat model: it takes a state plus a schema of typed
questions and returns one probability per option, with no text generation. So ember
plugs into agents as a tool, while their reasoning stays on their normal LLM:
- The model server (
ember/serving/server.py) loads the model once and stays warm across agent sessions. - The MCP server (
ember/mcp/mcp_server.py) never loads the model; it starts the model server on the first tool call, so the MCP handshake stays instant.
Names
"Clef" always refers to Cloudflare's upstream model, never to this product. Select a model
with ember model pull <name> / EMBER_MODEL=<name>; run ember model list to see every
registered entry (flash, full).
A hosted deployment (e.g. Outerbounds) that supplies the model's S3 location directly at start time doesn't use this registry at all — see "Hosted deployment: a model location supplied at start time" below.
Verified
On a MacBook Pro M4 Max / 128 GB, torch 2.14.1, transformers 5.18.0, mcp 2.3:
Install (Apple Silicon)
Requirements
- Apple Silicon Mac (M-series) on macOS for local use (MPS). Intel Macs remain out of
scope. For a hosted deployment on NVIDIA/CUDA compute (e.g. Outerbounds), see "Hosted
deployment" below and
deployment/README.md. - Unified memory above the model's size: 32 GB or more for
flash(9B), 64 GB or more forfull(27B). Only 128 GB has been verified. - Disk: about 18 GiB for
flashor 55 GiB forfull, in Hugging Face's shared cache (~/.cache/huggingface). Config, state, and logs live in~/Library/Application Support/ember. - Python 3.12, managed by uv.
Keep --python 3.12: uv otherwise picks your newest interpreter, which the pinned
torch/transformers stack is not tested on.
This installs the latest released wheel from PyPI.
To pin a specific release tag from GitHub instead (no PyPI required):
The repository is public, so the git install needs no credentials. To use SSH instead,
install from git+ssh://[email protected]/shapeandshare/[email protected].
Installed ember before the rename, as gut? Run uv tool uninstall gut first: both
distributions provide the same commands.
Restart opencode and every repo on this machine gains ember_advise. The model server
stays lazy — it starts on the first tool call (or with ember start).
Per-project skill and policy (run once in each repo you want ember-aware agents):
The skill teaches the agent when to consult ember and how to ask; the AGENTS.md snippet adds a project-specific policy you customize (e.g. "check change risk before every push"). See Agent onboarding below for all harnesses and options.
flash targets the commit verified on MPS (17f0b0a) and full its release commit
(2f3de3d, not yet verified locally) as a download convenience, and torch/torchvision are
pinned to the tested minor series — but ember does not restrict itself to a hand-maintained
allowlist of individually verified weights; it runs any model that fits its loader contract
(constitution Article V, "Model Loading"). Set EMBER_MODEL_DIR to run another weights
directory.
Cloud
ember is Apple-Silicon-first locally, and supports NVIDIA GPUs for hosted deployment
(constitution Article VI, "Apple Silicon and CUDA") — three devices total: run it on an
Apple Silicon host with the memory above (MPS), on an NVIDIA GPU host (EMBER_DEVICE=cuda,
float16 — see "Hosted deployment: a model location supplied at start time" and
deployment/README.md for a full Outerbounds example), or on any host with the CPU fallback
(EMBER_DEVICE=cpu — float32, roughly twice the memory, much slower). The HTTP server binds
to loopback by default; to serve other machines set EMBER_HOST and set
EMBER_SERVER_AUTH_TOKEN to require Authorization: Bearer <token> (or any header name via
EMBER_AUTH_HEADER, e.g. x-api-key, on the client side). Ember does not terminate TLS —
front a remote-serving
deployment with a proxy. Clients point at it with EMBER_SERVER_URL. See
COMPATIBILITY.md and SECURITY.md.
Agent onboarding
Installing the tool is half the job; the other half is making agents want to consult it
at the right moments and read its answers sensibly. ember/agent_kit/ ships that
guidance through every channel each agent actually reads:
The skill is a playbook, not a reference card: the decision points worth consulting ember about, copy-paste question sets for intent and readiness, failure triage, change risk, routing, and effort, and starting confidence thresholds calibrated from observed model output (re-measured whenever the pinned model revision changes). The snippet ends with a project policy — edit it to wire ember into your own workflow, e.g. "check change risk before every push".
Agent-specific notes:
-
opencode reads skills from
.opencode/skills,.claude/skills, and.agents/skills, so install one copy per project to avoid duplicate listings. -
Kilo Code:
ember init --kilocodewrites themcp.emberentry tokilo.json(same config shape as opencode) and installs the skill to.kilo/skills/ember-advise/SKILL.md; the tool appears asember_advise. Kilo Code also readsAGENTS.mdautomatically. -
Claude Code: register the server with
claude mcp add --scope user ember -- ember-mcp(every project; drop--scope userto register it for the current project only); the tool appears asmcp__ember__advise. Or install the ember plugin, which bundles the MCP entry and theember-adviseskill (so skipember agents install --agent claude):The plugin still runs the
ember-mcpyou installed withuv tool install, and its tool appears asmcp__plugin_ember_ember__advise. Use one route, not both, or the tool is listed twice. -
Codex CLI:
ember init --codexwrites[mcp_servers.ember]to.codex/config.toml(--global:~/.codex/config.toml, or$CODEX_HOME/config.toml) and installs the skill to.agents/skills/ember-advise/SKILL.md. Codex loads a project.codex/config.tomlonly for a trusted project, andember init --codextells you whether this one is. The entry forwardsEMBER_AUTH_TOKENby name (env_vars), never by value. Support for MCP server instructions and resources is unconfirmed, so rely on the skill and the AGENTS.md snippet there.
Check what's registered
ember doctor closes with two lines per harness: its binary on PATH, and where ember is
registered for it — in this directory and globally — or the command that registers it. It
only reads each harness's config files; it never runs a harness CLI.
Each harness has its own check too: opencode mcp list, codex mcp list, claude mcp list.
Older ember init versions also added this repository's vault MCP server to every config
they wrote; if you ran ember init --global, doctor points at the leftover entry in your
global opencode config so you can delete it.
ember init records the ember-mcp it finds on your PATH, falling back to the Python
that runs it, so run ember init --global from the installed tool (uv tool install …)
rather than a checkout's .venv. It merges into existing configs and refuses to rewrite one
it can't parse (an opencode.json with comments, say), leaving that file untouched.
CLI
gut is a short alias for every command below.
Usage
Ask the agent in natural language — "Is this bug report urgent, and which team should own
it?" — and it calls ember_advise with a state and typed questions:
Question types: noul (yes/no → P(true)), choice (named options), score (ordered options →
expected score plus legend). The model field echoes the upstream model label.
When the evidence is visual, attach images or video frames inline: images is a list of
data:image/png;base64,... URIs (or {"content_type": "image/png", "base64": "..."}
objects), and videos is a list of videos, each a list of frame refs. Remote URLs and local
paths are rejected — the model server never reads host files or fetches URLs for an agent.
Configuration
Settings resolve as CLI flag > environment variable > config file > default. The config file
is JSON at ember config path (keys model, host, port, device, max_length,
server_url, auth_token, auth_header, allow_insecure_transport, request_timeout,
server_auth_token; a max_length of 0 means the model's own maximum).
Hosted deployment: a model location supplied at start time
Some deployments (e.g. an Outerbounds app) supply the model's S3
location directly at start time, rather than selecting a REGISTRY entry. Set:
No AWS credentials need to be set explicitly when the compute already has an IAM role
attached (the expected case on Outerbounds) — boto3's default credential chain picks it up automatically. Set
EMBER_S3_ACCESS_KEY_ID/_SECRET_ACCESS_KEY explicitly only where no role is
attached.
This is checked once at startup, before the server begins serving, and takes priority over
EMBER_MODEL/the model config key entirely — there is no REGISTRY lookup for this path.
ember supports any model that can run under its loader contract (constitution Article V,
"Model Loading"); it is not a hand-maintained allowlist of individually hash-verified weights,
so a URI-supplied location loads directly, with no pin or hash to check against.
A complete Outerbounds deployment example — deploy.yaml, a generated requirements.txt,
and a make deployment-requirements target — lives in deployment/.
Remote inference
Local inference is the default. To serve another machine — for example, a second Mac that lacks the memory for the model — run the server on the host and point the client at it.
On the host machine (the one with the model)
Ember does not terminate TLS. On a trusted LAN, plain HTTP is fine (see the client note
below). Over the internet, put a TLS proxy (nginx, Caddy) in front and use https://.
On the client machine (the one running the agent)
Plain http to a non-loopback host is refused by default; set
EMBER_ALLOW_INSECURE_TRANSPORT=1 to allow it on a trusted LAN. Credentials may also be
sent as a custom header (EMBER_AUTH_HEADER=X-API-KEY). ember status / ember doctor
report the endpoint kind, reachability, the remote's advertised auth_required, and whether
a credential is configured; reverting to local is a single change (unset these variables).
Bootstrapping a client against a hosted endpoint
ember init can generate the mcp.ember entry for a remote endpoint directly, instead of
hand-editing the generated opencode.json / kilo.json / .codex/config.toml:
This writes EMBER_SERVER_URL (and EMBER_AUTH_HEADER, if given) into the config file and
forces EMBER_AUTOSTART=0 — there is nothing local to autostart. The credential itself is
never written to the file: export EMBER_AUTH_TOKEN in the shell that launches the agent
(or your GUI app's environment) instead, so a secret never lands in a config file that might
be project-committed. Without --server-url, ember init behaves exactly as before (a local
loopback entry).
Metrics
While the model server is running it exposes Prometheus metrics at
http://127.0.0.1:8765/metrics (the EMBER_HOST/EMBER_PORT address):
The endpoint binds to the same address as the rest of the API (loopback by default), so
it is reachable only there unless you change EMBER_HOST. Ember's metrics live in a
dedicated Prometheus registry, so the standard python_*/process_* collectors are not
included — /metrics shows only the table above.
Development
opencode.json, .opencode/plugins/ember.js, and .opencode/skills/ember-advise/ (like
kilo.json, .codex/config.toml, and the other harnesses' ember-advise skill copies)
embed this clone's absolute paths or copy packaged files, so they are gitignored — regenerate
them with make init / make opencode after cloning. .opencode/opencode.json is shared and
committed: it registers the vault MCP server that agents use to read and write vault/
(launch opencode from the repository root).
Make targets
Make re-syncs the environment automatically when pyproject.toml or uv.lock changes.
Tests
tests/ covers the supported call paths:
GET /healthandPOST /v1/systemoneacrossnoul/choice/score, probability normalization, and request-validation422s- MCP tool discovery,
adviseover stdio, actionable errors (server down, model missing, malformed questions), and autostart on a configured port with pid cleanup - the agent kit: instructions size, skill frontmatter, the
ember://guideresource - the CLI:
agents,init --opencode,status/stopon unused ports,doctor,uninstall - process safety:
stopnever signals a pid that is not an ember server it started
Safe on a host running other opencode instances. The suite binds random free ports (never
8765), keeps state in temporary directories, stops only the processes it started, and never
invokes the opencode CLI or touches global opencode config.
CI (.github/workflows/ci.yml) runs uv sync --locked and then, as separate jobs, format
and lint, mypy --strict, bandit, make check, uv build plus an install smoke of the built
wheel on a hosted Apple Silicon runner, a SonarCloud scan (needs the SONAR_TOKEN repository
secret), and a zizmor audit of the workflows. Model-backed tests are
not run in CI: the ~19 GB fp16 model does not fit the available runners (hosted or the
org's 8 GiB self-hosted VMs), so run make test locally for model-affecting changes.
Coverage is tracked and gated (constitution Article XI). Run make test-cov to see the
full report. The enforced floor (fail_under in pyproject.toml) is the current measured
level and may only increase — currently 81 %. Lowering it requires explicit, recorded
approval per Article XI §11.2.
Benchmark
evals/clef-flash.jsonl is a 264-item benchmark covering the five agent-kit recipes (intent
and readiness, failure triage, change risk, routing, effort and approach) plus four vision
recipes (vision_noul, vision_choice, vision_score, vision_video): 456 scored questions
(184 choice, 152 noul, 88 score, 32 across the four vision recipes), split into
dev (130 items) and test (134). Vision items carry an images or videos field
(base64 data: URIs) alongside state and questions. Each line is
self-contained (state, the recipe's fixed question set, a gold label for every question, and
a rationale), so other implementations can score it without this harness.
ember eval export writes results/<run_id>_report/ for external reviewers:
report.html (one self-contained file with inline charts, light and dark themes, a print
layout, and item filters; it loads nothing from the network), report.md with its
figures/*.svg, and data/ with the run's results, trace, and the exact dataset scored.
The report covers context, the system under test, benchmark design, metric definitions with
references, results, calibration, the agent kit's decision rules replayed on every item, a
review card for every miss, limitations, and reproduction pins.
The report gives accuracy with a 95% bootstrap interval, macro-F1, top-label ECE, and Brier
score (choice, noul); ranked probability score and MAE (score); and coverage and accuracy
at the agent kit's act-on-it thresholds. Each run records the git hash, the engine, and the
dataset's SHA-256, and --compare warns when two runs scored different item sets.
Gold labels are judged from state alone against the recipe's option descriptions, and they
are fixed before any run: needs_review is true exactly when risk is Medium or High, and
retry only for flaky failures. A person reviews disagreements; labels are never changed to
match the model. make check validates the dataset (tests/test_eval_benchmark.py).
ember eval reads evals/ and scripts/ from the checkout, so it works only in a clone.
Agent in the loop
The benchmark above scores the model. ember eval agent scores ember the way it is used:
through a coding agent. It gives opencode 54 scripted requests (vague and precise asks,
failing tests, risky and trivial commits, routing and effort questions, and controls), each
in a throwaway git repo, and judges the agent only by what it did (files, commits, tests,
its reply). Each scenario runs under four conditions: none (no ember), mcp (the MCP
server and its instructions), skill (plus the ember-advise skill), and full (plus the
AGENTS.md policy).
Every session runs opencode run --pure with a private HOME and XDG directories and no TCP
port, so your opencode config, plugins, sessions, and running instances are untouched. The
provider key is read from the environment or opencode's auth file and passed only to the child
process. The run reports, per model and condition: gold-action accuracy, its paired change
against none, how often the agent consulted ember at decision points (and on controls),
whether it consulted before acting, asked the kit's recipe questions, passed the evidence, and
followed ember's answer, plus cost and time per session. It spends real provider credit and is
never part of make check, make test, or CI.
Caveats
- Gated DeltaNet fallback. Qwen3.5's backbone is a hybrid linear-attention model. The
optimized CUDA kernels (
causal_conv1d,flash-linear-attention) don't exist for Apple Silicon, so it uses the pure-PyTorch reference path: correct but slower. You'll see two "falling back" log lines — expected. - fp16 on MPS.
bfloat16works but is emulated and less battle-tested on MPS, so the loader usesfloat16; CPU mode usesfloat32. - Vision works on MPS. Images and video frames run through the same fp16 path as text.
Inline them as base64
data:URIs inimages/videos; remote URLs and local paths are rejected on purpose, so an agent cannot make the server read host files or fetch URLs. Text-only inputs skip the vision tower entirely. - Numerics. MPS can differ slightly from CUDA/CPU. If calibrated probabilities matter,
cross-check with
EMBER_DEVICE=cpu ember restart. device_map={"": "mps"}segfaults with the pinned stack. The loader loads on CPU and then moves the model to MPS — don't "simplify" that away.
Troubleshooting
- Start with
ember doctor, thenember logs(~/Library/Application Support/ember/logs/server.log). - The agent can't see the tool:
ember doctorshows where ember is registered for each harness (and flags an untrusted Codex project or a Claude Code.mcp.jsonawaiting approval);opencode mcp listshould show✓ ember connected. - "model … is not pulled": run
ember model pull, or pointEMBER_MODEL_DIRat the weights. Qwen3VLVideoProcessor requires Torchvision→ torchvision is a pinned dependency; runuv sync(or reinstall the tool).- Any single unimplemented MPS op falls back to CPU (
PYTORCH_ENABLE_MPS_FALLBACK=1, set by the runtime).
License
MIT (see LICENSE). Cloudflare's Clef weights and joint_schema_model.py are Apache-2.0; they
are downloaded from Hugging Face at runtime, not redistributed here.
Brand assets and palette · Provenance and licensing · Project site
来源:README.md,提交 c65e705
工具
0版本历史
1- v0.8.0最新Oct 9, 2026

