
jevjam
io.github.beremaranv0.3.2更新于 Oct 2, 2026
Typed decisions (choice, score, yes/no) from Laya, Julia-1 and clef-flash on your own GPU
概览
自托管的 MCP 服务器,在你自己的 GPU 上运行小型决策模型,对文本、JSON、图像或视频回答选择、评分和是/否类问题。
- 功能
- jevjam 提供四个 MCP 工具:jevjam_predict 用于类型化提问,jevjam_preset 提供 guard、moderation、triage 和 model_router 预设,另有 jevjam_route 和 jevjam_status。它让模型对给定的文本、JSON、图像或视频选择标签、按量表评分或回答是/否,并在毫秒级返回校准后的概率。同一进程还提供兼容 Jev 的 HTTP 端点 /v1/systemone,共用同一队列和常驻模型。
- 适用场景
- 适合用于快速、低成本、不值得调用大型 LLM 的分类决策,例如工单分流、滥用标记、工具调用防护,或为提示词挑选模型。也适合希望数据留在自有硬件上的本地或自托管场景。
- 运行要求
- 需要 Docker、NVIDIA GPU 和 NVIDIA Container Toolkit,镜像以 --gpus all 运行。模型在首次请求时下载到挂载卷中,首次需数分钟。可选密钥:HF_TOKEN 用于提高 Hugging Face 下载速率限制,JEVJAM_API_KEY 用于要求 /mcp 和 /v1/systemone 携带 bearer 密钥。JEVJAM_IDLE_TIMEOUT 控制常驻检查点何时释放。
安装
在 SourceWeft 中
- 打开 控制台中的 jevjam,将其添加到工作区。
- 为需要使用其工具的对话启用该服务。
Desktop only,通过 STDIO。 STDIO 服务会启动本地进程,因此需要 SourceWeft 桌面宿主。
其他 MCP 客户端
参照 仓库 中的启动说明。
README
[jevjam: typed decisions from small models, as MCP tools and a Jev-compatible HTTP API]
[CI] [Latest release] [Container image] [MCP] [License]
jevjam
Self-hosted MCP server and Jev-compatible HTTP API for small decision models, on one GPU, in Docker.
Ask a model typed questions about a piece of text, JSON, an image or a video, and get
calibrated answers back in milliseconds: pick a label (choice), rate on a scale
(score), or answer yes or no (noul). Agents call it as MCP
tools; services call POST /v1/systemone, the same protocol as TypeSafe Jev, so a Jev
client only needs a new base URL.
Use it to route tickets, flag abuse, guard tool calls, pick a model for a prompt, or any other decision you would rather not spend a large LLM call on.
Models
One checkpoint stays in VRAM at a time, and it is freed after five idle minutes. See docs/models.md for sizes, quantization and limits.
Quick start
You need Docker, an NVIDIA GPU, and the NVIDIA Container Toolkit.
Or, from a clone, docker compose up -d. No model downloads at boot; the first request
fetches what it needs into the jevjam-models volume, which takes minutes once.
Ask over HTTP:
The answer, trimmed:
Or connect an agent over MCP, at http://127.0.0.1:8000/mcp:
Agents get four tools: jevjam_predict, jevjam_preset (guard, moderation,
triage, model_router), jevjam_route and jevjam_status. The
MCP guide covers OpenCode, Pi, remote access and reverse proxies.
Features
- One process, two doors. The HTTP API and MCP share one queue and one resident model, so neither starves the other of VRAM.
- Sleeps when idle. After
JEVJAM_IDLE_TIMEOUTseconds (300 by default) every checkpoint is freed and the GPU memory goes back to the driver. The next request loads only what it needs; Laya wakes in 0.6 s on an RTX 4070 Ti SUPER. - Fits the card it finds. clef-flash loads in BF16, 8-bit, 4-bit, or split across GPU and CPU, whichever fits.
- Drop-in for Jev. Same request and response shapes; unknown fields are ignored.
- Locked down by default. Runs as non-root, binds to loopback in Compose, and
takes an optional bearer key (
JEVJAM_API_KEY) for both endpoints.
Docs
Moving from laya-docker
This repo used to be laya-docker. The old image, ghcr.io/beremaran/laya-docker,
gets no more updates; switch to ghcr.io/beremaran/jevjam. Old LAYA_* settings
still work and log a warning; see Configuration.
Contributing
Bug reports and pull requests are welcome; see CONTRIBUTING.md. Report security problems privately, as SECURITY.md describes.
License
jevjam is licensed under Apache-2.0. The image also contains the
Apache-2.0 Laya package and checkpoints by
Convai Innovations, the Apache-2.0 Julia-1 code and checkpoint by Supersonic Labs, and
the Apache-2.0 clef-flash code and checkpoint by Cloudflare. The clef-flash code is
copied into src/jevjam/vendor/ with its license.
来源:README.md,提交 99406ff
工具
0版本历史
1- v0.3.2最新Oct 2, 2026


