jevjam

io.github.beremaranv0.3.2更新於 Oct 2, 2026

Typed decisions (choice, score, yes/no) from Laya, Julia-1 and clef-flash on your own GPU

已驗證STDIO僅桌面Developer ToolsAI & ML

概覽

AI 產生的概覽

自架的 MCP 伺服器,在你自己的 GPU 上執行小型決策模型,針對文字、JSON、影像或影片回答選擇、評分與是/否類問題。

功能
jevjam 提供四個 MCP 工具:jevjam_predict 用於型別化提問,jevjam_preset 提供 guard、moderation、triage 與 model_router 預設集,另有 jevjam_route 與 jevjam_status。它讓模型針對給定的文字、JSON、影像或影片挑選標籤、依量表評分或回答是/否,並在毫秒內回傳校準後的概率。同一個行程也提供相容於 Jev 的 HTTP 端點 /v1/systemone,共用同一個佇列與常駐模型。
適用情境
適合用於快速、低成本、不值得呼叫大型 LLM 的分類決策,例如工單分流、濫用標記、工具呼叫防護,或為提示詞挑選模型。也適合希望資料留在自有硬體上的本機或自架情境。
執行需求
需要 Docker、NVIDIA GPU 與 NVIDIA Container Toolkit,映像以 --gpus all 執行。模型會在首次請求時下載到掛載的磁碟區,首次需數分鐘。選用密鑰:HF_TOKEN 可提高 Hugging Face 下載速率限制,JEVJAM_API_KEY 則要求 /mcp 與 /v1/systemone 攜帶 bearer 密鑰。JEVJAM_IDLE_TIMEOUT 控制常駐檢查點何時釋放。
安裝前請注意
它以本機容器方式執行並可存取 GPU,首次使用時會從 Hugging Face 下載模型權重。JEVJAM_API_KEY 為選用,但在把端點暴露到回送位址之外前建議設定,否則 /mcp 與 /v1/systemone 會接受未驗證的請求。你提交的文字、JSON、影像或影片由本機模型處理;若內容敏感,請先確認要送出的內容。

安裝

在 SourceWeft 中

  1. 開啟 儀表板中的 jevjam,將其新增到工作區。
  2. 為需要使用其工具的對話啟用該服務。

Desktop only,透過 STDIO。 STDIO 服務會啟動本機處理程序,因此需要 SourceWeft 桌面主機。

其他 MCP 客戶端

參照 儲存庫 中的啟動說明。

README

[jevjam: typed decisions from small models, as MCP tools and a Jev-compatible HTTP API]

[CI] [Latest release] [Container image] [MCP] [License]

jevjam

Self-hosted MCP server and Jev-compatible HTTP API for small decision models, on one GPU, in Docker.

Ask a model typed questions about a piece of text, JSON, an image or a video, and get calibrated answers back in milliseconds: pick a label (choice), rate on a scale (score), or answer yes or no (noul). Agents call it as MCP tools; services call POST /v1/systemone, the same protocol as TypeSafe Jev, so a Jev client only needs a new base URL.

Use it to route tickets, flag abuse, guard tool calls, pick a model for a prompt, or any other decision you would rather not spend a large LLM call on.

Models

ModelBySizeReadsPicked when
LayaConvai Innovations3 checkpoints, ~1.2B in alltext, JSONby default; its router picks English, multilingual or typed-decisions
Julia-1Supersonic Labs144Mtext, JSONthe request names julia-1
clef-flashCloudflare9Btext, JSON, images, videothe request names clef-flash

One checkpoint stays in VRAM at a time, and it is freed after five idle minutes. See docs/models.md for sizes, quantization and limits.

Quick start

You need Docker, an NVIDIA GPU, and the NVIDIA Container Toolkit.

bash
docker run -d --name jevjam --gpus all -p 127.0.0.1:8000:8000 \  -v jevjam-models:/models ghcr.io/beremaran/jevjam:latest

Or, from a clone, docker compose up -d. No model downloads at boot; the first request fetches what it needs into the jevjam-models volume, which takes minutes once.

Ask over HTTP:

bash
curl -s http://127.0.0.1:8000/v1/systemone -H 'Content-Type: application/json' -d '{  "state": "We were billed twice for March. Refund it today or we cancel.",  "questions": {    "department": {"type": "choice", "instructions": "Who should handle this?",                   "criteria": {"billing": "payments, refunds", "technical": "bugs, outages"}},    "churn_risk": {"type": "noul", "instructions": "Does the user threaten to leave?"}  }}'

The answer, trimmed:

json
{  "answers": {    "department": {"type": "choice", "choice": "billing", "probabilities": {"billing": 0.97, "technical": 0.03}, ...},    "churn_risk": {"type": "noul", "noul": 0.82, ...}  },  "routing": {"model": "english", "reason": "English Latin text", ...}}

Or connect an agent over MCP, at http://127.0.0.1:8000/mcp:

bash
claude mcp add --transport http jevjam http://127.0.0.1:8000/mcp   # Claude Codecodex mcp add jevjam --url http://127.0.0.1:8000/mcp                 # Codex

Agents get four tools: jevjam_predict, jevjam_preset (guard, moderation, triage, model_router), jevjam_route and jevjam_status. The MCP guide covers OpenCode, Pi, remote access and reverse proxies.

Features

  • One process, two doors. The HTTP API and MCP share one queue and one resident model, so neither starves the other of VRAM.
  • Sleeps when idle. After JEVJAM_IDLE_TIMEOUT seconds (300 by default) every checkpoint is freed and the GPU memory goes back to the driver. The next request loads only what it needs; Laya wakes in 0.6 s on an RTX 4070 Ti SUPER.
  • Fits the card it finds. clef-flash loads in BF16, 8-bit, 4-bit, or split across GPU and CPU, whichever fits.
  • Drop-in for Jev. Same request and response shapes; unknown fields are ignored.
  • Locked down by default. Runs as non-root, binds to loopback in Compose, and takes an optional bearer key (JEVJAM_API_KEY) for both endpoints.

Docs

GuideWhat is in it
ConfigurationRunning, settings, the model cache, sleeping on idle
MCP serverTools, auth, client setup, reverse proxies
HTTP API/health, /v1/systemone, question types, JSON Schema, errors
ModelsLaya, Julia-1 and clef-flash: sizes, VRAM, limits

Moving from laya-docker

This repo used to be laya-docker. The old image, ghcr.io/beremaran/laya-docker, gets no more updates; switch to ghcr.io/beremaran/jevjam. Old LAYA_* settings still work and log a warning; see Configuration.

Contributing

Bug reports and pull requests are welcome; see CONTRIBUTING.md. Report security problems privately, as SECURITY.md describes.

License

jevjam is licensed under Apache-2.0. The image also contains the Apache-2.0 Laya package and checkpoints by Convai Innovations, the Apache-2.0 Julia-1 code and checkpoint by Supersonic Labs, and the Apache-2.0 clef-flash code and checkpoint by Cloudflare. The clef-flash code is copied into src/jevjam/vendor/ with its license.

來源:README.md,提交 99406ff

工具

0
工具後設資料尚未被收錄。

版本歷史

1
  1. v0.3.2最新Oct 2, 2026