Local Llm Worker

io.github.JoblessJoev0.1.0更新于 Oct 4, 2026

Local LLM does Claude's bulk work: reads logs and files, researches the web, writes test-gated code.

概览

AI 生成的概览

让本地大模型承担批量阅读、网络调研和以测试为门槛的文件编写,助手只收到简短结果。

功能
提供 offload、research 和 delegate 三个工具。offload 在本地执行 shell 命令或读取文件并返回简短答案;research 通过 SearXNG 实例搜索网页、完整读取页面并返回带引用的答案;delegate 在隔离的 git worktree 中编写一个目标文件并反复重试直到测试通过,只返回一行结论。configure 工具用于查看和修改设置。
适用场景
适合大型日志、大文件或文档页面会占用前沿模型上下文的情况,也适合在本地生成由测试约束的小型辅助代码。已运行本地模型服务、希望降低批量处理成本的用户会受益。
运行要求
需要 Node 18 或更高版本、git,以及正在运行的 Ollama 或 OpenAI 兼容模型服务。环境变量 LLW_BASE_URL(默认 LLW_MODEL,可选 LLW_SEARCH_URL 指向启用 JSON 输出的 SearXNG 实例。以本地 stdio 进程运行,仅限桌面端。
安装前请注意
delegate 会写入文件并运行 shell 测试命令,默认会把通过测试的文件复制回检出目录,合并前应人工检查。offload 也可执行任意命令。可选的 api_key 和 headers 设置保存凭据并发送到所配置的后端。research 会把查询发送到你的 SearXNG 并下载公开页面;不会发送你的文件。

安装

在 SourceWeft 中

  1. 打开 控制台中的 Local Llm Worker,将其添加到工作区。
  2. 为需要使用其工具的对话启用该服务。

Desktop only,通过 STDIO。 STDIO 服务会启动本地进程,因此需要 SourceWeft 桌面宿主。

其他 MCP 客户端

参照 仓库 中的启动说明。

README

local-llm-worker

Let Claude hand the bulk reading to your local LLM: test logs, big files, web pages. Claude only gets the answer.

[test] [License: MIT] [Node ≥ 18] [Zero dependencies] [Claude Code plugin] [MCP server]

Works with Ollama · llama.cpp · LM Studio · vLLM · LocalAI · any OpenAI-compatible endpoint


Reading a 5,000-line test log or three docs pages costs the same frontier-model tokens as hard architectural work. local-llm-worker is an MCP server and Claude Code plugin that moves that bulk work onto the GPU (or CPU) you already own:

  • offload: your local model runs the noisy command or reads the big file. Claude gets the answer.
  • research: your local model searches the web, reads the pages in full and returns a cited answer. Claude never sees the pages.
  • delegate (opt-in): your local model writes one file in an isolated git worktree and retries until a test Claude wrote first passes. Claude gets a one-line verdict, not the code.
text
> offload  npm test 2>&1   "Which test fails, and why?"`fixed coupon never goes below zero` (test/cart.test.js:404): applyCoupon(5, {type:'fixed',value:10}) returns -5, expected 0.— qwen3-coder:30b-a3b-q4_K_M · 19,132 input tokens read locally

Why

Claude does itClaude delegates it
Find one failure in a 50 KB test log~19,000 tokens of log in contexta ~130-token answer
Answer from three docs pagesthe pages, or a summary of thema cited paragraph + source list, from the full pages
Write a helper that tests can pin downoutput tokens for the code, then re-reading itwrites the test, reads PASS
Several independent helperssequential editsparallel delegates, each in its own worktree
  • Any model, any hardware. No hardcoded models, no GPU assumptions. CPU-only works, just slower.
  • Parallel by default. Several calls in one turn, or from several subagents, run at once. The concurrency limit is a config value, not a hardcoded lock.
  • Fits the context automatically. Material is sized to the model's context window per text (number-heavy logs need more tokens than prose). Anything cut is reported.
  • Zero dependencies. Plain Node ≥ 18. Installing from git needs no npm install.
  • Agent-configurable. One configure call shows the config, the backend and its models, and changes any setting. It takes effect on the next call, with no restart.

Install

Prerequisites: Node ≥ 18, git, and a running local model server (e.g. ollama pull qwen3-coder:30b).

Claude Code plugin

text
/plugin marketplace add JoblessJoe/local-llm-worker/plugin install local-llm-worker@local-llm-worker

Then just ask: "set up local-llm-worker for my machine". Claude finds your backend, picks a model, and asks which tools it should use on its own (auto_use). By default that's offload and research; delegate is opt-in. Every tool also works whenever you ask for it.

Make Claude use it every time

A session-start reminder nudges Claude toward the tools you chose. For dependable use, add this to your CLAUDE.md; setup offers to do it for you:

markdown
## Local LLM (local-llm-worker)- Run test suites, builds and other noisy commands through `offload` (`command`), and read logs or files over ~300 lines through it, instead of reading the output yourself.- Use `research` for web lookups instead of WebSearch/WebFetch.

With this in place, Claude ran a noisy failing test suite through offload every time we tried. That session cost 40 % less than one that read the output itself.

Use it in any MCP client

Claude Desktop, Cursor, Windsurf, and others:

json
{  "mcpServers": {    "local-llm-worker": {      "command": "node",      "args": ["/path/to/local-llm-worker/src/index.js"],      "env": { "LLW_MODEL": "qwen3-coder:30b" }    }  }}

Claude Code without the plugin: claude mcp add local-llm-worker -- node /path/to/local-llm-worker/src/index.js

How delegate works

mermaid
flowchart LR    A[Claude writes the test<br/>+ spec + invariants] --> B[delegate]    B --> C[fresh git worktree<br/>+ your uncommitted changes]    C --> D[local LLM writes<br/>target file]    D --> E{run test}    E -- fail --> F[parse failures<br/>test source stays hidden]    F --> D    E -- pass --> G[copy file into checkout<br/>unless it changed meanwhile]    G --> H[Claude gets a one-line verdict]
  • Isolated. Each call gets its own git worktree with your uncommitted changes mirrored in, plus symlinked node_modules / .venv, so parallel calls never collide.
  • Locked scope. Exactly one target file, jailed to the repo, and never one of the test_files you name.
  • Useful feedback. Retries get the parsed failing assertions (TAP, pytest, jest, vitest, go, cargo), not a stack-trace tail.
  • Safe apply. If you or another delegate touched the target meanwhile, nothing is overwritten.
  • Honest failure. After N attempts you get the last failure and the kept worktree. Transport errors are reported as errors, never as a model FAIL.
  • Self-cleaning. Each delegate prunes llw-* worktrees left behind by a crashed server, and kept ones older than 24 h. Untracked files over 10 MB are not copied into the worktree (the verdict says how many were skipped).

The bundled skill teaches Claude how to write specs that pass: one invariant per test, exact API surface in context, properties instead of examples, and never delegating auth or money code.

Results

On one 24 GB GPU (Tesla P40) with qwen3-coder:30b-a3b-q4_K_M on Ollama:

TaskRead locallyClaude received
offload: find the failing test in a 50 KB test log~19,100 tokens~130 tokens
offload: find one ERROR line in a 90 KB log~30,400 tokensthe line, quoted
research: "Latest Node.js LTS and its end of life?" (3 pages)7,025 tokensa cited answer, ~250 tokens

Several calls run in parallel: three in one turn took as long as the slowest one.

Which model?

From a reproducible benchmark (bench/) of 120 delegate runs across six task types:

ModelGood forSpeed per call
qwen3-coder:30b-a3bthe best default: parsers, pure functions, Python15–55 s
devstral-small-2:24bedits to existing files, parsers2–9 min
granite4.1:8bsmall, well-specified functions on modest hardware25–90 s

Retries matter: with failure feedback, qwen3-coder's pass rate rose by half from the first attempt to the third. offload and research work well with any of these. Every run is logged (see Stats), so you can measure your own setup.

Configuration

Agents should use configure. Humans can edit JSON. Layers (later wins):

  1. built-in defaults
  2. user: ~/.config/local-llm-worker/config.json (respects $XDG_CONFIG_HOME)
  3. project: <git root>/.local-llm-worker.json
  4. env: LLW_<KEY>, e.g. LLW_BASE_URL, LLW_MODEL. Integers are digits only, booleans true/false/1/0/yes/no, link_dirs a comma list or JSON array, headers a JSON object. An invalid value is ignored and configure lists it under warnings.

A config file with invalid JSON is an error that names the file.

KeyDefaultMeaning
base_urlhttp://localhost:11434Backend root. A trailing /v1 is stripped.
apiautoauto · ollama · openai. Auto probes /api/version.
api_key""Sent as Authorization: Bearer. Shown as (set), never printed. Plugin users can set it under the plugin's settings instead, which keeps it in the OS keychain.
headers{}Extra HTTP headers for a gateway, e.g. {"X-Api-Key": "..."}. Values shown as (set).
model""Default model for both tools.
offload_model""Override for offload (e.g. long-context).
delegate_model""Override for delegate (e.g. a coder).
research_model""Override for research (e.g. long-context).
search_url""SearXNG instance for research (JSON output enabled).
research_sources3Pages research reads per search (max 10).
page_timeout_ms30000Per page download and per search.
allow_private_urlsfalseLet research read pages on loopback, link-local or private addresses (the search_url itself is always allowed).
num_ctx32768Context window (Ollama).
max_tokens8192Output limit per answer, sent on both APIs (num_predict on Ollama). A cut-off answer is reported, not tested.
keep_alive""Ollama only: how long the model stays loaded, e.g. "30m" or -1 (forever). "" keeps the server default (5 min).
concurrency4Max in-flight LLM requests per server process.
timeout_ms600000Per LLM request, counted from slot acquisition.
test_timeout_ms600000Per test / command run.
max_attempts3Delegate rounds (cap 10).
link_dirsnode_modules, .venv, venv, vendorSymlinked into each worktree.
log_path~/.local-llm-worker/runs.jsonlRun log; "" disables.
auto_useoffload, researchTools Claude uses on its own, via a short session-start reminder. delegate is opt-in. Every tool still works when you ask for it; [] turns the reminder off.
jsonc
// configure — no args: effective config, value sources, backend, model list// configure — change settings (null removes a key):{ "set": { "model": "qwen3-coder:30b", "num_ctx": 65536 }, "scope": "user" }

With no model set, it uses the backend's only model, or fails with the list. It never guesses, because the guess could be an embedding model.

Backends

Backendbase_urlAPI
Ollamahttp://localhost:11434native /api/chat. Its OpenAI endpoint ignores num_ctx and silently truncates long prompts.
llama.cpp llama-serverhttp://localhost:8080/v1/chat/completions
LM Studiohttp://localhost:1234/v1/chat/completions
vLLMhttp://localhost:8000/v1/chat/completions
LocalAI, othersyour endpoint/v1/chat/completions

Verified end-to-end on Ollama (GPU) and llama.cpp llama-server (CPU only). <think> blocks from reasoning models are stripped. If a backend reads less of the prompt than was sent, the result says so.

Tool reference

offload: read locally, answer briefly
Arg
taskrequired. The question.
filesPaths relative to the git root; must stay inside it.
commandShell command whose stdout + stderr is the material, e.g. npm test 2>&1.
modelOverride.

Material that doesn't fit next to the answer is cut in the middle (head and tail kept, sized per text), and the cut is reported.

research: search and read the web locally, answer with citations
Arg
questionrequired. The research question.
querySearch query, if it should differ from the question.
urlsRead these pages instead of searching (no search_url needed).
max_sourcesPages to read when searching (1–10).
modelOverride.

Page URLs that resolve to loopback, link-local or private addresses are refused, on every redirect hop too, unless allow_private_urls is true. Fetches pages in parallel and walks further down the results when a page fails (404, PDF, JavaScript-only). Each page is reduced to readable text, with <main>/<article> preferred and nav, footer and scripts dropped, then the pages share the num_ctx budget. Searching needs a SearXNG instance: docker run -d -p 8888:8080 searxng/searxng, add json under search.formats in its settings.yml, then set search_url to http://localhost:8888.

delegate: write one file until the test passes
Arg
specrequired. What the file must do.
target_filerequired. The one file to create or modify.
test_commandrequired. Run in the worktree by the platform shell (sh -c, or cmd.exe /d /s /c on Windows); exit 0 = pass. On timeout the whole process tree is killed.
propertiesInvariants, one per gating test.
contextThe exact API surface the file may use.
test_filesThe target may not be one of these.
max_attempts, modelOverrides.
show_codeAlso return the code (default false, which is where the saving comes from).
applyCopy a passing file into the checkout (default true).
configure: inspect and change settings

No args returns the report. { "set": {...}, "scope": "user" | "project" } validates and writes. Unknown keys are rejected with the list of valid ones.

Stats

Every run appends one line to log_path:

json
{"ts":"2026-10-04T17:51:28.476Z","tool":"delegate","model":"qwen3-coder:30b-a3b-q4_K_M","api":"ollama","ok":true,"attempts":1,"in_tok":201,"out_tok":195,"ms":57909}
sh
jq -s 'group_by(.tool) | map({tool: .[0].tool, runs: length, ok: (map(select(.ok)) | length), out_tok: (map(.out_tok) | add)})' ~/.local-llm-worker/runs.jsonl

FAQ

Does Claude ever see the generated code? Only if it asks (show_code: true) or reads the file. Review before merging anything that matters. A passing test is a gate, not a code review.

What should I not delegate? Auth, money, permissions, security, and anything whose spec has no single right answer. The skill tells Claude this too.

Can the local model cheat the test? It never sees the test source. It can write only the one target file, and that file can't be one of the test_files you list.

A long delegate or research call was cut off. The server sends MCP progress notifications every 15 s when the client asks for them (progressToken), but whether that helps depends on the client:

ClientTool-call limitKnob
Claude CodeNo practical wall-clock limit by default (~28 h); a stdio call is aborted after 30 min with no response and no progress, which the heartbeat preventsMCP_TOOL_TIMEOUT (ms, hard limit, progress does not extend it), CLAUDE_CODE_MCP_TOOL_IDLE_TIMEOUT (ms, 0 disables), or "timeout" in .mcp.json
Claude Desktop~60 s for local servers; progress does not reset itnone known: run long jobs from Claude Code, or lower max_attempts and research_sources
Cursor~60 s; progress reportedly does not reset itnone known

Does it send my code anywhere? Only to the base_url you configure, which is localhost by default. research sends your search query to your own SearXNG and downloads public pages. It never sends your files.

Development

sh
npm test   # 88 tests, fake backend + fake web, no LLM or network needed

src/index.js is the stdio JSON-RPC server. src/worker.js holds the config, backend client, tools and failure extractor.

Releases are automatic: every push to main runs the tests, bumps the patch version everywhere it appears (npm run bump does the same locally; pass minor, major or x.y.z for more), tags it and creates a GitHub Release.

Roadmap

  • more search backends (Brave, Tavily) alongside SearXNG
  • multi-file delegates
  • per-task model routing
  • an async job API for clients with a 60 s tool limit (Claude Desktop, Cursor)

License

MIT © Johannes Tebbert

来源:README.md,提交 12ef462

工具

0
工具元数据尚未被收录。

版本历史

1
  1. v0.1.0最新Oct 4, 2026