Nodegrove VRAM: can I run it?
io.nodegrovev1.0.0更新于 Oct 3, 2026
Can this LLM run on my GPU? VRAM, speed ceiling and what fits instead, for any model and GPU.
概览
估算某个开放权重的大语言模型能否装进指定显卡的显存,并给出速度上限和可行的替代方案。
- 功能
- 提供只读工具,用于计算在特定显卡上运行开放权重大语言模型所需的内存。工具包括 can_i_run(是否装得下、内存分配、速度上限、可容纳的最长上下文)、what_fits(单张显卡能装下的模型)、estimate_vram(各量化下的权重、KV 缓存与开销)、estimate_from_hf_repo(读取任意 Hugging Face 仓库的 config.json),以及 list_models、list_gpus。回答以字段形式给出数值,并附模型与显卡页面链接和各项假设。
- 适用场景
- 适合在规划本地大模型推理时使用:想知道某个模型能否装进自己的显卡、能支持多长上下文,或该改选哪个模型或显卡。也适合询问量化方式、显存预算和单流速度的大致预期。
- 运行要求
- 远程方式:连接 streamable HTTP 端点,无需账号或密钥。本地方式:需要 Node.js 20 或更高版本,通过 npx 运行 npm 包。未声明任何环境变量或请求头。远程端点和 Hugging Face 仓库查询都需要网络访问。
安装
在 SourceWeft 中
- 打开 控制台中的 Nodegrove VRAM: can I run it?,将其添加到工作区。
- 为需要使用其工具的对话启用该服务。
Web executable,通过 Streamable HTTP。 远程服务在工作区中配置后即可从网页运行时运行。
其他 MCP 客户端
把它添加到你客户端的 mcpServers 配置中。
{
"mcpServers": {
"vram-mcp": {
"type": "http",
"url": "https://mcp.nodegrove.io/mcp"
}
}
}README
Can I run it?
An MCP server that tells your AI assistant whether an open-weight LLM fits your GPU, and what would fit instead.
The arithmetic behind nodegrove.io, for any model on Hugging Face and any card.
Can my RTX 4090 run Llama 3.3 70B?
No: Llama 3.3 70B at Q4_K_M with 8,192 tokens of context needs 45.8 GB, and the RTX 4090 holds 22.8 GB after headroom. Short by 23 GB. What would work instead:
- No context length helps: the weights alone are 43.1 GB before a single token of conversation.
- RTX 6000 Ada (48 GB) is within a whisker: 45.8 GB against 45.6 GB after headroom. With the KV cache at Q8 it needs 44.4 GB and fits.
- The smallest card here that runs it exactly as asked: A100, 80 GB usable.
- The biggest model your card does run at these settings, counting mixture-of-experts models at their dense equivalent: Qwen3 32B, 22.4 GB at ~37 tokens/s.
That is the server's own answer, word for word. Every figure is a stated formula over the model's config.json and the card's published specs: no benchmarks, no guesses. It is free, read-only, and needs no account or key.
Connect
Remote: nothing to install
[Install in Cursor] [Install in VS Code] [Add to LM Studio]
Claude (claude.ai and the desktop app)
Customize → Connectors → Add → Add custom connector. Name it Nodegrove VRAM, paste the URL, choose No sign-in, then turn it on in a chat from + → Connectors.
Claude Code
Add --scope user to have it in every project.
Cursor
~/.cursor/mcp.json:
VS Code
.vscode/mcp.json:
Devin Desktop (formerly Windsurf)
LM Studio
Program → Install → Edit mcp.json:
Open WebUI
Admin Settings → Integrations → External Tool Servers → Add Connection. Type MCP (Streamable HTTP), the URL above, authentication None.
Cline
The type must be stated, or Cline treats a URL as the older SSE transport:
Codex CLI
Gemini CLI
Zed
settings.json:
ChatGPT
Settings → Security and login → turn on Developer mode. At chatgpt.com/plugins, add one with the URL and No Authentication, then pick it in a chat from + → Developer mode.
Local: over stdio
Needs Node.js 20 or newer:
Tools
Models can be named the way people type them ("llama 3.3 70b", Llama-3.3-70B-Instruct) or given as any Hugging Face repo id. Cards can be named ("4090", "M4 Max") or described by their memory and bandwidth. Every answer carries the numbers as fields, links to the model and card pages on nodegrove.io, and the assumptions behind each figure. All six tools are read-only.
How the numbers are made
- Memory is weights + KV cache + overhead. Weights are parameters × bytes per parameter at the quantisation: FP16 2.00, Q8_0 1.06, Q6_K 0.82, Q5_K_M 0.71, Q4_K_M 0.58, Q3_K_M 0.47. Overhead is 0.5 GB plus 4% of the weights.
- The KV cache is counted the way each model caches. Sliding-window layers stop at their window, hybrid models (linear attention or Mamba) grow a cache only on their few full-attention layers, and latent attention stores one compressed vector per layer. A standard-transformer formula would overstate these models several times over at long context.
- Fit means at most 95% of the memory a runtime can address; above 85% it is tight. Apple silicon gives the GPU about 75% of unified memory.
- Speed is 0.7 × memory bandwidth ÷ bytes of active weights read per token. It is a single-stream ceiling, not a measurement, and the faster the figure, the further real runtimes fall below it.
- The model table was read from each model's
config.jsonand checked against Hugging Face. Any other repo is read live by the same rules, which reproduce every row of the table; anything the reader cannot model is named in the answer, never guessed.
The full method, with every constant, is at nodegrove.io/data. The same figures are an open dataset under CC BY 4.0, DOI 10.5281/zenodo.22966137.
The math as a library
The server is a thin layer over @nodegrove/llm-math, the package nodegrove.io's pages and calculators are built on:
Privacy
The remote server runs on Cloudflare Workers. Nodegrove keeps no request logs and no record of what you ask: each request is answered by a fresh, stateless instance and forgotten. Cloudflare keeps standard edge logs for a short period, as for any website. Requests are rate-limited to 120 a minute per address. When you ask about a Hugging Face repo, the server fetches that repo's public config.json and metadata; only the repo name is sent. The local version contacts nothing but huggingface.co, and only when you ask about a repo there. Details: nodegrove.io/privacy.
Run your own
The remote endpoint is mcp/src/worker.ts. To deploy a copy to your Cloudflare account, change the route in mcp/wrangler.jsonc, then:
Development
llm-math/ is the package of formulas and data; mcp/ is the server, with the stdio entry (src/stdio.ts), the Worker (src/worker.ts) and the tools (src/tools.ts). This repository is published from Nodegrove's main repository, where the site uses the same package. Issues and pull requests are welcome here; accepted changes are applied upstream and credited.
Found a figure that disagrees with its source? Open an issue with the model or card and the link. To report a security problem, see SECURITY.md.
Licence
The code is MIT. The model and GPU data, llm-math/src/models.ts and llm-math/src/gpus.ts, is CC BY 4.0: use it for anything, and credit Nodegrove (nodegrove.io).
Model and GPU names are trademarks of their owners.
来源:README.md,提交 98c4f60
工具
0版本历史
1- v1.0.0最新Oct 3, 2026


