Nodegrove VRAM: can I run it?

io.nodegrovev1.0.0更新於 Oct 3, 2026

Can this LLM run on my GPU? VRAM, speed ceiling and what fits instead, for any model and GPU.

已驗證Streamable HTTP可網頁執行Developer ToolsAI & ML

概覽

AI 產生的概覽

估算某個開放權重的大型語言模型能否放進指定顯示卡的 VRAM,並給出速度上限與可行的替代方案。

功能
提供唯讀工具,用來計算在特定顯示卡上執行開放權重大型語言模型所需的記憶體。工具包括 can_i_run(是否放得下、記憶體分配、速度上限、可容納的最長上下文)、what_fits(單張顯示卡能放下的模型)、estimate_vram(各量化下的權重、KV 快取與額外開銷)、estimate_from_hf_repo(讀取任何 Hugging Face 儲存庫的 config.json),以及 list_models、list_gpus。回答會以欄位形式給出數值,並附上模型與顯示卡頁面連結及各项假設。
適用情境
適合在規劃本機大型語言模型推論時使用:想知道某個模型能否放進自己的顯示卡、能支援多長上下文,或該改選哪個模型或顯示卡。也適合詢問量化方式、VRAM 預算與單流速度的大致預期。
執行需求
遠端方式:連線至 streamable HTTP 端點,不需要帳號或金鑰。本機方式:需要 Node.js 20 或更新版本,透過 npx 執行 npm 套件。未宣告任何環境變數或標頭。遠端端點與 Hugging Face 儲存庫查詢都需要網路連線。
安裝前請注意
唯讀且免費,不索取任何憑證。遠端伺服器為無狀態、不保留請求記錄,但託管方會保留短期的邊緣記錄,且會依位址限流。查詢 Hugging Face 儲存庫時會把儲存庫名稱傳送給伺服器,由其取得該儲存庫公開的 config.json。速度數值是單流理論上限,並非實測結果。

安裝

在 SourceWeft 中

  1. 開啟 儀表板中的 Nodegrove VRAM: can I run it?,將其新增到工作區。
  2. 為需要使用其工具的對話啟用該服務。

Web executable,透過 Streamable HTTP。 遠端服務在工作區中設定後即可從網頁執行環境執行。

其他 MCP 客戶端

把它新增到你客戶端的 mcpServers 設定中。

{
  "mcpServers": {
    "vram-mcp": {
      "type": "http",
      "url": "https://mcp.nodegrove.io/mcp"
    }
  }
}

README

[nodegrove]

Can I run it?

An MCP server that tells your AI assistant whether an open-weight LLM fits your GPU, and what would fit instead.
The arithmetic behind nodegrove.io, for any model on Hugging Face and any card.

[npm] [CI]


Can my RTX 4090 run Llama 3.3 70B?

No: Llama 3.3 70B at Q4_K_M with 8,192 tokens of context needs 45.8 GB, and the RTX 4090 holds 22.8 GB after headroom. Short by 23 GB. What would work instead:

  • No context length helps: the weights alone are 43.1 GB before a single token of conversation.
  • RTX 6000 Ada (48 GB) is within a whisker: 45.8 GB against 45.6 GB after headroom. With the KV cache at Q8 it needs 44.4 GB and fits.
  • The smallest card here that runs it exactly as asked: A100, 80 GB usable.
  • The biggest model your card does run at these settings, counting mixture-of-experts models at their dense equivalent: Qwen3 32B, 22.4 GB at ~37 tokens/s.

That is the server's own answer, word for word. Every figure is a stated formula over the model's config.json and the card's published specs: no benchmarks, no guesses. It is free, read-only, and needs no account or key.

Connect

Remote: nothing to install

https://mcp.nodegrove.io/mcp

[Install in Cursor] [Install in VS Code] [Add to LM Studio]

Claude (claude.ai and the desktop app)

Customize → Connectors → Add → Add custom connector. Name it Nodegrove VRAM, paste the URL, choose No sign-in, then turn it on in a chat from + → Connectors.

Claude Code
sh
claude mcp add --transport http nodegrove-vram https://mcp.nodegrove.io/mcp

Add --scope user to have it in every project.

Cursor

~/.cursor/mcp.json:

json
{ "mcpServers": { "nodegrove-vram": { "url": "https://mcp.nodegrove.io/mcp" } } }
VS Code

.vscode/mcp.json:

json
{ "servers": { "nodegrove-vram": { "type": "http", "url": "https://mcp.nodegrove.io/mcp" } } }
Devin Desktop (formerly Windsurf)
sh
devin mcp add -s user nodegrove-vram https://mcp.nodegrove.io/mcp
LM Studio

Program → Install → Edit mcp.json:

json
{ "mcpServers": { "nodegrove-vram": { "url": "https://mcp.nodegrove.io/mcp" } } }
Open WebUI

Admin Settings → Integrations → External Tool Servers → Add Connection. Type MCP (Streamable HTTP), the URL above, authentication None.

Cline

The type must be stated, or Cline treats a URL as the older SSE transport:

json
{ "mcpServers": { "nodegrove-vram": { "type": "streamableHttp", "url": "https://mcp.nodegrove.io/mcp" } } }
Codex CLI
sh
codex mcp add nodegrove-vram --url https://mcp.nodegrove.io/mcp
Gemini CLI
sh
gemini mcp add -s user --transport http nodegrove-vram https://mcp.nodegrove.io/mcp
Zed

settings.json:

json
{ "context_servers": { "nodegrove-vram": { "url": "https://mcp.nodegrove.io/mcp" } } }
ChatGPT

Settings → Security and login → turn on Developer mode. At chatgpt.com/plugins, add one with the URL and No Authentication, then pick it in a chat from + → Developer mode.

Local: over stdio

Needs Node.js 20 or newer:

json
{  "mcpServers": {    "nodegrove-vram": {      "command": "npx",      "args": ["-y", "@nodegrove/vram-mcp"]    }  }}

Tools

ToolAsk itIt answers with
can_i_runCan my RTX 4090 run Llama 3.3 70B with 32k of context?Fits, tight or no; the memory split; a speed ceiling; the longest context that fits; and on a no, every change that would make it fit
what_fitsWhat is the best model for my 16 GB card?Every model checked on one card, with a recommended everyday pick, the largest that fits, the best at Q8 and the first out of reach
estimate_vramHow much VRAM does Qwen3 32B need at 64k?Weights, KV cache and overhead at each quantisation, and the smallest common card that holds each
estimate_from_hf_repoHow much memory does Qwen/Qwen3-Next-80B-A3B-Instruct need?Any Hugging Face repo, read from its config.json: the attention layout, the cost of each 1,000 tokens of context, and memory at every quantisation
list_models, list_gpusWhich GPUs do you know?The models and cards with their specs, ids and pages

Models can be named the way people type them ("llama 3.3 70b", Llama-3.3-70B-Instruct) or given as any Hugging Face repo id. Cards can be named ("4090", "M4 Max") or described by their memory and bandwidth. Every answer carries the numbers as fields, links to the model and card pages on nodegrove.io, and the assumptions behind each figure. All six tools are read-only.

How the numbers are made

  • Memory is weights + KV cache + overhead. Weights are parameters × bytes per parameter at the quantisation: FP16 2.00, Q8_0 1.06, Q6_K 0.82, Q5_K_M 0.71, Q4_K_M 0.58, Q3_K_M 0.47. Overhead is 0.5 GB plus 4% of the weights.
  • The KV cache is counted the way each model caches. Sliding-window layers stop at their window, hybrid models (linear attention or Mamba) grow a cache only on their few full-attention layers, and latent attention stores one compressed vector per layer. A standard-transformer formula would overstate these models several times over at long context.
  • Fit means at most 95% of the memory a runtime can address; above 85% it is tight. Apple silicon gives the GPU about 75% of unified memory.
  • Speed is 0.7 × memory bandwidth ÷ bytes of active weights read per token. It is a single-stream ceiling, not a measurement, and the faster the figure, the further real runtimes fall below it.
  • The model table was read from each model's config.json and checked against Hugging Face. Any other repo is read live by the same rules, which reproduce every row of the table; anything the reader cannot model is named in the answer, never guessed.

The full method, with every constant, is at nodegrove.io/data. The same figures are an open dataset under CC BY 4.0, DOI 10.5281/zenodo.22966137.

The math as a library

The server is a thin layer over @nodegrove/llm-math, the package nodegrove.io's pages and calculators are built on:

sh
npm install @nodegrove/llm-math
ts
import { modelById, estimate } from '@nodegrove/llm-math';
estimate({ ...modelById('llama-3.3-70b')!, quant: 'q4', context: 8192 }).totalGb; // 45.77

Privacy

The remote server runs on Cloudflare Workers. Nodegrove keeps no request logs and no record of what you ask: each request is answered by a fresh, stateless instance and forgotten. Cloudflare keeps standard edge logs for a short period, as for any website. Requests are rate-limited to 120 a minute per address. When you ask about a Hugging Face repo, the server fetches that repo's public config.json and metadata; only the repo name is sent. The local version contacts nothing but huggingface.co, and only when you ask about a repo there. Details: nodegrove.io/privacy.

Run your own

The remote endpoint is mcp/src/worker.ts. To deploy a copy to your Cloudflare account, change the route in mcp/wrangler.jsonc, then:

sh
cd mcp && npm install && npx wrangler deploy

Development

sh
cd llm-math && npm install && npm testcd mcp && npm install && npm test && npm run build

llm-math/ is the package of formulas and data; mcp/ is the server, with the stdio entry (src/stdio.ts), the Worker (src/worker.ts) and the tools (src/tools.ts). This repository is published from Nodegrove's main repository, where the site uses the same package. Issues and pull requests are welcome here; accepted changes are applied upstream and credited.

Found a figure that disagrees with its source? Open an issue with the model or card and the link. To report a security problem, see SECURITY.md.

Licence

The code is MIT. The model and GPU data, llm-math/src/models.ts and llm-math/src/gpus.ts, is CC BY 4.0: use it for anything, and credit Nodegrove (nodegrove.io).

Model and GPU names are trademarks of their owners.

來源:README.md,提交 98c4f60

工具

0
工具後設資料尚未被收錄。

版本歷史

1
  1. v1.0.0最新Oct 3, 2026