
TokenMark
io.github.sypherinv0.2.1更新于 Oct 5, 2026
Measured local-LLM speeds and hardware-aware model picks for Strix Halo, DGX Spark and Mac.
概览
让助手查询 TokenMark 上实测的本地大模型推理速度,并获得与硬件匹配的模型推荐。
- 功能
- 提供多个工具:按硬件平台和任务给出排序后的模型与配置推荐;列出已收录的基准配置及实测解码速度;浏览硬件目录中的芯片规格与价格;搜索模型和配置;以及提交带基准数据的代码仓库供人工审核。数据从 TokenMark 网站实时获取,并声称不会编造数字。
- 适用场景
- 适合在为 Strix Halo、DGX Spark 或 Mac 等特定硬件挑选本地模型或量化方案时使用,也适合用实测吞吐数据替代估算值进行比较。
- 运行要求
- 以 stdio 方式在本地运行,通常通过 npx 启动,因此需要 Node.js 以及访问 TokenMark 网站的网络连接。未声明账号、API 密钥或环境变量;数据源地址可用环境变量 TOKENMARK_URL 覆盖。
安装
在 SourceWeft 中
- 打开 控制台中的 TokenMark,将其添加到工作区。
- 为需要使用其工具的对话启用该服务。
Desktop only,通过 STDIO。 STDIO 服务会启动本地进程,因此需要 SourceWeft 桌面宿主。
其他 MCP 客户端
参照 仓库 中的启动说明。
README
@altronis/tokenmark-mcp
An MCP server that gives Claude (and other agents) real local-LLM benchmark data + hardware-aware model recommendations from TokenMark.
Ask "what should I run on a Strix Halo for coding?" and the agent answers from measured configs (decode tok/s, quant, backend), each with a source link. It never invents numbers.
Add to Claude Code
Add to any MCP client
Run the server over stdio:
Or in a client config:
Tools
tokenmark_recommend:{ hardware, tasks?, prefer?, limit? }→ ranked model + best-config picks (5 by default, 10 max) with measured tok/s + why.tokenmark_configs:{ model?, hardware?, limit? }→ tracked benchmark configs, fastest first (25 by default, 50 max;matchedgives the full count), with the run mode (speculative: truefor MTP/DFlash/draft-model runs).tokenmark_hardware:{ platform? }→ the hardware catalogue (Strix Halo, Gorgon Halo, DGX Spark, Mac Max/Ultra). With a platform: chip specs with a source per value, the boxes that ship it, Singapore prices.tokenmark_search:{ term }→ matching models/configs.tokenmark_submit:{ repo, source?, note? }→ queues a GitHub/Hugging Face repo with benchmark numbers for human review.tokenmark_submission_status:{ id }→ where a submission is.
A speed someone measured themselves, or hardware missing from the catalogue, goes through the signed-in form at https://tokenmark.app/submit.
Data is pulled live from https://tokenmark.app (override with TOKENMARK_URL). Zero runtime dependencies.
The same data in your terminal
The CLI lives in cli/ of this repo:
The files in bin/ are the published builds, made from the TokenMark tracker that runs https://tokenmark.app.
MIT licensed. Data aggregated from public community benchmarks with attribution.
来源:README.md,提交 19596dc
工具
0版本历史
1- v0.2.1最新Oct 5, 2026


