TokenMark

io.github.sypherinv0.2.1更新于 Oct 5, 2026

Measured local-LLM speeds and hardware-aware model picks for Strix Halo, DGX Spark and Mac.

已验证STDIO仅桌面AI & MLData & Analytics

概览

AI 生成的概览

让助手查询 TokenMark 上实测的本地大模型推理速度,并获得与硬件匹配的模型推荐。

功能
提供多个工具:按硬件平台和任务给出排序后的模型与配置推荐;列出已收录的基准配置及实测解码速度;浏览硬件目录中的芯片规格与价格;搜索模型和配置;以及提交带基准数据的代码仓库供人工审核。数据从 TokenMark 网站实时获取,并声称不会编造数字。
适用场景
适合在为 Strix Halo、DGX Spark 或 Mac 等特定硬件挑选本地模型或量化方案时使用,也适合用实测吞吐数据替代估算值进行比较。
运行要求
以 stdio 方式在本地运行,通常通过 npx 启动,因此需要 Node.js 以及访问 TokenMark 网站的网络连接。未声明账号、API 密钥或环境变量;数据源地址可用环境变量 TOKENMARK_URL 覆盖。
安装前请注意
该服务器从第三方网站读取实时数据,因此查询内容以及提交的仓库信息会离开本机。提交工具会把第三方仓库链接和备注排队等待人工审核,README 还指向一个需登录的网页表单用于提交个人实测数据。

安装

在 SourceWeft 中

  1. 打开 控制台中的 TokenMark,将其添加到工作区。
  2. 为需要使用其工具的对话启用该服务。

Desktop only,通过 STDIO。 STDIO 服务会启动本地进程,因此需要 SourceWeft 桌面宿主。

其他 MCP 客户端

参照 仓库 中的启动说明。

README

@altronis/tokenmark-mcp

An MCP server that gives Claude (and other agents) real local-LLM benchmark data + hardware-aware model recommendations from TokenMark.

Ask "what should I run on a Strix Halo for coding?" and the agent answers from measured configs (decode tok/s, quant, backend), each with a source link. It never invents numbers.

Add to Claude Code

sh
claude mcp add tokenmark -- npx -y @altronis/tokenmark-mcp

Add to any MCP client

Run the server over stdio:

sh
npx -y @altronis/tokenmark-mcp

Or in a client config:

json
{  "mcpServers": {    "tokenmark": {      "command": "npx",      "args": ["-y", "@altronis/tokenmark-mcp"]    }  }}

Tools

  • tokenmark_recommend: { hardware, tasks?, prefer?, limit? } → ranked model + best-config picks (5 by default, 10 max) with measured tok/s + why.
  • tokenmark_configs: { model?, hardware?, limit? } → tracked benchmark configs, fastest first (25 by default, 50 max; matched gives the full count), with the run mode (speculative: true for MTP/DFlash/draft-model runs).
  • tokenmark_hardware: { platform? } → the hardware catalogue (Strix Halo, Gorgon Halo, DGX Spark, Mac Max/Ultra). With a platform: chip specs with a source per value, the boxes that ship it, Singapore prices.
  • tokenmark_search: { term } → matching models/configs.
  • tokenmark_submit: { repo, source?, note? } → queues a GitHub/Hugging Face repo with benchmark numbers for human review.
  • tokenmark_submission_status: { id } → where a submission is.

A speed someone measured themselves, or hardware missing from the catalogue, goes through the signed-in form at https://tokenmark.app/submit.

Data is pulled live from https://tokenmark.app (override with TOKENMARK_URL). Zero runtime dependencies.

The same data in your terminal

The CLI lives in cli/ of this repo:

sh
npx @altronis/tokenmark-cli recommend --hardware "Strix Halo" --tasks coding

The files in bin/ are the published builds, made from the TokenMark tracker that runs https://tokenmark.app.

MIT licensed. Data aggregated from public community benchmarks with attribution.

来源:README.md,提交 19596dc

工具

0
工具元数据尚未被收录。

版本历史

1
  1. v0.2.1最新Oct 5, 2026