TokenMark

io.github.sypherinv0.2.1更新於 Oct 5, 2026

Measured local-LLM speeds and hardware-aware model picks for Strix Halo, DGX Spark and Mac.

已驗證STDIO僅桌面AI & MLData & Analytics

概覽

AI 產生的概覽

讓助理查詢 TokenMark 上實測的本地 LLM 推論速度,並取得符合硬體的模型推薦。

功能
提供多項工具:依硬體平台與任務給出排序後的模型與設定推薦;列出已收錄的基準設定與實測解碼速度;瀏覽硬體目錄中的晶片規格與價格;搜尋模型與設定;以及提交附基準數據的程式碼倉庫供人工審核。資料從 TokenMark 網站即時取得,並聲稱不會捏造數字。
適用情境
適合在為 Strix Halo、DGX Spark 或 Mac 等特定硬體挑選本地模型或量化方案時使用,也適合以實測吞吐量數據取代估算值進行比較。
執行需求
以 stdio 方式在本機執行,通常透過 npx 啟動,因此需要 Node.js 以及連線至 TokenMark 網站的網路存取。未宣告帳號、API 金鑰或環境變數;資料來源網址可用環境變數 TOKENMARK_URL 覆寫。
安裝前請注意
此伺服器從第三方網站讀取即時資料,因此查詢內容以及提交的倉庫資訊會離開本機。提交工具會把第三方倉庫連結與備註排入佇列等待人工審核,README 另指向一個需登入的網頁表單用於提交個人實測資料。

安裝

在 SourceWeft 中

  1. 開啟 儀表板中的 TokenMark,將其新增到工作區。
  2. 為需要使用其工具的對話啟用該服務。

Desktop only,透過 STDIO。 STDIO 服務會啟動本機處理程序,因此需要 SourceWeft 桌面主機。

其他 MCP 客戶端

參照 儲存庫 中的啟動說明。

README

@altronis/tokenmark-mcp

An MCP server that gives Claude (and other agents) real local-LLM benchmark data + hardware-aware model recommendations from TokenMark.

Ask "what should I run on a Strix Halo for coding?" and the agent answers from measured configs (decode tok/s, quant, backend), each with a source link. It never invents numbers.

Add to Claude Code

sh
claude mcp add tokenmark -- npx -y @altronis/tokenmark-mcp

Add to any MCP client

Run the server over stdio:

sh
npx -y @altronis/tokenmark-mcp

Or in a client config:

json
{  "mcpServers": {    "tokenmark": {      "command": "npx",      "args": ["-y", "@altronis/tokenmark-mcp"]    }  }}

Tools

  • tokenmark_recommend: { hardware, tasks?, prefer?, limit? } → ranked model + best-config picks (5 by default, 10 max) with measured tok/s + why.
  • tokenmark_configs: { model?, hardware?, limit? } → tracked benchmark configs, fastest first (25 by default, 50 max; matched gives the full count), with the run mode (speculative: true for MTP/DFlash/draft-model runs).
  • tokenmark_hardware: { platform? } → the hardware catalogue (Strix Halo, Gorgon Halo, DGX Spark, Mac Max/Ultra). With a platform: chip specs with a source per value, the boxes that ship it, Singapore prices.
  • tokenmark_search: { term } → matching models/configs.
  • tokenmark_submit: { repo, source?, note? } → queues a GitHub/Hugging Face repo with benchmark numbers for human review.
  • tokenmark_submission_status: { id } → where a submission is.

A speed someone measured themselves, or hardware missing from the catalogue, goes through the signed-in form at https://tokenmark.app/submit.

Data is pulled live from https://tokenmark.app (override with TOKENMARK_URL). Zero runtime dependencies.

The same data in your terminal

The CLI lives in cli/ of this repo:

sh
npx @altronis/tokenmark-cli recommend --hardware "Strix Halo" --tasks coding

The files in bin/ are the published builds, made from the TokenMark tracker that runs https://tokenmark.app.

MIT licensed. Data aggregated from public community benchmarks with attribution.

來源:README.md,提交 19596dc

工具

0
工具後設資料尚未被收錄。

版本歷史

1
  1. v0.2.1最新Oct 5, 2026