
TokenMark
io.github.sypherinv0.2.1Updated Oct 5, 2026
Measured local-LLM speeds and hardware-aware model picks for Strix Halo, DGX Spark and Mac.
Overview
Lets an assistant look up measured local-LLM benchmark speeds and hardware-aware model recommendations from TokenMark.
- What it does
- Provides tools to recommend ranked model and configuration picks for a given hardware platform and task, list tracked benchmark configurations with measured decode tokens per second, browse a hardware catalogue with chip specs and prices, search models and configs, and submit a repository with benchmark numbers for human review. Data is fetched live from the TokenMark site, and the server claims it never invents numbers.
- When to use it
- Useful when choosing which local model or quantization to run on specific hardware such as Strix Halo, DGX Spark or Mac, or when comparing measured throughput figures instead of estimates.
- Requirements
- Runs locally over stdio, typically launched with npx, so Node.js and network access to the TokenMark site are needed. No accounts, API keys or environment variables are declared; the data source URL can be overridden with the TOKENMARK_URL environment variable.
Installation
In SourceWeft
- Open TokenMark in the dashboard and add it to a workspace.
- Enable the server for the chats that should use its tools.
Desktop only via STDIO. STDIO servers start a local process, so they need the SourceWeft desktop host.
Other MCP clients
Follow the launch instructions in the repository.
README
@altronis/tokenmark-mcp
An MCP server that gives Claude (and other agents) real local-LLM benchmark data + hardware-aware model recommendations from TokenMark.
Ask "what should I run on a Strix Halo for coding?" and the agent answers from measured configs (decode tok/s, quant, backend), each with a source link. It never invents numbers.
Add to Claude Code
Add to any MCP client
Run the server over stdio:
Or in a client config:
Tools
tokenmark_recommend:{ hardware, tasks?, prefer?, limit? }→ ranked model + best-config picks (5 by default, 10 max) with measured tok/s + why.tokenmark_configs:{ model?, hardware?, limit? }→ tracked benchmark configs, fastest first (25 by default, 50 max;matchedgives the full count), with the run mode (speculative: truefor MTP/DFlash/draft-model runs).tokenmark_hardware:{ platform? }→ the hardware catalogue (Strix Halo, Gorgon Halo, DGX Spark, Mac Max/Ultra). With a platform: chip specs with a source per value, the boxes that ship it, Singapore prices.tokenmark_search:{ term }→ matching models/configs.tokenmark_submit:{ repo, source?, note? }→ queues a GitHub/Hugging Face repo with benchmark numbers for human review.tokenmark_submission_status:{ id }→ where a submission is.
A speed someone measured themselves, or hardware missing from the catalogue, goes through the signed-in form at https://tokenmark.app/submit.
Data is pulled live from https://tokenmark.app (override with TOKENMARK_URL). Zero runtime dependencies.
The same data in your terminal
The CLI lives in cli/ of this repo:
The files in bin/ are the published builds, made from the TokenMark tracker that runs https://tokenmark.app.
MIT licensed. Data aggregated from public community benchmarks with attribution.
Source: README.md at commit 19596dc
Tools
0Version history
1- v0.2.1LatestOct 5, 2026


