Benchmark Model

modular/skills/benchmark-model

by modularb9b3a8e86700No license200 starsListed Oct 8, 2026Updated Oct 8, 2026Repository updated today

Benchmark a model served on MAX with the `max benchmark` command: measure throughput (tokens/sec), latency (TTFT, TPOT, inter-token latency), and GPU utilization by driving load against a running `max serve` endpoint. Use this whenever the user wants to benchmark, load-test, or measure the performance of a MAX model, get tokens-per-second / TTFT / TPOT numbers, run a concurrency or request-rate sweep, compare latency vs throughput, size a deployment, or produce benchmark JSON, even if they don't say "benchmark" by name. Also use when a `max benchmark` run fails to connect or reports zero/garbage numbers.

Only the file list is public. File contents are available once the skill is installed in a workspace.

PathSizeType
references/flags.md6.3 KBtext/markdown
references/metrics.md4.4 KBtext/markdown
references/troubleshooting.md2.4 KBtext/markdown
SKILL.md8 KBtext/markdown

Source and attribution

Source:modular/skillsinbenchmark-modelat commitb9b3a8e

License: No license

Content belongs to its original authors. SourceWeft indexes it from a public repository.

Report or request removal