Benchmark Model

modular/skills/benchmark-model

作者 modularb9b3a8e86700无许可证200 个星标收录于 2026年10月8日更新于 2026年10月8日仓库今天更新

Benchmark a model served on MAX with the `max benchmark` command: measure throughput (tokens/sec), latency (TTFT, TPOT, inter-token latency), and GPU utilization by driving load against a running `max serve` endpoint. Use this whenever the user wants to benchmark, load-test, or measure the performance of a MAX model, get tokens-per-second / TTFT / TPOT numbers, run a concurrency or request-rate sweep, compare latency vs throughput, size a deployment, or produce benchmark JSON, even if they don't say "benchmark" by name. Also use when a `max benchmark` run fails to connect or reports zero/garbage numbers.

仅公开文件列表。将技能安装到工作区后即可查看文件内容。

路径大小类型
references/flags.md6.3 KBtext/markdown
references/metrics.md4.4 KBtext/markdown
references/troubleshooting.md2.4 KBtext/markdown
SKILL.md8 KBtext/markdown

来源与署名

来源:modular/skills位于benchmark-model提交b9b3a8e

许可证: 无许可证

内容归原作者所有。SourceWeft 从公开仓库中收录这些内容。

举报或申请下架