Benchmark Model

modular/skills/benchmark-model

作者 modularb9b3a8e86700無授權條款200 個星標收錄於 2026年10月8日更新於 2026年10月8日儲存庫今天更新

Benchmark a model served on MAX with the `max benchmark` command: measure throughput (tokens/sec), latency (TTFT, TPOT, inter-token latency), and GPU utilization by driving load against a running `max serve` endpoint. Use this whenever the user wants to benchmark, load-test, or measure the performance of a MAX model, get tokens-per-second / TTFT / TPOT numbers, run a concurrency or request-rate sweep, compare latency vs throughput, size a deployment, or produce benchmark JSON, even if they don't say "benchmark" by name. Also use when a `max benchmark` run fails to connect or reports zero/garbage numbers.

  1. b9b3a8e86700目前提交 b9b3a8e發布於 2026年10月8日

來源與署名

來源:modular/skills位於benchmark-model提交b9b3a8e

授權條款: 無授權條款

內容歸原作者所有。SourceWeft 從公開儲存庫中收錄這些內容。

檢舉或申請下架