vllm-benchmarkinglisted
Install: claude install-skill air-gapped/skills
# vLLM benchmarking
Target audience: operators producing defensible latency/throughput numbers against production or pre-production vLLM deployments, on datacenter GPUs, often in containerized or air-gapped environments.
**This skill measures; it does not tune.** Once a number is trusted and the
verdict is "too slow", the knobs live elsewhere in the `vllm` plugin:
**`vllm-performance-tuning`** (scheduler, MoE kernels, CUDA graphs, parallelism),
**`vllm-caching`** (KV tiering when the bottleneck is prefill or cache capacity),
**`vllm-nvidia-hardware`** (the SKU's own ceiling). Measure → change one thing →
re-measure with the same methodology; a tuning change compared against a
differently-shaped benchmark run is not evidence.
## Why this matters
Bad benchmarks are worse than no benchmarks — they drive the wrong decisions with false confidence. The three common failure modes:
1. **Wrong methodology.** `--request-rate inf` answers "saturation throughput," not "TTFT my users see." Mixing those up leads to buying GPUs to solve a latency problem, or shipping a latency regression because total throughput looked fine.
2. **Wrong workload.** `--dataset-name random` has zero prefix structure. Real coding-agent or RAG traffic has heavy prefix reuse. Benchmarking caching wins on random produces numbers that don't survive contact with prod.
3. **No warmup / wrong tokenizer.** First N requests hit cold CUDA graphs. Token counts are fiction unless `--tokenizer` matches the served model