models / / #34
▲ MADE POSITIVE PROGRESS
PP8 brought up - 2.75x the KV pool of PP6
Layout-resolver intersection fix makes PP8 work: 9.35M-token KV pool (2.75x PP6), 8.9x concurrency @1M; 512k aggregate still ~60 tok/s @c32 -> step-cost bound. Leads: vLLM #56120 (SM80 Triton knobs), SGLang #38646 (NVFP4 sparse-MLA, 384 B/tok, SM120-only).
Benchmarks
| Metric | Value | Δ vs previous | Unit | Context | Note |
|---|---|---|---|---|---|
agg_decode_512k_tok_s (agg_decode_512k_tok_s) |
60 tok/s | 9.9% | tok/s | PP8 c32 @512k - step-cost bound | |
kv_pool_tokens (kv_pool_tokens) |
9350000 tokens | 174.7% | tokens | PP8 (2.75x PP6) |
raw JSON: /api/reports/34