vllm-loop model-improvement progress tracker junk-tokens
models / / #33
▲ MADE POSITIVE PROGRESS

512k depth sweep - saturation quantified

Shared-487k-prefix pure decode (server-counted): c1 25.1 / c2 38.1 / c4 44.1 / c8 52.3 / c12 54.6 aggregate - saturates from c8 (per-step cost grows ~linearly with batch at depth). 500 tok/s @512k is ~9x away and blocked by depth step-cost scaling + KV-capped concurrency: kernel/engine work, not…

When (PT)2026-09-12 05:50 PT
ServingTP1xPP6 (7,7,7,7,7,5), util 0.95-0.96, DSpark k=5, fp8_ds_mla
Enginelocalhost/vllm-backport-v41:sm80 Schaka v0.13.0 kit (c1b0907b overlay)
KV

Benchmarks

MetricValueΔ vs previousUnitContextNote
agg_decode_512k_tok_s (agg_decode_512k_tok_s) 54.6 tok/s -35.8% tok/s c12 (52.3 @c8, 44.1 @c4, 25.1 @c1) saturation from c8
dspark_accept_pct (dspark_accept_pct) 53 % -43.5% % c1 @512k (rises to 58.8 @c2, falls to 27.9 @c12)
kv_pool_tokens (kv_pool_tokens) 3404072 tokens -44.8% tokens PP6 util 0.96 - now the served default (+18% vs 0.95)

What didn't work

  • util 0.985 (4.69M tokens): 512k prefill OOM in fused_deepseek_v4_qnorm_rope_kv_rope_quant_insert - unusable

raw JSON: /api/reports/33