{"model":"deepseek-4.1-flash","title":"512k depth sweep - saturation quantified","status":"positive","summary":"Shared-487k-prefix pure decode (server-counted): c1 25.1 / c2 38.1 / c4 44.1 / c8 52.3 / c12 54.6 aggregate - saturates from c8 (per-step cost grows ~linearly with batch at depth). 500 tok/s @512k is ~9x away and blocked by depth step-cost scaling + KV-capped concurrency: kernel/engine work, not…","occurred_at":"2026-09-12T12:50:00Z","config":{"serving":"TP1xPP6 (7,7,7,7,7,5), util 0.95-0.96, DSpark k=5, fp8_ds_mla","engine":"localhost/vllm-backport-v41:sm80 Schaka v0.13.0 kit (c1b0907b overlay)"},"benchmarks":[{"metric":"agg_decode_512k_tok_s","value":54.6,"unit":"tok/s","context":"c12 (52.3 @c8, 44.1 @c4, 25.1 @c1)","note":"saturation from c8"},{"metric":"kv_pool_tokens","value":3404072,"unit":"tokens","context":"PP6 util 0.96 - now the served default (+18% vs 0.95)"},{"metric":"dspark_accept_pct","value":53,"unit":"%","context":"c1 @512k (rises to 58.8 @c2, falls to 27.9 @c12)"}],"failed":["util 0.985 (4.69M tokens): 512k prefill OOM in fused_deepseek_v4_qnorm_rope_kv_rope_quant_insert - unusable"]}