{"model":"deepseek-4.1-flash","title":"Perf finding - the prefill bottleneck is the TP all-reduce","status":"positive","summary":"Got Schaka's TP1xPP8 profile booting (2 patches: intersect per-rank KV layouts; guard relay warmup index-K write): PP8 prefill 5,738 tok/s @32k / 8,264 @128k = ~3.7x TP4xPP2 (link-bound all-reduce over Gen2 x4) - but the relay was numerically wrong, reverted.","occurred_at":"2026-09-11T13:32:00Z","config":{"serving":"TP4xPP2, DSpark k=5, fp8_ds_mla, 1M ctx, CUDA graphs","engine":"localhost/vllm-backport-v41:sm80 (vLLM 0.12.0-sm80 + PR#56201 + Ampere shims + PP relay)"},"benchmarks":[{"metric":"prefill_tok_s","value":5738,"unit":"tok/s","context":"TP1xPP8 @32k (~3.7x TP4xPP2)"},{"metric":"prefill_tok_s","value":8264,"unit":"tok/s","context":"TP1xPP8 @128k"}],"failed":["PP8 accuracy gate rejected it: needle 0/3, coding 1/5 with garbled identifiers - relay does not reproduce indexer-K cache write on every stage"]}