{"model":"deepseek-4.1-flash","title":"README: PP8/pipeline-parallel numbers documented","status":"positive","summary":"Published measured layout comparison: PP layouts prefill ~2x (no all-reduce; Gen2 x4 caps TP4 prefill ~1,700 tok/s) but lose deep decode vs TP4xPP2 (92 vs 130 @c64/512k). Third-party PP8 claims (117 single / 532 agg) use tokens/(last-first) + staggered arrivals and read higher than wall-clock.","occurred_at":"2026-09-12T02:21:00Z","found":["PP8 per-token KV is 2x TP4xPP2 because its pp_share relay replicates caches across ranks"]}