{"model":"glm-5.3-flash","title":"AutoRound W4A16 correctness gate FAIL","status":"stuck","summary":"Current gate FAIL (correctness): PP4 eager AR loads with the SM80 Triton fallback (BF16 KV, no cache/speculation) but five identical cold 1K-token greedy requests diverge as early as output position 18; forced single-pass sparse-KV also diverges (position 19). No performance claim admitted.","occurred_at":"2026-09-08T22:30:00Z","config":{"serving":"TP1/PP4 eager AR (GPUs 4,5,7,8; GPU6 never selected)","engine":"wtdcode/vllm-backport, DeepGEMM SM80 fix @ ffa38541; Intel/GLM-5.3-Flash-W4A16-AutoRound rev 5eee1846"},"failed":["Single-pass sparse-KV (VLLM_TRITON_MLA_FORCE_KV_SPLITS=1): still diverges @token 19","Persistent CUDA top-k disabled (portable top_k_per_row_decode): still diverges @tokens 18-40","MoE auto-\u003eMarlin reproduces the same nondeterministic token family; Triton MoE cannot load AutoGPTQ W4A16 expert bias (rejected before load)"],"next":["Capture exact per-step SM80 Triton indexer-logit + selected-top-k hashes and PP boundary state to locate the first divergent tensor","Scheduler/KV/MTP/PP6/DFlash changes are not justified before that evidence"]}