models / / #41
● AM STUCK
AutoRound W4A16 correctness gate FAIL
Current gate FAIL (correctness): PP4 eager AR loads with the SM80 Triton fallback (BF16 KV, no cache/speculation) but five identical cold 1K-token greedy requests diverge as early as output position 18; forced single-pass sparse-KV also diverges (position 19). No performance claim admitted.
What didn't work
- Single-pass sparse-KV (VLLM_TRITON_MLA_FORCE_KV_SPLITS=1): still diverges @token 19
- Persistent CUDA top-k disabled (portable top_k_per_row_decode): still diverges @tokens 18-40
- MoE auto->Marlin reproduces the same nondeterministic token family; Triton MoE cannot load AutoGPTQ W4A16 expert bias (rejected before load)
What's next
- Capture exact per-step SM80 Triton indexer-logit + selected-top-k hashes and PP boundary state to locate the first divergent tensor
- Scheduler/KV/MTP/PP6/DFlash changes are not justified before that evidence
raw JSON: /api/reports/41