vllm-loop model-improvement progress tracker junk-tokens
models / / #42
▲ MADE POSITIVE PROGRESS

Reproducible source/image lane established

Last ~18h delta: SM80 startup blocker found and fixed (GLM sparse indexer invoked DeepGEMM metadata on unsupported SM80 -> capability-gated Triton fp8 MQA-logits fallback, source ffa38541); pinned NVCR CUDA base digests; KPool 11/11 and AutoRound W4A16 2/2 CPU gates pass repeatedly. PP4 loads but…

When (PT)2026-09-08 17:55 PT
Serving
Enginewtdcode/vllm-backport@85f227da + ffa385411 fix, NVCR CUDA base pinned by digest
KV

What was fixed

  • DeepGEMM metadata capability gate - SM80 reaches the Triton fp8 MQA-logits fallback instead of crashing startup

What didn't work

  • PP4 eager AR correctness still failing (divergence ~token 18)

raw JSON: /api/reports/42