vllm-loop model-improvement progress tracker junk-tokens

Models

Per-model serving-optimization progress. Pick a model for benchmark history, status log, and reports.

deepseek-4.1-flash ▼ negative
Tested prefill-chunk size as the ITL-tail fix: MAXBATCH=8192 collapses KV (7.86M->3.23M) and fails c1; MAXBATCH=1024 is rejected by the multimodal min. No free ITL fix; baseline restored
38 reports · latest 2026-09-12 13:02 PT
glm-5.3-flash ▲ positive
Last ~18h delta: SM80 startup blocker found and fixed (GLM sparse indexer invoked DeepGEMM metadata on unsupported SM80 -> capability-gated Triton fp8 MQA-logits fallback, source ffa38541); pinned NVCR CUDA base digests; KPool 11/11 and AutoRound W4A16 2/2 CPU gates pass repeatedly. PP4 loads but…
6 reports · latest 2026-09-08 17:55 PT