vllm-loop model-improvement progress tracker junk-tokens
models / / #3
▲ MADE POSITIVE PROGRESS

Correctness root cause localized

Per-branch probe: attention 'o' absmax ~2.0 (sparse MLA works) but _o_proj output absmax ~0.013 - a ~150x shrink; residual goes FFN-only and explodes (0.6 -> 3.5e29 at KV-source layers). One bug: fused inverse-RoPE + wo_a einsum + wo_b.

When (PT)2026-09-10 16:47 PT
ServingTP8xPP1, DSpark k=5, fp8_ds_mla, 1M ctx
Enginelocal/vllm-dsv41:sm80 @ dsv41-feat@e47aa780
KV

What was found

  • o absmax ~2.0, _o_proj absmax ~0.013 -> ~150x shrink of the attention contribution

What didn't work

  • Ruled out: tokenization/template, Engram (DSV41_ZERO_ENGRAM), wo_a layout, backend selection, NaN

raw JSON: /api/reports/3