vllm-loop model-improvement progress tracker junk-tokens
models / / #4
▲ MADE POSITIVE PROGRESS

Schaka publishes first working recipe - pivoted

Schaka 170hx-journey (TP4xPP2) is the first working V4.1 on SM80: 43.8/72.9/104.2 tok/s @ c1/c4/c8, prefill 1,584/1,650/1,389 tok/s @ 39k/119k/319k, KV pool 14,058,003 tokens. His kit matches our bug: wo_a wrong-shaped under Marlin, and V4.0's 9-arg qnorm op dropping V4.1's apply_q_norm…

When (PT)2026-09-10 17:15 PT
ServingTP4xPP2 (Schaka 170hx-journey)
EngineSchaka/170hx-journey patch kit
KV

Benchmarks

MetricValueΔ vs previousUnitContextNote
agg_decode_c4_tok_s (agg_decode_c4_tok_s) 72.9 tok/s first point tok/s Schaka box (author-reported)
agg_decode_c8_tok_s (agg_decode_c8_tok_s) 104.2 tok/s first point tok/s Schaka box (author-reported)
c1_decode_tok_s (c1_decode_tok_s) 43.8 tok/s first point tok/s Schaka box, TP4xPP2 (author-reported)
kv_pool_tokens (kv_pool_tokens) 14058003 tokens first point tokens Schaka TP4xPP2 pool
prefill_tok_s (prefill_tok_s) 1584 tok/s first point tok/s Schaka box @39k (1,650 @119k, 1,389 @319k)

What was fixed

  • wo_a wrong-shaped under Marlin -> emulation kernel
  • V4.0 9-arg qnorm op dropped V4.1 apply_q_norm -> Q normalized twice -> wrong output, no error
  • pp_kv_group_relay.py unlocks PP (previously thought impossible)

raw JSON: /api/reports/4