{"model":"deepseek-4.1-flash","title":"LMCache c1 cost A/B - ~11% decode cost","status":"negative","summary":"Plain PP6 65.1 tok/s (512-tok gen, thinking off) vs PP6 + LMCacheMPConnector 58.0 - the LMCache path costs ~11% c1 decode; real-OpenCode c1 median drops ~30 -\u003e 26.5. Decision: plain PP6 stays the c1 served default; LMCache is a validated optional arm for cross-restart/evicted-prefix reuse (vLLM…","occurred_at":"2026-09-12T11:16:00Z","config":{"serving":"TP1xPP6 (7,7,7,7,7,5), util 0.95-0.96, DSpark k=5, fp8_ds_mla","engine":"localhost/vllm-backport-v41:sm80 Schaka v0.13.0 kit (c1b0907b overlay)"},"benchmarks":[{"metric":"c1_decode_tok_s","value":65.1,"unit":"tok/s","context":"plain PP6, 512-tok generation"},{"metric":"c1_decode_tok_s","value":58,"unit":"tok/s","context":"PP6 + LMCacheMPConnector","note":"-11%"}]}