models / / #32
▼ MADE NEGATIVE PROGRESS
LMCache c1 cost A/B - ~11% decode cost
Plain PP6 65.1 tok/s (512-tok gen, thinking off) vs PP6 + LMCacheMPConnector 58.0 - the LMCache path costs ~11% c1 decode; real-OpenCode c1 median drops ~30 -> 26.5. Decision: plain PP6 stays the c1 served default; LMCache is a validated optional arm for cross-restart/evicted-prefix reuse (vLLM…
Benchmarks
| Metric | Value | Δ vs previous | Unit | Context | Note |
|---|---|---|---|---|---|
c1_decode_tok_s (c1_decode_tok_s) |
65.1 tok/s | -38.1% | tok/s | plain PP6, 512-tok generation | |
c1_decode_tok_s (c1_decode_tok_s) |
58 tok/s | -44.8% | tok/s | PP6 + LMCacheMPConnector | -11% |
raw JSON: /api/reports/32