{"model":"deepseek-4.1-flash","title":"Benchmark #4 - live OpenCode sessions","status":"positive","summary":"5 real headless OpenCode sessions (qa + coding + read-tool): wall 18.4-21.1 s per prompt, output 48-120 tok, end-to-end 2.5-6.5 tok/s (median 5.07). Decode ~3.5x slower at 23k depth than shallow.","occurred_at":"2026-09-11T05:25:00Z","config":{"serving":"TP4xPP2, DSpark k=5, fp8_ds_mla, 1M ctx, CUDA graphs","engine":"localhost/vllm-backport-v41:sm80 (vLLM 0.12.0-sm80 + PR#56201 + Ampere shims + PP relay)"},"benchmarks":[{"metric":"opencode_c1_e2e_tok_s","value":5.07,"unit":"tok/s","context":"median of 5 real sessions (2.5-6.5 range)"}],"found":["OpenCode system prompt is ~23,462 tokens - every single-turn request pays ~10-11 s prefill (agent harness context floor)"]}