models / / #39
▲ MADE POSITIVE PROGRESS
PP8/c11 live deployment receipt
NVFP4 target + DFlash2 K=7 live on TP1/PP8 (GPUs 1,0,2,3,4,5,7,8; GPU6 excluded): 256K max request, max-num-seqs 11, GPU cache 3,200 blocks = 2,964,172 tokens (~11.3x 256K concurrency); FULL_DECODE_ONLY graphs with DFlash buckets 8-88; healthy LMCache sidecar. Explicitly not a claim that c11…
Benchmarks
| Metric | Value | Δ vs previous | Unit | Context | Note |
|---|---|---|---|---|---|
kv_pool_tokens (kv_pool_tokens) |
2964172 tokens | first point | tokens | PP8, 3,200 blocks, 256K max request |
raw JSON: /api/reports/39