This hand-maintained index is navigation for actual public submissions. The chronology below is the immutable audit history; model packets remain the source for verified results that have not been submitted.
Historical fixed-first-100-token rows used the repository’s legacy 100-event/99-interval calculation unless explicitly qualified otherwise. Those approved receipts remain immutable. New submissions are required to use the conventional 99-interval field.
| Model / lane | Hardware | Representative submitted result | LocalMaxxing ID | Evidence |
|---|---|---|---|---|
| Poolside Laguna S 2.1 INT4, TP4+EP4 DFlash11 | 4x Arc Pro B70 | 102.971 submitted legacy-event median; 101.942 conventional interval median; exact cold width-12 draft-FP8 graph gate | cms2ccv2d00lps201rej94pjy |
qualified packet; correction |
| Qwen3.6 27B AutoRound INT4, TP2 | 2x Arc Pro B70 | 95.385 median tok/s, fixed cold realistic gate | cmrh35ct50092mj01h7jgydqj |
packet |
| Gemma 4 26B A4B Q8 | 1x Arc Pro B70 | 124.977 median tok/s, fixed cold realistic gate | cmr1u77na01k2ld01kalwzs1e |
packet |
| Qwen3.6 35B Quark INT8, TP4 | 4x Arc Pro B70 | 93.551 output tok/s, strict deep gate | cmqq4mw4c00yfqo01gb2ucgxj |
packet |
| Qwen3.6 27B GGUF Q4_0, native DFlash5 + Xe2 M6 | 1x Arc Pro B70 | 47.819 median tok/s, fixed cold realistic gate | cmrjbx8bc02g8mj01yzz2v701 |
evidence |
| MiniMax M2.7 AutoRound INT4 | 4x Arc Pro B70 | 65.752 output tok/s, quality-gated public row | cmp6a5c1o00mpo3011hg8ncyp |
packet |
| DeepSeek V4 Flash uniform-K160, TP4+EP | 4x Arc Pro B70 | 80.820 median tok/s, target-verified DSpark7 sharded target argmax | cmrquta9905w3lg013m5vxoqx |
packet; repro |
| DeepSeek V4 Flash uniform-K160, TP4+EP nonspec | 4x Arc Pro B70 | 43.767 median tok/s, direct routed-MoE + wide-epoch oneCCL | cmrmnp7h81nntmj01lfenydgj |
ledger |
| Rapid model snapshots | 1x Arc Pro B70 | Multiple fixed cold realistic references | see packet | performance index |
Current measured-but-unsubmitted work belongs in its model packet, not this public-submission index.
Date: 2026-07-18
Model: 0xSero/DeepSeek-V4-Flash-180B, experimental uniform-K160 checkpoint,
vLLM/XPU TP4+EP on four Intel Arc Pro B70 GPUs.
| Label | LocalMaxxing ID | GPUs | Input | Output | tok/s out | tok/s total | Validation |
|---|---|---|---|---|---|---|---|
deepseek-v4-flash-k160-b70-tp4-dspark7-sharded-target-argmax-realistic-80.820tok-20260718 |
cmrquta9905w3lg013m5vxoqx |
4 | suite median 62 | 128 | 80.820 median 1-100 after TTFT | 67.763 median wall full128 | current target-verified record: guarded greedy-only target verification projects local 32,320-token vocabulary shards, gathers tiny top-1 pairs, and commits target IDs through native SYCL without full 129,280-token logits all-gather or FP32 sampler materialization. Independent strict medians are 80.820 / 76.900 / 78.287 tok/s; all 36 realistic requests are cache-zero, and four ordered canary suites pass 24/24. One active generation; unchanged K160 target verifies accepted tokens at M=8. Evidence /mnt/fast-ai/bench-results/deepseek-v4-flash-xpu/dspark7-sharded-target-argmax-candidate-20260718T2100Z. |
deepseek-v4-flash-k160-b70-tp4-dspark7-m8-router-realistic-80.164tok-20260718 |
cmrqp2uoa05ublg01lh6yluj8 |
4 | suite median 62 | 128 | 80.164 median 1-100 after TTFT | 62.997 median wall full128 | superseded target-verified record: one native M=8 submission fuses biased top-k, gather, normalization, and routed scaling with bit-exact IDs and FP32 weights. Independent strict medians are 75.846 / 77.573 / 80.164 tok/s; all 36 realistic requests are cache-zero, and four ordered canary suites pass 24/24. One active generation; unchanged K160 target verifies accepted tokens at M=8. Evidence /mnt/fast-ai/bench-results/deepseek-v4-flash-xpu/dspark7-m8-router-fused-candidate-20260718T1815Z. |
deepseek-v4-flash-k160-b70-tp4-dspark7-m8-w8a16-n128-realistic-78.288tok-20260718 |
cmrqlp9je05thlg01q4igkk0x |
4 | suite median 62 | 128 | 78.288 median 1-100 after TTFT | 63.060 median wall full128 | superseded target-verified record: selective oneDNN W8A16 removes activation quantization for four M=8 dense families and N128 improves the Xe2 routed-MXFP4 tile. Independent strict medians are 78.288 / 74.410 / 76.938 tok/s; all 36 realistic requests are unique and cache-zero, and four ordered canary suites pass 24/24. W8A16 is quality-gated rather than claimed bitwise row-invariant (maximum observed BF16 difference 0.0078125). One active generation; unchanged K160 target verifies accepted tokens at M=8. Evidence /mnt/fast-ai/bench-results/deepseek-v4-flash-xpu/dspark7-m8-w8a16-n128-candidate-20260718T2130Z. |
deepseek-v4-flash-k160-b70-tp4-dspark7-m8-compressor-realistic-71.507tok-20260718 |
cmrql07qs05t4lg01p86jjybx |
4 | suite median 62 | 128 | 71.507 median 1-100 after TTFT | 57.567 median wall full128 | current target-verified record: exact M=8 strided-batch compressor projections preserve independent-row accumulation while deleting seven GEMM submissions per compressor. Independent strict medians are 69.344 / 71.507 / 70.249 tok/s; all 36 realistic requests are fresh and cache-zero, and four exact suites pass 24/24. One active generation; unchanged K160 target verifies accepted tokens at M=8. Evidence /mnt/fast-ai/bench-results/deepseek-v4-flash-xpu/dspark7-xpu-compressor-m8-candidate-20260718T2030Z. |
deepseek-v4-flash-k160-b70-tp4-dspark7-replicated-w1-realistic-67.501tok-20260718 |
cmrqjhpmz05snlg01ujiehc0u |
4 | suite median 62 | 128 | 67.501 median 1-100 after TTFT | 54.563 median wall full128 | current target-verified record: only the small Markov W1 embedding is replicated, removing seven all-reduces while expensive W2 remains sharded and persistent. Independent strict medians are 65.657 / 67.501 / 67.182 tok/s, all 36 realistic requests are fresh and cache-zero, and four six-case exact suites pass. One active generation; unchanged K160 target verifies accepted tokens at M=8. Evidence /mnt/fast-ai/bench-results/deepseek-v4-flash-xpu/dspark7-xpu-replicated-w1-candidate-20260718T1800Z. |
deepseek-v4-flash-k160-b70-tp4-dspark7-persistent-markov-realistic-66.479tok-20260718 |
cmrqiovsv05s6lg012d8v5nz8 |
4 | suite median 62 | 128 | 66.479 median 1-100 after TTFT | 54.806 median wall full128 | current target-verified record: fixed device buffers and direct-output operations remove allocation/cat/copy overhead from the seven-step Markov transaction while keeping W1/W2 sharded. The exact four-card component saves 0.786613 ms/cycle at the slowest rank. Independent strict medians are 65.674 / 66.479 / 63.559 tok/s, all 36 realistic requests are fresh and cache-zero, and four six-case exact suites pass. One active generation; unchanged K160 target verifies accepted tokens at M=8. Evidence /mnt/fast-ai/bench-results/deepseek-v4-flash-xpu/dspark7-xpu-persistent-markov-bundle-20260718T1640Z. |
deepseek-v4-flash-k160-b70-tp4-dspark7-piecewise-exactm7-realistic-64.661tok-20260718 |
cmrpymqh505mxlg01tzg3e0yl |
4 | suite median 62 | 128 | 64.661 median 1-100 after TTFT | 51.581 median wall full128 | current target-verified record: one active generation uses the official three-stage DSpark7 draft in a private breakable PIECEWISE graph captured at exact M=7; the unchanged K160 target verifies accepted tokens at M=8. Independent strict medians are 64.661 / 61.725 / 64.275 tok/s, all 36 realistic requests are fresh and cache-zero, and three six-case exact suites pass. Full draft graph was correctness-rejected and fixed DSpark5 regressed. Evidence /mnt/fast-ai/bench-results/deepseek-v4-flash-xpu/dspark7-xpu-targetpw-draftpw-exactm7-20260718T0556Z. |
Date: 2026-07-16
Model: 0xSero/DeepSeek-V4-Flash-180B, experimental uniform-K160 checkpoint,
vLLM/XPU TP4+EP on four Intel Arc Pro B70 GPUs.
| Label | LocalMaxxing ID | GPUs | Input | Output | tok/s out | tok/s total | Validation |
|---|---|---|---|---|---|---|---|
deepseek-v4-flash-k160-b70-tp4-mtp1-qnorm-routeportfolio-realistic-63.851tok-20260716 |
cmrocpuhq029hlg01g3yzglko |
4 | suite median 62 | 128 | 63.851 median 1-100 after TTFT | 53.224 median wall full128 | current target-verified record: a predeclared portfolio extends fused QNorm/RoPE/FP8-KV insertion to verifier M=2 and selects direct remap plus two fixed 12-slot N64 compact MXFP4 GEMMs with canonical clamp-SwiGLU and generic gather. The standalone route component remains below its unchanged 0.50 ms gate and is not claimed independently. Four physical B70s pass 336/336 changed graph cases bitwise; the guarded production wrapper passes another 84/84. Same-binary B-A-B medians are 62.516 / 61.718 / 63.851 tok/s; 70/70 exact captures pass across positions 28/58 and every qualifying request is cached-zero. Evidence /mnt/fast-ai/bench-results/deepseek-v4-flash-xpu/qnorm-routeportfolio-candidate-b2-20260716T2255Z. |
deepseek-v4-flash-k160-b70-tp4-mtp1-m2-router-realistic-63.350tok-20260716 |
cmrncv39w003ylg01hogleazo |
4 | suite median 62 | 128 | 63.350 median 1-100 after TTFT | 52.972 median wall full128 | superseded target-verified record: exact trace arguments identified 40 target-router [2,160], K6 calls per cycle. One native submission selects, normalizes, and scales both verifier rows. Four physical B70s pass 160/160 changed eager and 128/128 graph-replay M=2 cases bitwise, with the generalized M=1 path passing the same counts. A same-build flag-off control reaches 59.108 versus independent candidate suites at 62.883/63.350; 70/70 exact captures pass across positions 28/58 and every request is cached-zero. Evidence /mnt/fast-ai/bench-results/deepseek-v4-flash-xpu/mtp1-m2-router-norm-candidate-20260716T0605Z. |
deepseek-v4-flash-k160-b70-tp4-mtp1-m2-mhc-realistic-60.264tok-20260716 |
cmrmvjbok1np3mj01p9il8486 |
4 | suite median 62 | 128 | 60.264 median 1-100 after TTFT | 50.329 median wall full128 | superseded target-verified record: one native command launches independent 256-thread workgroups for both M=2 verifier rows while preserving the proven M=1 reduction and BF16 arithmetic boundaries. Four-card changed-input eager/graph gates are bitwise exact and save 0.962-0.971 ms across 85 boundaries. Independent cold suites reach 60.264/59.292 versus 57.412; 70/70 exact captures pass across positions 28/58 and every request is cached-zero. Evidence /mnt/fast-ai/bench-results/deepseek-v4-flash-xpu/mtp1-m2-mhc-single-kernel-candidate-20260716T0210Z. |
deepseek-v4-flash-k160-b70-tp4-mtp1-m2-fusion-realistic-57.412tok-20260716 |
cmrmrgce51nojmj01bbxoruuu |
4 | suite median 62 | 128 | 57.412 median 1-100 after TTFT | 48.498 median wall full128 | superseded target-verified record: exact shared clamped-SwiGLU/dynamic-FP8 quantization covers verifier M=2, and generic M=2 routed MoE uses exact fused clamp/SiLU without losing expert grouping or cross-row weight reuse. Independent strict suites reach 57.412/56.952 versus the 55.704 record; 70/70 exact capture suites pass after both suites across former rollover positions 28/58, and all 444 requests are cached-zero. Evidence /mnt/fast-ai/bench-results/deepseek-v4-flash-xpu/mtp1-m2-fusion-candidate-20260716T000928Z. |
deepseek-v4-flash-k160-b70-tp4-mtp1-direct-wideepoch-realistic-55.704tok-20260715 |
cmrmoyenp1no3mj01fz2gjzo6 |
4 | suite median 62 | 128 | 55.704 median 1-100 after TTFT | 46.899 median wall full128 | current target-verified record: exact direct M=1 routed MoE accelerates the attached draft layer while row-exact batched compressors and selective M=2 W8A16 retain the target verifier; wide-epoch oneCCL removes the readiness rollover. Matching strict suites reach 55.704/55.668, 70/70 exact captures pass including 50 after both suites and former failure positions 28/58, all requests are cached-zero, and all worker maps contain the selected libccl. Evidence /mnt/fast-ai/bench-results/deepseek-v4-flash-xpu/mtp1-direct-moe-wideepoch-candidate-20260715T2315Z. |
deepseek-v4-flash-k160-b70-tp4-direct-moe-wideepoch-realistic-43.767tok-20260715 |
cmrmnp7h81nntmj01lfenydgj |
4 | suite median 62 | 128 | 43.767 median 1-100 after TTFT | 39.281 median wall full128 | current trustworthy nonspeculative record: exact router normalization plus direct M=1 routed-MoE gather removes generic intermediates in 40 normal MoE layers. A same-build direct-off control reaches 41.991/42.155; four strict candidate suites reach 43.767/43.699/43.694/43.668 with all 48 rows cached-zero. The oneCCL Arc readiness identity is widened from an 11-bit reused counter to a 24-bit collective epoch plus 7-bit communicator tag; 70/70 exact captures pass across former failures 28 and 58. Evidence /mnt/fast-ai/bench-results/deepseek-v4-flash-xpu/nospec-direct-moe-wideepoch-candidate-20260715T2220Z. |
deepseek-v4-flash-k160-b70-tp4-native-m1-router-realistic-41.733tok-20260715 |
cmrmjd3io1nn1mj013stqoe4b |
4 | suite median 62 | 128 | 41.733 median 1-100 after TTFT | 37.625 median wall full128 | current trustworthy nonspeculative record: a SIMD16 Xe2 kernel replaces generic M=1 bias-add/radix-top-k/gather in 40 normal MoE layers while preserving raw weights and existing normalization. Four cards pass 40/40 changed-input epochs; two strict suites reach 41.733/41.514 versus a same-commit flag-off control at 40.068. Twenty exact captures pass, including ten after both suites, and all requests are cached-zero. Evidence /mnt/fast-ai/bench-results/deepseek-v4-flash-xpu/nospec-m1-router-candidate-20260715T2021Z. |
deepseek-v4-flash-k160-b70-tp4-mtp1-bmm-compressor-realistic-55.524tok-20260715 |
cmrmgacdq1nmimj01i4sfqytp |
4 | suite median 62 | 128 | 55.524 median 1-100 after TTFT | 47.475 median wall full128 | current target-verified record: a strided-batch FP32 compressor keeps the verifier rows independent while sharing one graph node and weight view. Both real K160 compressor shapes pass 40/40 changing eager and graph replays on all four cards; two strict suites reach 55.524/54.709, 20/20 sustained exact captures pass, all requests are cached-zero, and acceptance is 77.96%. Evidence /mnt/fast-ai/bench-results/deepseek-v4-flash-xpu/mtp1-rowexact-bmm-w8a16-m2-candidate-20260715T2000Z. |
deepseek-v4-flash-k160-b70-tp4-mtp1-rowexact-w8a16m2-realistic-54.465tok-20260715 |
cmrmfivhg1nmamj012e3138my |
4 | suite median 62 | 128 | 54.465 median 1-100 after TTFT | 46.485 median wall full128 | current target-verified record: four selective target projection families use one row-exact M=2 W8A16 call instead of the verifier’s W8A8 fallback. Four cards × four shapes × 40 changing epochs pass bit-for-bit; two strict suites reach 54.465/54.445, 20/20 sustained exact captures pass, every request is cached-zero, and acceptance is 77.68%. Evidence /mnt/fast-ai/bench-results/deepseek-v4-flash-xpu/mtp1-rowexact-w8a16-m2-candidate-20260715T1945Z. |
deepseek-v4-flash-k160-b70-tp4-mtp1-rowexact-realistic-50.017tok-20260715 |
cmrmetch81nm3mj01w1pidsyt |
4 | suite median 62 | 128 | 50.017 median 1-100 after TTFT | 43.623 median wall full128 | current target-verified record: attached MTP drafts one token; exact M=1-per-row FP32 compressor projections repair a sustained M=2 graph-state failure. Twenty ordered exact captures pass, including ten after two strict suites; support is 49.420 tok/s, every request cached-zero, and measured acceptance is 77.42%. Evidence /mnt/fast-ai/bench-results/deepseek-v4-flash-xpu/mtp1-compressor-rowexact-graph-oneccl1712-20260715T1840Z. |
deepseek-v4-flash-k160-b70-tp4-repeatability-correct-allreduce-route-realistic-40.170tok-20260715 |
cmrmebmzg1nm0mj01k30nv6vw |
4 | suite median 62 | 128 | 40.170 median 1-100 after TTFT | 36.390 median wall full128 | current trustworthy record: vLLM 93fde4186 makes BF16 KV/RoPE write regions disjoint; exact-version oneCCL 2021.17.2 routes only all-reduces above 128 KiB around the corrupt prefill SYCL path while retaining fast decode collectives. Ten exact captures pass 10/10; strict support is 40.096 tok/s and all 24 rows are cold/cached-zero. Worker maps confirm the selected hashed libccl and matching SYCL/UR runtime. Evidence /mnt/fast-ai/bench-results/deepseek-v4-flash-xpu/nospec-graph-oneccl1712-bf16-allreduce128k-preload094-cache-fix-20260715T1530Z. |
deepseek-v4-flash-k160-b70-tp4-fused-qnorm-rope-kv-insert-realistic-40.136tok-20260715 |
cmrm601ig1hsmmj017npoivfd |
4 | suite median 62 | 128 | 40.136 median 1-100 after TTFT | 37.968 median wall full128 | historical speed evidence, not repeatability-certified: exact focused QNorm/RoPE/direct-KV gates passed before submission, but later consecutive changed-prompt testing exposed deterministic wrong arithmetic on the unmodified large-SYCL-allreduce path. Retain the row; use cmrmebmzg1nm0mj01k30nv6vw as the current quality authority. Evidence /mnt/fast-ai/bench-results/deepseek-v4-flash-xpu/fused-qnorm-rope-kv-insert-candidate-20260715T1040Z. |
deepseek-v4-flash-k160-b70-tp4-split-fp8-b4-qk16-realistic-40.021tok-20260715 |
cmrlnp01l12q4mj01p58ynsyd |
4 | suite median 62 | 128 | 40.021 median 1-100 after TTFT | 37.843 median wall full128 | superseded trustworthy record: 4-head/16-warp QK geometry exposes more Xe2 parallelism while retaining four-warp tiled PV. Focused output is bitwise exact; changed-input graph replay is 1073 -> 437 -> 1073; all 12 strict rows are cold/cached-zero. p10 39.608. Evidence /mnt/fast-ai/bench-results/deepseek-v4-flash-xpu/split-fp8-geometry-b4-qk16-recordidentity-20260715T0144Z. |
deepseek-v4-flash-k160-b70-tp4-shared-expert-fused-act-quant-realistic-34.067tok-20260715 |
cmrlf1hn609glmj019rsjdl4r |
4 | suite median 62 | 128 | 34.067 median 1-100 after TTFT | 30.653 median wall full128 | superseded trustworthy record: exact clamp-at-10 SwiGLU + dynamic E4M3FN quant feeds canonical W8A8 shared-down while retaining selective W8A16 elsewhere. Paired strict runs reached 34.067/34.050, all cold/cached-zero; sequential replay, exact canaries, executable gates, and frozen 101! - 1 pass. Evidence /mnt/fast-ai/bench-results/deepseek-v4-flash-xpu/shared-expert-fused-act-quant-20260715T0140Z. |
deepseek-v4-flash-k160-b70-tp4-w8a16-high4-realistic-33.434tok-20260714 |
cmrlb675r0705mj01k9psoub0 |
4 | suite median 62 | 128 | 33.434 median 1-100 after TTFT | 30.218 median wall full128 | current trustworthy record: W8A16 only for fused WQA/WKV, Q-B, O-B, and shared gate/up; shared-down stays W8A8. Two strict runs reached 33.434/33.363, all cold/cached-zero. Sequential replay, exact canaries, frozen 101! - 1, and executable quality gates pass; the known intermittent K160 CJK-corruption floor remains documented. Evidence /mnt/fast-ai/bench-results/deepseek-v4-flash-xpu/w8a16-high4-no-shared-down-20260714T2346Z. |
deepseek-v4-flash-k160-b70-tp4-mhc-post-pre-m1-single-realistic-30.340tok-20260714 |
cmrl9xiwe06zzmj01cof0k38p |
4 | suite median 62 | 128 | 30.340 median 1-100 after TTFT | 27.663 median wall full128 | current trustworthy record: one 256-thread Xe2 kernel preserves the BF16 producer boundary while fusing M=1 MHC post/pre; standard oneCCL remains unchanged. Three strict runs reached 30.340/30.214/30.240, combined 36-prompt median 30.271. Forty changing microcases, graph replay, sequential exact canaries, and every cold/cached-zero row pass. Evidence /mnt/fast-ai/bench-results/deepseek-v4-flash-xpu/mhc-post-pre-m1-single-gmem95-20260714T1921Z. |
deepseek-v4-flash-k160-b70-tp4-w8a8-woa-corrected-realistic-30.239tok-20260714 |
cmrl2619q06hwmj011j5rtnbt |
4 | suite median 62 | 128 | 30.239 median 1-100 after TTFT | 28.880 median wall full128 | superseded trustworthy record: standard dense scale grids are prepacked once while special wo_a BMM scales remain canonical. Two strict serial suites reached 30.230/30.239; sequential exact canaries and all 12 cold/cached-zero rows pass. Evidence /mnt/fast-ai/bench-results/deepseek-v4-flash-xpu/tp4-w8a8-woa-corrected-n64-20260714T1940Z. |
deepseek-v4-flash-k160-b70-tp4-block-fp8-w8a16-realistic-33.887tok-20260714 |
cmrl12pke06ehmj01i9a1f0gu |
4 | suite median 62 | 128 | 33.887 median 1-100 after TTFT | 32.187 median wall full128 | invalid historical submission: generic prepack transposed special wo_a BMM scales. Corrected W8A16 is also quality-rejected by the frozen long-math gate. Do not cite as a record. |
deepseek-v4-flash-k160-b70-tp4-fp8-scale-prepack-realistic-30.295tok-20260714 |
cmrl0rf5u06b6mj01y1s9ew2u |
4 | suite median 62 | 128 | 30.295 median 1-100 after TTFT | 28.904 median wall full128 | invalid historical submission: generic prepack transposed special wo_a BMM scales. Do not cite as a record. |
deepseek-v4-flash-k160-b70-tp4-inplace-allreduce-realistic-29.913tok-20260714 |
cmrkz6vo7061fmj01v09nwz72 |
4 | suite median 62 | 128 | 29.913 median 1-100 after TTFT | 28.578 median wall full128 | superseded policy-compliant record: mutation-declared TP-only in-place reduction for 87 contiguous BF16 [1,4096] outputs; changed-input graph replay and exact canaries passed. Two valid strict suites reached 29.913 and 29.911 with all 12 prompts cold/cached-zero and no speculation/reuse. Evidence /mnt/fast-ai/bench-results/deepseek-v4-flash-xpu/tp4-inplace-ar-candidate-20260714T1920Z; queue experiments/deepseek-v4-flash-reap-xpu-b70/localmaxxing/deepseek-v4-flash-k160-tp4-inplace-allreduce-realistic-20260714.queue.json. |
deepseek-v4-flash-k160-b70-tp4-split-fp8-attn-realistic-29.822tok-20260714 |
cmrkxoavs05uimj01p9ix2dtk |
4 | suite median 69 | 128 | 29.822 median 1-100 after TTFT | 28.460 median wall full128 | superseded policy-compliant record: one active generation, fixed realistic suite, 12 unique cold prompts, cached_tokens=0 throughout, no speculation/reuse. Split QK/LSE plus 8-by-64 tiled PV reads paged UE8M0 FP8 KV directly and preserves runtime lengths and learned sinks. p10 29.426, mean 29.764, exact arithmetic/copy/Paris/JSON canaries passed. Evidence /mnt/fast-ai/bench-results/deepseek-v4-flash-xpu/tp4-split-fp8-attn-recordidentity-20260714T1733Z; queue experiments/deepseek-v4-flash-reap-xpu-b70/localmaxxing/deepseek-v4-flash-k160-tp4-split-fp8-attn-realistic-20260714.queue.json. K160 is a hash-pruned performance artifact, not a quality-certified final pack. |
deepseek-v4-flash-k160-b70-tp4-direct-fp8-attn-realistic-21.545tok-20260714 |
cmrkw8db205nymj01gl5bbmsc |
4 | suite median 69 | 128 | 21.545 median 1-100 after TTFT | see queue | superseded direct-attention record: graph-captured Triton reads paged UE8M0 FP8 KV directly, uses device runtime lengths, and applies learned sinks. All cold/cached-zero and exact canary gates passed. Evidence /mnt/fast-ai/bench-results/deepseek-v4-flash-xpu/tp4-direct-fp8-attn-20260714T1500Z; queue experiments/deepseek-v4-flash-reap-xpu-b70/localmaxxing/deepseek-v4-flash-k160-tp4-direct-fp8-attn-realistic-20260714.queue.json. |
Earlier approved DeepSeek K160 progression is preserved in the experiment
queues and ledger: native mHC 10.0161 (cmrkt8m9g053umj01vgw9zlyw), direct
mHC 10.09033 (cmrkts5bg0578mj018fmka2ue), no-copy/no-MTP 10.11537
(cmrku8z0l05ahmj01raa1794f), context-bounded sparse 14.03864
(cmrkuloel05almj015vfgot4o), and C4 full-selection bypass 14.37840
(cmrkuv0is05dwmj01e6mdztv7). Consult the queue JSON for exact identities;
do not compare rows whose native-mHC, context, graph, or prompt-detail fields
differ.
Date: 2026-06-27
Model: unsloth/gemma-4-26B-A4B-it-GGUF, Gemma 4 26B A4B Q8 lane.
Status: active Gemma 4 Q8 B70 optimization. As of 2026-06-27, Gemma rows in this ledger that were submitted from synthetic, repeated, or filled-long benchmarks are classified as diagnostic / pre-final-gate unless they link a passing fixed realistic prompt-suite result. Do not cite them as representative real-world decode throughput. Keep single-replica records separate from four independent replica aggregate capacity, and keep natural-stop, short-prompt sustained, and filled-long sustained shapes separate.
Hardware note for the Gemma 4 26B submissions: these were run on a headless
Supermicro AMD Threadripper PRO 5955WX platform with 128 GB DDR4 and Intel Arc
Pro B70 32 GB GPUs. The cmqwkedg303jeqr013z753j62 submission used one B70
replica on GPU0, but is now diagnostic/pre-final-gate until revalidated; the
host has four B70s available for parallel single-replica experiments.
Required current submission gate: fixed realistic prompt suite, one cold
response per prompt, cached_tokens=0 every row, no prompt/KV/context/response
reuse or n-gram/history acceleration, target model/quantization unchanged,
verified speculation only, and primary metric
median_tok_s_1_100_after_ttft.
The Gemma table below is a historical submission ledger. Rows that do not link
a passing fixed-suite realistic-gate artifact are diagnostic / pre-final-gate
only, even if their original labels used fresh or the first measured row had
cached_tokens=0. A single synthetic/filled-long row0 is not enough for a
current headline claim or a new LocalMaxxing submission.
Date: 2026-07-13
Model: Qwen/Qwen3.6-27B, GGUF Q4_0 target with native Q8_0 DFlash draft,
llama.cpp/SYCL on one Intel Arc Pro B70.
| Label | LocalMaxxing ID | GPUs | Input | Output | tok/s out | tok/s total | Validation |
|---|---|---|---|---|---|---|---|
qwen36-27b-q4_0-b70-llamacpp-xe2m6-q6top1-dflash5-realistic-47tok-20260713 |
cmrjbx8bc02g8mj01yzz2v701 |
1 | suite median 69 | 128 | 47.819 median 1-100 after TTFT | 32.923 median wall full128 | policy-compliant fused draft-head record: fixed realistic suite, 12 unique cold prompts, cached_tokens=0 throughout, target-verified native DFlash5. The exact Q6_K x Q8_1 M=6 draft LM-head plus raw-greedy top-1 boundary removes full draft-logit materialization/readback while retaining ordinary-logit rollback on compact-read failure. A matching AOT control reproduced 44.221 tok/s; the candidate confirmed at 47.819 tok/s, an 8.14% end-to-end fusion gain and 8.05% over the prior public record. Strict p10 39.870, mean 46.639, TTFT 1156.511 ms; a first independent strict candidate passed at 47.114. Evidence data/qwen36-27b-mtp-gguf-q4-b70-baselines/q6top1-aot-realistic128-r2-20260713.json, queue experiments/qwen27-dflash-sycl-b70/localmaxxing/qwen36-27b-q4_0-xe2-m6-q6top1-dflash5-realistic-47tok-20260713.queue.json, approved response data/localmaxxing-responses/qwen36-27b-q4_0-xe2-m6-q6top1-aot-20260713.submit.log. |
qwen36-27b-q4_0-b70-llamacpp-xe2m6-gateup-down-gdncache-dflash5-realistic-44tok-20260713 |
cmrj8s2sy02a4mj01f18hanvc |
1 | suite median 69 | 128 | 44.255 median 1-100 after TTFT | 31.762 median wall full128 | policy-compliant compositional fusion record: fixed realistic suite, 12 unique cold prompts, cached_tokens=0 throughout, target-verified native DFlash5. This BMG-AOT row stacks the exact GDN snapshot-cache commit fusion onto the 187-projection Xe2 M6 path, removing the recurrent-state copy tail while retaining joint gate/up and canonical-metadata down DPAS. It improves the matching 42.641 record by 3.79%; strict p10 38.147, mean 44.348, TTFT 1155.477 ms. Evidence data/qwen36-27b-mtp-gguf-q4-b70-baselines/xe2-m6-full187-joint-gdncache-aot-realistic128-20260713T130908Z.json, queue experiments/qwen27-dflash-sycl-b70/localmaxxing/qwen36-27b-q4_0-xe2-m6-gateup-down-gdncache-dflash5-20260713.queue.json, approved response data/localmaxxing-responses/qwen36-27b-q4_0-xe2-m6-gateup-down-gdncache-aot-20260713.submit.log. |
qwen36-27b-q4_0-b70-llamacpp-xe2m6-gateup-down-dflash5-realistic-42tok-20260713 |
cmrj8fygq029ymj01e2404psy |
1 | suite median 69 | 128 | 42.641 median 1-100 after TTFT | 30.166 median wall full128 | policy-compliant joint gate/up plus canonical-down Xe2 verifier record: fixed realistic suite, 12 unique cold prompts, cached_tokens=0 throughout, target-verified native DFlash5. BMG-AOT llama.cpp 9976 (e3546c794), graph off, Q8_0 target KV and F16 draft KV. The guarded pack set contains 130 Q4_0 gate/up plus 57 Q4_0 down tensors; same-layer gate/up shares one quantization and ESIMD submission, while down consumes canonical Q8_1 metadata to remove the earlier sum error. Real down shadow max error 1.01e-7; strict p10 37.012, mean 42.957, TTFT 1162.638 ms. The supporting JIT row reached 45.484. Evidence data/qwen36-27b-mtp-gguf-q4-b70-baselines/xe2-m6-hybridquant-full187-joint-aot-realistic128-20260713T1305Z.json, queue experiments/qwen27-dflash-sycl-b70/localmaxxing/qwen36-27b-q4_0-xe2-m6-gateup-down-dflash5-20260713.queue.json, approved response data/localmaxxing-responses/qwen36-27b-q4_0-xe2-m6-gateup-down-aot-20260713.submit.log. |
qwen36-27b-q4_0-b70-llamacpp-xe2m6-dflash5-realistic-39tok-20260713 |
cmriq995z0210mj01fl13xmuc |
1 | suite median 69 | 128 | 39.249 median 1-100 after TTFT | 28.697 median wall full128 | policy-compliant realistic suite and first integrated Xe2 DPAS verifier win: fixed qwen36-27b-autoround-int4-b70-realistic-v1, 12 unique prompts each once cold, cached_tokens=0 every row, no prompt/KV/context/response/history reuse, native DFlash5 accepted tokens verified by the unchanged Q4_0 target. BMG-AOT llama.cpp 9976 (e3546c794), graph off, Q8_0 target KV, F16 draft KV, all 130 Q4_0 gate/up weights offline-packed into the signed-s4 N16/K32 layout, M=6 INT4xINT8 DPAS with one-workgroup SLM reduction. Real AOT shadow oracle max error 0.000363; strict p10 33.790, mean 39.726, TTFT 1168.469 ms. A separate JIT support row measured 40.338; the conservative AOT result is submitted. Evidence data/qwen36-27b-mtp-gguf-q4-b70-baselines/xe2-m6-full130-aot-native-dflash5-realistic128-20260713T043137Z.json, queue experiments/qwen27-dflash-sycl-b70/localmaxxing/qwen36-27b-q4_0-xe2-m6-dflash5-20260713.queue.json, approved response data/localmaxxing-responses/qwen36-27b-q4_0-xe2-m6-dflash5-aot-20260713.submit.log. |
Date: 2026-07-11
Model: webhie/Qwen3.6-27B-int4-AutoRound, AutoRound INT4 W4A16 plus runtime
INT8 target LM-head BF16 scales and runtime INT4 draft LM-head BF16 scales,
vLLM/XPU TP2 on two Intel Arc Pro B70 GPUs.
| Label | LocalMaxxing ID | GPUs | Input | Output | tok/s out | tok/s total | Validation |
|---|---|---|---|---|---|---|---|
qwen36-27b-webhie-int4-autoround-b70-vllm-tp2-public-oneccl-mtp3-draftgraph-82tok-20260711 |
cmrgjjw8n004qmj01cp91qxl0 |
2 | suite median 69 | 512 | 82.894 median 1-100 after TTFT | 73.290 median wall full output | policy-compliant realistic suite and graph-boundary win: fixed qwen36-27b-autoround-int4-b70-realistic-v1, 12 unique prompts each once cold, cached_tokens=0 every row, no prompt/KV/context/response/history reuse, target-verified MTP3. Pinned public oneCCL fixes target all-reduce graph replay; a default-off opaque compiled all-gather boundary avoids Inductor’s XPU-graph-incompatible functional wait_tensor and enables exact intrinsic-MTP draft graph capture. Conservative p10 72.752, mean 83.101, TTFT 748.908 ms; integrated full-quality high 85.394 passed exact cases, repeat128, baseline parity, and the 1K needle. A swapped four-GPU crossover measured graph 81.580/79.637 versus eager 75.664/77.308, +5.39% on average. Result packet results/qwen36-27b-autoround-int4-b70/tp2-public-oneccl-draftgraph-20260711.json, queue experiments/qwen36-27b-autoround-int4-b70/localmaxxing/qwen36-27b-webhie-int4-tp2-draftgraph-20260711.queue.json, approved response data/localmaxxing-responses/qwen36-27b-webhie-int4-tp2-draftgraph-20260711.submit.log. |
qwen36-27b-webhie-int4-autoround-b70-vllm-tp2-public-oneccl-mtp3-cg8-78tok-20260711 |
cmrghhs27004cmj01dijk9r9f |
2 | suite median 69 | 512 | 78.226 median 1-100 after TTFT | 69.879 median wall full output | policy-compliant realistic suite and new mechanism: fixed qwen36-27b-autoround-int4-b70-realistic-v1, 12 unique prompts each once cold, cached_tokens=0 every row, no prompt/KV/context/response/history reuse, target-verified MTP3. Pinned public oneCCL parent b52f40c / libccl 4ceafd1 fixes the installed runtime’s packed-verifier XPUGraph collective corruption; direct 256/256 and graph 512/512 oracles passed on both ranks. Conservative isolated p10 69.963, mean 78.598, TTFT 750.031 ms; separate full-quality high 81.341 passed exact cases, repeat128, baseline parity, and the 1K needle. The 3.98% run delta is inside the established 4.4% variance band, so 78.226 is the submitted headline. Result packet results/qwen36-27b-autoround-int4-b70/tp2-public-oneccl-4ceafd1-20260711.json, queue experiments/qwen36-27b-autoround-int4-b70/localmaxxing/qwen36-27b-webhie-int4-tp2-public-oneccl-20260711.queue.json, approved response data/localmaxxing-responses/qwen36-27b-webhie-int4-tp2-public-oneccl-20260711.submit.log. |
Date: 2026-07-03
Model: Intel/Qwen3.6-27B-int4-AutoRound, AutoRound INT4 W4A16, vLLM/XPU on
one Intel Arc Pro B70.
| Label | LocalMaxxing ID | GPUs | Input | Output | tok/s out | tok/s total | Validation |
|---|---|---|---|---|---|---|---|
qwen36-27b-webhie-int4-autoround-b70-vllm-realistic-int8lmhead-bf16scale-draftint4-replayssm-mtp3-cg8-68tok-current-confirm-20260706 |
cmr9atqb800msqr01u760xh0t |
1 | 69 | 512 | 68.236 median 1-100 after TTFT | 61.551 median wall full output | policy-compliant realistic suite, current Qwen27 best measured valid row: fixed qwen36-27b-autoround-int4-b70-realistic-v1, each prompt once, cached_tokens=0 every row, no prompt/KV/context/response reuse, no n-gram/history acceleration, qwen3_next_mtp accepted tokens verified by the target model. Same recipe as the approved 67.519 ReplaySSM target-INT8/draft-INT4 row, with repeat64 quality passed and matched baseline. Treat as a small variance-sensitive current confirm, not a new mechanism. Primary p10 62.317, mean 67.830, TTFT median 479.146 ms. Result packet results/qwen36-27b-autoround-int4-b70/webhie-int8lmhead-bf16scale-draftint4-replayssm-current-confirm-20260706.json, queue experiments/qwen36-27b-autoround-int4-b70/localmaxxing/qwen36-27b-webhie-int4-int8lmhead-bf16scale-draftint4-replayssm-current-confirm-20260706.queue.json, approved response data/localmaxxing-responses/qwen36-27b-webhie-int4-int8lmhead-bf16scale-draftint4-replayssm-current-confirm-20260706.submit.log. |
qwen36-27b-webhie-int4-autoround-b70-vllm-realistic-int8lmhead-bf16scale-draftint4-replayssm-mtp3-cg8-67tok-20260706 |
cmr8rg5d900glqr01g4fesy6i |
1 | 69 | 512 | 67.519 median 1-100 after TTFT | 61.272 median wall full output | policy-compliant realistic suite, superseded Qwen27 row: fixed qwen36-27b-autoround-int4-b70-realistic-v1, each prompt once, cached_tokens=0 every row, no prompt/KV/context/response reuse, no n-gram/history acceleration, qwen3_next_mtp accepted tokens verified by the target model. Config builds on the webhie BF16-scale INT8 target LM-head lane with ReplaySSM exact GDN state handling, commit-in-forward, and runtime INT4 draft LM-head BF16 scales. Headline uses the conservative solo confirmation (67.519) rather than the one-off 68.481 high; same-window native slot-copy vs PyTorch slot-management controls showed the new native slot-copy op was not the speed source (66.871 native vs 67.300 fallback). Repeat64 quality passed and matched baseline. Primary p10 62.663, mean 68.154, TTFT median 477.851 ms. Result packet results/qwen36-27b-autoround-int4-b70/webhie-int8lmhead-bf16scale-draftint4-replayssm-20260706.json, queue experiments/qwen36-27b-autoround-int4-b70/localmaxxing/qwen36-27b-webhie-int4-int8lmhead-bf16scale-draftint4-replayssm-20260706.queue.json, approved response data/localmaxxing-responses/qwen36-27b-webhie-int4-int8lmhead-bf16scale-draftint4-replayssm-20260706.submit2.log. First POST failed only because top-level quantization exceeded the API’s length limit; queue was shortened while retaining full details in notes/engineFlags. |
qwen36-27b-webhie-int4-autoround-b70-vllm-realistic-int8lmhead-bf16scale-mtp3-cg8-65tok-20260703 |
cmr5iu3gk00bfq901nidgcana |
1 | 69 | 128 | 65.276 median 1-100 after TTFT | 49.172 median wall full128 | policy-compliant realistic suite, superseded Qwen27 record: fixed qwen36-27b-autoround-int4-b70-realistic-v1, each prompt once, cached_tokens=0 every row, no prompt/KV/context/response reuse, no n-gram/history acceleration, qwen3_next_mtp accepted tokens verified by the target model. Config matches the promoted webhie MTP3/cg8 runtime INT8 LM-head lane plus BF16 scale storage (VLLM_XPU_LM_HEAD_INT8_BF16_SCALES=1), preserving INT8 LM-head weights while reducing scale bandwidth/format cost. Full quality gate passed and matched the prior webhie INT8-LM-head baseline, including 32-repeat stability and 1K long-context needle with cached tokens zero. Strict fresh BF16-scale support rows: 65.005, 64.864; same-window FP32-scale controls: 64.234, 64.090; baseline reconfirm: 64.431; primary p10 59.609, mean 65.077, TTFT median 603.580 ms. Result packet results/qwen36-27b-autoround-int4-b70/webhie-int8-lmhead-bf16scale-20260703.json, queue experiments/qwen36-27b-autoround-int4-b70/localmaxxing/qwen36-27b-webhie-int4-int8lmhead-bf16scale-20260703.queue.json, approved response data/localmaxxing-responses/qwen36-27b-webhie-int4-int8lmhead-bf16scale-20260703.submit.log. |
qwen36-27b-webhie-int4-autoround-b70-vllm-realistic-int8lmhead-mtp3-cg8-64tok-20260703 |
cmr576apv0079q901i6dvsh0l |
1 | 69 | 128 | 64.306 median 1-100 after TTFT | 48.194 median wall full128 | policy-compliant realistic suite, webhie AutoRound variant: fixed qwen36-27b-autoround-int4-b70-realistic-v1, each prompt once, cached_tokens=0 every row, no prompt/KV/context/response reuse, no n-gram/history acceleration, qwen3_next_mtp accepted tokens verified by the target model. Config matches the promoted MTP3/cg8 lane plus VLLM_XPU_LM_HEAD_INT8=1; label as webhie/Qwen3.6-27B-int4-AutoRound + runtime INT8 LM-head, separate from the Intel checkpoint row. Full quality gate passed and matched the prior Intel INT8-LM-head baseline, including 32-repeat stability and 1K long-context needle. Initial webhie support row: 63.336; same-window Intel INT8-LM-head control: 62.366; primary p10 59.496, mean 63.615, TTFT median 605.938 ms. Result packet results/qwen36-27b-autoround-int4-b70/webhie-int8-lmhead-20260703.json, queue experiments/qwen36-27b-autoround-int4-b70/localmaxxing/qwen36-27b-webhie-int4-int8lmhead-20260703.queue.json, approved response data/localmaxxing-responses/qwen36-27b-webhie-int4-int8lmhead-20260703.submit.log. |
qwen36-27b-int4-autoround-b70-vllm-realistic-int8lmhead-mtp3-cg8-62tok-20260703 |
cmr4zkcxb003yq9018408i1pn |
1 | 69 | 128 | 62.628 median 1-100 after TTFT | 47.656 median wall full128 | policy-compliant realistic suite, separate runtime-quantized variant: fixed qwen36-27b-autoround-int4-b70-realistic-v1, each prompt once, cached_tokens=0 every row, no prompt/KV/context/response reuse, no n-gram/history acceleration, target model unchanged, qwen3_next_mtp accepted tokens verified by the target model. Config matches the promoted MTP3/cg8 lane plus VLLM_XPU_LM_HEAD_INT8=1, which keeps the BF16 lm_head resident but uses a default-off transient per-output-channel INT8 LM-head projection; label as AutoRound INT4 W4A16 + runtime INT8 LM-head, not the original BF16-LM-head AutoRound identity. Full quality gate passed and matched baseline, including 32-repeat stability and 1K long-context needle. Same-window BF16-LM-head control: 53.332; INT8 repeat: 62.276; primary run p10 58.104, mean 62.998, TTFT median 606.575 ms. Result packet results/qwen36-27b-autoround-int4-b70/int8-lmhead-20260703.json, patch patches/qwen36-27b-autoround-int4-b70/vllm-xpu-lm-head-int8-quality-pass-20260703.patch, queue experiments/qwen36-27b-autoround-int4-b70/localmaxxing/qwen36-27b-int4-int8lmhead-20260703.queue.json, approved response data/localmaxxing-responses/qwen36-27b-int4-int8lmhead-20260703.submit.log. |
qwen36-27b-int4-autoround-b70-vllm-realistic-promotesource-mtp3-cg8-53tok-20260703 |
cmr4gokx90061nv01lhoe3ft8 |
1 | 69 | 128 | 53.522 median 1-100 after TTFT | 42.545 median wall full128 | policy-compliant realistic suite: fixed qwen36-27b-autoround-int4-b70-realistic-v1, each prompt once, cached_tokens=0 every row, no prompt/KV/context/response reuse, no n-gram/history acceleration, target model/quant unchanged, qwen3_next_mtp accepted tokens verified by the target model. Config: TP1, XPU graph on, num_speculative_tokens=3, max_cudagraph_capture_size=8, MAX_NUM_BATCHED_TOKENS=1024, VLLM_XPU_GDN_PROMOTE_ACCEPTED_SPEC_STATE=1, VLLM_XPU_GDN_NONSPEC_POSTPROCESS_ACCEPTED_STATE=0. Quality suite passed and matched baseline. Support rows: 54.861 and 53.992; same-window plain-MTP3/cg8 control: 48.345. Result packet results/qwen36-27b-autoround-int4-b70/promote-source-noacceptedpost-20260703.md, queue experiments/qwen36-27b-autoround-int4-b70/localmaxxing/qwen36-27b-int4-promotesource-20260703.queue.json, approved response data/localmaxxing-responses/qwen36-27b-int4-promotesource-20260703.submit2.log. First POST failed only because top-level promptTokens was 68.5; queue was corrected to integer 69 while preserving per-prompt token counts in engineFlags. |
Date: 2026-07-04
Model: unsloth/Qwen3-30B-A3B-Instruct-2507-GGUF, GGUF
UD-Q4_K_XL, llama.cpp/SYCL on one Intel Arc Pro B70.
| Label | LocalMaxxing ID | GPUs | Input | Output | tok/s out | tok/s total | Validation |
|---|---|---|---|---|---|---|---|
qwen3-30b-a3b-instruct-2507-udq4-llamacpp-realistic128 |
cmr6rr2kv008imn019frg0x3m |
1 | suite median 65 | 128 | 107.484 median 1-100 after TTFT | 94.118 median wall full128 | policy-compliant rapid realistic suite: fixed rapid-model-snapshots-b70-realistic-v1, 12 unique prompts, each prompt once, llama.cpp cache_prompt=false, cached_tokens=0 every row, no prompt/KV/context checkpoint/response reuse, no n-gram/history acceleration, no speculation. Config: Qwen3-30B-A3B-Instruct-2507-UD-Q4_K_XL.gguf, HF revision eea7b2be5805a5f151f8847ede8e5f9a9284bf77, llama.cpp/SYCL on one B70, ctx=4096, batch=1024, ubatch=256, FlashAttention on, f16 KV. Primary p10 106.898, mean 104.993, full-output after-TTFT median 107.421, TTFT median 166.953 ms. Result packet results/rapid-model-snapshots-b70/qwen3-30b-a3b-instruct-2507-udq4/README.md, evidence data/rapid-model-snapshots-b70/qwen3-30b-a3b-instruct-2507-udq4-llamacpp-faon-nocacheprompt-realistic128-20260704T193409Z.json, queue experiments/rapid-model-snapshots-b70/localmaxxing/qwen3-30b-a3b-instruct-2507-udq4-llamacpp-realistic128-20260704.queue.json, approved response data/localmaxxing-responses/qwen3-30b-a3b-instruct-2507-udq4-llamacpp-realistic128-20260704.submit.log. A quick four-GPU runtime sweep found no reproducible sub-percent knob win; this is a first-pass rapid snapshot, not a deep per-model optimization lane. |
Date: 2026-07-04
Model: unsloth/Qwen3-Coder-30B-A3B-Instruct-GGUF, GGUF
UD-Q4_K_XL, llama.cpp/SYCL on one Intel Arc Pro B70.
| Label | LocalMaxxing ID | GPUs | Input | Output | tok/s out | tok/s total | Validation |
|---|---|---|---|---|---|---|---|
qwen3-coder-30b-a3b-instruct-udq4-llamacpp-realistic128 |
cmr6w2ekt00gimn01orbith22 |
1 | suite median 65 | 128 | 108.117 median 1-100 after TTFT | 94.754 median wall full128 | policy-compliant rapid realistic suite: fixed rapid-model-snapshots-b70-realistic-v1, 12 unique prompts, each prompt once, llama.cpp server prompt cache disabled with --cache-ram 0, per-request cache_prompt=false, cached_tokens=0 every row, no prompt/KV/context checkpoint/response reuse, no n-gram/history acceleration, no speculation. Config: Qwen3-Coder-30B-A3B-Instruct-UD-Q4_K_XL.gguf, HF revision b17cb02dd882d5b6ab62fc777ad2995f19668350, llama.cpp/SYCL on one B70, ctx=4096, batch=1024, ubatch=256, POLL=100, FlashAttention on, f16 KV. Primary p10 106.573, mean 105.328, full-output after-TTFT median 107.897, TTFT median 164.129 ms. Result packet results/rapid-model-snapshots-b70/qwen3-coder-30b-a3b-instruct-udq4/README.md, evidence data/rapid-model-snapshots-b70/qwen3-coder-30b-a3b-instruct-udq4-llamacpp-faon-cacheoff-poll100-confirm-ctx4096-realistic128-20260704T214053Z.json, queue experiments/rapid-model-snapshots-b70/localmaxxing/qwen3-coder-30b-a3b-instruct-udq4-llamacpp-realistic128-20260704.queue.json, approved response data/localmaxxing-responses/qwen3-coder-30b-a3b-instruct-udq4-llamacpp-realistic128-20260704.submit.log. Same-window quick screen found only sub-percent movement, so treat this as a useful first-pass model snapshot, not a deep optimization lane. |
Date: 2026-07-04
Model: bartowski/DeepSeek-Coder-V2-Lite-Instruct-GGUF, GGUF Q4_K_M,
llama.cpp/SYCL on one Intel Arc Pro B70.
| Label | LocalMaxxing ID | GPUs | Input | Output | tok/s out | tok/s total | Validation |
|---|---|---|---|---|---|---|---|
deepseek-coder-v2-lite-q4km-llamacpp-realistic128 |
cmr6zbkbw00hpmn01nq858vcg |
1 | suite median 64 | 128 | 57.097 median 1-100 after TTFT | 53.414 median wall full128 | policy-compliant rapid realistic suite, coder-model reference: fixed rapid-model-snapshots-b70-realistic-v1, 12 unique prompts, each prompt once, llama.cpp server prompt cache disabled with --cache-ram 0, per-request cache_prompt=false, cached_tokens=0 every row, no prompt/KV/context checkpoint/response reuse, no n-gram/history acceleration, no speculation. Config: DeepSeek-Coder-V2-Lite-Instruct-Q4_K_M.gguf, HF revision 8f248fa2072348f77a8bc37754e470de1f61866e, llama.cpp/SYCL on one B70, ctx=2048, batch=1024, ubatch=256, FlashAttention on, f16 KV. Primary p10 56.932, mean 57.083, full-output after-TTFT median 56.711, TTFT median 139.827 ms; same-recipe support row 57.212 tok/s; ctx=4096 baseline 56.033 tok/s. Result packet results/rapid-model-snapshots-b70/deepseek-coder-v2-lite-q4km/README.md, evidence data/rapid-model-snapshots-b70/deepseek-coder-v2-lite-q4km-llamacpp-faon-cacheoff-ctx2048-confirm-realistic128-20260704T231049Z.json, queue experiments/rapid-model-snapshots-b70/localmaxxing/deepseek-coder-v2-lite-q4km-llamacpp-realistic128-20260704.queue.json, approved response data/localmaxxing-responses/deepseek-coder-v2-lite-q4km-llamacpp-realistic128-20260704.submit.log. Concurrent four-GPU screen rows underreported and were not used as headline. |
Date: 2026-07-04
Model: bartowski/nvidia_Nemotron-Cascade-2-30B-A3B-GGUF, GGUF Q4_K_M,
llama.cpp/SYCL on one Intel Arc Pro B70.
| Label | LocalMaxxing ID | GPUs | Input | Output | tok/s out | tok/s total | Validation |
|---|---|---|---|---|---|---|---|
nemotron-cascade-2-30b-a3b-q4km-llamacpp-realistic128 |
cmr7128uq00jdmn01dn0uttm7 |
1 | suite median 65 | 128 | 50.904 median 1-100 after TTFT | 43.119 median wall full128 | policy-compliant rapid realistic suite, Nemotron-family reference: fixed rapid-model-snapshots-b70-realistic-v1, 12 unique prompts, each prompt once, llama.cpp server prompt cache disabled with --cache-ram 0, per-request cache_prompt=false, cached_tokens=0 every row, no prompt/KV/context checkpoint/response reuse, no n-gram/history acceleration, no speculation. Config: nvidia_Nemotron-Cascade-2-30B-A3B-Q4_K_M.gguf, HF revision 931b595fc71b7ca14fb9d935af011f69f7c0434c, llama.cpp/SYCL on one B70, ctx=2048, batch=1024, ubatch=256, poll=50, FlashAttention on, f16 KV, --jinja, --reasoning off. Primary p10 50.877, mean 50.896, full-output after-TTFT median 50.789, TTFT median 449.159 ms; output preview check showed normal prose, zero reasoning deltas, and no visible <think> leakage. Quick ctx/batch/ubatch/poll/thread probes all landed around 50.7-50.9 tok/s, so this is an expected-performance snapshot, not a frontier optimization lane. Result packet results/rapid-model-snapshots-b70/nemotron-cascade-2-30b-a3b-q4km/README.md, evidence data/rapid-model-snapshots-b70/nemotron-cascade-2-30b-a3b-q4km-llamacpp-faon-cacheoff-ctx2048-confirm-realistic128-20260704T235714Z.json, queue experiments/rapid-model-snapshots-b70/localmaxxing/nemotron-cascade-2-30b-a3b-q4km-llamacpp-realistic128-20260704.queue.json, approved response data/localmaxxing-responses/nemotron-cascade-2-30b-a3b-q4km-llamacpp-realistic128-20260704.submit2.log. First POST failed only because the payload used engineName=llama.cpp-sycl; the accepted queue uses LocalMaxxing’s llama.cpp enum and keeps SYCL details in notes/runtime metadata. |
Date: 2026-07-04
Model: unsloth/GLM-4.7-Flash-GGUF, GGUF UD-Q4_K_XL,
llama.cpp/SYCL on one Intel Arc Pro B70.
| Label | LocalMaxxing ID | GPUs | Input | Output | tok/s out | tok/s total | Validation |
|---|---|---|---|---|---|---|---|
glm-4.7-flash-udq4-llamacpp-realistic128 |
cmr6xkr2f00gomn01k4u2dua8 |
1 | suite median 62 | 128 | 40.769 median 1-100 after TTFT | 38.165 median wall full128 | policy-compliant rapid realistic suite, valid/modest GLM-4.7-Flash reference: fixed rapid-model-snapshots-b70-realistic-v1, 12 unique prompts, each prompt once, llama.cpp server prompt cache disabled with --cache-ram 0, per-request cache_prompt=false, cached_tokens=0 every row, no prompt/KV/context checkpoint/response reuse, no n-gram/history acceleration, no speculation. Config: GLM-4.7-Flash-UD-Q4_K_XL.gguf, HF revision 0d32489ecb9db6d2a4fc93bd27ef01519f95474d, llama.cpp/SYCL on one B70, ctx=4096, batch=1024, ubatch=256, POLL=100, FlashAttention on, f16 KV. Primary p10 40.019, mean 40.261, full-output after-TTFT median 40.678, TTFT median 206.206 ms. Result packet results/rapid-model-snapshots-b70/glm-4.7-flash-udq4/README.md, evidence data/rapid-model-snapshots-b70/glm-4.7-flash-udq4-llamacpp-faon-cacheoff-poll100-confirm-ctx4096-realistic128-20260704T221455Z.json, queue experiments/rapid-model-snapshots-b70/localmaxxing/glm-4.7-flash-udq4-llamacpp-realistic128-20260704.queue.json, approved response data/localmaxxing-responses/glm-4.7-flash-udq4-llamacpp-realistic128-20260704.submit.log. Faster ~44 tok/s rows appeared only in concurrent four-GPU screens, so the promoted row uses the conservative standalone confirmation. |
Date: 2026-07-04
Model: bartowski/microsoft_Phi-4-mini-instruct-GGUF, GGUF Q4_K_M and Q8_0,
llama.cpp/SYCL on one Intel Arc Pro B70.
| Label | LocalMaxxing ID | GPUs | Input | Output | tok/s out | tok/s total | Validation |
|---|---|---|---|---|---|---|---|
phi4-mini-instruct-q4km-llamacpp-realistic128 |
cmr6yazhe00hcmn01i5gz2xe0 |
1 | suite median 60 | 128 | 96.548 median 1-100 after TTFT | 91.750 median wall full128 | policy-compliant rapid realistic suite, Q4_K_M compact reference: fixed rapid-model-snapshots-b70-realistic-v1, 12 unique prompts, each prompt once, llama.cpp server prompt cache disabled with --cache-ram 0, per-request cache_prompt=false, cached_tokens=0 every row, no prompt/KV/context checkpoint/response reuse, no n-gram/history acceleration, no speculation. Config: microsoft_Phi-4-mini-instruct-Q4_K_M.gguf, HF revision 7ff82c2aaa4dde30121698a973765f39be5288c0, llama.cpp/SYCL on one B70, ctx=4096, batch=1024, ubatch=256, FlashAttention on, f16 KV. Primary p10 96.351, mean 97.461, full-output after-TTFT median 96.580, TTFT median 69.937 ms; same-recipe standalone repeat 96.574 tok/s. Result packet results/rapid-model-snapshots-b70/phi4-mini-instruct-gguf/README.md, evidence data/rapid-model-snapshots-b70/phi4-mini-instruct-q4km-llamacpp-faon-cacheoff-confirm-ctx4096-realistic128-20260704T224303Z.json, queue experiments/rapid-model-snapshots-b70/localmaxxing/phi4-mini-instruct-q4km-llamacpp-realistic128-20260704.queue.json, approved response data/localmaxxing-responses/phi4-mini-instruct-q4km-llamacpp-realistic128-20260704.submit.log. Concurrent four-GPU screen rows were lower and are support-only. |
phi4-mini-instruct-q8-llamacpp-realistic128 |
cmr6yazvy00hgmn01s5rtowwa |
1 | suite median 60 | 128 | 72.246 median 1-100 after TTFT | 67.592 median wall full128 | policy-compliant rapid realistic suite, Q8_0 compact higher-quality reference: same strict suite and no-cache policy as the Q4 row, no speculation or history acceleration. Config: microsoft_Phi-4-mini-instruct-Q8_0.gguf, HF revision 7ff82c2aaa4dde30121698a973765f39be5288c0, llama.cpp/SYCL on one B70, ctx=4096, batch=1024, ubatch=256, FlashAttention on, f16 KV. Primary p10 71.908, mean 72.588, full-output after-TTFT median 72.159, TTFT median 119.457 ms; same-recipe support row 72.884 tok/s. Result packet results/rapid-model-snapshots-b70/phi4-mini-instruct-gguf/README.md, evidence data/rapid-model-snapshots-b70/phi4-mini-instruct-q8-llamacpp-faon-cacheoff-confirm2-ctx4096-realistic128-20260704T224430Z.json, queue experiments/rapid-model-snapshots-b70/localmaxxing/phi4-mini-instruct-q8-llamacpp-realistic128-20260704.queue.json, approved response data/localmaxxing-responses/phi4-mini-instruct-q8-llamacpp-realistic128-20260704.submit.log. |
Date: 2026-07-04
Model: unsloth/Mistral-Small-3.2-24B-Instruct-2506-GGUF, GGUF
UD-Q4_K_XL, llama.cpp/SYCL on one Intel Arc Pro B70.
| Label | LocalMaxxing ID | GPUs | Input | Output | tok/s out | tok/s total | Validation |
|---|---|---|---|---|---|---|---|
mistral-small-3.2-24b-instruct-2506-udq4-llamacpp-realistic128 |
cmr6ura7300e4mn01yrdw7wto |
1 | suite median 616 | 128 | 27.297 median 1-100 after TTFT | 20.634 median wall full128 | policy-compliant rapid realistic suite, valid/modest dense-model reference: fixed rapid-model-snapshots-b70-realistic-v1, 12 unique prompts, each prompt once, llama.cpp server prompt cache disabled with --cache-ram 0, per-request cache_prompt=false, cached_tokens=0 every row, no prompt/KV/context checkpoint/response reuse, no n-gram/history acceleration, no speculation. Config: Mistral-Small-3.2-24B-Instruct-2506-UD-Q4_K_XL.gguf, HF revision b750ec2299225e492f1bd27cab88a0a595fa848f, llama.cpp/SYCL on one B70, ctx=4096, batch=1024, ubatch=256, FlashAttention on, f16 KV. Primary p10 27.126, mean 27.356, full-output after-TTFT median 27.224, TTFT median 1501.774 ms. Result packet results/rapid-model-snapshots-b70/mistral-small-3.2-24b-instruct-2506-udq4/README.md, evidence data/rapid-model-snapshots-b70/mistral-small-3.2-24b-instruct-2506-udq4-llamacpp-faon-cacheoff-v2-ctx4096-realistic128-20260704T205443Z.json, queue experiments/rapid-model-snapshots-b70/localmaxxing/mistral-small-3.2-24b-instruct-2506-udq4-llamacpp-realistic128-20260704.queue.json, approved response data/localmaxxing-responses/mistral-small-3.2-24b-instruct-2506-udq4-llamacpp-realistic128-20260704.submit.log. Q8 fit check passed but was only 16.380 tok/s; quick Q4 knob screen found no easy win, so this is a useful expected-performance snapshot rather than a frontier lane. |
Date: 2026-07-04
Model: unsloth/Qwen3.6-27B-MTP-GGUF, GGUF UD-Q4_K_XL, llama.cpp/SYCL on
one Intel Arc Pro B70.
| Label | LocalMaxxing ID | GPUs | Input | Output | tok/s out | tok/s total | Validation |
|---|---|---|---|---|---|---|---|
qwen36-27b-mtp-gguf-q4-b70-llamacpp-realistic-mtp3-30tok-20260703 |
cmr6mn5ct0076mn01on3dnpyn |
1 | 69 | 128 | 30.679 median 1-100 after TTFT | 26.581 median wall full128 | policy-compliant realistic suite, non-competitive model/runtime reference: fixed qwen36-27b-autoround-int4-b70-realistic-v1, 12 unique prompts, each prompt once, cached_tokens=0 every row, no prompt/KV/context checkpoint/response reuse, no n-gram/history acceleration, draft-MTP3 accepted tokens verified by the target GGUF model. Config: Qwen3.6-27B-UD-Q4_K_XL.gguf, llama.cpp/SYCL 9860 (fdb1db877), one B70, ctx=4096, batch=1024, ubatch=256, FlashAttention on, f16 KV, n_max=3, n_min=0, p_min=0.00, --ctx-checkpoints 0. Primary p10 27.589, mean 30.405, full-output after-TTFT median 29.860, TTFT median 499.824 ms. Result packet results/qwen36-27b-mtp-gguf-q4-b70/README.md, evidence data/qwen36-27b-mtp-gguf-q4-b70-baselines/llamacpp-mtp3-aot-np1-realistic128-20260703T060748Z.json, queue experiments/qwen36-27b-mtp-gguf-q4-b70/localmaxxing/qwen36-27b-gguf-q4-mtp3-20260703.queue.json, approved response data/localmaxxing-responses/qwen36-27b-gguf-q4-mtp3-20260703.submit.log. This row is useful for expected-performance comparison but does not displace the faster AutoRound vLLM Qwen27 record. |
Date: 2026-07-03
Model: unsloth/gemma-4-26B-A4B-it-GGUF, Gemma 4 26B A4B Q8 lane,
llama.cpp/SYCL on one Intel Arc Pro B70.
| Label | LocalMaxxing ID | GPUs | Input | Output | tok/s out | tok/s total | Validation |
|---|---|---|---|---|---|---|---|
gemma4-26b-a4b-q8-b70-llamacpp-service-32k-smoke-20260703 |
cmr47ivql0045nv011pfdjlaa |
1 | 32571 | 76 | 115.179 after TTFT | 979.156 prompt+output wall | approved long-context service / prompt-processing result, not the short-decode headline: one cold near-32K request from lc-24000-late, cached_tokens=0, unique prompt, exact JSON retrieval passed, UD-Q8_K_XL target/verifier with Q4_0 MTP draft verified by target; tokSPrefill=996.600, TTFT 32682 ms; supporting service ladder passed 32/32 long-context rows and 64/64 canary rows across four B70 lanes with average lane median prefill 1192.965 tok/s and long-context decode 131.786 tok/s; payload data/localmaxxing-gemma4-26b-a4b-q8-b70-llamacpp-service-32k-20260703.payload.json, response data/localmaxxing-responses/gemma4-26b-a4b-q8-b70-llamacpp-service-32k-20260703.submit.log |
gemma4-26b-a4b-q8-b70-llamacpp-realistic-finalpostnorm-faon-vmm-ctx32768-full512-124tok-20260701 |
cmr1u77na01k2ld01kalwzs1e |
1 | suite median 69 | 512 | 124.977 median 1-100 after TTFT | 108.581 median wall full512 | policy-compliant realistic suite: exact promoted full512 recipe rerun after the LM-head/Q8 subgroup experiment; fixed gemma4-26b-a4b-q8-b70-realistic-v1, each prompt once, cached_tokens=0 every row, no prompt/KV/context/response reuse, no n-gram/history acceleration, UD-Q8_K_XL target/verifier with Q4_0 MTP draft verified by target; FA-on 32K/VMM VDR2 selected-down baseline plus LLAMA_GEMMA4_FUSED_FINAL_POST_NORM_RESIDUAL=1; reordered-Q8 VDR2, VDR2 selected-down fused weighted-sum, LLAMA_SYCL_F16_P021_SMALL_NCOLS=1, LLAMA_SPEC_VERIFY_BULK_SAMPLED_IDS=1, FLASH_ATTN=on, GGML_SYCL_ENABLE_VMM=1, CTX_SIZE=32768, n_max=3, n_min=2, p_min=0.0475, UBATCH_SIZE=1024, LM-head subgroup unset; p10 103.836, mean 122.474, full512 after-TTFT 114.871, TTFT median 178.694 ms; valid but high-variance, with same exact batch support 121.591, 119.264, 113.633; supersedes cmr01nnet000mld01x2tt6qds |
gemma4-26b-a4b-q8-b70-llamacpp-realistic-finalpostnorm-faon-vmm-ctx32768-full512-123tok-20260630 |
cmr01nnet000mld01x2tt6qds |
1 | suite median 69 | 512 | 123.677 median 1-100 after TTFT | 106.441 median wall full512 | policy-compliant realistic suite: fixed gemma4-26b-a4b-q8-b70-realistic-v1, each prompt once, cached_tokens=0 every row, no prompt/KV/context/response reuse, no n-gram/history acceleration, UD-Q8_K_XL target/verifier with Q4_0 MTP draft verified by target; FA-on 32K/VMM VDR2 selected-down baseline plus LLAMA_GEMMA4_FUSED_FINAL_POST_NORM_RESIDUAL=1; reordered-Q8 VDR2, VDR2 selected-down fused weighted-sum, LLAMA_SYCL_F16_P021_SMALL_NCOLS=1, LLAMA_SPEC_VERIFY_BULK_SAMPLED_IDS=1, FLASH_ATTN=on, GGML_SYCL_ENABLE_VMM=1, CTX_SIZE=32768, n_max=3, n_min=2, p_min=0.0475, UBATCH_SIZE=1024; p10 105.673, mean 120.825, full512 after-TTFT 110.683, TTFT median 179.125 ms; valid but high-variance, with second finalpost 116.551 and controls 117.873 / 114.709; supersedes cmqztiqdn02vnoe01egox6q3f |
gemma4-26b-a4b-q8-b70-llamacpp-realistic-vdr2-selecteddown-faon-vmm-ctx32768-full512-121tok-20260629 |
cmqztiqdn02vnoe01egox6q3f |
1 | suite median 69 | 512 | 121.414 median 1-100 after TTFT | 105.881 median wall full512 | policy-compliant realistic suite: same FA-on 32K/VMM VDR2 selected-down baseline/control identity as the previous 117.914 row; fixed gemma4-26b-a4b-q8-b70-realistic-v1, each prompt once, cached_tokens=0 every row, no prompt/KV/context/response reuse, no n-gram/history acceleration, UD-Q8_K_XL target/verifier with Q4_0 MTP draft verified by target; reordered-Q8 VDR2, VDR2 selected-down fused weighted-sum, LLAMA_SYCL_F16_P021_SMALL_NCOLS=1, LLAMA_SPEC_VERIFY_BULK_SAMPLED_IDS=1, FLASH_ATTN=on, GGML_SYCL_ENABLE_VMM=1, CTX_SIZE=32768, n_max=3, n_min=2, p_min=0.0475, UBATCH_SIZE=1024; LM-head experiment flags unset (LLAMA_SYCL_Q8_0_LM_HEAD_1COL_DMMV=0, LLAMA_SYCL_Q8_0_LM_HEAD_1COL_NO_REORDER=0); p10 107.032, mean 120.136, full512 after-TTFT 110.391, TTFT median 179.118 ms; supporting same-family confirmation 119.948 plus lower variance rows 113.572, 114.088, 111.988; supersedes the FA-on 32K/VMM 117.914 row |
gemma4-26b-a4b-q8-b70-llamacpp-realistic-vdr2-selecteddown-faon-vmm-ctx32768-full512-20260629 |
cmqzq5zu402troe01t774uyox |
1 | suite median 69 | 512 | 117.915 median 1-100 after TTFT | 106.807 median wall full512 | policy-compliant realistic suite, superseded by same-family baseline row: fixed gemma4-26b-a4b-q8-b70-realistic-v1, each prompt once, cached_tokens=0 every row, no prompt/KV/context/response reuse, no n-gram/history acceleration, UD-Q8_K_XL target/verifier with Q4_0 MTP draft verified by target; reordered-Q8 VDR2, VDR2 selected-down fused weighted-sum, FLASH_ATTN=on, GGML_SYCL_ENABLE_VMM=1, CTX_SIZE=32768, LLAMA_SYCL_F16_P021_SMALL_NCOLS=1, LLAMA_SPEC_VERIFY_BULK_SAMPLED_IDS=1, n_max=3, n_min=2, p_min=0.0475, UBATCH_SIZE=1024; p10 107.807, mean 118.881, full512 after-TTFT 110.958, TTFT median 180.169 ms; superseded by cmqztiqdn02vnoe01egox6q3f at 121.414 and then cmr01nnet000mld01x2tt6qds at 123.677 |
gemma4-26b-a4b-q8-b70-llamacpp-realistic-vdr2-selecteddown-reordervdr2-full512-repeat-20260629 |
cmqyrpox4021dqk01co5o4fcw |
1 | suite median 69 | 512 | 115.847 median 1-100 after TTFT | 100.640 median wall full512 | policy-compliant realistic suite: same VDR2 selected-down recipe as the prior 115.728 row; fixed gemma4-26b-a4b-q8-b70-realistic-v1, each prompt once, cached_tokens=0 every row, no prompt/KV/context/response reuse, no n-gram/history acceleration, UD-Q8_K_XL target/verifier with Q4_0 MTP draft verified by target; reordered-Q8 VDR2, VDR2 selected-down fused weighted-sum, LLAMA_SYCL_F16_P021_SMALL_NCOLS=1, LLAMA_SPEC_VERIFY_BULK_SAMPLED_IDS=1, n_max=3, n_min=2, p_min=0.0475, UBATCH_SIZE=1024; p10 102.573, mean 114.574, full512 after-TTFT 104.661, TTFT median 181.167 ms; BF16-direct lanes in the adjacent retest did not beat controls; supersedes the selected-down 115.728 row |
gemma4-26b-a4b-q8-b70-llamacpp-realistic-vdr2-selecteddown-reordervdr2-full512-20260629 |
cmqyo0jyt08ippk01vhiobdnm |
1 | suite median 69 | 512 | 115.728 median 1-100 after TTFT | 100.228 median wall full512 | policy-compliant realistic suite, superseded by same-recipe repeat: fixed gemma4-26b-a4b-q8-b70-realistic-v1, each prompt once, cached_tokens=0 every row, no prompt/KV/context/response reuse, no n-gram/history acceleration, UD-Q8_K_XL target/verifier with Q4_0 MTP draft verified by target; reordered-Q8 VDR2, VDR2 selected-down fused weighted-sum, LLAMA_SYCL_F16_P021_SMALL_NCOLS=1, LLAMA_SPEC_VERIFY_BULK_SAMPLED_IDS=1, n_max=3, n_min=2, p_min=0.0475, UBATCH_SIZE=1024; p10 101.449, mean 113.158, full512 after-TTFT 104.602, TTFT median 181.348 ms; supporting full512 confirmations measured 113.471, 113.815, and 114.811; supersedes the bulk sampled-ID verifier row 98.340 |
gemma4-26b-a4b-q8-b70-llamacpp-realistic-vdr2-mtp-n3-nmin2-p00475-ub1024-f16p021-bulksampled-full512-20260628 |
cmqxchyra03xmqr01b963gmi1 |
1 | suite median 69 | 512 | 98.340 median 1-100 after TTFT | 87.737 median wall full512 | policy-compliant realistic suite: fixed gemma4-26b-a4b-q8-b70-realistic-v1, each prompt once, cached_tokens=0 every row, no prompt/KV/context/response reuse, no n-gram/history acceleration, UD-Q8_K_XL target/verifier with Q4_0 MTP draft verified by target; reordered-Q8 VDR2, LLAMA_SYCL_F16_P021_SMALL_NCOLS=1, LLAMA_SPEC_VERIFY_BULK_SAMPLED_IDS=1, n_max=3, n_min=2, p_min=0.0475, UBATCH_SIZE=1024; p10 85.979, mean 95.953, full512 after-TTFT 91.174, TTFT median 180.211 ms; supporting full512 confirmations measured 96.015, 95.903, and 94.941; supersedes the F16-p021 95.825 row |
gemma4-26b-a4b-q8-b70-llamacpp-realistic-vdr2-mtp-n3-nmin2-p00475-ub1024-f16p021-smallncols-full512-20260628T010121 |
cmqx3687103v4qr01ace1ft3m |
1 | suite median 69 | 512 | 95.825 median 1-100 after TTFT | 88.262 median wall full512 | policy-compliant realistic suite, superseded: fixed gemma4-26b-a4b-q8-b70-realistic-v1, each prompt once, cached_tokens=0 every row, no prompt/KV/context/response reuse, no n-gram/history acceleration, UD-Q8_K_XL target/verifier with Q4_0 MTP draft verified by target; reordered-Q8 VDR2, LLAMA_SYCL_F16_P021_SMALL_NCOLS=1, n_max=3, n_min=2, p_min=0.0475, UBATCH_SIZE=1024; p10 85.504, mean 95.605, full512 after-TTFT 91.142, TTFT median 179.723 ms; supporting full512 confirmations measured 95.817, 93.422, and 95.566; superseded by bulk sampled-ID verifier row 98.340 |
gemma4-26b-a4b-q8-b70-llamacpp-realistic-vdr2-mtp-n3-nmin2-p00475-ub1024-recordconfirm-20260627T221722 |
cmqwxep4a03qiqr010chjn93s |
1 | suite median 69 | 512 | 90.983 median 1-100 after TTFT | 82.897 median wall full512 | policy-compliant realistic suite: fixed gemma4-26b-a4b-q8-b70-realistic-v1, each prompt once, cached_tokens=0 every row, no prompt/KV/context/response reuse, no n-gram/history acceleration, UD-Q8_K_XL target/verifier with Q4_0 MTP draft verified by target; reordered-Q8 VDR2, n_max=3, n_min=2, p_min=0.0475, UBATCH_SIZE=1024; p10 80.120, mean 90.184, full512 after-TTFT 85.919, TTFT median 179.287 ms; supporting strict repeats in the same batch measured 88.571, 89.873, and 87.300, with a prior same-identity high observation at 91.393; submitted conservatively from the repeated-confirmation batch and supersedes the VDR2 90.322 row |
gemma4-26b-a4b-q8-b70-llamacpp-realistic-vdr2-mtp-n3-nmin2-p00475-ub1024-v21-20260627T201757 |
cmqwt1zk803ozqr01hctqss2z |
1 | suite median 69 | 512 | 90.322 median 1-100 after TTFT | 83.212 median wall full512 | policy-compliant realistic suite: fixed gemma4-26b-a4b-q8-b70-realistic-v1, each prompt once, cached_tokens=0 every row, no prompt/KV/context/response reuse, no n-gram/history acceleration, UD-Q8_K_XL target/verifier with Q4_0 MTP draft verified by target; reordered-Q8 VDR2, n_max=3, n_min=2, p_min=0.0475, UBATCH_SIZE=1024; p10 86.029, mean 92.180, full512 after-TTFT 86.217, TTFT median 179.681 ms; supporting strict VDR2 rows 89.455, 89.437, 88.063, and 85.906; supersedes the VDR2 89.455 row and VDR4 87.611 row |
gemma4-26b-a4b-q8-b70-llamacpp-realistic-vdr2-mtp-n3-nmin2-p00475-ub1024-v19-20260627T191931 |
cmqwqzayr03o8qr01j6lgx93n |
1 | suite median 69 | 512 | 89.455 median 1-100 after TTFT | 80.625 median wall full512 | policy-compliant realistic suite, superseded: fixed gemma4-26b-a4b-q8-b70-realistic-v1, each prompt once, cached_tokens=0 every row, no prompt/KV/context/response reuse, no n-gram/history acceleration, UD-Q8_K_XL target/verifier with Q4_0 MTP draft verified by target; reordered-Q8 VDR2, n_max=3, n_min=2, p_min=0.0475, UBATCH_SIZE=1024; p10 77.556, mean 87.849, full512 after-TTFT 84.452; supporting strict VDR2 rows 87.308, 87.240, 87.274, and 88.906; superseded by the VDR2 90.322 row |
gemma4-26b-a4b-q8-b70-llamacpp-realistic-mtp-n3-nmin2-p005-ub1024-v8-20260627T174753 |
cmqwnl2ag03lgqr01ch5bxknq |
1 | suite median 69 | 512 | 87.611 median 1-100 after TTFT | 77.865 median wall full512 | policy-compliant realistic suite, superseded: fixed gemma4-26b-a4b-q8-b70-realistic-v1, each prompt once, cached_tokens=0 every row, no prompt/KV/context/response reuse, no n-gram/history acceleration, UD-Q8_K_XL target/verifier with Q4_0 MTP draft verified by target; prior VDR4 high in the confirmed n3/p0.05/UB1024 family, with prior strict rows 84.825, 83.836, and 84.527; superseded by the VDR2 90.322 strict row |
gemma4-26b-a4b-q8-b70-llamacpp-realistic-mtp-n3-nmin2-p005-ub1024-20260627T171157 |
cmqwn5wq703l3qr01ilxrw6p2 |
1 | suite median 69 | 512 | 84.825 median 1-100 after TTFT | 78.321 median wall full512 | policy-compliant realistic suite: fixed gemma4-26b-a4b-q8-b70-realistic-v1, each prompt once, cached_tokens=0 every row, no prompt/KV/context/response reuse, no n-gram/history acceleration, UD-Q8_K_XL target/verifier with Q4_0 MTP draft verified by target; confirmations 83.836 and 84.527 in same n3/p0.05/UB1024 family |
gemma4-26b-a4b-q8-b70-llamacpp-mtp-n7-q8target-q40draft-q8reorder-vdr2-ub720-nmin3-pmin010-fresh-20260627T155347 |
cmqwkedg303jeqr013z753j62 |
1 | 588 | 512 | 176.216 first / 176.403 mean | 139.317 | 1536 repeats / 6144 canary rows passed; Q8 target/verifier with Q4_0 MTP draft only; same Q8 MoE-ID reorder broad verifier path as the prior UB720 record, with reordered Q8_0 MMVQ compile knob GGML_SYCL_REORDER_Q8_0_VDR_MMVQ=2; diagnostic pre-final-gate row0 metric only, all benchmark rows cached_tokens=0; not warmed/history/ngram accelerated |
gemma4-26b-a4b-q8-b70-llamacpp-mtp-n7-q8target-q40draft-q8reorder-ub720-nmin3-pmin010-fresh-20260627T144855 |
cmqwi45d803gyqr01td3vf9ka |
1 | 588 | 512 | 171.108 first / 170.129 mean | 135.666 | 1536 repeats / 6144 canary rows passed; Q8 target/verifier with Q4_0 MTP draft only; same Q8 MoE-ID reorder broad verifier path as the superseded UB704 record, with UBATCH_SIZE=720; diagnostic pre-final-gate row0 metric only, all benchmark rows cached_tokens=0; not warmed/history/ngram accelerated |
gemma4-26b-a4b-q8-b70-llamacpp-mtp-n7-q8target-q40draft-q8reorder-ub704-nmin3-pmin010-fresh-20260627T143126 |
cmqwhkbzj03guqr01h00c8n04 |
1 | 588 | 512 | 170.112 first / 169.876 mean | 134.896 | 1536 repeats / 6144 canary rows passed; Q8 target/verifier with Q4_0 MTP draft only; same Q8 MoE-ID reorder broad verifier path as the superseded UB768 record, with UBATCH_SIZE=704; diagnostic pre-final-gate row0 metric only, all benchmark rows cached_tokens=0; not warmed/history/ngram accelerated |
gemma4-26b-a4b-q8-b70-llamacpp-mtp-n7-q8target-q40draft-q8reorder-ub768-nmin3-pmin010-fresh-20260627T142318 |
cmqwh8du403gfqr01d6ut1ddo |
1 | 588 | 512 | 169.949 first / 169.550 mean | 135.182 | 1536 repeats / 6144 canary rows passed; Q8 target/verifier with Q4_0 MTP draft only; adds LLAMA_SYCL_MUL_MAT_ID_MULTI_TOKEN_FAST=1 + LLAMA_SYCL_MUL_MAT_ID_Q8_0_REORDER=1 to the prior route-cache/fused-output/fused-selected-softmax/RMS-reuse stack; diagnostic pre-final-gate row0 metric only, all benchmark rows cached_tokens=0; not warmed/history/ngram accelerated |
gemma4-26b-a4b-q8-b70-llamacpp-syclopt0-faoff-20260623T0715 |
cmqq8phxt0103qo01afcgyjq8 |
1 | 574 | 156 | 41.806 | n/a | 384/384 chat canary |
gemma4-26b-a4b-q8-b70-llamacpp-syclopt0-faoff-parallel1-cache0-20260623T0915 |
cmqq9nqbh010gqo01a9jnzl6r |
1 | 574 | 146 | 42.154 | n/a | 384/384 chat canary |
gemma4-26b-a4b-q8-b70-llamacpp-syclopt0-faoff-parallel1-cache0-long512-20260623T0945 |
cmqqa6zbx010xqo01cdtfn8e0 |
1 | 75 | 512 | 42.716 | 41.351 | 384/384 chat canary |
gemma4-26b-a4b-q8-b70-llamacpp-mtp-n3-aot-repeat-long512-20260623T0353 |
cmqqctk4w014kqo011gyyks7r |
1 | 75 | 512 | 48.347 | 46.602 | 384/384 chat canary |
gemma4-26b-a4b-q8-b70-llamacpp-mtp-n3-aot-filledlong512-20260623T0853 |
cmqqexo5x0151qo0154xsie7s |
1 | 588 | 512 | 68.192 | 63.428 | 384/384 chat canary |
gemma4-26b-a4b-q8-b70-llamacpp-mtp-n3-aot-psplit020-filledlong512-20260623T0858 |
cmqqf759s0154qo01gwqa14uc |
1 | 588 | 512 | 68.515 | 63.666 | 384/384 chat canary |
gemma4-26b-a4b-q8-b70-llamacpp-mtp-n4-aot-filledlong512-20260623T0858 |
cmqqf75p70157qo018fsavf0g |
1 | 588 | 512 | 74.395 | 68.797 | 384/384 chat canary |
gemma4-26b-a4b-q8-b70-llamacpp-mtp-n4-aot-psplit020-filledlong512-20260623T0907 |
cmqqfe75s015aqo01xr94yxh0 |
1 | 588 | 512 | 74.498 | 68.900 | 384/384 chat canary |
gemma4-26b-a4b-q8-b70-llamacpp-mtp-n6-aot-nmin2-pmin015-filledlong512-20260623T0912 |
cmqqfnilo015lqo011nm0q2tn |
1 | 588 | 512 | 83.520 | 76.569 | 384/384 chat canary |
gemma4-26b-a4b-q8-b70-llamacpp-mtp-n7-aot-nmin2-pmin015-filledlong512-20260623T0919 |
cmqqfv296015sqo0126mym3ko |
1 | 588 | 512 | 87.878 | 80.252 | 384/384 chat canary |
gemma4-26b-a4b-q8-b70-llamacpp-mtp-n7-aot-nmin2-pmin010-filledlong512-20260623T0925 |
cmqqg1r0l015xqo01e6d696mx |
1 | 588 | 512 | 88.345 | 80.553 | 384/384 chat canary |
gemma4-26b-a4b-q8-b70-llamacpp-mtp-n7-aot-nmin2-pmin010-nobs-filledlong512-20260623T0936 |
cmqqgftv50160qo01km3s7lkt |
1 | 588 | 512 | 90.243 | 82.243 | 384/384 chat canary |
gemma4-26b-a4b-q8-b70-llamacpp-mtp-n7-aot-nmin2-pmin010-nobs-dthreads32-filledlong512-20260623T0941 |
cmqqgn3cm0163qo010optg91u |
1 | 588 | 512 | 90.419 | 82.342 | 384/384 chat canary |
gemma4-26b-a4b-q8-b70-llamacpp-mtp-n7-aot-nmin2-pmin012-nobs-dthreads32-dtb32-filledlong512-20260623T1018 |
cmqqi1p2c016jqo01vndau1y9 |
1 | 588 | 512 | 91.050 | 82.970 | 384/384 chat canary |
gemma4-26b-a4b-q8-b70-llamacpp-mtp-n7-c926ad098-ctxcp0-nmin2-pmin012-nobs-dthreads32-dtb32-filledlong512-20260623 |
cmqqkmbhr017oqo017rdfxqh2 |
1 | 588 | 512 | 91.157 | 71.057 | 384/384 chat canary |
gemma4-26b-a4b-q8-b70-llamacpp-mtp-n7-c926ad098-fasttopk10-filledlong512-20260623T1508 |
cmqqsecuk01azqo018ahv0i1s |
1 | 588 | 512 | 91.619 | 71.287 | 384/384 chat canary |
gemma4-26b-a4b-q8-b70-llamacpp-mtp-n7-c926ad098-fasttopk10-cpucleanup-filledlong512-20260623T2217 |
cmqr7ni7u01gxqo01wtqsrn3u |
1 | 588 | 512 | 91.877 first / 91.899 mean | 71.485 | 384/384 chat canary |
gemma4-26b-a4b-q8-b70-llamacpp-mtp-n7-c926ad098-fastargmax-cpucleanup-vmm0-ub512-poll100-filledlong512-20260623T2228 |
cmqr82niq01hgqo01v42y7ue8 |
1 | 588 | 512 | 92.397 first / 92.767 mean | 83.289 | 384/384 chat canary; conservative diagnostic pre-final-gate first-request metric only |
gemma4-26b-a4b-q8-b70-llamacpp-mtp-n7-q8target-q40draft-fresh-20260624T0812 |
cmqrsupdk000jqr01af3eu6vu |
1 | 588 | 512 | 95.264 first / 95.386 mean | 81.285 | 384/384 chat canary; Q8 target/verifier with Q4_0 MTP draft only; diagnostic pre-final-gate first-request metric only |
gemma4-26b-a4b-q8-b70-llamacpp-mtp-n7-q8target-q40draft-directunroll7-qonly-fresh-20260624T1432 |
cmqs4jnx100k6qr01d1iy78kl |
1 | 588 | 512 | 96.822 first / 97.226 mean | 82.462 | 384/384 chat canary; Q8 target/verifier with Q4_0 MTP draft only; direct argmax-ID unroll + q-only assistant inputs; diagnostic pre-final-gate first-request metric only |
gemma4-26b-a4b-q8-b70-llamacpp-mtp-n7-q8target-q40draft-directunroll7-qonly-b1024u1024-th8-fresh-20260624T1357 |
cmqs56wv100kjqr01de3fdspd |
1 | 588 | 512 | 98.491 first / 97.886 mean | 86.194 | 384/384 chat canary; Q8 target/verifier with Q4_0 MTP draft only; direct argmax-ID unroll + q-only assistant inputs; BATCH_SIZE=1024, UBATCH_SIZE=1024, THREADS=8; diagnostic pre-final-gate row0 metric only |
gemma4-26b-a4b-q8-b70-llamacpp-mtp-n7-q8target-q40draft-directunroll7-qonly-b1024u1024-th8-syclgraph0-fresh-20260624T1447 |
cmqs7uyqb00lnqr01u9dtv63r |
1 | 588 | 512 | 98.617 first / 97.956 mean | 86.262 | 384/384 chat canary; Q8 target/verifier with Q4_0 MTP draft only; direct argmax-ID unroll + q-only assistant inputs; BATCH_SIZE=1024, UBATCH_SIZE=1024, THREADS=8, GGML_SYCL_DISABLE_GRAPH=0; diagnostic pre-final-gate row0 metric only |
gemma4-26b-a4b-q8-b70-llamacpp-mtp-n7-q8target-q40draft-rowargmax-deferh-pmin014-fresh-20260624T1735 |
cmqsd2jpn00pwqr017fq21akz |
1 | 588 | 512 | 101.428 first / 100.769 mean | 88.374 | 384/384 chat canary; Q8 target/verifier with Q4_0 MTP draft only; verifier row-argmax IDs + deferred target h_nextn + MTP_P_MIN=0.14; superseded by safer verifier row-argmax result; diagnostic pre-final-gate row0 metric only |
gemma4-26b-a4b-q8-b70-llamacpp-mtp-n7-q8target-q40draft-rowargmax-safer-deferh-pmin014-fresh-20260624T1830 |
cmqsf630x00r1qr01d1usfo2d |
1 | 588 | 512 | 101.482 first / 101.249 mean | 88.582 | 1536/1536 chat canary; Q8 target/verifier with Q4_0 MTP draft only; stricter verifier row-argmax shape guard + deferred target h_nextn + MTP_P_MIN=0.14; superseded by immediate-command-list result; diagnostic pre-final-gate row0 metric only, all benchmark rows cached_tokens=0 |
gemma4-26b-a4b-q8-b70-llamacpp-mtp-n7-q8target-q40draft-rowargmax-safer-deferh-pmin014-immediatecl1-fresh-20260624T1932 |
cmqshlz8j00s0qr01f7lr24oh |
1 | 588 | 512 | 101.602 first / 100.835 mean | 88.508 | 1536/1536 chat canary; Q8 target/verifier with Q4_0 MTP draft only; safer verifier row-argmax + deferred target h_nextn + MTP_P_MIN=0.14 plus UR_L0_USE_IMMEDIATE_COMMANDLISTS=1; superseded by selected-softmax/weighted-sum result; diagnostic pre-final-gate row0 metric only, all benchmark rows cached_tokens=0 |
gemma4-26b-a4b-q8-b70-llamacpp-mtp-n7-q8target-q40draft-selectedsoftmax-weightedsum-pmin0136-fresh-20260625T0315 |
cmqsylo2l011nqr011yydjvne |
1 | 588 | 512 | 103.299 first / 102.193 mean | 89.849 | 1536/1536 chat canary; Q8 target/verifier with Q4_0 MTP draft only; selected-softmax + weighted-sum MoE source guards, safer verifier row-argmax, deferred target h_nextn, MTP_P_MIN=0.136, and UR_L0_USE_IMMEDIATE_COMMANDLISTS=1; superseded by route-cache micro-record; diagnostic pre-final-gate row0 metric only, all benchmark rows cached_tokens=0 |
gemma4-26b-a4b-q8-b70-llamacpp-mtp-n7-q8target-q40draft-routecache-mtpfusedoutargmax-selfusedweights-rmsreuse-ub768-nmin3-pmin010-fresh-20260627T070421 |
cmqw1tgzx0366qr01g4lkv7f1 |
1 | 588 | 512 | 104.309 first / 103.934 mean | 90.851 | 1536 repeats / 6144 canary rows passed; Q8 target/verifier with Q4_0 MTP draft only; same route-cache/fused-output/fused-selected-softmax recipe with UBATCH_SIZE=768, MTP_N_MIN=3, MTP_P_MIN=0.10 on GPU0 plus LLAMA_GEMMA4_MOE_REUSE_ATTN_RMS=1; superseded within the diagnostic/pre-final-gate lane by later Q8 MoE-ID reorder rows; all benchmark rows cached_tokens=0 |
gemma4-26b-a4b-q8-b70-llamacpp-mtp-n7-q8target-q40draft-routecache-mtpfusedoutargmax-selfusedweights-ub768-nmin3-pmin010-fresh-20260627T035307 |
cmqvv3kop0309qr013ekr8apu |
1 | 588 | 512 | 104.226 first / 104.174 mean | 90.741 | 1536 repeats / 6144 canary rows passed; Q8 target/verifier with Q4_0 MTP draft only; same route-cache/fused-output/fused-selected-softmax recipe with UBATCH_SIZE=768, MTP_N_MIN=3, MTP_P_MIN=0.10 on GPU0; diagnostic pre-final-gate row0 metric only, all benchmark rows cached_tokens=0; support mean also improves over the previous 104.071 record, but this remains a small variance-class row, not material progress toward >150 |
gemma4-26b-a4b-q8-b70-llamacpp-mtp-n7-q8target-q40draft-routecache-mtpfusedoutargmax-selfusedweights-ub768-fresh-20260627T002926 |
cmqvmjvzx02qvqr01qh9jikow |
1 | 588 | 512 | 104.071 first / 103.589 mean | 90.487 | 1536 repeats / 6144 canary rows passed; Q8 target/verifier with Q4_0 MTP draft only; same route-cache/fused-output/fused-selected-softmax recipe with UBATCH_SIZE=768 on GPU3; superseded by cmqvv3kop0309qr013ekr8apu; diagnostic pre-final-gate row0 metric only, all benchmark rows cached_tokens=0; support mean is lower than prior record, so this was a row0 variance-class micro-record over 103.983, not material progress toward >150 |
gemma4-26b-a4b-q8-b70-llamacpp-mtp-n7-q8target-q40draft-routecache-mtpfusedoutargmax-selfusedweights-repeat-fresh-20260626T230510 |
cmqvjupek02pgqr01d46algvg |
1 | 588 | 512 | 103.983 first / 104.096 mean | 90.479 | 1536/1536 chat canary; Q8 target/verifier with Q4_0 MTP draft only; exact same route-cache/fused-output recipe as prior row, repeated on GPU0; superseded by the 104.226 UBATCH_SIZE=768, n_min=3, p_min=0.10 micro-record; diagnostic pre-final-gate row0 metric only, all benchmark rows cached_tokens=0; variance-class micro-record over 103.954 |
gemma4-26b-a4b-q8-b70-llamacpp-mtp-n7-q8target-q40draft-routecache-mtpfusedoutargmax-selfusedweights-fresh-20260626T222525 |
cmqviful602p0qr01vp27jw5i |
1 | 588 | 512 | 103.954 first / 104.135 mean | 90.686 | 1536/1536 chat canary; Q8 target/verifier with Q4_0 MTP draft only; route-cache recipe plus LLAMA_GEMMA4_MTP_FUSED_OUTPUT_ARGMAX=1 and LLAMA_GEMMA4_MOE_SELECTED_SOFTMAX_FUSED=1; superseded same-stack diagnostic pre-final-gate row0 metric only, all benchmark rows cached_tokens=0; small validated micro-record over 103.515 |
gemma4-26b-a4b-q8-b70-llamacpp-mtp-n7-q8target-q40draft-routecache-ctx8192-gpu2-pmin0136-fresh-20260626T191746 |
cmqvbq8tf02m1qr010dom0vu1 |
1 | 588 | 512 | 103.515 first / 103.193 mean | 90.220 | 1536/1536 chat canary; Q8 target/verifier with Q4_0 MTP draft only; same route-cache recipe, validated after a four-GPU CTX screen on GPU2/ctx8192; superseded by fused-output-argmax/fused-selected-softmax record; diagnostic pre-final-gate row0 metric only, all benchmark rows cached_tokens=0; small validated micro-record over 103.301 |
gemma4-26b-a4b-q8-b70-llamacpp-mtp-n7-q8target-q40draft-routecache-pmin0136-fresh-20260626T184617 |
cmqvalync02lhqr01h76rnti3 |
1 | 588 | 512 | 103.301 first / 103.063 mean | 89.977 | 1536/1536 chat canary; Q8 target/verifier with Q4_0 MTP draft only; same selected-softmax + weighted-sum recipe plus default-off LLAMA_SYCL_MUL_MAT_ID_ROUTE_CACHE=1; superseded by GPU2/ctx8192 route-cache validation; diagnostic pre-final-gate row0 metric only, all benchmark rows cached_tokens=0; micro-record only (+0.001884 tok/s) |
These four rows were submitted before the fresh/warmed policy was clarified.
They are valid Q8 verification of a repeated continuation, but not valid
realistic cold-suite speed claims because the draftless n-gram source had already
seen the benchmark output. Local queue artifacts were corrected on 2026-06-26
so top-level tokSOut records the cold row0 rate and warmed means live under
diagnostic engineFlags.
| Label | LocalMaxxing ID | GPUs | Input | Output | row0 fresh tok/s | warmed tok/s | Validation |
|---|---|---|---|---|---|---|---|
gemma4-26b-a4b-q8-b70-llamacpp-ngrammod-24-48-64-filledlong512-20260623T1745 |
cmqqxbkzx01cxqo01j8p97627 |
1 | 588 | 512 | 41.138 | 245.980 | 384/384 chat canary; warmed/history artifact, retraction-needed |
gemma4-26b-a4b-q8-b70-llamacpp-ngrammod-20-32-64-filledlong512-20260623T1750 |
cmqqxjnif01d0qo01ix4oeixo |
1 | 588 | 512 | 41.097 | 255.041 | 384/384 chat canary; warmed/history artifact, retraction-needed |
gemma4-26b-a4b-q8-b70-llamacpp-ngrammod-20-32-64-filledlong512-20260623T1815 |
cmqqxx7bp01dbqo012d2qiiw6 |
1 | 588 | 512 | 41.364 | 280.040 | 384/384 chat canary; warmed/history artifact, retraction-needed |
gemma4-26b-a4b-q8-b70-llamacpp-ngrammod-20-32-64-filledlong512-20260623T1855 |
cmqqyby6801dvqo01as3wenz2 |
1 | 588 | 512 | 41.308 | 280.642 | 384/384 chat canary; warmed/history artifact, retraction-needed |
Required packet: see
results/gemma4-26b-a4b-q8-b70/localmaxxing-and-targets.md.
Submit artifacts:
data/localmaxxing-gemma4-26b-a4b-q8-b70-syclopt0-faoff-20260623.queue.jsondata/localmaxxing-responses/gemma4-26b-a4b-q8-b70-syclopt0-faoff-20260623.submit.logdata/localmaxxing-responses/gemma4-26b-a4b-q8-b70-syclopt0-faoff-20260623.submit2.logdata/localmaxxing-responses/gemma4-26b-a4b-q8-b70-syclopt0-faoff-parallel1-cache0-20260623.submit.logdata/localmaxxing-gemma4-26b-a4b-q8-b70-long512-20260623.queue.jsondata/localmaxxing-responses/gemma4-26b-a4b-q8-b70-long512-20260623.submit.logdata/localmaxxing-responses/gemma4-26b-a4b-q8-b70-mtp-n3-aot-repeat-long512-20260623.submit.logdata/localmaxxing-responses/gemma4-26b-a4b-q8-b70-mtp-n3-aot-filledlong512-20260623.submit.log,
data/localmaxxing-responses/gemma4-26b-a4b-q8-b70-mtp-n3-aot-psplit020-filledlong512-20260623.submit.log,
data/localmaxxing-responses/gemma4-26b-a4b-q8-b70-mtp-n4-aot-filledlong512-20260623.submit.log,
data/localmaxxing-responses/gemma4-26b-a4b-q8-b70-mtp-n4-aot-psplit020-filledlong512-20260623.submit.log,
data/localmaxxing-responses/gemma4-26b-a4b-q8-b70-mtp-n6-aot-nmin2-pmin015-filledlong512-20260623.submit.log,
data/localmaxxing-responses/gemma4-26b-a4b-q8-b70-mtp-n7-aot-nmin2-pmin015-filledlong512-20260623.submit.log,
data/localmaxxing-responses/gemma4-26b-a4b-q8-b70-mtp-n7-aot-nmin2-pmin010-filledlong512-20260623.submit.log,
data/localmaxxing-responses/gemma4-26b-a4b-q8-b70-mtp-n7-aot-nmin2-pmin010-nobs-filledlong512-20260623.submit.log,
data/localmaxxing-responses/gemma4-26b-a4b-q8-b70-mtp-n7-aot-nmin2-pmin010-nobs-dthreads32-filledlong512-20260623.submit.log,
data/localmaxxing-responses/gemma4-26b-a4b-q8-b70-mtp-n7-aot-nmin2-pmin012-nobs-dthreads32-dtb32-filledlong512-20260623.submit.log,
data/localmaxxing-responses/gemma4-26b-a4b-q8-b70-mtp-n7-c926ad098-ctxcp0-nmin2-pmin012-nobs-dthreads32-dtb32-filledlong512-20260623.submit.log,
data/localmaxxing-responses/gemma4-26b-a4b-q8-b70-mtp-n7-c926ad098-fasttopk10-filledlong512-20260623.submit.log,
data/localmaxxing-responses/gemma4-26b-a4b-q8-b70-mtp-n7-c926ad098-fasttopk10-cpucleanup-filledlong512-20260623.submit.log,
data/localmaxxing-responses/gemma4-26b-a4b-q8-b70-mtp-n7-c926ad098-fastargmax-cpucleanup-vmm0-ub512-poll100-filledlong512-20260623.submit.logdata/localmaxxing-responses/gemma4-26b-a4b-q8-b70-llamacpp-mtp-n7-q8target-q40draft-fresh-20260624.submit.logdata/localmaxxing-responses/gemma4-26b-a4b-q8-b70-llamacpp-mtp-n7-q8target-q40draft-directunroll7-qonly-b1024u1024-th8-syclgraph0-fresh-20260624.submit.logdata/localmaxxing-responses/gemma4-26b-a4b-q8-b70-llamacpp-mtp-n7-q8target-q40draft-rowargmax-deferh-pmin014-fresh-20260624.submit.logdata/localmaxxing-responses/gemma4-26b-a4b-q8-b70-llamacpp-mtp-n7-q8target-q40draft-selectedsoftmax-weightedsum-pmin0136-fresh-20260625.submit.logUBATCH_SIZE=768, MTP_N_MIN=3, MTP_P_MIN=0.10
Q8-target/Q4_0-draft pre-final-gate approved response:
data/localmaxxing-responses/gemma4-26b-a4b-q8-b70-llamacpp-mtp-n7-q8target-q40draft-routecache-mtpfusedoutargmax-selfusedweights-rmsreuse-ub768-nmin3-pmin010-fresh-20260627.submit.logUBATCH_SIZE=768, MTP_N_MIN=3, MTP_P_MIN=0.10
Q8-target/Q4_0-draft pre-final-gate approved response:
data/localmaxxing-responses/gemma4-26b-a4b-q8-b70-llamacpp-mtp-n7-q8target-q40draft-routecache-mtpfusedoutargmax-selfusedweights-ub768-nmin3-pmin010-fresh-20260627.submit.logdata/localmaxxing-responses/gemma4-26b-a4b-q8-b70-llamacpp-mtp-n7-q8target-q40draft-routecache-mtpfusedoutargmax-selfusedweights-fresh-20260626.submit.logdata/localmaxxing-responses/gemma4-26b-a4b-q8-b70-ngrammod-24-48-64-filledlong512-20260623.submit.log,
data/localmaxxing-responses/gemma4-26b-a4b-q8-b70-ngrammod-20-32-64-filledlong512-20260623.submit.log,
data/localmaxxing-responses/gemma4-26b-a4b-q8-b70-ngrammod-20-32-64-ctx4096ub512-filledlong512-20260623.submit2.log,
data/localmaxxing-responses/gemma4-26b-a4b-q8-b70-ngrammod-20-32-64-poll100-filledlong512-20260623.submit.logdata/localmaxxing-responses/gemma4-26b-a4b-q8-b70-ngrammod-20-32-64-ctx4096ub512-filledlong512-20260623.submit.log
(empty because the first command filtered on a non-matching label)Correction on 2026-06-23: the four draftless ngram-mod submissions above are
not valid fresh-response headline throughput because the speedup depends on
repeated-output continuation history. They remain useful warmed/history
artifacts, but should be retracted from any public headline leaderboard view.
Their first synthetic measured rows were only about 41 tok/s after TTFT:
41.138 (cmqqxbkzx01cxqo01j8p97627), 41.097
(cmqqxjnif01d0qo01ix4oeixo), 41.364
(cmqqxx7bp01dbqo012d2qiiw6), and 41.308
(cmqqyby6801dvqo01as3wenz2). Do not average warmed repeated rows into a
fresh-response claim.
API deletion was attempted for all four IDs and returned 404 because
LocalMaxxing currently exposes only GET/POST /api/benchmarks and
POST /api/benchmarks/dry-run; see
data/localmaxxing-responses/gemma4-ngram-history-accelerated-delete-attempts-20260623.json
and the OpenAPI method snapshot at
data/localmaxxing-responses/localmaxxing-openapi-benchmark-methods-20260623.json.
The first attempt failed only because the payload used backend="SYCL/Level Zero".
The accepted payload uses LocalMaxxing’s enum backend="xpu" and stores
SYCL/Level Zero as engineFlags.backendDetail.
Date: 2026-06-23
Model: nameistoken/Qwen3.6-35B-A3B-Quark-W8A8-INT8, Quark W8A8 INT8,
vLLM/XPU on Intel Arc Pro B70.
| Label | LocalMaxxing ID | GPUs | Input | Output | tok/s out | tok/s total |
|---|---|---|---|---|---|---|
qwen36-35b-quark-int8-b70-tp4-strict-deep-gate-20260615a13deep2 |
cmqq4mw4c00yfqo01gb2ucgxj |
4 | 512 | 512 | 93.551 | 178.773 |
qwen36-35b-quark-int8-b70-tp2-safe-smoke-20260615tp2safe1 |
cmqq4mwgm00yiqo0133bj962q |
2 | 512 | 512 | 85.869 | 162.283 |
Note: the TP4 submission is the current strict-valid deep gate: JSON 128/128,
color 256/256, and quality suite pass. The TP2 submission is the best safer
reference smoke with JSON 16/16 and color 16/16; quality suite was skipped,
so it is labeled as a TP2 reference rather than a stronger deep-gate result.
Payload queue and response log:
data/localmaxxing-qwen36-35b-quark-int8-b70-valid-2x4x-20260623.queue.json
and
data/localmaxxing-responses/qwen36-35b-quark-int8-b70-valid-2x4x-20260623.submit.log.
Date: 2026-05-15
Model: Lasimeri/MiniMax-M2.7-int4-AutoRound, AutoRound W4A16 safetensors,
vLLM/XPU TP4.
| Label | LocalMaxxing ID | GPUs | Input | Output | tok/s out | tok/s total |
|---|---|---|---|---|---|---|
vllm-minimax-m27-clean-weight-piecewise-aot-p512-n1536 |
cmp6a5c1o00mpo3011hg8ncyp |
4 | 512 | 1536 | 65.752 | 87.670 |
Note: repaired piecewise/AOT compiled path with the default-off MiniMax Q/K
RMSNorm clean-weight guard enabled. Three p512/n1536 repeats were 64.622,
66.659, and 65.976 output tok/s. Raw-prompt quality canaries at 64 and
256 generated tokens both passed with 0 NUL tokens, 0 non-space control
chars, and nontrivial token diversity. This supersedes the earlier quality-
corrected ~61 tok/s TP4 baseline, but the older ~73 tok/s AOT diagnostic
remains invalid because it failed the raw corruption gate.
Date: 2026-05-09
Model: Lasimeri/MiniMax-M2.7-int4-AutoRound, AutoRound W4A16 safetensors, vLLM/XPU TP4.
| Label | LocalMaxxing ID | GPUs | Input | Output | tok/s out | tok/s total |
|---|---|---|---|---|---|---|
vllm-minimax-m27-autoround-u4-decode-p512-n128 |
cmoxptkfd00hsml01hf2ajhhp |
4 | 512 | 128 | 29.748 | 148.742 |
vllm-minimax-m27-autoround-u4-decode-p512-n256 |
cmoxq7cww00i8ml019ihbeqc9 |
4 | 512 | 256 | 33.034 | 99.101 |
vllm-minimax-m27-autoround-u4-fp32-route-p512-n256 |
cmoy8hs3n002smk01ksgcpavr |
4 | 512 | 256 | 34.158 | 102.474 |
vllm-minimax-m27-autoround-u4-pp2tp2-negative-p512-n256 |
cmoy9exmf003lmk01d3it9cz2 |
4 | 512 | 256 | 17.550 | 52.651 |
vllm-minimax-m27-autoround-u4-default-ipc-p512-n256 |
cmoy9qat60040mk01l5y8n3al |
4 | 512 | 256 | 34.578 | 103.734 |
vllm-minimax-m27-autoround-u4-default-ipc-p512-n512 |
cmoyagit0004dmk014gk25e2k |
4 | 512 | 512 | 37.136 | 74.272 |
vllm-minimax-m27-autoround-xpu-graph-fixedkv-p512-n256 |
cmoyfl7cm0057mk01suxo0glp |
4 | 512 | 256 | 32.723 | 98.169 |
Note: unsigned llm-scaler u4 decode-only MoE path, no speculative decode, no expert dropping, no sampling changes, and no power-limit changes. The XPU graph fixed-KV result is a negative/diagnostic run: PIECEWISE graph capture succeeded with local vLLM patches, but it was slower than the non-graph default-IPC path.
Date: 2026-05-03
Model: Lorbus/Qwen3.6-27B-int4-AutoRound
All submitted results returned APPROVED.
| Label | LocalMaxxing ID | GPUs | Input | Output | tok/s out | tok/s total |
| — | — | —: | —: | —: | —: | —: |
| vllm-int4-single-b70-mtp-500-256 | cmoq41b9d001alg043wsnthz2 | 1 | 500 | 256 | 45.2 | 133.44 |
| vllm-int4-single-b70-mtp-500-512 | cmoq47sll0005l104v3i0f9l3 | 1 | 500 | 512 | 41.3 | 81.60 |
| vllm-int4-tp2-b70-nonmtp-500-256 | cmoq4e9dw0002js04ledqyycn | 2 | 500 | 256 | 49.1 | 144.88 |
| vllm-int4-tp2-b70-nonmtp-500-512 | cmoq4krfb000cl40456wobg7e | 2 | 500 | 512 | 48.3 | 95.56 |
| vllm-int4-single-b70-nonmtp-500-256 | cmoq4r8rc0001l804tocgibus | 1 | 500 | 256 | 31.8 | 93.80 |
| vllm-int4-tp2-b70-mtp-500-256 | cmoq4xppt0003ky04xidngli9 | 2 | 500 | 256 | 35.6 | 105.03 |
Date: 2026-07-11
Model: webhie/Qwen3.6-27B-int4-AutoRound, AutoRound INT4 W4A16 target,
runtime INT8 target LM-head, runtime INT4 intrinsic-MTP draft LM-head,
vLLM/XPU TP2 on two Intel Arc Pro B70 GPUs.
| Label | LocalMaxxing ID | GPUs | Output | tok/s out | tok/s wall |
|---|---|---|---|---|---|
qwen36-27b-webhie-int4-autoround-b70-vllm-tp2-fp16-fullgraph-transaction-95tok-20260711 |
cmrh35ct50092mj01h7jgydqj |
2 | 512 | 95.385 | 80.405 |
qwen36-27b-webhie-int4-autoround-b70-vllm-tp2-fp16-graphsafe-fa-fullgraph-93tok-20260711 |
cmrgue7kl007pmj01yrkcyqmv |
2 | 512 | 93.036 | 79.837 |
qwen36-27b-webhie-int4-autoround-b70-vllm-tp2-fp16-capturegdn-91tok-20260711 |
cmrgojixq005rmj0141e9fjj2 |
2 | 512 | 91.714 | 76.670 |
qwen36-27b-webhie-int4-autoround-b70-vllm-tp2-capturegdn-87tok-20260711 |
cmrgn3szj005dmj01u8tel6yd |
2 | 512 | 87.029 | 75.780 |
Strict fresh-response record: 12 unique realistic prompts once, every request
cached_tokens=0, no cache/history/response reuse, target-verified MTP3,
exact cases + repeat128 + baseline parity + 1K needle passed. Graph-safe
FlashAttention reduces target graph calls from 33 PIECEWISE segments to one
full graph; exact ReplaySSM pending/direct-output transaction fusions raise the
current record to 95.385 tok/s. The route is short-context-only pending
graph-safe paged decode.
Current packet:
results/qwen36-27b-autoround-int4-b70/tp2-fp16-fullgraph-transaction-20260711.json.
2026-07-22: cmrw7cn1k006jnz01gq2z981v APPROVED — DFlash batched-exact, 33.086 tok/s (median_tok_s_1_100_after_ttft, full-512 contract), 13/13 exact vs deterministic q=1 teacher on 2 fresh starts, cached_tokens=0, TP4+EP4 INT4 W4A16, one active generation. Payload: data/localmaxxing-laguna-s-2.1-int4-b70-dflash-bulletproof-33.086tok-20260722.queue.json
2026-07-22: cmrwlyxez00f4nz01zefturuv APPROVED — Laguna S 2.1 INT4 (4x B70), 33.268 tok/s (fused W1+SiLU + route-parallel W2), lower of 2 fresh starts (33.303/33.268, spread 0.108%), +0.55% over prior record cmrw7cn1k. 13/13 exact vs q1 teacher x2 starts, cross-req+rollover exact, cached_tokens=0. Payload: data/localmaxxing-laguna-s-2.1-int4-b70-dflash-fused-w1-route-w2-33.268tok-20260722.queue.json
2026-07-22: cmrwot89400gqnz014oodtlbp APPROVED — Laguna S 2.1 INT4 (4x B70), 33.439 tok/s (M8 route-interleave expert GEMM occupancy: W1 EU 43.8->47.7%, W2 46.5->49.1%), lower of 2 starts (33.439/33.546), +0.51% over cmrwlyxez. 13/13 exact x2, rollover exact, cached_tokens=0. Superseded by cmrx6p5dv001bo4017hb7sixz.
2026-07-23: cmrx6p5dv001bo4017hb7sixz APPROVED — Laguna S 2.1 INT4 (4x B70), 33.895 tok/s (exact shared-elementwise + QKNorm/RoPE launch-reduction stack on the route-interleaved MoE record base), conservative lower candidate of a preregistered A-B-B-A endpoint (34.551/33.895 candidate versus 32.827/33.273 adjacent controls), +1.364% over cmrwot894. Both candidate/control pairs passed; candidate won 12/13 and 13/13 prompt rows and saved 3.490/4.015 ms per target cycle. All 52 requests were exact vs canonical q1 and cached_tokens=0; long-next 8/8 and rollover 4/4 passed. Queue: data/localmaxxing-laguna-s-2.1-int4-b70-dflash-shared-elementwise-qknorm-33.895tok-20260723.queue.json; response: data/localmaxxing-responses/laguna-s-2.1-int4-b70-dflash-shared-elementwise-qknorm-33.895tok-20260723.response.json.
2026-07-24: cmrzjb7i906x4o401egrnm05m APPROVED — Laguna S 2.1 INT4 (4x B70), 92.164 tok/s (validated exact Breakable M8 PIECEWISE graph on the DFlash7 route-interleaved/shared-elementwise/QKNorm-RoPE stack), conservative lower graph start of a preregistered fresh A1-B1-B2-A2 campaign (92.761/92.164 graph versus 34.491/34.591 eager). Graph won 13/13 and 12/13 rows and saved 55.049/54.220 ms per target cycle with acceptance drift below 0.000308. All 52 requests were bitwise exact vs canonical q1 and cached_tokens=0; long-next 8/8 and rollover 4/4 passed. Each graph start captured/replayed exactly once on all four ranks with audited 146/145 topology. Queue: data/localmaxxing-laguna-s-2.1-int4-b70-dflash-breakable-graph-92.164tok-20260724.queue.json; response: data/localmaxxing-responses/laguna-s-2.1-int4-b70-dflash-breakable-graph-92.164tok-20260724.response.json.
2026-07-25: cmrzrd4tf001ipa013xpx4kid APPROVED — Laguna S 2.1 INT4 (4x B70), 94.920 tok/s (persistent exact-attention metadata on the validated Breakable M8 PIECEWISE graph stack), conservative lower candidate of a preregistered graph-vs-graph A1-B1-B2-A2 campaign (94.920/95.067 metadata-on versus 92.550/92.878 metadata-off). Candidate won 13/13 rows in both adjacent pairs, improved headline throughput by 2.561%/2.356%, and saved 0.911/1.648 ms per aggregate target cycle with acceptance drift below 0.000308. All 52 requests were bitwise exact vs canonical q1 and cached_tokens=0; cross-leg exactness was 39/39, long-next 8/8, and rollover 4/4. All four starts captured/replayed the audited 146/145 topology on all ranks. Queue: data/localmaxxing-laguna-s-2.1-int4-b70-dflash-persistent-metadata-94.920tok-20260725.queue.json; response: data/localmaxxing-responses/laguna-s-2.1-int4-b70-dflash-persistent-metadata-94.920tok-20260725.response.json.
2026-07-26: cms2ccv2d00lps201rej94pjy APPROVED — Laguna S 2.1 INT4 (4x B70), 102.971435596 tok/s under the submitted legacy 100 events / 99-interval span convention; a later reproduction audit recomputed 101.941721240 tok/s under conventional interval accounting. Width 12 / DFlash depth 11 ran on the audited 146/145 Breakable PIECEWISE graph with 31 runtime E4M3FN W8A16 draft-projection conversions per rank. This was the first valid score from one preregistered cold service: every fixed prompt once, no warmup generation or retry, 13/13 token IDs and output-text hashes exact vs canonical q1, all cached_tokens=0, 512-output-then-next 2/2, rollover 1/1, clean teardown, and 73-second pre/post idle gates. The relative improvement over the prior row remains 8.482% because both used the same convention. No gain is attributed to the intended FP8 draft LM head because its runtime preparation log was absent. Preserve this approved receipt and do not duplicate-submit. Correction: experiments/laguna-s-2.1-xpu-b70/notes/2026-07-26-throughput-window-accounting-correction.md; queue: data/localmaxxing-laguna-s-2.1-int4-b70-width12-dflash-fp8-102.971tok-20260726.queue.json; response: data/localmaxxing-responses/laguna-s-2.1-int4-b70-width12-dflash-fp8-102.971tok-20260726.response.json.