The primary Q8/INT8-quality lane appears plateaued:
data/gemma4-q8-gpu3-mtp-n7-directunroll7-qonly-b1024u1024-th8-full-20260624T135701Z/;98.4913689785019 tok/s after TTFT on row 0,
cached_tokens=0;384/384;gemma-4-26B-A4B-it-UD-Q8_K_XL.gguf;MTP/gemma-4-26B-A4B-it-Q4_0-MTP.gguf.Small Q8 MTP knobs, draft quant variants, Q8_0 target, Vulkan, vLLM INT8 smoke,
direct argmax IDs, device hidden-state handoff, and direct unroll were explored
in 20260624T0235-mtp-fused-unroll-feasibility.md; direct-unroll7 plus
q-only assistant inputs advanced the Q8 fresh record from 95.263 to
96.822 tok/s, and a later BATCH_SIZE=1024 / UBATCH_SIZE=1024 /
THREADS=8 tune advanced it again to 98.491 tok/s; this remains far below
the >150 tok/s research target.
Later update: the Q8 target/verifier lane subsequently advanced to
103.2992004295621 tok/s fresh row0 with selected-softmax/weighted-sum MoE
guards (cmqsylo2l011nqr011yydjvne). Keep the QAT/Q4XL results in this note
as lower-precision side-lane evidence, not as the Q8 headline.
This side lane tests whether the QAT-trained 4-bit family can deliver a major fresh-response speed jump while preserving acceptable quality. It is not the default Q8/INT8 headline lane.
Only row 0 of a no-cache benchmark can be used as a headline unless every row is a distinct fresh prompt with no usable prior generated continuation. Repeated prompt warmed/history throughput, n-gram continuation learning, prefix/cache reuse, response reuse, and averaged repeated-output speedups are not valid fresh-response headline numbers.
For this lane, every candidate must record:
tok_s_after_ttft;cached_tokens from the benchmark response usage;Repository:
unsloth/gemma-4-26B-A4B-it-qat-GGUF
Files being downloaded to:
/mnt/fast-ai/llm-models/gemma4-26b-a4b-it-qat-gguf/gemma-4-26B-A4B-it-qat-UD-Q4_K_XL.gguf
/mnt/fast-ai/llm-models/gemma4-26b-a4b-it-qat-gguf/MTP/gemma-4-26B-A4B-it-Q4_0-MTP.gguf
Checksums after download:
dcf179a91153e3a7ece792e48ef872180d9d6ef9b7677f0a0bd3e83cfe624d5e gemma-4-26B-A4B-it-qat-UD-Q4_K_XL.gguf
62bd3af7f66c9308de9a5454233852f8c7324c93767e8dfb824ed45b9179864a MTP/gemma-4-26B-A4B-it-Q4_0-MTP.gguf
Planned first smoke, once the files are complete:
UD-Q4_K_XL;Q4_0-MTP;n_max=7, n_min=2, p_min=0.12, backend sampling off;GGML_SYCL_ENABLE_VMM=0, UBATCH_SIZE=512, POLL=100, flash attention off,
--ctx-checkpoints 0;BENCH_PROMPT_MODE=filled-long, PROMPT_TOKENS=512, MAX_TOKENS=512;CANARY_REPEATS=128, BENCH_REPEATS=1 for the first fresh smoke.Promotion rule: even if this exceeds the Q8 record, treat it as QAT/Q4 side-lane throughput until a separate quality decision says it is acceptable to publish as a quality-equivalent result. Do not submit it to LocalMaxxing as the Q8/INT8 lane.
For this QAT lane, each run is pinned with ONEAPI_DEVICE_SELECTOR=level_zero:N.
After that selector is applied, llama.cpp sees the selected physical GPU as
SYCL0. Therefore replica runs on physical GPU 1/2/3 must still use:
LLAMA_DEVICES=SYCL0
--spec-draft-device SYCL0
Using SYCL1, SYCL2, or SYCL3 inside the filtered process produced
invalid device failures.
All results below are fresh-response row-0 numbers unless noted. cached_tokens
was 0 for promoted comparisons. Output SHA matched the current deterministic
reference (d3236ebed08dda8f19a0fec78622967b3622704da06819c98a7e0e63f90d982b)
for passing canary runs.
Baseline target-only:
data/gemma4-qatq4xl-gpu0-baseline-smoke-20260624T110634Z/128/128tok_s_after_ttft: 60.25923650182368654.348825368384645, TTFT: 0.9240040310251061MTP first screen:
data/gemma4-qatq4xl-gpu1-mtp-n7-q40-smoke2-20260624T110741Z/
128/128tok_s_after_ttft: 131.58998767789097107.9229990859438, TTFT: 0.8532496340048965>150 tok/s targetdata/gemma4-qatq4xl-gpu2-mtp-n7-directids-deviceh-smoke2-20260624T110741Z/
128/128tok_s_after_ttft: 128.91546669424767data/gemma4-qatq4xl-gpu3-mtp-n10-q40-smoke2-20260624T110741Z/
128/128tok_s_after_ttft: 66.8693327163563n=10 costs too much on this targetAcceptance detail from server logs:
446/454, mean length 7.86;
cumulative mean accepted length 6.58463/480, mean length 10.65;
cumulative mean accepted length 7.59, but draft duration rose enough to
erase the acceptance gainCANARY_REPEATS=64, one fresh p512/o512 row:
data/gemma4-qatq4xl-gpu0-mtp-n5-q40-depthsmoke-20260624T111316Z/
256/256, fresh 119.81185100696993data/gemma4-qatq4xl-gpu1-mtp-n6-q40-depthsmoke-20260624T111316Z/
256/256, fresh 124.94807934521532data/gemma4-qatq4xl-gpu2-mtp-n8-q40-depthsmoke-20260624T111316Z/
256/256, fresh 58.307855554435925data/gemma4-qatq4xl-gpu3-mtp-n9-q40-depthsmoke-20260624T111316Z/
256/256, fresh 62.6342644516168Conclusion: n7 remains the best depth. More depth improves nominal acceptance on some requests, but the extra draft work dominates.
data/gemma4-qatq4xl-gpu0-mtp-n7-q8kv-smoke-20260624T111646Z/
V cache quantization requires flash_attn, followed by
segfault/core dumpdata/gemma4-qatq4xl-gpu1-mtp-n7-faon-smoke-20260624T111647Z/
256/256, fresh 128.95397328567188BATCH_SIZE=1024, UBATCH_SIZE=1024:
data/gemma4-qatq4xl-gpu2-mtp-n7-b1024ub1024-smoke-20260624T111647Z/
256/256, fresh 130.72860282579887data/gemma4-qatq4xl-gpu3-mtp-n7-q8kv-faon-smoke-20260624T111647Z/
256/256, fresh 117.58276876994238h_nextn PatchThis used LLAMA_MTP_DRAFT_DIRECT_ARGMAX_IDS=1 and
LLAMA_MTP_DRAFT_DIRECT_ARGMAX_UNROLL=N on the QAT n7 config.
data/gemma4-qatq4xl-gpu0-mtp-n7-directunroll2-smoke-20260624T111929Z/
256/256, fresh 88.79800191427907data/gemma4-qatq4xl-gpu1-mtp-n7-directunroll3-smoke-20260624T111929Z/
256/256, fresh 98.52502086073136data/gemma4-qatq4xl-gpu2-mtp-n7-directunroll4-smoke-20260624T111929Z/
256/256, fresh 107.72326757594458data/gemma4-qatq4xl-gpu3-mtp-n7-directunroll5-smoke-20260624T111929Z/
256/256, fresh 116.44839253254389Conclusion: direct unroll had the right shape but stayed below scalar QAT n7.
The next source patch under validation skips the final unused h_nextn
projection/export inside the direct-unroll graph. This is intended to reduce
one nextn_proj_post matmul on the terminal unroll step without changing
sampled token outputs.
h_nextn Patch ResultPatch artifact:
patches/gemma4-26b-a4b-q8-b70/20260624T1120-llamacpp-gemma4-assistant-directunroll-skip-final-hnext.patch
sha256: 184323a276a892d2750124bcbd92b5ba6d982e13d530bacb67337b193538bd5f
The patch avoids computing/exporting h_nextn for the terminal direct-unroll
step in src/models/gemma4-assistant.cpp. It is semantically scoped to
LLAMA_MTP_DRAFT_DIRECT_ARGMAX_UNROLL > 1.
Post-patch results:
data/gemma4-qatq4xl-gpu0-mtp-n7-directunroll4-skipfinalh-smoke/
108.6551909635287data/gemma4-qatq4xl-gpu1-mtp-n7-directunroll5-skipfinalh-smoke/
116.84653883707952data/gemma4-qatq4xl-gpu2-mtp-n7-directunroll6-skipfinalh-smoke/
123.54119385339989data/gemma4-qatq4xl-gpu3-mtp-n7-directunroll7-skipfinalh-smoke/
ggml_backend_sched_reserveConclusion: valid loss. The patch is directionally positive but too small and
direct-unroll remains below scalar QAT n7 (131.58998767789097). Preserve the
patch artifact for reference, but do not promote it as the active optimization.
After direct-unroll lost, a scalar QAT n7 control and three low-risk
argmax/verify variants were tested. CANARY_REPEATS=64, one fresh
p512/o512 row.
data/gemma4-qatq4xl-gpu0-mtp-n7-scalar-control-argmaxscreen-20260624T112850Z/
256/256131.70696381526648LLAMA_MTP_DRAFT_BACKEND_ARGMAX=1:
data/gemma4-qatq4xl-gpu1-mtp-n7-draftbackendargmax-argmaxscreen-20260624T112850Z/
256/256130.28549558899817LLAMA_SPEC_VERIFY_GREEDY_ARGMAX=1:
data/gemma4-qatq4xl-gpu2-mtp-n7-verifygreedy-argmaxscreen-20260624T112850Z/
256/256130.0338863476161LLAMA_MTP_DRAFT_DIRECT_ARGMAX_IDS=1:
data/gemma4-qatq4xl-gpu3-mtp-n7-directids-argmaxscreen-20260624T112850Z/
256/256128.1626864461717Conclusion: no argmax/verify shortcut beats scalar QAT n7. The scalar control
reproduced the current best at 131.7 tok/s, but the >150 tok/s target needs
a real reduction in draft/verify work, not another sampler knob.
A second patch increased graph_max_nodes() only for Gemma4 assistant MTP direct
argmax unroll (ctx_type == MTP, direct argmax IDs enabled, unroll > 1). This
tested the hypothesis that unroll7 was failing because scheduler graph reserve
was undersized, not because the optimization was semantically invalid.
Result: the patch fixed the crash, but did not improve throughput enough.
data/gemma4-qatq4xl-gpu0-mtp-n7-scalar-control-reservepatch-20260624T113221Z/
256/256 rows132.08433147783174data/gemma4-qatq4xl-gpu1-mtp-n7-directunroll6-reservepatch-20260624T113221Z/
256/256 rows123.31885614515221data/gemma4-qatq4xl-gpu2-mtp-n7-directunroll7-reservepatch-20260624T113221Z/
256/256 rows129.0966695597713data/gemma4-qatq4xl-gpu3-mtp-n7-directunroll7-deviceh-reservepatch-20260624T113221Z/
256/256 rows128.61442742132576Conclusion: graph-reserve patch is useful as a bug fix for testing unroll7, but
direct unroll still loses to scalar QAT n7. Do not promote it as a performance
win. The next useful step is phase profiling around scalar QAT n7 to identify
where the ~132 -> >150 gap remains.
Profile run:
data/gemma4-qatq4xl-gpu0-mtp-n7-scalar-profile-20260624T113430Z/
server log: /mnt/fast-ai/bench-results/gemma4-26b-a4b-q8/servers/gemma4-qatq4xl-gpu0-mtp-n7-scalar-profile-20260624T113430Z.server.log
Canary passed (32 rows in the short profile run) and fresh throughput was
133.1792477533177, so profiling did not materially distort the result.
Profile totals at the end of the run:
mtp decode phase profile:
calls=904, tokens=905, total_ms=864.874, per_call_ms=0.957
process_ubatch_ms=418.935
post_extract_ms=430.189
balloc_ms=2.367, sched_reserve_ms=0.713, init_batch_ms=11.280
draft-mtp:
draft_decode_ms=854.781
fast_scan_ms=79.482
hidden_get_ms=0.459
handoff_ms=0.755
draft_decodes=903
Interpretation: the scalar QAT n7 draft path is split roughly in half between
assistant graph compute/submit (process_ubatch) and output extraction
(post_extract). Sampler/top-k and hidden-row bookkeeping are already small.
The remaining >150 tok/s gap likely needs lower-level output extraction or
draft graph work reduction, not another sampler knob.
After adding finer extraction counters, scalar QAT n7 showed:
data/gemma4-qatq4xl-gpu0-mtp-n7-scalar-profile-split-20260624T114353Z/
fresh: 134.46824800804353
post_extract_ms=344.677
logits_extract_ms=338.839
h_nextn_extract_ms=5.695
sampled_extract_ms=0.021
So full raw-logit extraction is the dominant post-extract cost. Direct IDs were profiled as a comparison:
data/gemma4-qatq4xl-gpu0-mtp-n7-directids-profile-split-20260624T114443Z/
fresh: 131.3893864046233
post_extract_ms=516.040
logits_extract_ms=0.009
h_nextn_extract_ms=512.904
sampled_extract_ms=3.007
This means the direct-ID graph path removes host logits transfer, but the graph work / first remaining extraction waits longer overall. Do not pursue direct IDs further until there is a better on-device argmax/top1 kernel.
Device hidden-state handoff alone was then tested without direct IDs:
data/gemma4-qatq4xl-gpu0-mtp-n7-deviceh-scalar-smoke-20260624T113647Z/
256/256 rows133.13947504649965This is a small valid win over prior scalar controls, but still far below
>150 tok/s. It suggests hidden-state host extraction is not the main blocker;
the expensive part is broader output extraction / decode envelope.
Split profile with device-H:
data/gemma4-qatq4xl-gpu0-mtp-n7-deviceh-profile-split-20260624T114555Z/
fresh: 133.70532508555132
post_extract_ms=334.508
logits_extract_ms=334.344
h_nextn_extract_ms=0.020
Device-H removes host h_nextn extraction, but logits extraction remains the
dominant cost.
Device-H depth sweep:
data/gemma4-qatq4xl-gpu0-mtp-n6-deviceh-depthsweep-20260624T113843Z/
256/256 rows126.29351651105031data/gemma4-qatq4xl-gpu1-mtp-n7-deviceh-depthsweep-20260624T113843Z/
256/256 rows130.92546265828116data/gemma4-qatq4xl-gpu2-mtp-n8-deviceh-depthsweep-20260624T113843Z/
256/256 rows58.27021144392839data/gemma4-qatq4xl-gpu3-mtp-n9-deviceh-depthsweep-20260624T113843Z/
256/256 rows62.641875750565234Conclusion: device-H does not make deeper MTP viable. n7 remains the only useful depth.
Draft quantization screen with QAT target + device-H:
data/gemma4-qatq4xl-gpu0-mtp-n7-deviceh-draftq2k-20260624T114712Z/
256/256 rows73.70665309458039data/gemma4-qatq4xl-gpu1-mtp-n7-deviceh-draftq3ks-20260624T114712Z/
256/256 rows108.75157605134396data/gemma4-qatq4xl-gpu2-mtp-n7-deviceh-draftq3km-20260624T114712Z/
256/256 rows102.94829765078032data/gemma4-qatq4xl-gpu3-mtp-n7-deviceh-draftqatq40-20260624T114712Z/
256/256 rows131.48596740753342Conclusion: smaller Q2/Q3 draft files do not produce faster draft kernels here;
acceptance/cost balance is much worse. Keep the QAT Q4_0-MTP draft.
Draft-only KV cache type screen with QAT target + device-H:
q8_0/q8_0, flash attention off:
data/gemma4-qatq4xl-gpu0-mtp-n7-deviceh-draftkvq8off-20260624T115007Z/
256/256 rows60.72740408065323q4_0/q4_0, flash attention off:
data/gemma4-qatq4xl-gpu1-mtp-n7-deviceh-draftkvq4off-20260624T115007Z/
256/256 rows60.63701450277927q8_0/q8_0, flash attention on:
data/gemma4-qatq4xl-gpu2-mtp-n7-deviceh-draftkvq8fa-20260624T115007Z/
256/256 rows129.62435631697284q4_0/q4_0, flash attention on:
data/gemma4-qatq4xl-gpu3-mtp-n7-deviceh-draftkvq4fa-20260624T115007Z/
256/256 rows129.7760268775364Conclusion: draft KV quantization is not useful here. Without FA it collapses; with FA it remains below f16 draft KV. Keep draft KV at f16.
Fast top-k confidence screen:
Earlier p_min sweeps with LLAMA_MTP_DRAFT_FAST_ARGMAX=1 were not real
confidence-stopping tests because the fast-argmax path assigns p=1.0, so
p_min never stops drafting. A real fast-top-k screen was run with device-H:
p_min=0.12:
data/gemma4-qatq4xl-gpu0-mtp-n7-deviceh-fasttopk2-pmin012-20260624T115253Z/
256/256 rows130.29429192620302p_min=0.12:
data/gemma4-qatq4xl-gpu1-mtp-n7-deviceh-fasttopk4-pmin012-20260624T115253Z/
256/256 rows129.39971464224635p_min=0.12:
data/gemma4-qatq4xl-gpu2-mtp-n7-deviceh-fasttopk10-pmin012-20260624T115253Z/
256/256 rows130.08465611052088p_min=0.20:
data/gemma4-qatq4xl-gpu3-mtp-n7-deviceh-fasttopk4-pmin020-20260624T115253Z/
256/256 rows129.30564777627023Conclusion: top-k confidence stopping is slower than fast argmax on this stack. The extra top-k/probability work costs more than it saves by stopping marginal draft positions.
Poll sweep on current scalar/device-H config:
POLL=25:
data/gemma4-qatq4xl-gpu0-mtp-n7-deviceh-poll25-20260624T115512Z/
256/256 rows131.1970643804725POLL=50:
data/gemma4-qatq4xl-gpu1-mtp-n7-deviceh-poll50-20260624T115512Z/
256/256 rows131.17397959625876POLL=75:
data/gemma4-qatq4xl-gpu2-mtp-n7-deviceh-poll75-20260624T115512Z/
256/256 rows132.2933221288594POLL=100:
data/gemma4-qatq4xl-gpu3-mtp-n7-deviceh-poll100-20260624T115512Z/
256/256 rows131.18864073147495Conclusion: flat/no win. Keep POLL=100 or 75; neither changes the plateau.
Thread sweep on current scalar/device-H config:
data/gemma4-qatq4xl-gpu0-mtp-n7-deviceh-th8-dt32-dtb32-20260624T115725Z/
256/256 rows132.2556766085762data/gemma4-qatq4xl-gpu1-mtp-n7-deviceh-th24-dt32-dtb32-20260624T115725Z/
256/256 rows131.94871682863848data/gemma4-qatq4xl-gpu2-mtp-n7-deviceh-th16-dt16-dtb16-20260624T115725Z/
256/256 rows130.50890944113308data/gemma4-qatq4xl-gpu3-mtp-n7-deviceh-th16-dt48-dtb48-20260624T115725Z/
256/256 rows132.05585283370584Conclusion: flat/no win. Keep the prior THREADS=16,
--spec-draft-threads 32, --spec-draft-threads-batch 32 recipe.
F16 draft-logit extraction patch screen:
data/gemma4-qatq4xl-gpu0-mtp-n7-scalarcontrol-20260624T120354Z/
256/256 rows131.99366752340984data/gemma4-qatq4xl-gpu1-mtp-n7-f16logits-20260624T120354Z/
256/256 rows124.33233152510599data/gemma4-qatq4xl-gpu2-mtp-n7-f16logits-deviceh-20260624T120354Z/
256/256 rows124.01156749684175POLL=75:
data/gemma4-qatq4xl-gpu3-mtp-n7-f16logits-deviceh-poll75-20260624T120354Z/
256/256 rows125.11164247060259Conclusion: failed optimization. The patch reduced the draft-logit transfer width, but the host-side fp16 conversion / scan path costs more than it saves. It is canary-clean but slower than the scalar/device-H baseline, so keep it as a negative artifact and do not promote it.
Patch artifact:
patches/gemma4-26b-a4b-q8-b70/20260624T1215-llamacpp-gemma4-mtp-f16-logits-extract-current.patch9fcc6d12e6cd4b55b541351243310bf53d6f0de9947a794e143b0ebbf1310c7aContext / batch shape sweep on current scalar/device-H config:
ctx=4096, batch=512, ubatch=512:
data/gemma4-qatq4xl-gpu0-mtp-n7-deviceh-ctx4096-b512u512-20260624T120810Z/
384/384 rows132.0164670122808, cached_tokens=0ctx=2048, batch=512, ubatch=512:
data/gemma4-qatq4xl-gpu1-mtp-n7-deviceh-ctx2048-b512u512-20260624T120810Z/
384/384 rows130.9805435084394, cached_tokens=0ctx=4096, batch=256, ubatch=256:
data/gemma4-qatq4xl-gpu2-mtp-n7-deviceh-ctx4096-b256u256-20260624T120810Z/
384/384 rows130.80932646728493, cached_tokens=0ctx=2048, batch=1024, ubatch=1024:
data/gemma4-qatq4xl-gpu3-mtp-n7-deviceh-ctx2048-b1024u1024-20260624T120810Z/
384/384 rows131.750767610216, cached_tokens=0Conclusion: no decode-throughput win. Smaller ctx can reduce TTFT in one
shape (ctx=2048, batch=1024, ubatch=1024 got ttft=0.665s), but the
fresh decode plateau remains ~131-132 tok/s. Keep ctx=8192, batch=512,
ubatch=512 for comparability unless optimizing TTFT separately.
Source-audit follow-up: redundant syncs in scalar fast argmax.
Explorer audit found the current scalar fast-argmax path was doing more synchronization than intended:
draft_fast_argmax_sample() explicitly calls llama_synchronize(ctx_dft)
after llama_decode();llama_get_logits_f16_ith(), whose public wrapper also
synchronizes, even when f16 draft-logit extraction is disabled;llama_get_logits_ith(), whose public wrapper synchronizes
again.Patch under test:
src/llama-ext.h /
src/llama-context.cpp;LLAMA_MTP_DRAFT_F16_LOGITS_EXTRACT, and skip the f16 probe
unless that flag is enabled;Correctness expectation: unchanged draft tokens and verifier behavior; this only removes redundant synchronization / dead f16 probing after the decode has already been synchronized.
Four-way screen launched:
Q4_0-MTP draft, no-sync patch, profile enabled;Q4_0-MTP draft;Q5_K_M-MTP draft;Q6_K-MTP draft.Result:
Q4_0-MTP draft, profile enabled:
data/gemma4-qatq4xl-gpu0-mtp-n7-nosync-qatq40-profile-20260624T122327Z/
256/256 rows, cached_tokens=0131.97589151993552logits_extract_ms=1665.110, fast_scan_ms=356.506,
fast_sync_ms=12.869Q4_0-MTP draft:
data/gemma4-qatq4xl-gpu1-mtp-n7-nosync-q8repo-q40-20260624T122327Z/
256/256 rows, cached_tokens=0114.17324125727583Q5_K_M-MTP draft:
data/gemma4-qatq4xl-gpu2-mtp-n7-nosync-q8repo-q5km-20260624T122327Z/
256/256 rows, cached_tokens=0108.13964245947918Q6_K-MTP draft:
data/gemma4-qatq4xl-gpu3-mtp-n7-nosync-q8repo-q6k-20260624T122327Z/
256/256 rows, cached_tokens=0108.95410818516535Conclusion: no-sync is canary-clean but not a win; non-QAT repo MTP drafts are significantly slower when paired with the QAT target. The redundant public wrapper syncs were not the bottleneck at the current plateau; draft logits extraction and host scan remain the dominant cost.
Patch artifact:
patches/gemma4-26b-a4b-q8-b70/20260624T1235-llamacpp-gemma4-mtp-nosync-argmax512-loss-current.patch82540145978fb4fca0e293ae281f49c763b5d5d320cd9e1cd2107dc1cf7a90a7Follow-up screen after rebuilding with the SYCL_ARGMAX_BLOCK_SIZE=512 change
and no-sync scalar accessors:
data/gemma4-qatq4xl-gpu0-mtp-n7-l0copy-20260624T-cont/
256/256 rows, cached_tokens=0132.9256806633963133.13947504649965);
the Level Zero copy knobs do not break the plateau.data/gemma4-qatq4xl-gpu1-mtp-n7-directids-argmax512-20260624T-cont/
256/256 rows, cached_tokens=0129.5679255672916data/gemma4-qatq4xl-gpu2-mtp-n7-directunroll7-argmax512-20260624T-cont/
GGML_ASSERT(i01 >= 0 && i01 < ne01) failed in
ggml_compute_forward_get_rows()data/gemma4-qatq4xl-gpu3-mtp-n7-directids-l0copy-20260624T-cont/
invalid token[0] = 1101530873 / Invalid input batchConclusion from this follow-up: the current plateau is not caused by redundant
public API syncs, Level Zero copy-engine policy, or the SYCL argmax block size.
The valid fresh-response QAT/Q4 path remains the scalar fast-argmax +
device-H-handoff lane at about 133 tok/s. The unstable direct-ID/unroll
routes should stay as negative artifacts unless a deeper sampled-token buffer
ownership/indexing bug is fixed.
Small scalar/device-H parameter sweep, keeping the stable path and changing only draft minimum length or draft thread counts:
--spec-draft-n-min 1:
data/gemma4-qatq4xl-gpu0-mtp-n7-nmin1-20260624T123353Z/
256/256 rows, cached_tokens=0132.43790195100183--spec-draft-n-min 3:
data/gemma4-qatq4xl-gpu1-mtp-n7-nmin3-20260624T123353Z/
256/256 rows, cached_tokens=0130.809314034844658/8:
data/gemma4-qatq4xl-gpu2-mtp-n7-dt8-20260624T123353Z/
256/256 rows, cached_tokens=0131.9048123904550264/64:
data/gemma4-qatq4xl-gpu3-mtp-n7-dt64-20260624T123353Z/
256/256 rows, cached_tokens=0131.8125776077952Conclusion: no scalar/runtime parameter win. n_min=2, draft threads
32/32, POLL=100, ctx=8192, batch=512, ubatch=512, FA off remains the
best validated scalar/device-H shape for this side lane.
Separate fresh draft-model screen with downloaded E4B Q4_0 GGUF:
/mnt/fast-ai/llm-models/gemma4-e4b-it-gguf/gemma-4-E4B-it-Q4_0.gguf
d51af14007fe16deb45107de2411660be3b59c9e1c89bbde3bfec8ece65eeedcunsloth/gemma-4-E4B-it-GGUF, file gemma-4-E4B-it-Q4_0.ggufUD-Q4_K_XL 26B, so output quality is still
verifier-bound; the E4B model is only a fresh draft source.draft-simple, E4B Q4_0, n_max=4:
data/gemma4-qatq4xl-gpu0-draftsimple-e4bq40-n4-20260624T124516Z/
128/128, cached_tokens=048.44968921986737draft-simple, E4B Q4_0, n_max=8:
data/gemma4-qatq4xl-gpu1-draftsimple-e4bq40-n8-20260624T124516Z/
128/128, cached_tokens=020.8685271235626draft-simple, E4B Q4_0, n_max=12:
data/gemma4-qatq4xl-gpu2-draftsimple-e4bq40-n12-20260624T124516Z/
128/128, cached_tokens=026.133640247024097draft-simple, E4B Q4_0, n_max=16:
data/gemma4-qatq4xl-gpu3-draftsimple-e4bq40-n16-20260624T124516Z/
128/128, cached_tokens=030.038289500818628Conclusion: true separate-draft speculation is valid fresh-response work, but
not useful here. llama.cpp draft-simple repeatedly decodes the full E4B
draft model and samples through the generic draft-simple path, so it is far
slower than the tiny Gemma4 MTP assistant. Do not pursue E4B/12B draft-simple
unless the draft-simple implementation itself is redesigned.
SYCL argmax512 consistency fix:
SYCL_ARGMAX_BLOCK_SIZE had been changed from 256 to 512, but
the SYCL argmax kernel still used hard-coded 256 for local memory, column
stride, and reduction stride. That made the direct-ID argmax512 tests
invalid/fragile and explained the earlier invalid sampled tokens.ggml/src/ggml-sycl/ggml-sycl.cpp now uses
SYCL_ARGMAX_BLOCK_SIZE consistently in the argmax kernel.patches/gemma4-26b-a4b-q8-b70/20260624T1305-llamacpp-gemma4-sycl-argmax512-consistency-current.patch
1e9ef3df311ee194540a8912f84963795baa0f56e23707c8d2edb3a0797f4f7dPost-fix retest:
data/gemma4-qatq4xl-gpu0-mtp-n7-scalar-argmaxfix-20260624T125911Z/
256/256, cached_tokens=0132.04618692363044data/gemma4-qatq4xl-gpu1-mtp-n7-directids-argmaxfix-20260624T125911Z/
256/256, cached_tokens=0131.70152035416316data/gemma4-qatq4xl-gpu2-mtp-n7-directids-l0copy-argmaxfix-20260624T125911Z/
256/256, cached_tokens=0130.44999768038858data/gemma4-qatq4xl-gpu3-mtp-n7-directunroll7-argmaxfix-20260624T125911Z/
256/256, cached_tokens=0132.36125747200137Conclusion: the argmax fix is a real correctness cleanup (direct-ID/unroll no longer crash), but not a throughput win. Backend argmax remains slightly slower than exporting logits and doing the scalar host scan on this workload.
Post-fix direct-unroll depth sweep:
data/gemma4-qatq4xl-gpu0-mtp-n7-directunroll2-argmaxfix-20260624T130153Z/
256/256, cached_tokens=090.48132775600303data/gemma4-qatq4xl-gpu1-mtp-n7-directunroll3-argmaxfix-20260624T130153Z/
256/256, cached_tokens=099.98989486911354data/gemma4-qatq4xl-gpu2-mtp-n7-directunroll4-argmaxfix-20260624T130153Z/
256/256, cached_tokens=0109.54124212289017data/gemma4-qatq4xl-gpu3-mtp-n7-directunroll5-argmaxfix-20260624T130153Z/
256/256, cached_tokens=0119.99443045776854Conclusion: no direct-unroll win. The deeper direct-unroll graph must amortize
more MTP steps before it approaches scalar, but even unroll7 post-fix only
reached 132.36125747200137, still below the scalar/device-H best.
SYCL argmax1024 retest:
After the argmax kernel consistency fix, SYCL_ARGMAX_BLOCK_SIZE was raised
from 512 to 1024 and the four most relevant lanes were rebuilt/retested.
This is a valid test of the larger-block argmax kernel, not a Q8 headline
result. CANARY_REPEATS=64; all throughput values below are fresh row-0
p512/o512 with cached_tokens=0.
data/gemma4-qatq4xl-gpu0-mtp-n7-scalar-argmax1024-20260624T1316Z/
256/256132.54422874725984, wall 108.48264126608782,
TTFT 0.8567875849839766data/gemma4-qatq4xl-gpu1-mtp-n7-directids-argmax1024-20260624T1316Z/
256/256131.32701664309184, wall 107.57920512289195,
TTFT 0.8606194260064512data/gemma4-qatq4xl-gpu2-mtp-n7-directids-l0copy-argmax1024-20260624T1316Z/
256/256132.84141637936202, wall 108.72100597565084,
TTFT 0.8550818619842175data/gemma4-qatq4xl-gpu3-mtp-n7-directunroll7-argmax1024-20260624T1316Z/
256/256133.27420695985995, wall 109.09064377070227,
TTFT 0.8516411530144978Conclusion: argmax1024 is canary-clean and produces the best measured QAT/Q4
side-lane row so far (133.27420695985995), but the delta over the previous
best (133.13947504649965) is only +0.1347 tok/s and should be treated as
noise-level/marginal unless repeated. It does not change the main conclusion:
the QAT/Q4 lane remains below the >150 tok/s target, and it is not the
primary Q8/INT8-quality headline lane.
Patch artifact:
patches/gemma4-26b-a4b-q8-b70/20260624T1325-llamacpp-gemma4-sycl-argmax1024-current.patch090d2118eeafd484a3215077d9957c2d20772dc7b7ff8b55a6235019ec40d7f0Gemma4Assistant q-only attention input patch:
Patch idea: Gemma4Assistant uses shared target K/V and calls
build_attn(..., k_cur=nullptr, v_cur=nullptr, ...), but the generic ISWA
input builder still creates and uploads K/V index tensors. An env-gated patch
(LLAMA_GEMMA4_MTP_QONLY_ATTN_INPUTS=1) nulls those unused K/V index inputs
for the assistant graph while preserving masks and RoPE rotation inputs.
Patch artifact:
patches/gemma4-26b-a4b-q8-b70/20260624T1400-llamacpp-gemma4-mtp-qonly-attn-inputs-current.patch7ebb524016b2a63a8cb8130f8fc861d24a4a35a164f588e26b5622d2ded0430dQ8-quality screen:
data/gemma4-q8-gpu0-mtp-n7-control-qonlyscreen-20260624T1344Z/
256/256, cached_tokens=095.1198102331274data/gemma4-q8-gpu1-mtp-n7-qonly-20260624T1344Z/
256/256, cached_tokens=095.04517249053589Early conclusion for Q8: q-only alone was not a win in this scalar screen.
Later, q-only combined with direct-unroll7 did become the current promoted Q8
record (98.491 tok/s, cmqs56wv100kjqr01de3fdspd). Keep this distinction:
the patch is useful as part of the direct-unroll stack, not as a standalone
scalar-path improvement.
QAT/Q4 screen:
data/gemma4-qatq4xl-gpu2-mtp-n7-qonly-20260624T1344Z/
256/256, cached_tokens=0131.76819466381565data/gemma4-qatq4xl-gpu3-mtp-n7-directunroll7-qonly-20260624T1344Z/
256/256, cached_tokens=0134.37693544373184Follow-up around direct-unroll7 q-only:
data/gemma4-qatq4xl-gpu0-mtp-n7-directunroll7-qonly-repeat-20260624T1351Z/
256/256, cached_tokens=0133.25860672997848data/gemma4-qatq4xl-gpu1-mtp-n7-directunroll7-qonly-l0copy-20260624T1351Z/
256/256, cached_tokens=0133.62711777209057POLL=75:
data/gemma4-qatq4xl-gpu2-mtp-n7-directunroll7-qonly-poll75-20260624T1351Z/
256/256, cached_tokens=0133.63996044381167BATCH_SIZE=1024, UBATCH_SIZE=1024:
data/gemma4-qatq4xl-gpu3-mtp-n7-directunroll7-qonly-b1024-20260624T1351Z/
256/256, cached_tokens=0134.21956263064533Conclusion: q-only attention inputs are canary-clean and give the best measured
QAT/Q4 side-lane row so far (134.37693544373184), but the identical repeat
fell to 133.25860672997848, so treat this as a marginal/noisy side-lane
improvement, not a stable breakthrough. It does not solve the >150 tok/s
fresh-response target and does not help the Q8-quality lane.
Single-server confirmation:
The four-way sweeps above run four llama-server processes concurrently and can
understate single-session throughput because each server requests substantial
CPU-side runtime support (THREADS=16, HTTP workers, and SYCL host work) on the
same 16-core host. A single-server confirmation of the best QAT/Q4 shape:
data/gemma4-qatq4xl-gpu0-mtp-n7-directunroll7-qonly-profile-20260624T1404Z/
64/64, cached_tokens=0136.3354126697164draft_decode_ms=1190.012, fast_scan_ms=0,
sampled_extract_ms=83.269, mean accepted length 6.96data/gemma4-qatq4xl-gpu0-mtp-n7-directunroll7-qonly-single-20260624T1410Z/
256/256, cached_tokens=0136.05802749417853, wall 111.29408723654275,
TTFT 0.8373238200147171data/gemma4-qatq4xl-gpu0-mtp-n7-scalar-qonly-single-20260624T1423Z/
256/256, cached_tokens=0133.93770311710102, wall 109.8640310483012,
TTFT 0.8376332809857558Conclusion update: use the single-server no-profile row (136.05802749417853)
as the best solid QAT/Q4 side-lane result. Use four-way sweeps for fast research
screening, but confirm any candidate alone before making a headline throughput
claim.
Runtime screen around q-only direct-unroll7:
data/gemma4-qatq4xl-gpu0-mtp-n7-directunroll7-qonly-faon-20260624T1417Z/
256/256, cached_tokens=0132.257925712836GGML_SYCL_DISABLE_OPT=1:
data/gemma4-qatq4xl-gpu1-mtp-n7-directunroll7-qonly-syclopt1-20260624T1417Z/
256/256, cached_tokens=0105.77855494409444GGML_SYCL_ENABLE_VMM=1:
data/gemma4-qatq4xl-gpu2-mtp-n7-directunroll7-qonly-vmm1-20260624T1417Z/
256/256, cached_tokens=0133.64486218684613THREADS=8:
data/gemma4-qatq4xl-gpu3-mtp-n7-directunroll7-qonly-th8-20260624T1417Z/
256/256, cached_tokens=0132.92761516323216Conclusion: no runtime knob win. The remaining gap is inside assistant graph
execution (process_ubatch_ms), not flash attention, VMM, Level Zero copy
policy, or CPU thread count.
Best QAT/Q4 side-lane result:
data/gemma4-qatq4xl-gpu0-mtp-n7-directunroll7-qonly-single-20260624T1410Z/
fresh tok_s_after_ttft: 136.05802749417853
canary: pass, 256 rows completed
cached_tokens: 0
target: QAT UD-Q4_K_XL
draft: Q4_0-MTP
This is a valid fresh-response result for the QAT/Q4 side lane, but it is not a
Q8/INT8-quality headline and should not be mixed with the primary Q8 record.
It also does not meet the user target of >150 tok/s.