b70-optimization-lab

Gemma 4 26B A4B LocalMaxxing Targets

Research snapshot: updated 2026-06-30.

This page separates public leaderboard context from this lane’s promoted result rules. The goal is a valid Q8 / INT8-or-better result on Intel Arc Pro B70, not a speed-only lower-precision entry.

Quality/label guardrail: promoted Gemma 26B B70 submissions in this lane must use the literal gemma-4-26B-A4B-it-UD-Q8_K_XL.gguf target/verifier unless the user explicitly accepts a different target quantization. Historical labels like q8target-q40draft are shorthand for a Q8-quality UD-Q8_K_XL target with a Q4_0 MTP draft; they are not claims that the target was gemma-4-26B-A4B-it-Q8_0.gguf. Literal Q8_0.gguf target runs are separate controls and are not LocalMaxxing-promotable under the no-quality-loss rule.

As of 2026-06-27, synthetic/repeated/filled-long benchmark scores are diagnostic only. Do not submit or advertise them as real-world throughput. Promotion and LocalMaxxing submission require the fixed realistic prompt suite, one cold response per prompt, cached_tokens=0 every row, no prompt/KV/context/response reuse or n-gram/history acceleration, verified speculation only, and primary metric median_tok_s_1_100_after_ttft.

Public Target Context

Current public pages are useful as a speed target, but not as direct quality-equivalent comparisons:

Page Current public top context Why it is not directly comparable
google/gemma-4-26B-A4B-it about 87.3 tok/s Rows include mixed engines, hardware, and quantization such as MXFP4/Q4.
unsloth/gemma-4-26B-A4B-it-GGUF about 94.3 tok/s GGUF page, but public top rows are still mixed precision/hardware.
Jackrong/Gemopus-4-26B-A4B-it-GGUF about 94.5 tok/s Fine-tune, useful idea source only; not the same checkpoint.

Interpretation for this lane:

Current policy-compliant LocalMaxxing submission:

Current service/prompt-processing LocalMaxxing submission, separate from the short-decode headline:

Previous policy-compliant LocalMaxxing submission, now superseded:

Previous policy-compliant LocalMaxxing submission, now superseded:

Previous policy-compliant LocalMaxxing submission, now superseded:

The adjacent BF16-direct retest did not beat controls and is recorded as a negative in ../../patches/gemma4-26b-a4b-q8-b70/20260629-selecteddown-bf16direct-currentstack-negative.md.

Previous policy-compliant LocalMaxxing submission, now superseded:

Earlier policy-compliant LocalMaxxing submission, now superseded:

Previous policy-compliant F16-p021 submission, now superseded:

Previous policy-compliant VDR2 submission, now superseded:

Previous policy-compliant VDR2 submission, now superseded:

Previous policy-compliant VDR2 submission, now superseded:

Previous policy-compliant VDR4 submission, now superseded:

Previous realistic-suite local Q8 observation, now superseded:

Current realistic-suite no-spec control:

Current local Q8 baseline:

Historical natural-stop local Q8 best:

Historical short-prompt sustained-decode Q8 best:

Historical short-prompt draft-MTP sustained-decode Q8 best:

Current filled-long draftless ngram-mod warmed/history artifact:

Historical filled-long draft-MTP Q8-target diagnostic best:

Superseded Q8 MoE-ID reorder pre-final-gate diagnostic:

Superseded Q8 MoE-ID reorder pre-final-gate diagnostic:

Earlier superseded Q8 MoE-ID reorder pre-final-gate diagnostic:

Superseded UBATCH_SIZE=768, n_min=3, p_min=0.10 pre-final-gate Q8-target diagnostic:

Superseded UBATCH_SIZE=768 draft-MTP pre-final-gate Q8-target best:

Superseded same-stack filled-long draft-MTP pre-final-gate Q8-target best:

Superseded same-stack filled-long draft-MTP pre-final-gate Q8-target best:

Superseded filled-long draft-MTP route-cache pre-final-gate Q8-target best:

Earlier superseded filled-long draft-MTP route-cache pre-final-gate Q8-target best:

Previous material filled-long draft-MTP pre-final-gate Q8-target best:

Superseded filled-long draft-MTP pre-final-gate Q8-target record:

Superseded safer row-argmax/defer-H pre-final-gate Q8-target record:

Superseded row-argmax/defer-H pre-final-gate Q8-target record:

Previous direct-unroll/q-only batch/thread filled-long draft-MTP Q8-target best:

Previous batch/thread filled-long draft-MTP Q8-target best:

Previous direct-unroll/q-only filled-long draft-MTP Q8-target best:

Previous filled-long draft-MTP pre-final-gate Q8-target best:

Previous filled-long draft-MTP pre-final-gate Q8-draft diagnostic:

Previous filled-long draft-MTP pre-final-gate Q8 diagnostic:

Previous filled-long draft-MTP pre-final-gate Q8 diagnostic:

Previous filled-long draft-MTP sustained-decode Q8 best:

Previous draft-MTP approved result:

Submission Packet

LocalMaxxing requires at minimum:

Useful optional fields for this repo’s records:

The API supports a dry-run endpoint before writing a real benchmark. The local helper reads the key from LMX_API_KEY or /home/steve/.config/localmaxxing/api_key; never put that key in a payload, note, shell history snippet, or commit.

Gemma 4 Payload Shape

For the primary GGUF lane:

hfId: unsloth/gemma-4-26B-A4B-it-GGUF
modelRevision: 3bb10d594514ef4edb7f3a65d41a7e4eb8c5767a
engineName: llama.cpp
backend: sycl/xpu
quantization: UD-Q8_K_XL
hardware.hwClass: DISCRETE_GPU
hardware.gpuName: Intel Arc Pro B70
hardware.vramGb: 32
hardware.gpuCount: 1 for a single-replica record, 4 only for aggregate service records

Engine flags should include the command snippet and the relevant values from the server log:

Do Not Submit If