Disabling Gemma 4 thinking mode should make direct-answer chat canaries return
normal message.content, preserving Q8 quality while establishing the first
valid single-B70 llama.cpp baseline.
/mnt/fast-ai/llm-models/gemma4-26b-a4b-it-q8-gguf/gemma-4-26B-A4B-it-UD-Q8_K_XL.gguf27,636,230,944unsloth/gemma-4-26B-A4B-it-GGUF@3bb10d594514ef4edb7f3a65d41a7e4eb8c5767adec5ca557, SYCL/Level Zerolevel_zero:0, port 182608192512 / 64-fa on, --poll 50, GGML_SYCL_DISABLE_OPT=1off/mnt/fast-ai/bench-results/gemma4-26b-a4b-q8/servers/gemma4-26b-q8-llamacpp-gpu0-ctx8192-20260623T052850Z.server.logn_ctx=8192;n_ctx_train=262144;n_params=25233142046;27620407416;thinking = 0.32 repeats x 4 cases).26.0997 tok/s mean after TTFT;24.2443 tok/s mean wall;1110 completion tokens across 8 requests;0.00028, so steady decode is stable.Win for correctness and fit; not a speed win. Promote as the conservative control baseline and begin four-at-a-time optimization from it.
Do not submit this as a LocalMaxxing record unless explicitly recording a low baseline. The public Gemma 4 family context is much higher, and this lane should try batch/ubatch, SYCL runtime, AOT, vLLM, and later MTP before publishing.
data/gemma4-26b-q8-llamacpp-gpu0-ctx8192-20260623T052850Z/models.jsondata/gemma4-26b-q8-llamacpp-gpu0-ctx8192-20260623T052850Z/chat-canary.jsondata/gemma4-26b-q8-llamacpp-gpu0-ctx8192-20260623T052850Z/p512o512.jsondata/gemma4-26b-q8-llamacpp-gpu0-ctx8192-20260623T052850Z/summary.json