This Qwen3.6 35B Quark INT8 lane should now be treated as a reference, not the main active optimization target. Move to another model unless doing a controlled upstream/runtime bakeoff or a deliberate speculative-state engineering project.
../../repro/minimax-m27-b70-110tps-ubuntu24-20260523/
and ../../docs/current-reproducibility-map.md.q4_0-gguf-2026-05-03-four-b70-sycl.md,
q4_0-gguf-2026-05-04-sycl-single-kernel-allreduce.md,
and fp8-vllm-xpu-qwen36-2026-05-04.md.../../experiments/gemma4-12b-int4-autoround-vllm/.The current repo has solid Gemma 4 12B TP4 material, not a validated Gemma 35B TP4 recipe. For Gemma-family or other large TP4 work, carry over these limits:
results-20260607-production-c8-xpugraph.json.780 tok/s class for Gemma 4
12B, with LocalMaxxing-approved c8 records.849.59 tok/s) but did not
improve near-32K throughput; it is research-only.UR_RESULT_ERROR_OUT_OF_RESOURCES and
UR_RESULT_ERROR_DEVICE_LOST. KV estimates alone were not sufficient for
production promotion.Useful Gemma references:
Gemma experiment READMEGemma production c8 XPU graph resultGemma c10/c12 32K boundary2026-06-06 Qwen35/Gemma candidates