Goal from user:
Initial local findings:
level_zero:0, level_zero:1, level_zero:2, level_zero:3./opt/intel/oneapi/compiler/2026.0/bin/icpx.llama-server, llama-cli, or llama-bench binary was on PATH./home/steve/src/llama.cpp was absent at lane start; a fresh upstream clone is
appropriate./mnt/fast-ai.Initial runtime decision:
UD-Q8_K_XL.gguf first.ONEAPI_DEVICE_SELECTOR=level_zero:N, not
llama.cpp multi-GPU split.google/gemma-4-26B-A4B-it --quantization int8_per_channel_weight_only.External evidence recorded in the result packet:
UD-Q8_K_XL GGUF for this model.--data-parallel-size > 1 crash report; this
supports the four independent replica plan.Setup progress:
/home/steve/src/llama.cpp, commit
dec5ca557.scripts/build-llama-cpp-sycl-b70.sh completed successfully with oneAPI
2026, GGML_SYCL=ON, GGML_SYCL_F16=ON, Level Zero support, oneDNN, and MKL.
Binaries:
/home/steve/src/llama.cpp/build-sycl-b70/bin/llama-server/home/steve/src/llama.cpp/build-sycl-b70/bin/llama-cli/home/steve/src/llama.cpp/build-sycl-b70/bin/llama-benchnpm install hit a local Node-version mismatch, then
upstream’s build fell back to the prebuilt UI archive and completed. This is
not blocking for serving./mnt/fast-ai/llm-models/gemma4-26b-a4b-it-q8-gguf/; no local baseline run
yet because the model file is still downloading.