b70-optimization-lab

Performance Index

This is a small, manually maintained index of useful performance expectations. It is not the identity of the project, a universal leaderboard, or a replacement for the linked result packets. Those packets and their JSON/log evidence are the source of truth.

The current local verification platform is Intel Arc Pro B70. Results and patches for other Intel Arc GPUs are welcome, but they remain clearly labeled until matching hardware is available for independent reproduction. Testing a contributor’s patch on B70 validates the resulting B70 behavior; it does not verify a score originally reported on different hardware.

Last manual review: 2026-07-26.

Read This Before Comparing Rows

B70-Verified Results

Model / revision Quantization Hardware / layout Engine / runtime Suite / benchmark shape Result Quality / status Patch, evidence, contributor
poolside/Laguna-S-2.1-INT4
4bbfc285f2f8b3b6b526274c133b7b17aae6c8cb
compressed-tensors INT4 group-32 W4A16 target; matched INT4 DFlash draft with 31 runtime E4M3FN W8A16 projections per rank; BF16 KV; target-verified speculation 4x Arc Pro B70 32 GB, TP4+EP4, one active generation vLLM/XPU e596ef154; XPU kernels 6f9dd3c3a; exact width-12 stack, persistent exact-attention metadata, validated 146/145-segment Breakable PIECEWISE graph Laguna fixed realistic suite; 13 unique cold prompts, output up to 512; first-to-100th token timestamps; greedy canonical-q1 exactness; DFlash depth 11 102.971436 tok/s submitted legacy 100-event/99-interval median; 101.941721 tok/s conventional interval median; legacy p10 71.148884; full wall 52.767621 B70-verified, sealed, metric-qualified; first valid preregistered cold score, 13/13 token-and-text exact and cache-zero, 512-output-then-next 2/2, rollover 1/1, no warmup generation or retry, clean 73-second pre/post idle gates; LocalMaxxing cms2ccv2d00lps201rej94pjy qualified packet; repro; correction; contributor: Steve Seguin
0xSero/DeepSeek-V4-Flash-180B
7c360e1cd4a5168099dbc54d16d929bf6df04990, experimental uniform-K160 artifact
FP8 block-scaled dense weights; FP4 experts; FP8 KV; unchanged K160 target with target-verified DSpark7 4x Arc Pro B70 32 GB, TP4+EP, one active generation vLLM/XPU 264c7f2f7; XPU kernels 313156737; oneCCL 48fda4f0e; target and draft PIECEWISE graphs Fixed realistic suite; 12 unique cold prompts, output 128; median tokens 1-100 after TTFT; target verifier M=8 80.820052 tok/s median high; p10 71.669556; three-run median 78.287226; wall full128 67.762818 B70-verified closed-lane record; 36/36 realistic rows cache-zero, 24/24 exact canaries, unchanged target verifies accepted tokens; LocalMaxxing cmrquta9905w3lg013m5vxoqx packet; repro; summary; contributor: Steve Seguin
unsloth/Qwen3.6-27B-MTP-GGUF
Q4_0 file identity recorded in the linked payload
GGUF Q4_0 target; Q8 target KV; native Q4_K_M DFlash draft with F16 draft KV; target-verified speculation 1x Arc Pro B70 32 GB, one active generation llama.cpp/SYCL base e3546c794 plus preserved BMG-AOT Xe2 M=6 verifier, GDN snapshot-cache, and fused Q6_K draft-head/top-1 stack Fixed realistic suite; 12 unique cold prompts, output 128; median tokens 1-100 after TTFT; native DFlash5 47.818818 tok/s median; p10 39.869534; mean 46.638647 B70-verified closed-lane record; all cached tokens zero; unchanged target verifies accepted draft tokens; matching AOT control 44.2205 tok/s; LocalMaxxing cmrjbx8bc02g8mj01yzz2v701 closure; evidence; record note; contributor: Steve Seguin
webhie/Qwen3.6-27B-int4-AutoRound
f5750c90b3776db658594df5fe8051098226dd8e
AutoRound INT4 W4A16 target; FP16 target compute; runtime INT8 target LM-head with BF16 scales; runtime INT4 group-128 draft LM-head with BF16 scales 2x Arc Pro B70 32 GB, TP2, concurrency 1 vLLM/XPU 0.20.2rc1.dev13 local stack; pinned public oneCCL parent b52f40c / libccl 4ceafd1; captured draft, graph-safe FlashAttention full target graph, ReplaySSM pending/direct-output transaction fusion Qwen fixed realistic suite; 12 unique cold prompts, output 512; median tokens 1-100 after TTFT; ctx 2048; target-verified MTP3 95.384868 tok/s median; p10 86.975415; mean 95.623050; full after-TTFT 91.698097 B70-verified; strict fresh gate and full quality pass, all cached tokens zero, exact cases + repeat128 + baseline parity + 1K needle pass; both swapped crossover assignments positive; short-context-only forced chunk-decode route packet; strict JSON; quality JSON; note; contributor: Steve Seguin
unsloth/gemma-4-26B-A4B-it-GGUF
HF revision not recorded in packet; exact byte-verified GGUF is identified there
UD-Q8_K_XL target/verifier; Q4_0 MTP draft; f16 KV 1x Arc Pro B70 32 GB, one full replica, concurrency 1 llama.cpp/SYCL c926ad098 plus the preserved local Gemma record stack gemma4-26b-a4b-q8-b70-realistic-v1; 12 unique cold prompts, output 512; median tokens 1-100 after TTFT; ctx 32768; target-verified MTP3 124.97714084813418 tok/s median; p10 103.836100; full-output after-TTFT median 114.871070 B70-verified; realistic final gate and fresh-response validity pass, all cached tokens zero, 512/512 canary rows packet; standalone repro; summary JSON; patch snapshot; contributor: Steve Seguin
Lasimeri/MiniMax-M2.7-int4-AutoRound
model revision not recorded in packet
AutoRound INT4 W4A16, FP16 activations / FP16-family KV; no speculation 4x Arc Pro B70 32 GB, TP4, batch 1 vLLM/XPU 0.20.1-local; base vLLM c51df430; llm-scaler 4bfc007; XPU graph / Level Zero Strict speed lane: p512/n1536, ctx 2048, max batched tokens 512; mean of four clean long repeats 89.314195 output tok/s; 119.085594 total tok/s B70-verified, historical strict-speed reference; exact n64/n256 hashes, semantic suite, arithmetic repeat, and extended sixpack passed repro; result JSON; patch snapshots; contributor: Steve Seguin
Lasimeri/MiniMax-M2.7-int4-AutoRound
model revision not recorded in packet
AutoRound INT4 W4A16, FP16 activations / FP16-family KV; no speculation 4x Arc Pro B70 32 GB, TP4, one active generation vLLM/XPU based on c51df430; llm-scaler 4bfc007; XPU kernels 28e1f5e; OpenAI-compatible endpoint Deployable 32K endpoint; comparable strict gate p512/n1536 at ctx 2048; mean of four repeats 83.172184 output tok/s; 110.896246 total tok/s; warm endpoint about 83.8 output tok/s B70-verified, deployable reference; strict gate passed; serves ctx 32768. Kept separate from the faster 2K strict lane fresh-install repro; summary JSON; applied patch snapshots; contributor: Steve Seguin
Qwen3.6-35B-A3B Quark W8A8 INT8 4x Arc Pro B70 32 GB, TP4 vLLM/XPU PIECEWISE forced-comm graph strict deep gate, p512/n512 93.550542 output tok/s B70-verified closed reference; strict quality gate passed; LocalMaxxing cmqq4mw4c00yfqo01gb2ucgxj packet; evidence
unsloth/Qwen3.6-27B-MTP-GGUF GGUF UD-Q4_K_XL target, MTP3 draft 1x Arc Pro B70 32 GB llama.cpp/SYCL fdb1db877 fixed 12-prompt cold realistic gate, output 128 30.678767 tok/s median 1-100 after TTFT B70-verified model/runtime reference; cached tokens zero; LocalMaxxing cmr6mn5ct0076mn01on3dnpyn packet
unsloth/Qwen3-30B-A3B-Instruct-2507-GGUF
eea7b2be5805a5f151f8847ede8e5f9a9284bf77
GGUF UD-Q4_K_XL; f16 KV; no speculation 1x Arc Pro B70 32 GB llama.cpp/SYCL server 9763 (dec5ca557); runner recorded clean source snapshot fdb1db877 rapid-model-snapshots-b70-realistic-v1; 12 unique cold prompts, output 128; median tokens 1-100 after TTFT; ctx 4096 107.483884 tok/s median; p10 106.897744; wall full-output median 94.118294 B70-verified rapid snapshot; realistic final gate passed, prompt cache disabled, all cached tokens zero; first-pass baseline packet; result JSON; no result-specific source patch; contributor: Steve Seguin
bartowski/microsoft_Phi-4-mini-instruct-GGUF
7ff82c2aaa4dde30121698a973765f39be5288c0
GGUF Q4_K_M; f16 KV; no speculation 1x Arc Pro B70 32 GB llama.cpp/SYCL; runner recorded clean source snapshot fdb1db877 rapid-model-snapshots-b70-realistic-v1; 12 unique cold prompts, output 128; median tokens 1-100 after TTFT; ctx 4096 96.548341 tok/s median; p10 96.350769; wall full-output median 91.749702 B70-verified rapid snapshot; prompt caches disabled, all cached tokens zero; standalone confirmation, first-pass baseline packet; result JSON; no result-specific source patch; contributor: Steve Seguin

Community-Reported Intel Arc Results

No rows have been added yet. Results from other Intel Arc configurations are welcome when they include the exact hardware, OS, model revision, quantization, runtime identity, command, benchmark shape, quality gate, and JSON/log evidence. A community-reported row remains distinct from a B70-verified B70 row unless it is independently reproduced on matching hardware.

Portability And Other-Hardware Observations

No rows have been added yet. Portable patches and observations from other hardware may be useful to Intel XPU work, but their original performance claims will be labeled as contributor-reported. If such a patch is tested locally, its B70 measurement belongs in a separate row with its own runtime identity and evidence.

Maintaining This Index

Update this page by hand only after reviewing the linked packet and evidence. Do not replace a row merely because a single run is faster. Preserve distinct rows when the model revision, quantization or quality class, runtime, hardware count, benchmark shape, cache policy, or metric definition changes. Superseded rows should remain discoverable in their model packet even when this compact index advances to a newer representative result.