This is a small, manually maintained index of useful performance expectations. It is not the identity of the project, a universal leaderboard, or a replacement for the linked result packets. Those packets and their JSON/log evidence are the source of truth.
The current local verification platform is Intel Arc Pro B70. Results and patches for other Intel Arc GPUs are welcome, but they remain clearly labeled until matching hardware is available for independent reproduction. Testing a contributor’s patch on B70 validates the resulting B70 behavior; it does not verify a score originally reported on different hardware.
Last manual review: 2026-07-26.
bench-openai-realistic-suite.py used a
100-event numerator over the first-to-100th timestamp span, which contains
99 intervals. Unless a row explicitly gives both values, treat its displayed
first-100-token score as the published legacy convention; multiply by 0.99
for conventional interval accounting. Relative comparisons wholly within
that historical convention are unchanged. Laguna is the first row audited
and displayed both ways.B70-verified means the row was produced and quality-gated on the
local B70 system. It does not imply upstream support, portability, warranty,
or identical results on another host.| Model / revision | Quantization | Hardware / layout | Engine / runtime | Suite / benchmark shape | Result | Quality / status | Patch, evidence, contributor |
|---|---|---|---|---|---|---|---|
poolside/Laguna-S-2.1-INT44bbfc285f2f8b3b6b526274c133b7b17aae6c8cb |
compressed-tensors INT4 group-32 W4A16 target; matched INT4 DFlash draft with 31 runtime E4M3FN W8A16 projections per rank; BF16 KV; target-verified speculation | 4x Arc Pro B70 32 GB, TP4+EP4, one active generation | vLLM/XPU e596ef154; XPU kernels 6f9dd3c3a; exact width-12 stack, persistent exact-attention metadata, validated 146/145-segment Breakable PIECEWISE graph |
Laguna fixed realistic suite; 13 unique cold prompts, output up to 512; first-to-100th token timestamps; greedy canonical-q1 exactness; DFlash depth 11 | 102.971436 tok/s submitted legacy 100-event/99-interval median; 101.941721 tok/s conventional interval median; legacy p10 71.148884; full wall 52.767621 |
B70-verified, sealed, metric-qualified; first valid preregistered cold score, 13/13 token-and-text exact and cache-zero, 512-output-then-next 2/2, rollover 1/1, no warmup generation or retry, clean 73-second pre/post idle gates; LocalMaxxing cms2ccv2d00lps201rej94pjy |
qualified packet; repro; correction; contributor: Steve Seguin |
0xSero/DeepSeek-V4-Flash-180B7c360e1cd4a5168099dbc54d16d929bf6df04990, experimental uniform-K160 artifact |
FP8 block-scaled dense weights; FP4 experts; FP8 KV; unchanged K160 target with target-verified DSpark7 | 4x Arc Pro B70 32 GB, TP4+EP, one active generation | vLLM/XPU 264c7f2f7; XPU kernels 313156737; oneCCL 48fda4f0e; target and draft PIECEWISE graphs |
Fixed realistic suite; 12 unique cold prompts, output 128; median tokens 1-100 after TTFT; target verifier M=8 | 80.820052 tok/s median high; p10 71.669556; three-run median 78.287226; wall full128 67.762818 |
B70-verified closed-lane record; 36/36 realistic rows cache-zero, 24/24 exact canaries, unchanged target verifies accepted tokens; LocalMaxxing cmrquta9905w3lg013m5vxoqx |
packet; repro; summary; contributor: Steve Seguin |
unsloth/Qwen3.6-27B-MTP-GGUFQ4_0 file identity recorded in the linked payload |
GGUF Q4_0 target; Q8 target KV; native Q4_K_M DFlash draft with F16 draft KV; target-verified speculation | 1x Arc Pro B70 32 GB, one active generation | llama.cpp/SYCL base e3546c794 plus preserved BMG-AOT Xe2 M=6 verifier, GDN snapshot-cache, and fused Q6_K draft-head/top-1 stack |
Fixed realistic suite; 12 unique cold prompts, output 128; median tokens 1-100 after TTFT; native DFlash5 | 47.818818 tok/s median; p10 39.869534; mean 46.638647 |
B70-verified closed-lane record; all cached tokens zero; unchanged target verifies accepted draft tokens; matching AOT control 44.2205 tok/s; LocalMaxxing cmrjbx8bc02g8mj01yzz2v701 |
closure; evidence; record note; contributor: Steve Seguin |
webhie/Qwen3.6-27B-int4-AutoRoundf5750c90b3776db658594df5fe8051098226dd8e |
AutoRound INT4 W4A16 target; FP16 target compute; runtime INT8 target LM-head with BF16 scales; runtime INT4 group-128 draft LM-head with BF16 scales | 2x Arc Pro B70 32 GB, TP2, concurrency 1 | vLLM/XPU 0.20.2rc1.dev13 local stack; pinned public oneCCL parent b52f40c / libccl 4ceafd1; captured draft, graph-safe FlashAttention full target graph, ReplaySSM pending/direct-output transaction fusion |
Qwen fixed realistic suite; 12 unique cold prompts, output 512; median tokens 1-100 after TTFT; ctx 2048; target-verified MTP3 | 95.384868 tok/s median; p10 86.975415; mean 95.623050; full after-TTFT 91.698097 |
B70-verified; strict fresh gate and full quality pass, all cached tokens zero, exact cases + repeat128 + baseline parity + 1K needle pass; both swapped crossover assignments positive; short-context-only forced chunk-decode route | packet; strict JSON; quality JSON; note; contributor: Steve Seguin |
unsloth/gemma-4-26B-A4B-it-GGUFHF revision not recorded in packet; exact byte-verified GGUF is identified there |
UD-Q8_K_XL target/verifier; Q4_0 MTP draft; f16 KV | 1x Arc Pro B70 32 GB, one full replica, concurrency 1 | llama.cpp/SYCL c926ad098 plus the preserved local Gemma record stack |
gemma4-26b-a4b-q8-b70-realistic-v1; 12 unique cold prompts, output 512; median tokens 1-100 after TTFT; ctx 32768; target-verified MTP3 |
124.97714084813418 tok/s median; p10 103.836100; full-output after-TTFT median 114.871070 |
B70-verified; realistic final gate and fresh-response validity pass, all cached tokens zero, 512/512 canary rows |
packet; standalone repro; summary JSON; patch snapshot; contributor: Steve Seguin |
Lasimeri/MiniMax-M2.7-int4-AutoRoundmodel revision not recorded in packet |
AutoRound INT4 W4A16, FP16 activations / FP16-family KV; no speculation | 4x Arc Pro B70 32 GB, TP4, batch 1 | vLLM/XPU 0.20.1-local; base vLLM c51df430; llm-scaler 4bfc007; XPU graph / Level Zero |
Strict speed lane: p512/n1536, ctx 2048, max batched tokens 512; mean of four clean long repeats | 89.314195 output tok/s; 119.085594 total tok/s |
B70-verified, historical strict-speed reference; exact n64/n256 hashes, semantic suite, arithmetic repeat, and extended sixpack passed | repro; result JSON; patch snapshots; contributor: Steve Seguin |
Lasimeri/MiniMax-M2.7-int4-AutoRoundmodel revision not recorded in packet |
AutoRound INT4 W4A16, FP16 activations / FP16-family KV; no speculation | 4x Arc Pro B70 32 GB, TP4, one active generation | vLLM/XPU based on c51df430; llm-scaler 4bfc007; XPU kernels 28e1f5e; OpenAI-compatible endpoint |
Deployable 32K endpoint; comparable strict gate p512/n1536 at ctx 2048; mean of four repeats | 83.172184 output tok/s; 110.896246 total tok/s; warm endpoint about 83.8 output tok/s |
B70-verified, deployable reference; strict gate passed; serves ctx 32768. Kept separate from the faster 2K strict lane | fresh-install repro; summary JSON; applied patch snapshots; contributor: Steve Seguin |
Qwen3.6-35B-A3B |
Quark W8A8 INT8 | 4x Arc Pro B70 32 GB, TP4 | vLLM/XPU PIECEWISE forced-comm graph | strict deep gate, p512/n512 | 93.550542 output tok/s | B70-verified closed reference; strict quality gate passed; LocalMaxxing cmqq4mw4c00yfqo01gb2ucgxj |
packet; evidence |
unsloth/Qwen3.6-27B-MTP-GGUF |
GGUF UD-Q4_K_XL target, MTP3 draft | 1x Arc Pro B70 32 GB | llama.cpp/SYCL fdb1db877 |
fixed 12-prompt cold realistic gate, output 128 | 30.678767 tok/s median 1-100 after TTFT | B70-verified model/runtime reference; cached tokens zero; LocalMaxxing cmr6mn5ct0076mn01on3dnpyn |
packet |
unsloth/Qwen3-30B-A3B-Instruct-2507-GGUFeea7b2be5805a5f151f8847ede8e5f9a9284bf77 |
GGUF UD-Q4_K_XL; f16 KV; no speculation | 1x Arc Pro B70 32 GB | llama.cpp/SYCL server 9763 (dec5ca557); runner recorded clean source snapshot fdb1db877 |
rapid-model-snapshots-b70-realistic-v1; 12 unique cold prompts, output 128; median tokens 1-100 after TTFT; ctx 4096 |
107.483884 tok/s median; p10 106.897744; wall full-output median 94.118294 |
B70-verified rapid snapshot; realistic final gate passed, prompt cache disabled, all cached tokens zero; first-pass baseline | packet; result JSON; no result-specific source patch; contributor: Steve Seguin |
bartowski/microsoft_Phi-4-mini-instruct-GGUF7ff82c2aaa4dde30121698a973765f39be5288c0 |
GGUF Q4_K_M; f16 KV; no speculation | 1x Arc Pro B70 32 GB | llama.cpp/SYCL; runner recorded clean source snapshot fdb1db877 |
rapid-model-snapshots-b70-realistic-v1; 12 unique cold prompts, output 128; median tokens 1-100 after TTFT; ctx 4096 |
96.548341 tok/s median; p10 96.350769; wall full-output median 91.749702 |
B70-verified rapid snapshot; prompt caches disabled, all cached tokens zero; standalone confirmation, first-pass baseline | packet; result JSON; no result-specific source patch; contributor: Steve Seguin |
No rows have been added yet. Results from other Intel Arc configurations are welcome when they include the exact hardware, OS, model revision, quantization, runtime identity, command, benchmark shape, quality gate, and JSON/log evidence. A community-reported row remains distinct from a B70-verified B70 row unless it is independently reproduced on matching hardware.
No rows have been added yet. Portable patches and observations from other hardware may be useful to Intel XPU work, but their original performance claims will be labeled as contributor-reported. If such a patch is tested locally, its B70 measurement belongs in a separate row with its own runtime identity and evidence.
Update this page by hand only after reviewing the linked packet and evidence. Do not replace a row merely because a single run is faster. Preserve distinct rows when the model revision, quantization or quality class, runtime, hardware count, benchmark shape, cache policy, or metric definition changes. Superseded rows should remain discoverable in their model packet even when this compact index advances to a newer representative result.