| 2026-07-20 |
Option 4 Phase 0b raw Level Zero mixed replay |
GO / Phase 1 unblocked |
../notes/2026-07-20-option4-phase0b-raw-level-zero-replay.md, ../data/option4-phase0b-raw-level-zero-20260720.json, ../../../option4-decoder/, raw /mnt/fast-ai/bench-results/deepseek-v4-flash-xpu/option4-phase0b-raw-lz-20260720T195000Z |
Borrowing PyTorch’s in-order immediate-list handle and the finalized graph’s regular-list handle allows direct zeCommandListImmediateAppendCommandListsExp replay. Changed-input parity passes 40/40 twice. The decisive PTI window has 1 boundary, 0 host syncs, versus Phase 0’s one zeEventHostSynchronize; 100 pending-input raw replays have 39.129 us median host enqueue time. EAGLE stayed on XPU 1/renderD131 while the gate used XPU 2/renderD128. Proceed to M1AttentionBoundaryV1; no model load or LocalMax submission occurred. |
| 2026-07-20 |
nonspec M=1 MHC exact-efficiency + inexact M8 quality measurement |
exact component below gate / inexact quality-rejected |
../notes/2026-07-20-nospec-m1-kernel-efficiency-iteration3-mhc.md, ../data/nospec-m1-kernel-efficiency-iteration3-mhc-20260720.json, ../scripts/bench-m1-mhc-rms-reuse.py, ../../../patches/deepseek-v4-flash-xpu-b70/20260720-mhc-reuse-rms-reduction.patch, XPU 5a1e9fa, raw /mnt/fast-ai/bench-results/deepseek-v4-flash-xpu/iter3-mhc-* |
Angle A is already promoted: current post→BF16→pre is one launch, and 85 counts semantic boundaries. Exact angle B reuses the canonical first-pass RMS reduction, passes 160/160 eager + 160/160 graph cases, but saves only 0.069904 ms/token on the slowest candidate card with launches unchanged at 85→85. Skip B-A-B. Measurement-only M8 TF32 DPAS saves 0.314281 ms/emitted-token in a one-pass same-binary public+DEV screen but changes 1,453/2,725 greedy token positions (46.678899% match, 17/22 prompts); additional DEV-only match is 62.405383%, 5/10 prompts changed. Keep both flags default-off; no record claim or LocalMax submission. |
| 2026-07-20 |
nonspec M=1 exact GEMM-efficiency bundle |
exact sub-gate / bundle rejected |
../notes/2026-07-20-nospec-m1-kernel-efficiency-iteration2.md, ../data/nospec-m1-kernel-efficiency-iteration2-20260720.json, ../scripts/bench-m1-mxfp4-grf-efficiency.py, ../scripts/bench-m1-dense-prepack-efficiency.py, XPU c9f20d2, vLLM bbb633a90, raw four-card evidence under /mnt/fast-ai/bench-results/deepseek-v4-flash-xpu/m1-{mxfp4-prefetch,dense-prepack}-* |
MXFP4 prefetch distance 3 passes 160/160 eager + 160/160 graph cases but saves only 0.026013 ms/token on the slowest candidate card. Removing duplicate A hints and distances 2/4 regress. Shared-down prepack is exact but slower; WQ_B and shared gate/up prepack are inexact and slower because layout changes JIT accumulation. Retain only default-off MXFP4 d3; combined bundle FAILS 0.50 ms/token, so no model load, B-A-B, or LocalMax submission. |
| 2026-07-20 |
nonspec M=1 per-kernel bandwidth profile + MXFP4 GRF128 occupancy |
exact/performance-rejected |
../notes/2026-07-20-nospec-m1-kernel-efficiency-iteration1.md, ../data/nospec-m1-kernel-efficiency-iteration1-20260720.json, ../scripts/bench-m1-mxfp4-grf-efficiency.py, ../../../patches/deepseek-v4-flash-xpu-b70/20260720-m1-mxfp4-grf128-occupancy.patch, XPU 790479d, raw /mnt/fast-ai/bench-results/deepseek-v4-flash-xpu/m1-mxfp4-grf128-efficiency-gate-20260720T154042Z-b |
Routed MXFP4 is 288.7 GB/s (54.8% peak) with 1.576 ms/token theoretical slack. GRF256->128 preserves identical M8xN64xK32 arithmetic and passes 160/160 changing eager + 160/160 fixed-address graph cases, but the slowest candidate card regresses 107.315 -> 205.147 us/layer, or -4.2067 ms/token. Fail the +0.30 gate; no service B-A-B and no LocalMax submission. |
| 2026-07-20 |
nonspec M=1 Intel PTI Level Zero attribution cross-check |
host trace confirms / device-timestamp mode rejected |
../notes/2026-07-20-nospec-unitrace-attribution-crosscheck.md, ../data/nospec-unitrace-host-crosscheck-20260720.json, ../scripts/vllm-unitrace-wrapper.sh, ../scripts/capture-unitrace-steady-decode.py, ../scripts/summarize-unitrace-host-crosscheck.py, valid raw host trace /mnt/fast-ai/bench-results/deepseek-v4-flash-xpu/nospec-unitrace-host-crosscheck-20260720T141855Z |
Intel PTI unitrace 2.4.0 host-only temporal tracing across 24 one-generation decode intervals measures 70.458 effective boundaries, 10.792 command-list host syncs, 2.510 ms/token no-Level-Zero CPU gaps, and only 0.430 ms/token nonblocking Level Zero API work at normal 21.863 ms interarrival. The inclusive oneCCL-associated wait is 18.925 ms and explicitly non-additive. Full kernel timestamps kill a worker and are rejected. This confirms the prior ~70/10 and ~0.9 ms native-recovery conclusion; no new >=0.5 ms removable bucket, no LocalMax submission. |
| 2026-07-18 |
DSpark7 sharded target argmax record |
promoted; final closed-lane record |
../notes/2026-07-18-sharded-target-argmax-record.md, ../data/dspark-sharded-target-argmax-record-20260718.json, vLLM 264c7f2f7, XPU kernels 313156737, oneCCL 48fda4f0e, /mnt/fast-ai/bench-results/deepseek-v4-flash-xpu/dspark7-sharded-target-argmax-candidate-20260718T2100Z, LocalMaxxing cmrquta9905w3lg013m5vxoqx, standalone ../../../repro/deepseek-v4-flash-k160-b70-80tps-20260718/README.md |
Guarded greedy verification projects only local vocabulary shards, gathers tiny top-1 pairs, and commits target IDs natively. Three strict medians are 80.820052 / 76.900178 / 78.287226 tok/s; all 36 realistic requests are cache-zero, four ordered exact suites pass 24/24, and the unchanged K160 target verifies accepted tokens at M=8. No later verified endpoint exceeded it. Promote and preserve exact source bundles. |
| 2026-07-18 |
fixed M8 MHC post/pre + RMSNorm |
correctness/performance rejected |
../notes/2026-07-18-m8-mhc-rms-fusion-closure.md, ../data/m8-mhc-rms-fusion-closure-20260718.json, ../scripts/bench-m8-mhc-rms-fusion.py, XPU 2cc25d0 |
The 512-lane fused geometry changes 113 post, 2,378 comb, and 10 normalized output bits across 40 changed inputs and regresses 21.4560 -> 22.6202 us/boundary, a projected 0.0990 ms/cycle loss. Reject after card 0; do not spend four-card, graph, model-load, endpoint, or LocalMax gates. |
| 2026-07-18 |
DSpark M7 local-base IPC bundle + Xe2 BF16 DPAS |
exact component / endpoint rejected |
../notes/2026-07-18-dspark-m7-ipc-dpas-bundle-closure.md, ../data/dspark-m7-ipc-dpas-bundle-closure-20260718.json, vLLM 80f1ad820, XPU kernels 585a4bc105, /mnt/fast-ai/bench-results/deepseek-v4-flash-xpu/dspark7-ipc-bundle-dpas-candidate-20260718T2040Z |
Real W2 BF16 DPAS is bit-exact and 1.679x faster. The combined seven-stage transaction is exact and saves 0.994 ms (1.520 -> 0.526 ms), but the endpoint is only 67.227723 tok/s versus the 80.820052 record. All 12 realistic requests are cache-zero and 12/12 pre/post canaries pass. Preserve DPAS, reject the one-shot event endpoint, no LocalMax submission. |
| 2026-07-18 |
fixed M8 target input/block/slot builder |
exact graph pass/performance rejected |
../notes/2026-07-18-fixed-m8-target-builder-closure.md, ../data/fixed-m8-target-builder-closure-20260718.json, ../scripts/bench-fixed-m8-target-builder.py, vLLM ec7d27e0c, /mnt/fast-ai/bench-results/deepseek-v4-flash-xpu/fixed-m8-target-builder-gate-20260718TpoststopZ |
One fixed transaction replaces position/length, token assembly, block gather, and slot mapping. Four B70s pass 16 changed eager schedules and 70 graph replays bit-for-bit. Eager improves 194.3145 -> 85.4355 us, but captured control/candidate are 33.974/33.718 us: only 0.256 us survives. Keep default-off, skip endpoint/LocalMax, and move to a transaction that deletes collective/device work. |
| 2026-07-18 |
padded Markov winner-pair all-gather |
exact communication gate/performance rejected |
../notes/2026-07-18-padded-markov-winner-exchange-closure.md, ../data/padded-markov-winner-exchange-closure-20260718.json, ../scripts/bench-tp4-padded-pair-allgather.py, /mnt/fast-ai/bench-results/deepseek-v4-flash-xpu/tp4-padded-pair-allgather-20260718T2350Z/sweep.json |
The unpadded eight-byte pair falls onto a pathological tiny route at 1,192.6645 us for seven exchanges. Padding repairs the route, but the best conservative 1-KiB payload takes 352.652 us versus a 371.347 us full-bias control and saves only 6.4315 us/cycle. All payloads are exact on all ranks. Reject before integration: the incumbent gather is already latency-bound and winner selection would erase the microscopic saving. |
| 2026-07-18 |
gathered target winner/commit and sharded Markov transport follow-ups |
target component retained / all Markov transports rejected |
../notes/2026-07-18-gathered-winner-fusion-and-markov-pair-closure.md, ../data/gathered-winner-and-markov-pair-closure-20260718.json, vLLM 35ce4e8a6/06c5ef710/6a77e5940, XPU 7936e0c4e/917a9398d/d10262ea7 |
Native gathered-winner plus greedy commit passes every four-card eager/graph gate and improves the three-suite center to 79.122226 tok/s, but does not beat the 80.820052 public high; retain without LocalMax submission. Native target-local pair packing is 28.935 us captured versus 25.530 us control and is reverted. Ordinary tiny oneCCL reaches only 73.458134 tok/s. Raw Level Zero IPC remote notification times out. A process-shared host barrier is exact and cuts the isolated seven-step transport 1.533740 -> 0.137513 ms, but endpoint medians regress to 74.840996/75.764457 tok/s because host synchronization destroys overlap. Keep every Markov transport default-off; require a proven device-resident protocol or a larger transaction that removes the exchange. |
| 2026-07-18 |
DSpark7 exact native M=8 router |
promoted target-verified record |
../notes/2026-07-18-m8-router-fusion-record-and-postrecord-closures.md, ../data/dspark-m8-router-record-20260718.json, vLLM db1863c799, XPU kernels 6cad2518d, /mnt/fast-ai/bench-results/deepseek-v4-flash-xpu/dspark7-m8-router-fused-candidate-20260718T1815Z, LocalMaxxing cmrqp2uoa05ublg01lh6yluj8 |
Matching the width-dependent XPU K=6 sum tree makes fused bias/top-k/gather/normalize/scale bit-exact at M=8. Four cards pass 160/160 changing eager and 128/128 changing graph cases, saving 1.205-1.222 ms/cycle. Strict medians are 75.845916 / 77.572536 / 80.163578 tok/s, 36/36 realistic requests are cache-zero, and 24/24 ordered canaries pass. Promote M=8; M=4 remains excluded. Route-direct compact and width-aware attention geometry are exact component negatives and remain reverted. |
| 2026-07-18 |
DSpark7 M=8 selective W8A16 + MXFP4 N128 |
promoted target-verified record |
../notes/2026-07-18-dspark-m8-w8a16-n128-record.md, ../data/dspark-m8-w8a16-n128-record-20260718.json, /mnt/fast-ai/bench-results/deepseek-v4-flash-xpu/dspark7-m8-w8a16-n128-candidate-20260718T2130Z, LocalMaxxing cmrqlp9je05thlg01q4igkk0x |
Four-card gates project 3.168-3.549 ms/cycle from bypassing activation quantization in four M=8 dense families and 0.464-0.562 ms/cycle from N128. W8A16 is not bitwise row-invariant (max BF16 difference 0.0078125), so promotion depends on the endpoint quality gate: strict medians 78.288267 / 74.410268 / 76.937587 tok/s, 36/36 realistic cache-zero requests, and 24/24 ordered exact-output canaries. Promote the bundle; keep N32 rejected. |
| 2026-07-18 |
DSpark target M=8 eager shape profile |
diagnostic complete |
../notes/2026-07-18-dspark-target-m8-eager-profile.md, ../data/dspark-target-m8-eager-profile-20260718.json, /mnt/fast-ai/bench-results/deepseek-v4-flash-xpu/dspark7-replicated-w1-target-eager-profile-20260718T1930Z |
Excluding distorted oneCCL durations, dense GEMM is 11.818190 ms/cycle and routed MXFP4 is 7.988468. Exact shape correlation finds 3.335443 ms/cycle in hundreds of independent M=1 compressor GEMMs. This directly selects the subsequently promoted M=8 batched-compressor boundary; remaining priorities are routed MXFP4 and M=8 FP8/BF16 dense families. |
| 2026-07-18 |
DSpark7 exact M=8 strided-batch compressor |
promoted target-verified record |
../notes/2026-07-18-dspark-m8-batched-compressor-record.md, ../data/dspark-m8-batched-compressor-record-20260718.json, vLLM 1f6d6be49, /mnt/fast-ai/bench-results/deepseek-v4-flash-xpu/dspark7-xpu-compressor-m8-candidate-20260718T2030Z, LocalMaxxing cmrql07qs05t4lg01p86jjybx |
Real K160 C4/C128 shapes pass 40/40 changed eager and 40/40 graph comparisons per shape per B70. Three strict medians are 69.343725 / 71.506808 / 70.249021 tok/s; 36/36 realistic requests are cache-zero and four exact suites pass 24/24. Promote over the 67.501117 W1-replication record. |
| 2026-07-18 |
post-W1 stage profile and greedy copy elision |
profile complete / exact endpoint negative |
../notes/2026-07-18-dspark-post-w1-profile-and-copy-closure.md, ../data/dspark-post-w1-profile-copy-closure-20260718.json, vLLM f7734caed, profile /mnt/fast-ai/bench-results/deepseek-v4-flash-xpu/dspark7-replicated-w1-stage-profile-20260718T1830Z, endpoint dspark7-xpu-copy-elision-candidate-20260718T1900Z |
Record-identity scopes put target verification at approximately 20.68-21.23 ms/cycle and DSpark proposal at 15.70-16.44 ms under instrumentation. Direct argmax-to-draft output saves only 0.076263 ms/cycle; adding unused greedy request-copy elision remains exact but reaches only 64.764976 tok/s versus the 67.501117 record. Keep both flags default-off and move to target-verifier attribution. |
| 2026-07-18 |
DSpark7 W1-only replication over persistent Markov |
promoted target-verified record |
../notes/2026-07-18-dspark-replicated-w1-record.md, ../data/dspark-replicated-w1-record-20260718.json, vLLM 019e6f0e2, XPU kernels 0b99fc536, /mnt/fast-ai/bench-results/deepseek-v4-flash-xpu/dspark7-xpu-replicated-w1-candidate-20260718T1800Z, LocalMaxxing cmrqjhpmz05snlg01ujiehc0u |
Replicating only W1 removes seven all-reduces while W2 stays sharded. The exact component saves 0.452403 ms/cycle, below the 0.50 ms standalone gate, but an explicitly authorized endpoint exception establishes strict medians of 65.656734 / 67.501117 / 67.182469 tok/s. All 36 requests are cache-zero and four six-case exact suites pass. Full replication, fused argmax, tiny-pair exchange, and pre-gather local add remain rejected. |
| 2026-07-18 |
DSpark7 persistent sharded Markov transaction |
promoted target-verified record |
../notes/2026-07-18-dspark-persistent-markov-record.md, ../data/dspark-persistent-markov-record-20260718.json, vLLM 0873ffa67, XPU kernels 0b99fc536, /mnt/fast-ai/bench-results/deepseek-v4-flash-xpu/dspark7-xpu-persistent-markov-bundle-20260718T1640Z, LocalMaxxing cmrqiovsv05s6lg012d8v5nz8 |
Fixed device buffers and direct-output operations remove allocation/cat/copy overhead while retaining sharded W1/W2 work and TP4 collectives. The exact component saves 0.786613 ms/cycle at the slowest rank. Strict fresh suite medians are 65.674202 / 66.479103 / 63.558530 tok/s; all 36 requests are cache-zero and four six-case exact suites pass. This is one active generation, not aggregate. Full W2/LM-head replication is rejected; next replicate W1 only while W2 stays sharded. |
| 2026-07-18 |
DSpark7 private PIECEWISE exact-M7 draft replay |
promoted target-verified record |
../notes/2026-07-18-dspark-piecewise-exact-m7-record.md, ../data/dspark-piecewise-exact-m7-record-20260718.json, vLLM 48401ed6a, XPU kernels 0b99fc536, /mnt/fast-ai/bench-results/deepseek-v4-flash-xpu/dspark7-xpu-targetpw-draftpw-exactm7-20260718T0556Z, LocalMaxxing cmrpymqh505mxlg01tzg3e0yl |
A private breakable PIECEWISE graph captures only the official three-stage DSpark draft’s exact M=7 query while target verification remains proven M=8. Strict fresh suite medians are 64.661411 / 61.724506 / 64.275173 tok/s; all 36 requests are cache-zero and three six-case exact suites pass. This is one active generation, not aggregate. Full draft graph is correctness-rejected, DSpark5 loses, and padded-M8 reaches only 60.518331. Promote exact-M7 and profile the remaining context-KV/draft/sampler cycle. |
| 2026-07-18 |
genuine sequential M=4/M=8 verifier and repeated-MTP screen |
verifier exact/predictor rejected |
../notes/2026-07-18-sequential-mwidth-verifier-and-predictor-pivot.md, ../data/mwidth-sequential-verifier-20260718.json, vLLM 57cfb6771, XPU kernels 50646a2, sequential corpora and replay results under /mnt/fast-ai/ |
Real consecutive M=4/M=8 tensors close the duplicated-row evidence gap. Fixed MHC saves 1.441370 ms (20.74%) at M=4 and 4.314321 ms (34.91%) at M=8 versus segmented M2, with 70/70 graph replays exact on every B70. Repeating K160’s one-layer MTP three times reaches only 46.247281 tok/s on eight eligible cold rows; proposal three accepts 0.0-3.2%. Reject repeated MTP2+ and use the official three-stage DSpark draft with the frozen held-out gate. |
| 2026-07-17 |
subgroup-split paired GEMM1 incremental upper bound |
closed before implementation |
../notes/2026-07-17-mtp1-sg-split-incremental-upper-bound-closure.md, prior XPU c069ed8, ../notes/2026-07-17-mtp1-postportfolio-eager-cycle-profile.md |
The old fused producer saved at most 0.520384 ms/cycle on all-remote versus generic, while promoted route-direct already owns 0.397-0.414 ms of that scope. The generous incremental ceiling is only 0.123114 ms/cycle, and fresh trace activation work is 0.079980 ms. Since all-remote has no local projection arithmetic, subgroup splitting cannot clear the unchanged 0.50 ms every-route gate. Preserve the design for a larger specialized-decoder package; do not build or load it standalone. |
| 2026-07-17 |
MTP1 post-portfolio eager cycle profile |
complete |
../notes/2026-07-17-mtp1-postportfolio-eager-cycle-profile.md, ../data/eager-cycle-postportfolio-20260717-summary.json, ../scripts/serve-k160-mtp1-postportfolio-eager-profile.sh, /mnt/fast-ai/bench-results/deepseek-v4-flash-xpu/mtp1-postportfolio-eager-profile-20260717T032135Z |
Exact record source/selectors rerun as an eager diagnostic twin. Cross-rank noncollective work falls from 19.4779 to 17.8497 ms/cycle, a measured 1.6283 ms portfolio reduction. Dense remains 6.5639 ms; compact routed MXFP4 remains the largest open kernel family at 3.9424 ms. Kineto oneCCL and host durations remain excluded. Proceed to the subgroup-split/SLM-exchange M=2 MXFP4 hardware gate. |
| 2026-07-16 |
MTP1 QNorm-M2 + route-direct N64 portfolio |
promoted |
../notes/2026-07-16-qnorm-routeportfolio-record.md, ../data/qnorm-routeportfolio-20260716/summary.json, vLLM 4a6fd8747, XPU kernels 18a44f440, /mnt/fast-ai/bench-results/deepseek-v4-flash-xpu/qnorm-routeportfolio-candidate-b2-20260716T2255Z, LocalMaxxing cmrocpuhq029hlg01g3yzglko |
The standalone route component keeps the frozen 0.50 ms gate false at a 0.397 ms/cycle worst-card floor and is admitted only with the independently proven, non-overlapping QNorm-M2 floor. Four cards pass 336/336 changed graph cases bitwise; the guarded production wrapper passes 84/84. Same-binary B-A-B medians are 62.515661 / 61.717893 / 63.851301 tok/s, 70/70 ordered exact suites pass across positions 28/58, and every qualifying request is cached-zero. One max-8 diagnostic truncated strict JSON and is preserved as invalid harness evidence; the frozen canary uses max-32. Promote as the target-verified record. |
| 2026-07-16 |
MTP1 M=2 MXFP4 N32/N128 policy |
exact isolated positive, not promoted |
../notes/2026-07-16-mtp1-m2-mxfp4-policy-closure.md, ../data/mtp1-m2-mxfp4-policy-closure-20260716.json, ../scripts/probe-mxfp4-m2-policy.py, XPU kernels 351a06a442, /mnt/fast-ai/bench-results/deepseek-v4-flash-xpu/mtp1-m2-mxfp4-n128-candidate-20260716T0630Z |
Ordered prelaunch scheduler reset repairs N32/N128 graph replay. N32 is exact but loses 0.287-0.300 ms/43 layers. N128 is 48/48 exact on each B70 and saves 0.247-0.283 ms/43, but strict suites reach 62.649706/63.628477 versus a 63.349928 record while same-binary N64 controls span 61.205692-63.101865. Close without promotion or LocalMax submission; keep N64 and require an architectural MXFP4 change above 0.50 ms/cycle. |
| 2026-07-16 |
native MTP1 M=2 router selection/normalization |
promoted |
../notes/2026-07-16-mtp1-m2-router-record.md, ../data/mtp1-m2-router-record-20260716.json, vLLM 4a6fd8747, XPU kernels d15ce87d0, /mnt/fast-ai/bench-results/deepseek-v4-flash-xpu/mtp1-m2-router-norm-candidate-20260716T0605Z, LocalMaxxing cmrncv39w003ylg01hogleazo |
Exact trace arguments corrected the apparent indexer radix hotspot to 40 target-router [2,160], K6 calls/cycle. One submission now selects, normalizes, and scales both verifier rows. Four B70s pass 160/160 changed eager plus 128/128 graph epochs bitwise and save 1.123-1.128 ms/cycle. A same-build flag-off control is 59.108299 tok/s; independent candidate suites reach 62.882999/63.349928 tok/s, 70/70 ordered exact captures pass, and every request is cached-zero. Promote as the target-verified record. |
| 2026-07-16 |
exact-identity eager MTP1 cycle profile |
complete |
../notes/2026-07-16-mtp1-eager-cycle-profile.md, ../data/eager-cycle-profile-20260716-summary.json, ../scripts/summarize-eager-cycle-trace.py, /mnt/fast-ai/bench-results/deepseek-v4-flash-xpu/mtp1-record-eager-cycle-profile-20260716T0550Z |
Host-submission timestamps repair the profiler’s GPU clock offset; distorted oneCCL durations remain excluded. Cross-rank noncollective means are dense GEMM 6.580 ms, MXFP4 MoE 4.151, MHC 2.843, QK/LSE 1.317, PV 0.505. Exact top-k arguments prove 40 [2,160], K6 target-router calls/cycle. Schema v2 further decomposes the dense path: its largest families are already optimized or closed, making M=2 routed MXFP4 the largest genuinely open verifier family. |
| 2026-07-16 |
native MTP1 M=2 MHC post/pre |
promoted |
../notes/2026-07-16-mtp1-m2-mhc-record.md, ../data/mtp1-m2-mhc-record-20260716.json, vLLM 9cf403e51, XPU kernels 46b95e64a, /mnt/fast-ai/bench-results/deepseek-v4-flash-xpu/mtp1-m2-mhc-single-kernel-candidate-20260716T0210Z, LocalMaxxing cmrmvjbok1np3mj01p9il8486 |
One command launches two independent 256-thread verifier-row workgroups while preserving the proven M=1 reduction and BF16 arithmetic boundaries. All four B70s are bitwise exact and save 0.962-0.971 ms across 85 boundaries. Independent cold suites reach 59.291531/60.264242 tok/s versus 57.412142; 70/70 ordered exact captures pass across positions 28/58 and all requests are cached-zero. Promote as the target-verified record. |
| 2026-07-16 |
XPU modular-MoE output alias |
exact/noise-floor negative |
../notes/2026-07-16-xpu-moe-output-alias-negative.md, vLLM 8007ee686, /mnt/fast-ai/bench-results/deepseek-v4-flash-xpu/mtp1-moe-output-alias-candidate-20260716T0130Z |
The default-off alias removes an apparent routed-output copy and passes 10/10 exact capture suites, but strict runs reach only 57.204014/56.198992 tok/s versus the qualified 57.412142/56.952065 record/support. Graph construction already amortizes allocation, and the changed workspace lifetime removes too little replay work to survive full-model variance. Keep default-off; no LocalMax submission. |
| 2026-07-16 |
MTP1 M=2 shared/routed fusion |
promoted |
../notes/2026-07-16-mtp1-m2-fusion-record.md, ../data/mtp1-m2-fusion-record-20260716.json, vLLM 068d6beb2, XPU kernels 84c10f4f1, /mnt/fast-ai/bench-results/deepseek-v4-flash-xpu/mtp1-m2-fusion-candidate-20260716T000928Z, LocalMaxxing cmrmrgce51nojmj01bbxoruuu |
Extending exact shared clamped-SwiGLU/dynamic-FP8 quantization through M=2 and selecting exact fused clamp/SiLU inside generic M=2 routed MoE preserves expert grouping and cross-row weight reuse. Independent strict suites reach 57.412142/56.952065 tok/s versus the 55.703731 record. Seventy ordered exact capture suites pass across positions 28/58 after both suites; all 444 requests are cached-zero. Two direct-M1 M=2 chains are explicitly rejected because overlap/duplicate routes regress 1.4-3.2 ms/cycle. |
| 2026-07-15 |
combined MTP1 + direct M1 routed MoE + wide collective epoch |
promoted |
../notes/2026-07-15-mtp1-direct-moe-wideepoch-record.md, ../data/mtp1-direct-moe-wideepoch-record-20260715.json, /mnt/fast-ai/bench-results/deepseek-v4-flash-xpu/mtp1-direct-moe-wideepoch-candidate-20260715T2315Z, LocalMaxxing cmrmoyenp1no3mj01fz2gjzo6 |
All required commits compose without cherry-picks. Matching strict 12-prompt suites reach 55.703731/55.668081 tok/s, all cached-zero. Seventy ordered exact captures pass, including 50 after both suites and former rollover positions 28 and 58; every worker maps wide-epoch libccl. Promote as the target-verified record. Direct-M1 helps only the attached draft layer because the M=2 target verifier falls back unchanged, explaining the small but repeatable gain. |
| 2026-07-15 |
late nonspec upper-bound closures |
performance-gate fail |
../notes/2026-07-15-late-nospec-upper-bound-closures.md, ../data/late-nospec-upper-bound-closures-20260715.json, ../scripts/bench-m1-direct-gather-upper-bound.py, ../scripts/bench-next-weight-l2-prefetch.py, XPU kernel experiment/revert 5a7f39e9/46bdf344 |
Four cards bound exact GEMM2+gather deletion to 0.151-0.168 ms/token at the typical route and 0.230 ms maximum. Finite L2 hints on real 4/6/8 MiB dense weights preserve exact output but expose no consumer gain; immediate warm-cache sensitivity is only 0.884 us. Close both before model integration. No measured nonspec backlog candidate now clears 0.50 ms/token. |
| 2026-07-15 |
direct M1 routed-MoE plus widened oneCCL epoch |
promoted |
../notes/2026-07-15-direct-routed-moe-wideepoch-record.md, ../data/m1-direct-routed-moe-wideepoch-record-20260715.json, vLLM a681dbb2b, XPU kernels 6522849b0, oneCCL 48fda4f0e, /mnt/fast-ai/bench-results/deepseek-v4-flash-xpu/nospec-direct-moe-wideepoch-candidate-20260715T2220Z, LocalMaxxing cmrmnp7h81nntmj01lfenydgj |
Direct routed-MoE gather and exact router normalization raise the nonspec record to 43.766673 tok/s, with three support suites at 43.699/43.694/43.668. The direct-off control is 41.991/42.155. A reused 11-bit oneCCL readiness counter caused deterministic corruption at graph captures 28 and 58 even with fusion off; a 24-bit collective epoch plus 7-bit communicator tag repairs it without added work. The final identity passes 70/70 exact captures and 48/48 cached-zero strict rows. Promote. |
| 2026-07-15 |
native SIMD16 M1 biased top-k |
promoted |
../notes/2026-07-15-m1-biased-topk-record.md, ../data/m1-biased-topk-record-20260715.json, vLLM a66f3486c, XPU kernels 2a07cf2e8, candidate /mnt/fast-ai/bench-results/deepseek-v4-flash-xpu/nospec-m1-router-candidate-20260715T2021Z, paired control nospec-m1-router-control-20260715T2027Z, LocalMaxxing cmrmjd3io1nn1mj013stqoe4b |
SIMD16 register selection replaces generic bias-add/radix-top-k/gather in 40 normal M=1 MoE layers. Four cards pass 40/40 bitwise ID/raw-weight epochs; the isolated boundary improves 77.128 -> 7.178 us. Two strict nonspec suites reach 41.513661/41.733256 tok/s versus a same-commit flag-off control at 40.067691, saving 0.87-1.00 ms/token. Twenty ordered exact captures pass before/after the suites and every request is cached-zero. Promote as the trustworthy nonspeculative record. |
| 2026-07-15 |
MTP1 exact strided-batch compressor |
promoted |
../notes/2026-07-15-mtp1-batched-compressor-record.md, ../data/mtp1-batched-compressor-record-20260715.json, four-card compressor-m2-bmm-exact-card*-20260715.json, vLLM 3bd0eb321, /mnt/fast-ai/bench-results/deepseek-v4-flash-xpu/mtp1-rowexact-bmm-w8a16-m2-candidate-20260715T2000Z, LocalMaxxing cmrmgacdq1nmimj01i4sfqytp |
Real K160 C4/C128 compressor shapes pass 40/40 changing eager and graph-replay comparisons on every B70. One strided-batch BMM is 1.84-1.91x faster than two M=1 calls plus concatenation. Strict suites reach 55.524496/54.708889 tok/s, 20/20 sustained exact captures pass, every request is cached-zero, and acceptance is 77.96%. Promote as the current target-verified record. |
| 2026-07-15 |
MTP1 target-verifier selective W8A16 M=2 |
promoted |
../notes/2026-07-15-mtp1-w8a16-m2-record.md, ../data/mtp1-rowexact-w8a16m2-record-20260715.json, four-card w8a16-m2-row-invariance-card*-20260715.json, /mnt/fast-ai/bench-results/deepseek-v4-flash-xpu/mtp1-rowexact-w8a16-m2-candidate-20260715T1945Z, LocalMaxxing cmrmfivhg1nmamj012e3138my |
All four production shapes pass 40/40 changing row-exact cases on every B70; one M=2 W8A16 call is 2.42-2.50x faster than two M=1 calls. Strict suites reach 54.464909/54.445287 tok/s, 20/20 sustained exact captures pass, all cached-zero, and acceptance remains 77.68%. Promote as the current target-verified record. |
| 2026-07-15 |
repeated single-layer MTP2 |
correctness-prepass then deadlock |
../notes/2026-07-15-mtp2-reuse-deadlock-closure.md, vLLM 4e47b18c9, /mnt/fast-ai/bench-results/deepseek-v4-flash-xpu/mtp2-rowexact-graph-oneccl1712-20260715T1900Z |
Generalized row-exact compressors let M=3 graph capture and 10/10 initial exact captures pass, but the reused layer’s second draft position accepted only about 0.5-2.2% on realistic prompts. One request then hung and the engine exhausted shared-memory broadcast blocks for 180 seconds. No valid suite or speed claim. Close MTP2+ and restore MTP1. |
| 2026-07-15 |
attached MTP1 with row-exact verifier compressors |
promoted |
../notes/2026-07-15-mtp1-rowexact-record.md, ../data/mtp1-rowexact-record-20260715.json, /mnt/fast-ai/bench-results/deepseek-v4-flash-xpu/mtp1-compressor-rowexact-graph-oneccl1712-20260715T1840Z, LocalMaxxing cmrmetch81nm3mj01w1pidsyt |
Plain MTP1 reached 50.74/50.10 tok/s but later leaked prompt text after 437, so it is rejected. Running the two-token FP32 compressor as two exact M=1 projections repairs sustained graph replay: 20/20 ordered exact captures pass, including ten after the two strict suites. Qualified medians are 50.016860/49.420459 tok/s, with 77.42% measured acceptance and every request cached-zero. Promote as the speculative record; retain 40.170350 as the base record. |
| 2026-07-15 |
KV-overlap and large-allreduce repeatability repair |
promoted |
../notes/2026-07-15-kv-repeatability-and-oneccl-allreduce-routing.md, ../data/oneccl-allreduce-routing-record-20260715.json, vLLM 93fde4186, oneCCL 6da44bc, /mnt/fast-ai/bench-results/deepseek-v4-flash-xpu/nospec-graph-oneccl1712-bf16-allreduce128k-preload094-cache-fix-20260715T1530Z, LocalMaxxing cmrmebmzg1nm0mj01k30nv6vw |
Removed overlapping BF16 KV/RoPE stores, localized deterministic graph-prefill corruption to large SYCL all-reduce, and routed only all-reduces above 128 KiB through the safe path in an exact-version oneCCL 2021.17.2 build. Ten exact captures pass 10/10; two cold suites pass at 40.096205/40.170350 tok/s. This is the current trustworthy base identity; the older 40.135724 submission remains historical speed evidence but is not repeatability-certified. |
| 2026-07-13 |
investment-red-team |
complete |
../data/fit-audit-20260713.json, ../../../plans/2026-07-13-deepseek-v4-flash-b70-investment-gated-plan.md |
Strategic go; reject direct K180 commitment. Run Stages 0-3.5 before download, build K160 first, and climb only after quality/warm-memory gates. |
| 2026-07-13 |
storage-runtime-download-start |
active |
../scripts/download-k160.sh, ../scripts/capture-stage0.sh, clean worktrees under /home/steve/src/deepseek-v4-* |
Archived 170 GiB of reviewed inactive artifacts with compatibility symlinks; internal free space rose from 11 to about 180 GiB. Prioritize frozen public uniform-K160 as the first runnable checkpoint. |
| 2026-07-13 |
k160-provenance-audit |
complete |
../quality/calibration-v1-plan.json, ../quality/suite-v1.json |
Public K160 is valid for smoke/performance bring-up but not quality-certified: hash layers are pruned and published observations are not true REAP. Preserve the official-source teacher and hash-preserved final-pack lanes. |
| 2026-07-13 |
exact-shape-test-scaffold |
implemented |
XPU-kernel commit 552c9ce, ../scripts/run-exact-shape-gates.sh |
Added low-level H4096/I2048/top-k6/M1,4,8 correctness coverage for MXFP4 and INT4 controls at E=40/64. This is not yet the Stage-1 performance, selector, replay, fallback, or TP4/EP gate. |
| 2026-07-13 |
runtime-build-j16 |
loss |
resumable build/cache under /home/steve/src/deepseek-v4-xpu-kernels-clean |
The Xe2 grouped-GEMM translation unit was killed under 16-way SYCL compilation. Preserve completed objects/ccache and resume at the durable eight-job default; this is host build-memory pressure, not a kernel result. |
| 2026-07-14 |
runtime-build-j8 |
pass |
../data/runtime-pin-20260714.json, /mnt/fast-ai/bench-results/deepseek-v4-flash-xpu/stage0-20260714T044033Z.txt |
Incremental eight-job Xe2/SYCL-TLA build completed in 10m40s; pinned vLLM/kernel imports resolve to clean worktrees and all four B70s enumerate. The selector fix prevents -k 40 from also matching E=64 via dimension 4096. |
| 2026-07-14 |
exact-shape-scaffold-preflight |
infrastructure-fail |
/mnt/fast-ai/bench-results/deepseek-v4-flash-xpu/stage1-20260714T044101Z |
Clean environment lacked pytest; no GPU case ran. Installed and pinned pytest 9.0.2 before retrying. |
| 2026-07-14 |
exact-shape-scaffold |
scaffold-pass |
../data/exact-shape-scaffold-20260714.json, /mnt/fast-ai/bench-results/deepseek-v4-flash-xpu/stage1-20260714T044200Z |
All 12 exact M=1/4/8, E=40/64 MXFP4/INT4 low-level reference cases passed on four B70s. This does not clear Stage 1A/1B; performance, metrics, selector, fallback, replay, and TP4+EP evidence remain. |
| 2026-07-14 |
four-card-xccl-preflight |
pass |
../data/runtime-pin-20260714.json, /mnt/fast-ai/bench-results/deepseek-v4-flash-xpu/xccl-preflight-20260714T044502Z.log |
All four per-device smokes, four-rank XCCL init/barrier, and allreduce passed over oneCCL/OFI on eno1; topology correctly reports PCIe/NODE rather than direct fabric. |
| 2026-07-14 |
k160-download-resume |
active |
../data/k160-download-start-20260713.json, ../scripts/download-k160.sh |
Five completed shards survived the 16-way host-memory event. The frozen revision resumed through the external Xet cache after constraining compilation to eight jobs; final HF/SHA-256 verification remains mandatory. |
| 2026-07-14 |
k160-artifact-promotion |
pass |
../../../data/deepseek-v4-k160-tp4-bringup-20260714.json, /mnt/fast-ai/llm-models/deepseek-v4-flash-xpu/current-k160 |
All 46 shards, 43,843 index entries, safetensor headers, HF metadata, and SHA-256 values passed. Archive and hot copy are verified; K160 remains an experimental smoke checkpoint, not a quality promotion. |
| 2026-07-14 |
k160-tp4-eager-construction |
pass |
/mnt/fast-ai/bench-results/deepseek-v4-flash-xpu/tp4-smoke-20260714T132607Z |
TP4+EP assigned 40/160 experts per rank and selected XPUExpertsMxFp4. The 2K/95% configuration fits at 24.95 GiB model plus 2.11 GiB KV per rank; the arithmetic canary returned 1073. Earlier 8K/90% and 2K/98% memory attempts remain preserved as failures. |
| 2026-07-14 |
k160-tp4-eager-decode |
performance-fail |
/mnt/fast-ai/bench-results/deepseek-v4-flash-xpu/tp4-smoke-20260714T132607Z/steady-long-128.json, ../../../data/deepseek-v4-k160-tp4-bringup-20260714.json |
The second warm 128-token diagnostic row reached only 2.616225 tok/s after TTFT. This is about 19x below the 50 tok/s gate; speculation remains prohibited. |
| 2026-07-14 |
k160-breakable-xpu-graph |
infrastructure-fail |
/mnt/fast-ai/bench-results/deepseek-v4-flash-xpu/tp4-smoke-20260714T133502Z/server.log |
Capture fails when sparse FP8 decode executes combined_lens.max().item(), causing a prohibited host wait on a Level Zero command-graph event. Remove or make that scalar graph-safe before another graph performance run. |
| 2026-07-14 |
k160-static-sparse-pack |
pass |
vLLM commit 0ed5ecc5, ../data/graph-recovery-20260714.json |
Replaced graph-unsafe host scalar reads and per-token packing with fixed-width device-only packing. Added finite masked-chunk sentinel. Exact graph replay with changed lengths and focused tests pass; the first one-kernel Triton attempt segfaulted and remains a preserved negative. |
| 2026-07-14 |
sycl8-jit-lane |
pass |
../notes/2026-07-14-xpu-graph-recovery-and-tp4-profile.md, ../scripts/serve-k160-tp4-smoke.sh |
Removed umbrella oneAPI setup and kept venv SYCL/UR libraries ahead of side-by-side oneAPI 2026. This fixes the libsycl.so.9/urDeviceWaitExp JIT ABI mismatch. Quarantined contaminated cache entries instead of deleting them. |
| 2026-07-14 |
k160-piecewise-graph |
performance-pass |
/mnt/fast-ai/bench-results/deepseek-v4-flash-xpu/tp4-graph-piecewise-20260714T1015Z |
Correct 1073 canary; fresh 128-token headline reached 8.512062 tok/s after TTFT versus 2.616225 warm eager. Graph capture is now reusable, but throughput remains far below the 40-50 tok/s base gate. |
| 2026-07-14 |
k160-full-narrow-indexer-break |
performance-pass |
vLLM commit 436298dcd, /mnt/fast-ai/bench-results/deepseek-v4-flash-xpu/tp4-full-narrow-indexer-20260714T1020Z |
Native paged-indexer scratch memory cannot enter SYCL graphs. A narrow FULL-mode eager break gives correct replay and 8.616232 tok/s fresh headline, only 1.22% over PIECEWISE. Attention graph coverage is not the dominant remaining boundary. |
| 2026-07-14 |
k160-tp4-full-profile |
diagnostic |
/mnt/fast-ai/bench-results/deepseek-v4-flash-xpu/tp4-full-profile-20260714T1024Z |
Profile observes 87 c10d all-reduces per decoder step, approximately two across 43 layers. Profiler collective time is distorted and not wall latency; call count plus normal 116 ms/token points to PCIe TP synchronization as the next high-value boundary. |
| 2026-07-14 |
k160-tp2-dp2-ep4 |
infrastructure-fail |
/mnt/fast-ai/bench-results/deepseek-v4-flash-xpu/tp2-dp2-ep4-eager-kv64m-20260714T1047Z, ../data/graph-recovery-20260714.json |
Model fits at 26.66 GiB/card with a manual 64 MiB KV cache. Both graph modes stall in the mixed dummy run; eager first request leaves all ranks waiting in DPEP all-gather. No performance claim. Close until eager XPU DPEP collectives work. |
| 2026-07-14 |
k160-native-mhc-and-c4-bounds |
performance-pass |
vLLM history through cff886b, LocalMaxxing cmrku8z0l05ahmj01raa1794f, cmrkuloel05almj015vfgot4o, cmrkuv0is05dwmj01e6mdztv7 |
Graph-safe native mHC plus context-bounded C4/C128 sparse work raised the strict cold suite from 10.0161 to 14.3784 tok/s. C4 full-selection bypass is valid only because a 1K context contains at most 256 compressed C4 candidates for top-512. |
| 2026-07-14 |
k160-direct-fp8-sparse-attention |
promoted |
vLLM cff886bb7165a7328f6412c199bb93c2fbdcfb98, /mnt/fast-ai/bench-results/deepseek-v4-flash-xpu/tp4-direct-fp8-attn-20260714T1500Z, LocalMaxxing cmrkw8db205nymj01gl5bbmsc |
A graph-captured Triton path reads paged UE8M0 FP8 KV directly and applies runtime lengths plus learned sinks. Strict cold median improved 49.84%, from 14.3784 to 21.5448 tok/s; canaries and all 12 cached-zero rows pass. |
| 2026-07-14 |
rank1-xe-device-loss-recovery |
recovered |
kernel journal for PCI 0000:27:00.0; four-device compute and XCCL post-recovery smokes |
Two launches hit UR_RESULT_ERROR_DEVICE_LOST after a Xe timeout storm. A PCI function reset left a GT PF self-configuration error; targeted Xe driver unbind/rebind restored the card. All four independent compute tests and TP4 barrier/allreduce passed. Prefer targeted rebind for this symptom and preserve the installed-vs-recommended GuC 70.44.1/70.49.4 firmware warning for maintenance. |
| 2026-07-14 |
split-attention-wrong-identity |
rejected |
/mnt/fast-ai/bench-results/deepseek-v4-flash-xpu/tp4-split-fp8-attn-recovery-20260714T1728Z |
The first post-recovery split run accidentally disabled native mHC, used 2K context/batch, and omitted prompt-token details. Its 19.74 tok/s and corrupted long outputs are confounded and are not kernel evidence. This is a benchmark-identity failure, not a split-attention conclusion. |
| 2026-07-14 |
k160-split-tiled-fp8-attention |
promoted |
vLLM b63b1f2b2106a26128a8a0dae55855493b3ada1d, /mnt/fast-ai/bench-results/deepseek-v4-flash-xpu/tp4-split-fp8-attn-recordidentity-20260714T1733Z, LocalMaxxing cmrkxoavs05uimj01p9ix2dtk |
QK/LSE plus 8x64 tiled PV cuts register pressure while retaining direct FP8 cache reads, runtime lengths, and learned sinks. Strict cold median is 29.82238 tok/s, p10 29.42632, +38.41% over direct attention; all exact canaries and 12 cached-zero rows pass. Current record. |
| 2026-07-14 |
k160-split-residual-profile |
diagnostic |
/mnt/fast-ai/bench-results/deepseek-v4-flash-xpu/tp4-split-eager-profiler-20260714T1737Z |
Steady eager decode reconciles to about 23.3 ms of non-collective kernels: dense GEMMs 9.1 ms, MXFP4 MoE 3.67 ms, native mHC 2.69 ms, split QK 1.94 ms, indexer select/sort 0.98 ms, activation quant/scale copies 0.91 ms, split PV 0.28 ms, and about 3.4 ms remaining. Normal graph token period is 33.53 ms; real collectives contribute roughly 8 ms and residual gaps/host about 2 ms. Profiler oneCCL durations are distorted and must not be quoted. |
| 2026-07-14 |
k160-fp8-wo-a |
performance-loss |
vLLM 4b2d2e0efaee9e08ebe9d77392cead912fdb6909, /mnt/fast-ai/bench-results/deepseek-v4-flash-xpu/tp4-fp8-woa-canary-20260714T1810Z |
The existing fused inverse-RoPE/UE8M0 quantizer plus true FP8 BMM passed M=1/4 correctness and graph replay. Exact G2/M1/K4096/N1024 microbench improved BF16 BMM 30.79→24.89 us and full isolated path 55.97→39.78 us, but the strict suite regressed from 29.82238 to 29.51408 tok/s (-1.03%, p10 28.96445). All cold/canary gates passed. Preserve default-off; do not submit. |
| 2026-07-14 |
tp-only-inplace-allreduce |
promoted |
vLLM bfac86e7d22418a88b0f50f4593aa1e2281d3f2d, /mnt/fast-ai/bench-results/deepseek-v4-flash-xpu/tp4-inplace-ar-candidate-20260714T1920Z, LocalMaxxing cmrkz6vo7061fmj01v09nwz72 |
A mutation-declared custom op removes clones only for TP-group contiguous BF16 4096-element decode outputs. Changed-input replay 1073 -> 437 -> 1073 and exact canaries passed. Two strict suites reached 29.9133 and 29.9113 tok/s versus 29.8224 prior (+0.30%); this is a valid record but proves XCCL wait dominates clone cost. |
| 2026-07-14 |
pp4-tp1-single-session |
performance-loss |
/mnt/fast-ai/bench-results/deepseek-v4-flash-xpu/pp4-tp1-split-attn-20260714T1945Z/steady-long-128.json |
PP4 partitioned the 43 layers as [11,11,11,10], fit at 22.62-24.99 GiB/card, captured reusable graphs, and passed changed-input plus exact canaries with zero cached tokens. Fresh 128-token decode reached only 16.7988 tok/s after TTFT versus 29.9133 on TP4. A batch-one token traverses pipeline stages serially, forfeiting four-way concurrent weight bandwidth; three small pipeline transfers cannot compensate. Reject PP4 and PP2/TP2 as base-decode speed lanes. |
| 2026-07-14 |
mhc-postpre-rmsnorm-fusion |
performance-loss |
vLLM 520df585c06c5227f2bdbafd3f4789b9bcca9dd2, XPU kernels 473a55e2a8b34da3c97c143401955d0c5746120b, /mnt/fast-ai/bench-results/deepseek-v4-flash-xpu/tp4-mhc-norm-fusion-20260714T1840Z |
The opt-in kernel preserves the BF16 producer boundary and standalone RMSNorm cast order; M=1/4 parity, command-graph replay, changed-input, and exact canaries pass. Two fresh screens reached 29.4080 and 29.4411 tok/s versus the 29.4844 comparison screen. The saved launch is offset by an in-kernel reduction/fence and reread. Preserve default-off; only revisit as part of producer + norm + consumer quantization fusion. |
| 2026-07-14 |
xccl-twoshots-8k |
microbench-win/end-to-end-loss |
oneCCL 4ceafd15c03ce46f11eeaf91781a92afebd3cecf, /mnt/fast-ai/bench-results/deepseek-v4-flash-xpu/xccl-8k-sweep-20260714T1845Z, /mnt/fast-ai/bench-results/deepseek-v4-flash-xpu/tp4-twoshots-20260714T1848Z |
CCL_SYCL_ALLREDUCE_LL=twoshots reduced synchronized standalone BF16 8 KiB rank-0 median from 128.6915 to 87.369 us and all ranks were correct. Exact graph decode nevertheless measured only 29.4212 and 29.3399 tok/s, below the 29.4844 screen. The host-synchronized microbench does not model reusable command-graph critical-path behavior. Keep the knob explicit in identity but retain ring as default; do not submit. |