b70-optimization-lab

DeepSeek V4 REAP/XPU Experiment Ledger

Preserve every meaningful attempt, including failures.

Date Label Status Evidence Decision
2026-07-20 Option 4 Phase 0b raw Level Zero mixed replay GO / Phase 1 unblocked ../notes/2026-07-20-option4-phase0b-raw-level-zero-replay.md, ../data/option4-phase0b-raw-level-zero-20260720.json, ../../../option4-decoder/, raw /mnt/fast-ai/bench-results/deepseek-v4-flash-xpu/option4-phase0b-raw-lz-20260720T195000Z Borrowing PyTorch’s in-order immediate-list handle and the finalized graph’s regular-list handle allows direct zeCommandListImmediateAppendCommandListsExp replay. Changed-input parity passes 40/40 twice. The decisive PTI window has 1 boundary, 0 host syncs, versus Phase 0’s one zeEventHostSynchronize; 100 pending-input raw replays have 39.129 us median host enqueue time. EAGLE stayed on XPU 1/renderD131 while the gate used XPU 2/renderD128. Proceed to M1AttentionBoundaryV1; no model load or LocalMax submission occurred.
2026-07-20 nonspec M=1 MHC exact-efficiency + inexact M8 quality measurement exact component below gate / inexact quality-rejected ../notes/2026-07-20-nospec-m1-kernel-efficiency-iteration3-mhc.md, ../data/nospec-m1-kernel-efficiency-iteration3-mhc-20260720.json, ../scripts/bench-m1-mhc-rms-reuse.py, ../../../patches/deepseek-v4-flash-xpu-b70/20260720-mhc-reuse-rms-reduction.patch, XPU 5a1e9fa, raw /mnt/fast-ai/bench-results/deepseek-v4-flash-xpu/iter3-mhc-* Angle A is already promoted: current post→BF16→pre is one launch, and 85 counts semantic boundaries. Exact angle B reuses the canonical first-pass RMS reduction, passes 160/160 eager + 160/160 graph cases, but saves only 0.069904 ms/token on the slowest candidate card with launches unchanged at 85→85. Skip B-A-B. Measurement-only M8 TF32 DPAS saves 0.314281 ms/emitted-token in a one-pass same-binary public+DEV screen but changes 1,453/2,725 greedy token positions (46.678899% match, 17/22 prompts); additional DEV-only match is 62.405383%, 5/10 prompts changed. Keep both flags default-off; no record claim or LocalMax submission.
2026-07-20 nonspec M=1 exact GEMM-efficiency bundle exact sub-gate / bundle rejected ../notes/2026-07-20-nospec-m1-kernel-efficiency-iteration2.md, ../data/nospec-m1-kernel-efficiency-iteration2-20260720.json, ../scripts/bench-m1-mxfp4-grf-efficiency.py, ../scripts/bench-m1-dense-prepack-efficiency.py, XPU c9f20d2, vLLM bbb633a90, raw four-card evidence under /mnt/fast-ai/bench-results/deepseek-v4-flash-xpu/m1-{mxfp4-prefetch,dense-prepack}-* MXFP4 prefetch distance 3 passes 160/160 eager + 160/160 graph cases but saves only 0.026013 ms/token on the slowest candidate card. Removing duplicate A hints and distances 2/4 regress. Shared-down prepack is exact but slower; WQ_B and shared gate/up prepack are inexact and slower because layout changes JIT accumulation. Retain only default-off MXFP4 d3; combined bundle FAILS 0.50 ms/token, so no model load, B-A-B, or LocalMax submission.
2026-07-20 nonspec M=1 per-kernel bandwidth profile + MXFP4 GRF128 occupancy exact/performance-rejected ../notes/2026-07-20-nospec-m1-kernel-efficiency-iteration1.md, ../data/nospec-m1-kernel-efficiency-iteration1-20260720.json, ../scripts/bench-m1-mxfp4-grf-efficiency.py, ../../../patches/deepseek-v4-flash-xpu-b70/20260720-m1-mxfp4-grf128-occupancy.patch, XPU 790479d, raw /mnt/fast-ai/bench-results/deepseek-v4-flash-xpu/m1-mxfp4-grf128-efficiency-gate-20260720T154042Z-b Routed MXFP4 is 288.7 GB/s (54.8% peak) with 1.576 ms/token theoretical slack. GRF256->128 preserves identical M8xN64xK32 arithmetic and passes 160/160 changing eager + 160/160 fixed-address graph cases, but the slowest candidate card regresses 107.315 -> 205.147 us/layer, or -4.2067 ms/token. Fail the +0.30 gate; no service B-A-B and no LocalMax submission.
2026-07-20 nonspec M=1 Intel PTI Level Zero attribution cross-check host trace confirms / device-timestamp mode rejected ../notes/2026-07-20-nospec-unitrace-attribution-crosscheck.md, ../data/nospec-unitrace-host-crosscheck-20260720.json, ../scripts/vllm-unitrace-wrapper.sh, ../scripts/capture-unitrace-steady-decode.py, ../scripts/summarize-unitrace-host-crosscheck.py, valid raw host trace /mnt/fast-ai/bench-results/deepseek-v4-flash-xpu/nospec-unitrace-host-crosscheck-20260720T141855Z Intel PTI unitrace 2.4.0 host-only temporal tracing across 24 one-generation decode intervals measures 70.458 effective boundaries, 10.792 command-list host syncs, 2.510 ms/token no-Level-Zero CPU gaps, and only 0.430 ms/token nonblocking Level Zero API work at normal 21.863 ms interarrival. The inclusive oneCCL-associated wait is 18.925 ms and explicitly non-additive. Full kernel timestamps kill a worker and are rejected. This confirms the prior ~70/10 and ~0.9 ms native-recovery conclusion; no new >=0.5 ms removable bucket, no LocalMax submission.
2026-07-18 DSpark7 sharded target argmax record promoted; final closed-lane record ../notes/2026-07-18-sharded-target-argmax-record.md, ../data/dspark-sharded-target-argmax-record-20260718.json, vLLM 264c7f2f7, XPU kernels 313156737, oneCCL 48fda4f0e, /mnt/fast-ai/bench-results/deepseek-v4-flash-xpu/dspark7-sharded-target-argmax-candidate-20260718T2100Z, LocalMaxxing cmrquta9905w3lg013m5vxoqx, standalone ../../../repro/deepseek-v4-flash-k160-b70-80tps-20260718/README.md Guarded greedy verification projects only local vocabulary shards, gathers tiny top-1 pairs, and commits target IDs natively. Three strict medians are 80.820052 / 76.900178 / 78.287226 tok/s; all 36 realistic requests are cache-zero, four ordered exact suites pass 24/24, and the unchanged K160 target verifies accepted tokens at M=8. No later verified endpoint exceeded it. Promote and preserve exact source bundles.
2026-07-18 fixed M8 MHC post/pre + RMSNorm correctness/performance rejected ../notes/2026-07-18-m8-mhc-rms-fusion-closure.md, ../data/m8-mhc-rms-fusion-closure-20260718.json, ../scripts/bench-m8-mhc-rms-fusion.py, XPU 2cc25d0 The 512-lane fused geometry changes 113 post, 2,378 comb, and 10 normalized output bits across 40 changed inputs and regresses 21.4560 -> 22.6202 us/boundary, a projected 0.0990 ms/cycle loss. Reject after card 0; do not spend four-card, graph, model-load, endpoint, or LocalMax gates.
2026-07-18 DSpark M7 local-base IPC bundle + Xe2 BF16 DPAS exact component / endpoint rejected ../notes/2026-07-18-dspark-m7-ipc-dpas-bundle-closure.md, ../data/dspark-m7-ipc-dpas-bundle-closure-20260718.json, vLLM 80f1ad820, XPU kernels 585a4bc105, /mnt/fast-ai/bench-results/deepseek-v4-flash-xpu/dspark7-ipc-bundle-dpas-candidate-20260718T2040Z Real W2 BF16 DPAS is bit-exact and 1.679x faster. The combined seven-stage transaction is exact and saves 0.994 ms (1.520 -> 0.526 ms), but the endpoint is only 67.227723 tok/s versus the 80.820052 record. All 12 realistic requests are cache-zero and 12/12 pre/post canaries pass. Preserve DPAS, reject the one-shot event endpoint, no LocalMax submission.
2026-07-18 fixed M8 target input/block/slot builder exact graph pass/performance rejected ../notes/2026-07-18-fixed-m8-target-builder-closure.md, ../data/fixed-m8-target-builder-closure-20260718.json, ../scripts/bench-fixed-m8-target-builder.py, vLLM ec7d27e0c, /mnt/fast-ai/bench-results/deepseek-v4-flash-xpu/fixed-m8-target-builder-gate-20260718TpoststopZ One fixed transaction replaces position/length, token assembly, block gather, and slot mapping. Four B70s pass 16 changed eager schedules and 70 graph replays bit-for-bit. Eager improves 194.3145 -> 85.4355 us, but captured control/candidate are 33.974/33.718 us: only 0.256 us survives. Keep default-off, skip endpoint/LocalMax, and move to a transaction that deletes collective/device work.
2026-07-18 padded Markov winner-pair all-gather exact communication gate/performance rejected ../notes/2026-07-18-padded-markov-winner-exchange-closure.md, ../data/padded-markov-winner-exchange-closure-20260718.json, ../scripts/bench-tp4-padded-pair-allgather.py, /mnt/fast-ai/bench-results/deepseek-v4-flash-xpu/tp4-padded-pair-allgather-20260718T2350Z/sweep.json The unpadded eight-byte pair falls onto a pathological tiny route at 1,192.6645 us for seven exchanges. Padding repairs the route, but the best conservative 1-KiB payload takes 352.652 us versus a 371.347 us full-bias control and saves only 6.4315 us/cycle. All payloads are exact on all ranks. Reject before integration: the incumbent gather is already latency-bound and winner selection would erase the microscopic saving.
2026-07-18 gathered target winner/commit and sharded Markov transport follow-ups target component retained / all Markov transports rejected ../notes/2026-07-18-gathered-winner-fusion-and-markov-pair-closure.md, ../data/gathered-winner-and-markov-pair-closure-20260718.json, vLLM 35ce4e8a6/06c5ef710/6a77e5940, XPU 7936e0c4e/917a9398d/d10262ea7 Native gathered-winner plus greedy commit passes every four-card eager/graph gate and improves the three-suite center to 79.122226 tok/s, but does not beat the 80.820052 public high; retain without LocalMax submission. Native target-local pair packing is 28.935 us captured versus 25.530 us control and is reverted. Ordinary tiny oneCCL reaches only 73.458134 tok/s. Raw Level Zero IPC remote notification times out. A process-shared host barrier is exact and cuts the isolated seven-step transport 1.533740 -> 0.137513 ms, but endpoint medians regress to 74.840996/75.764457 tok/s because host synchronization destroys overlap. Keep every Markov transport default-off; require a proven device-resident protocol or a larger transaction that removes the exchange.
2026-07-18 DSpark7 exact native M=8 router promoted target-verified record ../notes/2026-07-18-m8-router-fusion-record-and-postrecord-closures.md, ../data/dspark-m8-router-record-20260718.json, vLLM db1863c799, XPU kernels 6cad2518d, /mnt/fast-ai/bench-results/deepseek-v4-flash-xpu/dspark7-m8-router-fused-candidate-20260718T1815Z, LocalMaxxing cmrqp2uoa05ublg01lh6yluj8 Matching the width-dependent XPU K=6 sum tree makes fused bias/top-k/gather/normalize/scale bit-exact at M=8. Four cards pass 160/160 changing eager and 128/128 changing graph cases, saving 1.205-1.222 ms/cycle. Strict medians are 75.845916 / 77.572536 / 80.163578 tok/s, 36/36 realistic requests are cache-zero, and 24/24 ordered canaries pass. Promote M=8; M=4 remains excluded. Route-direct compact and width-aware attention geometry are exact component negatives and remain reverted.
2026-07-18 DSpark7 M=8 selective W8A16 + MXFP4 N128 promoted target-verified record ../notes/2026-07-18-dspark-m8-w8a16-n128-record.md, ../data/dspark-m8-w8a16-n128-record-20260718.json, /mnt/fast-ai/bench-results/deepseek-v4-flash-xpu/dspark7-m8-w8a16-n128-candidate-20260718T2130Z, LocalMaxxing cmrqlp9je05thlg01q4igkk0x Four-card gates project 3.168-3.549 ms/cycle from bypassing activation quantization in four M=8 dense families and 0.464-0.562 ms/cycle from N128. W8A16 is not bitwise row-invariant (max BF16 difference 0.0078125), so promotion depends on the endpoint quality gate: strict medians 78.288267 / 74.410268 / 76.937587 tok/s, 36/36 realistic cache-zero requests, and 24/24 ordered exact-output canaries. Promote the bundle; keep N32 rejected.
2026-07-18 DSpark target M=8 eager shape profile diagnostic complete ../notes/2026-07-18-dspark-target-m8-eager-profile.md, ../data/dspark-target-m8-eager-profile-20260718.json, /mnt/fast-ai/bench-results/deepseek-v4-flash-xpu/dspark7-replicated-w1-target-eager-profile-20260718T1930Z Excluding distorted oneCCL durations, dense GEMM is 11.818190 ms/cycle and routed MXFP4 is 7.988468. Exact shape correlation finds 3.335443 ms/cycle in hundreds of independent M=1 compressor GEMMs. This directly selects the subsequently promoted M=8 batched-compressor boundary; remaining priorities are routed MXFP4 and M=8 FP8/BF16 dense families.
2026-07-18 DSpark7 exact M=8 strided-batch compressor promoted target-verified record ../notes/2026-07-18-dspark-m8-batched-compressor-record.md, ../data/dspark-m8-batched-compressor-record-20260718.json, vLLM 1f6d6be49, /mnt/fast-ai/bench-results/deepseek-v4-flash-xpu/dspark7-xpu-compressor-m8-candidate-20260718T2030Z, LocalMaxxing cmrql07qs05t4lg01p86jjybx Real K160 C4/C128 shapes pass 40/40 changed eager and 40/40 graph comparisons per shape per B70. Three strict medians are 69.343725 / 71.506808 / 70.249021 tok/s; 36/36 realistic requests are cache-zero and four exact suites pass 24/24. Promote over the 67.501117 W1-replication record.
2026-07-18 post-W1 stage profile and greedy copy elision profile complete / exact endpoint negative ../notes/2026-07-18-dspark-post-w1-profile-and-copy-closure.md, ../data/dspark-post-w1-profile-copy-closure-20260718.json, vLLM f7734caed, profile /mnt/fast-ai/bench-results/deepseek-v4-flash-xpu/dspark7-replicated-w1-stage-profile-20260718T1830Z, endpoint dspark7-xpu-copy-elision-candidate-20260718T1900Z Record-identity scopes put target verification at approximately 20.68-21.23 ms/cycle and DSpark proposal at 15.70-16.44 ms under instrumentation. Direct argmax-to-draft output saves only 0.076263 ms/cycle; adding unused greedy request-copy elision remains exact but reaches only 64.764976 tok/s versus the 67.501117 record. Keep both flags default-off and move to target-verifier attribution.
2026-07-18 DSpark7 W1-only replication over persistent Markov promoted target-verified record ../notes/2026-07-18-dspark-replicated-w1-record.md, ../data/dspark-replicated-w1-record-20260718.json, vLLM 019e6f0e2, XPU kernels 0b99fc536, /mnt/fast-ai/bench-results/deepseek-v4-flash-xpu/dspark7-xpu-replicated-w1-candidate-20260718T1800Z, LocalMaxxing cmrqjhpmz05snlg01ujiehc0u Replicating only W1 removes seven all-reduces while W2 stays sharded. The exact component saves 0.452403 ms/cycle, below the 0.50 ms standalone gate, but an explicitly authorized endpoint exception establishes strict medians of 65.656734 / 67.501117 / 67.182469 tok/s. All 36 requests are cache-zero and four six-case exact suites pass. Full replication, fused argmax, tiny-pair exchange, and pre-gather local add remain rejected.
2026-07-18 DSpark7 persistent sharded Markov transaction promoted target-verified record ../notes/2026-07-18-dspark-persistent-markov-record.md, ../data/dspark-persistent-markov-record-20260718.json, vLLM 0873ffa67, XPU kernels 0b99fc536, /mnt/fast-ai/bench-results/deepseek-v4-flash-xpu/dspark7-xpu-persistent-markov-bundle-20260718T1640Z, LocalMaxxing cmrqiovsv05s6lg012d8v5nz8 Fixed device buffers and direct-output operations remove allocation/cat/copy overhead while retaining sharded W1/W2 work and TP4 collectives. The exact component saves 0.786613 ms/cycle at the slowest rank. Strict fresh suite medians are 65.674202 / 66.479103 / 63.558530 tok/s; all 36 requests are cache-zero and four six-case exact suites pass. This is one active generation, not aggregate. Full W2/LM-head replication is rejected; next replicate W1 only while W2 stays sharded.
2026-07-18 DSpark7 private PIECEWISE exact-M7 draft replay promoted target-verified record ../notes/2026-07-18-dspark-piecewise-exact-m7-record.md, ../data/dspark-piecewise-exact-m7-record-20260718.json, vLLM 48401ed6a, XPU kernels 0b99fc536, /mnt/fast-ai/bench-results/deepseek-v4-flash-xpu/dspark7-xpu-targetpw-draftpw-exactm7-20260718T0556Z, LocalMaxxing cmrpymqh505mxlg01tzg3e0yl A private breakable PIECEWISE graph captures only the official three-stage DSpark draft’s exact M=7 query while target verification remains proven M=8. Strict fresh suite medians are 64.661411 / 61.724506 / 64.275173 tok/s; all 36 requests are cache-zero and three six-case exact suites pass. This is one active generation, not aggregate. Full draft graph is correctness-rejected, DSpark5 loses, and padded-M8 reaches only 60.518331. Promote exact-M7 and profile the remaining context-KV/draft/sampler cycle.
2026-07-18 genuine sequential M=4/M=8 verifier and repeated-MTP screen verifier exact/predictor rejected ../notes/2026-07-18-sequential-mwidth-verifier-and-predictor-pivot.md, ../data/mwidth-sequential-verifier-20260718.json, vLLM 57cfb6771, XPU kernels 50646a2, sequential corpora and replay results under /mnt/fast-ai/ Real consecutive M=4/M=8 tensors close the duplicated-row evidence gap. Fixed MHC saves 1.441370 ms (20.74%) at M=4 and 4.314321 ms (34.91%) at M=8 versus segmented M2, with 70/70 graph replays exact on every B70. Repeating K160’s one-layer MTP three times reaches only 46.247281 tok/s on eight eligible cold rows; proposal three accepts 0.0-3.2%. Reject repeated MTP2+ and use the official three-stage DSpark draft with the frozen held-out gate.
2026-07-17 subgroup-split paired GEMM1 incremental upper bound closed before implementation ../notes/2026-07-17-mtp1-sg-split-incremental-upper-bound-closure.md, prior XPU c069ed8, ../notes/2026-07-17-mtp1-postportfolio-eager-cycle-profile.md The old fused producer saved at most 0.520384 ms/cycle on all-remote versus generic, while promoted route-direct already owns 0.397-0.414 ms of that scope. The generous incremental ceiling is only 0.123114 ms/cycle, and fresh trace activation work is 0.079980 ms. Since all-remote has no local projection arithmetic, subgroup splitting cannot clear the unchanged 0.50 ms every-route gate. Preserve the design for a larger specialized-decoder package; do not build or load it standalone.
2026-07-17 MTP1 post-portfolio eager cycle profile complete ../notes/2026-07-17-mtp1-postportfolio-eager-cycle-profile.md, ../data/eager-cycle-postportfolio-20260717-summary.json, ../scripts/serve-k160-mtp1-postportfolio-eager-profile.sh, /mnt/fast-ai/bench-results/deepseek-v4-flash-xpu/mtp1-postportfolio-eager-profile-20260717T032135Z Exact record source/selectors rerun as an eager diagnostic twin. Cross-rank noncollective work falls from 19.4779 to 17.8497 ms/cycle, a measured 1.6283 ms portfolio reduction. Dense remains 6.5639 ms; compact routed MXFP4 remains the largest open kernel family at 3.9424 ms. Kineto oneCCL and host durations remain excluded. Proceed to the subgroup-split/SLM-exchange M=2 MXFP4 hardware gate.
2026-07-16 MTP1 QNorm-M2 + route-direct N64 portfolio promoted ../notes/2026-07-16-qnorm-routeportfolio-record.md, ../data/qnorm-routeportfolio-20260716/summary.json, vLLM 4a6fd8747, XPU kernels 18a44f440, /mnt/fast-ai/bench-results/deepseek-v4-flash-xpu/qnorm-routeportfolio-candidate-b2-20260716T2255Z, LocalMaxxing cmrocpuhq029hlg01g3yzglko The standalone route component keeps the frozen 0.50 ms gate false at a 0.397 ms/cycle worst-card floor and is admitted only with the independently proven, non-overlapping QNorm-M2 floor. Four cards pass 336/336 changed graph cases bitwise; the guarded production wrapper passes 84/84. Same-binary B-A-B medians are 62.515661 / 61.717893 / 63.851301 tok/s, 70/70 ordered exact suites pass across positions 28/58, and every qualifying request is cached-zero. One max-8 diagnostic truncated strict JSON and is preserved as invalid harness evidence; the frozen canary uses max-32. Promote as the target-verified record.
2026-07-16 MTP1 M=2 MXFP4 N32/N128 policy exact isolated positive, not promoted ../notes/2026-07-16-mtp1-m2-mxfp4-policy-closure.md, ../data/mtp1-m2-mxfp4-policy-closure-20260716.json, ../scripts/probe-mxfp4-m2-policy.py, XPU kernels 351a06a442, /mnt/fast-ai/bench-results/deepseek-v4-flash-xpu/mtp1-m2-mxfp4-n128-candidate-20260716T0630Z Ordered prelaunch scheduler reset repairs N32/N128 graph replay. N32 is exact but loses 0.287-0.300 ms/43 layers. N128 is 48/48 exact on each B70 and saves 0.247-0.283 ms/43, but strict suites reach 62.649706/63.628477 versus a 63.349928 record while same-binary N64 controls span 61.205692-63.101865. Close without promotion or LocalMax submission; keep N64 and require an architectural MXFP4 change above 0.50 ms/cycle.
2026-07-16 native MTP1 M=2 router selection/normalization promoted ../notes/2026-07-16-mtp1-m2-router-record.md, ../data/mtp1-m2-router-record-20260716.json, vLLM 4a6fd8747, XPU kernels d15ce87d0, /mnt/fast-ai/bench-results/deepseek-v4-flash-xpu/mtp1-m2-router-norm-candidate-20260716T0605Z, LocalMaxxing cmrncv39w003ylg01hogleazo Exact trace arguments corrected the apparent indexer radix hotspot to 40 target-router [2,160], K6 calls/cycle. One submission now selects, normalizes, and scales both verifier rows. Four B70s pass 160/160 changed eager plus 128/128 graph epochs bitwise and save 1.123-1.128 ms/cycle. A same-build flag-off control is 59.108299 tok/s; independent candidate suites reach 62.882999/63.349928 tok/s, 70/70 ordered exact captures pass, and every request is cached-zero. Promote as the target-verified record.
2026-07-16 exact-identity eager MTP1 cycle profile complete ../notes/2026-07-16-mtp1-eager-cycle-profile.md, ../data/eager-cycle-profile-20260716-summary.json, ../scripts/summarize-eager-cycle-trace.py, /mnt/fast-ai/bench-results/deepseek-v4-flash-xpu/mtp1-record-eager-cycle-profile-20260716T0550Z Host-submission timestamps repair the profiler’s GPU clock offset; distorted oneCCL durations remain excluded. Cross-rank noncollective means are dense GEMM 6.580 ms, MXFP4 MoE 4.151, MHC 2.843, QK/LSE 1.317, PV 0.505. Exact top-k arguments prove 40 [2,160], K6 target-router calls/cycle. Schema v2 further decomposes the dense path: its largest families are already optimized or closed, making M=2 routed MXFP4 the largest genuinely open verifier family.
2026-07-16 native MTP1 M=2 MHC post/pre promoted ../notes/2026-07-16-mtp1-m2-mhc-record.md, ../data/mtp1-m2-mhc-record-20260716.json, vLLM 9cf403e51, XPU kernels 46b95e64a, /mnt/fast-ai/bench-results/deepseek-v4-flash-xpu/mtp1-m2-mhc-single-kernel-candidate-20260716T0210Z, LocalMaxxing cmrmvjbok1np3mj01p9il8486 One command launches two independent 256-thread verifier-row workgroups while preserving the proven M=1 reduction and BF16 arithmetic boundaries. All four B70s are bitwise exact and save 0.962-0.971 ms across 85 boundaries. Independent cold suites reach 59.291531/60.264242 tok/s versus 57.412142; 70/70 ordered exact captures pass across positions 28/58 and all requests are cached-zero. Promote as the target-verified record.
2026-07-16 XPU modular-MoE output alias exact/noise-floor negative ../notes/2026-07-16-xpu-moe-output-alias-negative.md, vLLM 8007ee686, /mnt/fast-ai/bench-results/deepseek-v4-flash-xpu/mtp1-moe-output-alias-candidate-20260716T0130Z The default-off alias removes an apparent routed-output copy and passes 10/10 exact capture suites, but strict runs reach only 57.204014/56.198992 tok/s versus the qualified 57.412142/56.952065 record/support. Graph construction already amortizes allocation, and the changed workspace lifetime removes too little replay work to survive full-model variance. Keep default-off; no LocalMax submission.
2026-07-16 MTP1 M=2 shared/routed fusion promoted ../notes/2026-07-16-mtp1-m2-fusion-record.md, ../data/mtp1-m2-fusion-record-20260716.json, vLLM 068d6beb2, XPU kernels 84c10f4f1, /mnt/fast-ai/bench-results/deepseek-v4-flash-xpu/mtp1-m2-fusion-candidate-20260716T000928Z, LocalMaxxing cmrmrgce51nojmj01bbxoruuu Extending exact shared clamped-SwiGLU/dynamic-FP8 quantization through M=2 and selecting exact fused clamp/SiLU inside generic M=2 routed MoE preserves expert grouping and cross-row weight reuse. Independent strict suites reach 57.412142/56.952065 tok/s versus the 55.703731 record. Seventy ordered exact capture suites pass across positions 28/58 after both suites; all 444 requests are cached-zero. Two direct-M1 M=2 chains are explicitly rejected because overlap/duplicate routes regress 1.4-3.2 ms/cycle.
2026-07-15 combined MTP1 + direct M1 routed MoE + wide collective epoch promoted ../notes/2026-07-15-mtp1-direct-moe-wideepoch-record.md, ../data/mtp1-direct-moe-wideepoch-record-20260715.json, /mnt/fast-ai/bench-results/deepseek-v4-flash-xpu/mtp1-direct-moe-wideepoch-candidate-20260715T2315Z, LocalMaxxing cmrmoyenp1no3mj01fz2gjzo6 All required commits compose without cherry-picks. Matching strict 12-prompt suites reach 55.703731/55.668081 tok/s, all cached-zero. Seventy ordered exact captures pass, including 50 after both suites and former rollover positions 28 and 58; every worker maps wide-epoch libccl. Promote as the target-verified record. Direct-M1 helps only the attached draft layer because the M=2 target verifier falls back unchanged, explaining the small but repeatable gain.
2026-07-15 late nonspec upper-bound closures performance-gate fail ../notes/2026-07-15-late-nospec-upper-bound-closures.md, ../data/late-nospec-upper-bound-closures-20260715.json, ../scripts/bench-m1-direct-gather-upper-bound.py, ../scripts/bench-next-weight-l2-prefetch.py, XPU kernel experiment/revert 5a7f39e9/46bdf344 Four cards bound exact GEMM2+gather deletion to 0.151-0.168 ms/token at the typical route and 0.230 ms maximum. Finite L2 hints on real 4/6/8 MiB dense weights preserve exact output but expose no consumer gain; immediate warm-cache sensitivity is only 0.884 us. Close both before model integration. No measured nonspec backlog candidate now clears 0.50 ms/token.
2026-07-15 direct M1 routed-MoE plus widened oneCCL epoch promoted ../notes/2026-07-15-direct-routed-moe-wideepoch-record.md, ../data/m1-direct-routed-moe-wideepoch-record-20260715.json, vLLM a681dbb2b, XPU kernels 6522849b0, oneCCL 48fda4f0e, /mnt/fast-ai/bench-results/deepseek-v4-flash-xpu/nospec-direct-moe-wideepoch-candidate-20260715T2220Z, LocalMaxxing cmrmnp7h81nntmj01lfenydgj Direct routed-MoE gather and exact router normalization raise the nonspec record to 43.766673 tok/s, with three support suites at 43.699/43.694/43.668. The direct-off control is 41.991/42.155. A reused 11-bit oneCCL readiness counter caused deterministic corruption at graph captures 28 and 58 even with fusion off; a 24-bit collective epoch plus 7-bit communicator tag repairs it without added work. The final identity passes 70/70 exact captures and 48/48 cached-zero strict rows. Promote.
2026-07-15 native SIMD16 M1 biased top-k promoted ../notes/2026-07-15-m1-biased-topk-record.md, ../data/m1-biased-topk-record-20260715.json, vLLM a66f3486c, XPU kernels 2a07cf2e8, candidate /mnt/fast-ai/bench-results/deepseek-v4-flash-xpu/nospec-m1-router-candidate-20260715T2021Z, paired control nospec-m1-router-control-20260715T2027Z, LocalMaxxing cmrmjd3io1nn1mj013stqoe4b SIMD16 register selection replaces generic bias-add/radix-top-k/gather in 40 normal M=1 MoE layers. Four cards pass 40/40 bitwise ID/raw-weight epochs; the isolated boundary improves 77.128 -> 7.178 us. Two strict nonspec suites reach 41.513661/41.733256 tok/s versus a same-commit flag-off control at 40.067691, saving 0.87-1.00 ms/token. Twenty ordered exact captures pass before/after the suites and every request is cached-zero. Promote as the trustworthy nonspeculative record.
2026-07-15 MTP1 exact strided-batch compressor promoted ../notes/2026-07-15-mtp1-batched-compressor-record.md, ../data/mtp1-batched-compressor-record-20260715.json, four-card compressor-m2-bmm-exact-card*-20260715.json, vLLM 3bd0eb321, /mnt/fast-ai/bench-results/deepseek-v4-flash-xpu/mtp1-rowexact-bmm-w8a16-m2-candidate-20260715T2000Z, LocalMaxxing cmrmgacdq1nmimj01i4sfqytp Real K160 C4/C128 compressor shapes pass 40/40 changing eager and graph-replay comparisons on every B70. One strided-batch BMM is 1.84-1.91x faster than two M=1 calls plus concatenation. Strict suites reach 55.524496/54.708889 tok/s, 20/20 sustained exact captures pass, every request is cached-zero, and acceptance is 77.96%. Promote as the current target-verified record.
2026-07-15 MTP1 target-verifier selective W8A16 M=2 promoted ../notes/2026-07-15-mtp1-w8a16-m2-record.md, ../data/mtp1-rowexact-w8a16m2-record-20260715.json, four-card w8a16-m2-row-invariance-card*-20260715.json, /mnt/fast-ai/bench-results/deepseek-v4-flash-xpu/mtp1-rowexact-w8a16-m2-candidate-20260715T1945Z, LocalMaxxing cmrmfivhg1nmamj012e3138my All four production shapes pass 40/40 changing row-exact cases on every B70; one M=2 W8A16 call is 2.42-2.50x faster than two M=1 calls. Strict suites reach 54.464909/54.445287 tok/s, 20/20 sustained exact captures pass, all cached-zero, and acceptance remains 77.68%. Promote as the current target-verified record.
2026-07-15 repeated single-layer MTP2 correctness-prepass then deadlock ../notes/2026-07-15-mtp2-reuse-deadlock-closure.md, vLLM 4e47b18c9, /mnt/fast-ai/bench-results/deepseek-v4-flash-xpu/mtp2-rowexact-graph-oneccl1712-20260715T1900Z Generalized row-exact compressors let M=3 graph capture and 10/10 initial exact captures pass, but the reused layer’s second draft position accepted only about 0.5-2.2% on realistic prompts. One request then hung and the engine exhausted shared-memory broadcast blocks for 180 seconds. No valid suite or speed claim. Close MTP2+ and restore MTP1.
2026-07-15 attached MTP1 with row-exact verifier compressors promoted ../notes/2026-07-15-mtp1-rowexact-record.md, ../data/mtp1-rowexact-record-20260715.json, /mnt/fast-ai/bench-results/deepseek-v4-flash-xpu/mtp1-compressor-rowexact-graph-oneccl1712-20260715T1840Z, LocalMaxxing cmrmetch81nm3mj01w1pidsyt Plain MTP1 reached 50.74/50.10 tok/s but later leaked prompt text after 437, so it is rejected. Running the two-token FP32 compressor as two exact M=1 projections repairs sustained graph replay: 20/20 ordered exact captures pass, including ten after the two strict suites. Qualified medians are 50.016860/49.420459 tok/s, with 77.42% measured acceptance and every request cached-zero. Promote as the speculative record; retain 40.170350 as the base record.
2026-07-15 KV-overlap and large-allreduce repeatability repair promoted ../notes/2026-07-15-kv-repeatability-and-oneccl-allreduce-routing.md, ../data/oneccl-allreduce-routing-record-20260715.json, vLLM 93fde4186, oneCCL 6da44bc, /mnt/fast-ai/bench-results/deepseek-v4-flash-xpu/nospec-graph-oneccl1712-bf16-allreduce128k-preload094-cache-fix-20260715T1530Z, LocalMaxxing cmrmebmzg1nm0mj01k30nv6vw Removed overlapping BF16 KV/RoPE stores, localized deterministic graph-prefill corruption to large SYCL all-reduce, and routed only all-reduces above 128 KiB through the safe path in an exact-version oneCCL 2021.17.2 build. Ten exact captures pass 10/10; two cold suites pass at 40.096205/40.170350 tok/s. This is the current trustworthy base identity; the older 40.135724 submission remains historical speed evidence but is not repeatability-certified.
2026-07-13 investment-red-team complete ../data/fit-audit-20260713.json, ../../../plans/2026-07-13-deepseek-v4-flash-b70-investment-gated-plan.md Strategic go; reject direct K180 commitment. Run Stages 0-3.5 before download, build K160 first, and climb only after quality/warm-memory gates.
2026-07-13 storage-runtime-download-start active ../scripts/download-k160.sh, ../scripts/capture-stage0.sh, clean worktrees under /home/steve/src/deepseek-v4-* Archived 170 GiB of reviewed inactive artifacts with compatibility symlinks; internal free space rose from 11 to about 180 GiB. Prioritize frozen public uniform-K160 as the first runnable checkpoint.
2026-07-13 k160-provenance-audit complete ../quality/calibration-v1-plan.json, ../quality/suite-v1.json Public K160 is valid for smoke/performance bring-up but not quality-certified: hash layers are pruned and published observations are not true REAP. Preserve the official-source teacher and hash-preserved final-pack lanes.
2026-07-13 exact-shape-test-scaffold implemented XPU-kernel commit 552c9ce, ../scripts/run-exact-shape-gates.sh Added low-level H4096/I2048/top-k6/M1,4,8 correctness coverage for MXFP4 and INT4 controls at E=40/64. This is not yet the Stage-1 performance, selector, replay, fallback, or TP4/EP gate.
2026-07-13 runtime-build-j16 loss resumable build/cache under /home/steve/src/deepseek-v4-xpu-kernels-clean The Xe2 grouped-GEMM translation unit was killed under 16-way SYCL compilation. Preserve completed objects/ccache and resume at the durable eight-job default; this is host build-memory pressure, not a kernel result.
2026-07-14 runtime-build-j8 pass ../data/runtime-pin-20260714.json, /mnt/fast-ai/bench-results/deepseek-v4-flash-xpu/stage0-20260714T044033Z.txt Incremental eight-job Xe2/SYCL-TLA build completed in 10m40s; pinned vLLM/kernel imports resolve to clean worktrees and all four B70s enumerate. The selector fix prevents -k 40 from also matching E=64 via dimension 4096.
2026-07-14 exact-shape-scaffold-preflight infrastructure-fail /mnt/fast-ai/bench-results/deepseek-v4-flash-xpu/stage1-20260714T044101Z Clean environment lacked pytest; no GPU case ran. Installed and pinned pytest 9.0.2 before retrying.
2026-07-14 exact-shape-scaffold scaffold-pass ../data/exact-shape-scaffold-20260714.json, /mnt/fast-ai/bench-results/deepseek-v4-flash-xpu/stage1-20260714T044200Z All 12 exact M=1/4/8, E=40/64 MXFP4/INT4 low-level reference cases passed on four B70s. This does not clear Stage 1A/1B; performance, metrics, selector, fallback, replay, and TP4+EP evidence remain.
2026-07-14 four-card-xccl-preflight pass ../data/runtime-pin-20260714.json, /mnt/fast-ai/bench-results/deepseek-v4-flash-xpu/xccl-preflight-20260714T044502Z.log All four per-device smokes, four-rank XCCL init/barrier, and allreduce passed over oneCCL/OFI on eno1; topology correctly reports PCIe/NODE rather than direct fabric.
2026-07-14 k160-download-resume active ../data/k160-download-start-20260713.json, ../scripts/download-k160.sh Five completed shards survived the 16-way host-memory event. The frozen revision resumed through the external Xet cache after constraining compilation to eight jobs; final HF/SHA-256 verification remains mandatory.
2026-07-14 k160-artifact-promotion pass ../../../data/deepseek-v4-k160-tp4-bringup-20260714.json, /mnt/fast-ai/llm-models/deepseek-v4-flash-xpu/current-k160 All 46 shards, 43,843 index entries, safetensor headers, HF metadata, and SHA-256 values passed. Archive and hot copy are verified; K160 remains an experimental smoke checkpoint, not a quality promotion.
2026-07-14 k160-tp4-eager-construction pass /mnt/fast-ai/bench-results/deepseek-v4-flash-xpu/tp4-smoke-20260714T132607Z TP4+EP assigned 40/160 experts per rank and selected XPUExpertsMxFp4. The 2K/95% configuration fits at 24.95 GiB model plus 2.11 GiB KV per rank; the arithmetic canary returned 1073. Earlier 8K/90% and 2K/98% memory attempts remain preserved as failures.
2026-07-14 k160-tp4-eager-decode performance-fail /mnt/fast-ai/bench-results/deepseek-v4-flash-xpu/tp4-smoke-20260714T132607Z/steady-long-128.json, ../../../data/deepseek-v4-k160-tp4-bringup-20260714.json The second warm 128-token diagnostic row reached only 2.616225 tok/s after TTFT. This is about 19x below the 50 tok/s gate; speculation remains prohibited.
2026-07-14 k160-breakable-xpu-graph infrastructure-fail /mnt/fast-ai/bench-results/deepseek-v4-flash-xpu/tp4-smoke-20260714T133502Z/server.log Capture fails when sparse FP8 decode executes combined_lens.max().item(), causing a prohibited host wait on a Level Zero command-graph event. Remove or make that scalar graph-safe before another graph performance run.
2026-07-14 k160-static-sparse-pack pass vLLM commit 0ed5ecc5, ../data/graph-recovery-20260714.json Replaced graph-unsafe host scalar reads and per-token packing with fixed-width device-only packing. Added finite masked-chunk sentinel. Exact graph replay with changed lengths and focused tests pass; the first one-kernel Triton attempt segfaulted and remains a preserved negative.
2026-07-14 sycl8-jit-lane pass ../notes/2026-07-14-xpu-graph-recovery-and-tp4-profile.md, ../scripts/serve-k160-tp4-smoke.sh Removed umbrella oneAPI setup and kept venv SYCL/UR libraries ahead of side-by-side oneAPI 2026. This fixes the libsycl.so.9/urDeviceWaitExp JIT ABI mismatch. Quarantined contaminated cache entries instead of deleting them.
2026-07-14 k160-piecewise-graph performance-pass /mnt/fast-ai/bench-results/deepseek-v4-flash-xpu/tp4-graph-piecewise-20260714T1015Z Correct 1073 canary; fresh 128-token headline reached 8.512062 tok/s after TTFT versus 2.616225 warm eager. Graph capture is now reusable, but throughput remains far below the 40-50 tok/s base gate.
2026-07-14 k160-full-narrow-indexer-break performance-pass vLLM commit 436298dcd, /mnt/fast-ai/bench-results/deepseek-v4-flash-xpu/tp4-full-narrow-indexer-20260714T1020Z Native paged-indexer scratch memory cannot enter SYCL graphs. A narrow FULL-mode eager break gives correct replay and 8.616232 tok/s fresh headline, only 1.22% over PIECEWISE. Attention graph coverage is not the dominant remaining boundary.
2026-07-14 k160-tp4-full-profile diagnostic /mnt/fast-ai/bench-results/deepseek-v4-flash-xpu/tp4-full-profile-20260714T1024Z Profile observes 87 c10d all-reduces per decoder step, approximately two across 43 layers. Profiler collective time is distorted and not wall latency; call count plus normal 116 ms/token points to PCIe TP synchronization as the next high-value boundary.
2026-07-14 k160-tp2-dp2-ep4 infrastructure-fail /mnt/fast-ai/bench-results/deepseek-v4-flash-xpu/tp2-dp2-ep4-eager-kv64m-20260714T1047Z, ../data/graph-recovery-20260714.json Model fits at 26.66 GiB/card with a manual 64 MiB KV cache. Both graph modes stall in the mixed dummy run; eager first request leaves all ranks waiting in DPEP all-gather. No performance claim. Close until eager XPU DPEP collectives work.
2026-07-14 k160-native-mhc-and-c4-bounds performance-pass vLLM history through cff886b, LocalMaxxing cmrku8z0l05ahmj01raa1794f, cmrkuloel05almj015vfgot4o, cmrkuv0is05dwmj01e6mdztv7 Graph-safe native mHC plus context-bounded C4/C128 sparse work raised the strict cold suite from 10.0161 to 14.3784 tok/s. C4 full-selection bypass is valid only because a 1K context contains at most 256 compressed C4 candidates for top-512.
2026-07-14 k160-direct-fp8-sparse-attention promoted vLLM cff886bb7165a7328f6412c199bb93c2fbdcfb98, /mnt/fast-ai/bench-results/deepseek-v4-flash-xpu/tp4-direct-fp8-attn-20260714T1500Z, LocalMaxxing cmrkw8db205nymj01gl5bbmsc A graph-captured Triton path reads paged UE8M0 FP8 KV directly and applies runtime lengths plus learned sinks. Strict cold median improved 49.84%, from 14.3784 to 21.5448 tok/s; canaries and all 12 cached-zero rows pass.
2026-07-14 rank1-xe-device-loss-recovery recovered kernel journal for PCI 0000:27:00.0; four-device compute and XCCL post-recovery smokes Two launches hit UR_RESULT_ERROR_DEVICE_LOST after a Xe timeout storm. A PCI function reset left a GT PF self-configuration error; targeted Xe driver unbind/rebind restored the card. All four independent compute tests and TP4 barrier/allreduce passed. Prefer targeted rebind for this symptom and preserve the installed-vs-recommended GuC 70.44.1/70.49.4 firmware warning for maintenance.
2026-07-14 split-attention-wrong-identity rejected /mnt/fast-ai/bench-results/deepseek-v4-flash-xpu/tp4-split-fp8-attn-recovery-20260714T1728Z The first post-recovery split run accidentally disabled native mHC, used 2K context/batch, and omitted prompt-token details. Its 19.74 tok/s and corrupted long outputs are confounded and are not kernel evidence. This is a benchmark-identity failure, not a split-attention conclusion.
2026-07-14 k160-split-tiled-fp8-attention promoted vLLM b63b1f2b2106a26128a8a0dae55855493b3ada1d, /mnt/fast-ai/bench-results/deepseek-v4-flash-xpu/tp4-split-fp8-attn-recordidentity-20260714T1733Z, LocalMaxxing cmrkxoavs05uimj01p9ix2dtk QK/LSE plus 8x64 tiled PV cuts register pressure while retaining direct FP8 cache reads, runtime lengths, and learned sinks. Strict cold median is 29.82238 tok/s, p10 29.42632, +38.41% over direct attention; all exact canaries and 12 cached-zero rows pass. Current record.
2026-07-14 k160-split-residual-profile diagnostic /mnt/fast-ai/bench-results/deepseek-v4-flash-xpu/tp4-split-eager-profiler-20260714T1737Z Steady eager decode reconciles to about 23.3 ms of non-collective kernels: dense GEMMs 9.1 ms, MXFP4 MoE 3.67 ms, native mHC 2.69 ms, split QK 1.94 ms, indexer select/sort 0.98 ms, activation quant/scale copies 0.91 ms, split PV 0.28 ms, and about 3.4 ms remaining. Normal graph token period is 33.53 ms; real collectives contribute roughly 8 ms and residual gaps/host about 2 ms. Profiler oneCCL durations are distorted and must not be quoted.
2026-07-14 k160-fp8-wo-a performance-loss vLLM 4b2d2e0efaee9e08ebe9d77392cead912fdb6909, /mnt/fast-ai/bench-results/deepseek-v4-flash-xpu/tp4-fp8-woa-canary-20260714T1810Z The existing fused inverse-RoPE/UE8M0 quantizer plus true FP8 BMM passed M=1/4 correctness and graph replay. Exact G2/M1/K4096/N1024 microbench improved BF16 BMM 30.79→24.89 us and full isolated path 55.97→39.78 us, but the strict suite regressed from 29.82238 to 29.51408 tok/s (-1.03%, p10 28.96445). All cold/canary gates passed. Preserve default-off; do not submit.
2026-07-14 tp-only-inplace-allreduce promoted vLLM bfac86e7d22418a88b0f50f4593aa1e2281d3f2d, /mnt/fast-ai/bench-results/deepseek-v4-flash-xpu/tp4-inplace-ar-candidate-20260714T1920Z, LocalMaxxing cmrkz6vo7061fmj01v09nwz72 A mutation-declared custom op removes clones only for TP-group contiguous BF16 4096-element decode outputs. Changed-input replay 1073 -> 437 -> 1073 and exact canaries passed. Two strict suites reached 29.9133 and 29.9113 tok/s versus 29.8224 prior (+0.30%); this is a valid record but proves XCCL wait dominates clone cost.
2026-07-14 pp4-tp1-single-session performance-loss /mnt/fast-ai/bench-results/deepseek-v4-flash-xpu/pp4-tp1-split-attn-20260714T1945Z/steady-long-128.json PP4 partitioned the 43 layers as [11,11,11,10], fit at 22.62-24.99 GiB/card, captured reusable graphs, and passed changed-input plus exact canaries with zero cached tokens. Fresh 128-token decode reached only 16.7988 tok/s after TTFT versus 29.9133 on TP4. A batch-one token traverses pipeline stages serially, forfeiting four-way concurrent weight bandwidth; three small pipeline transfers cannot compensate. Reject PP4 and PP2/TP2 as base-decode speed lanes.
2026-07-14 mhc-postpre-rmsnorm-fusion performance-loss vLLM 520df585c06c5227f2bdbafd3f4789b9bcca9dd2, XPU kernels 473a55e2a8b34da3c97c143401955d0c5746120b, /mnt/fast-ai/bench-results/deepseek-v4-flash-xpu/tp4-mhc-norm-fusion-20260714T1840Z The opt-in kernel preserves the BF16 producer boundary and standalone RMSNorm cast order; M=1/4 parity, command-graph replay, changed-input, and exact canaries pass. Two fresh screens reached 29.4080 and 29.4411 tok/s versus the 29.4844 comparison screen. The saved launch is offset by an in-kernel reduction/fence and reread. Preserve default-off; only revisit as part of producer + norm + consumer quantization fusion.
2026-07-14 xccl-twoshots-8k microbench-win/end-to-end-loss oneCCL 4ceafd15c03ce46f11eeaf91781a92afebd3cecf, /mnt/fast-ai/bench-results/deepseek-v4-flash-xpu/xccl-8k-sweep-20260714T1845Z, /mnt/fast-ai/bench-results/deepseek-v4-flash-xpu/tp4-twoshots-20260714T1848Z CCL_SYCL_ALLREDUCE_LL=twoshots reduced synchronized standalone BF16 8 KiB rank-0 median from 128.6915 to 87.369 us and all ranks were correct. Exact graph decode nevertheless measured only 29.4212 and 29.3399 tok/s, below the 29.4844 screen. The host-synchronized microbench does not model reusable command-graph critical-path behavior. Keep the knob explicit in identity but retain ring as default; do not submit.
2026-07-14 fp8-static-scale-prepack invalid/superseded vLLM a77aadb26ac27699aa0537915547cf67c7ab3281, /mnt/fast-ai/bench-results/deepseek-v4-flash-xpu/tp4-fp8-scale-prepack-20260714T2110Z, LocalMaxxing cmrl0rf5u06b6mj01y1s9ew2u The generic prepack also transposed DeepSeek’s special wo_a BMM scales, whose BF16 cache expects canonical layout. Short canaries missed the numerical corruption. Preserve the public ID as invalid history; do not cite 30.2953 as a record.
2026-07-14 exact-fp8-dense-shape-microbench diagnostic scripts/bench-fp8-dense-shapes.py, /mnt/fast-ai/bench-results/deepseek-v4-flash-xpu/fp8-dense-shapes-w8a16-20260714T1920Z.json Rank-0 tracing identified five repeated M=1 TP4 shapes per layer: (N,K)=(1536,4096),(8192,1024),(4096,2048),(1024,4096),(4096,512). Direct BF16-activation/block-FP8-weight oneDNN was 1.23-1.88x faster than quantized W8A8 GEMM on every shape and removed the separate quantizer; the weighted eager microbench fell from about 8.16 ms/token for quant+W8A8 to 4.01 ms/token for W8A16. Use graph end-to-end results, not standalone event totals, for promotion.
2026-07-14 block-fp8-w8a16-small-m invalid/superseded vLLM 9fe91a6d6c36806b0428b6c3487bd10b05eee20c, /mnt/fast-ai/bench-results/deepseek-v4-flash-xpu/tp4-block-fp8-w8a16-20260714T1930Z, LocalMaxxing cmrl12pke06ehmj01i9a1f0gu This inherited the corrupted wo_a scale layout. Preserve 33.8866 and its public ID as invalid history; it is not the frontier.
2026-07-14 corrected-w8a8-scale-prepack promoted vLLM 61c87db645c256651b5a366f538898485077ad32, XPU kernels d553fd2ac0cfc86edbb4fe9c65d567318931fe91, /mnt/fast-ai/bench-results/deepseek-v4-flash-xpu/tp4-w8a8-woa-corrected-n64-20260714T1940Z, LocalMaxxing cmrl2619q06hwmj011j5rtnbt Preserve canonical wo_a BMM scales and prepack only standard dense scales. Two strict serial suites reached 30.2304 and 30.2390 tok/s, all 12 rows cached-zero; sequential changed-input and exact canaries pass. This is the trustworthy frontier.
2026-07-14 corrected-w8a16-small-m quality-rejected vLLM 61c87db645c256651b5a366f538898485077ad32, /mnt/fast-ai/bench-results/deepseek-v4-flash-xpu/tp4-w8a16-woa-corrected-n64-20260714T1935Z, /mnt/fast-ai/bench-results/deepseek-v4-flash-xpu/tp4-w8a16-woa-corrected-n64-quality-20260714T2020Z Corrected speed is 34.0145/33.9236 tok/s, but early greedy parity with W8A8 is 83.3% and the frozen long math-invariant case corrupts/terminates while W8A8 returns 101! - 1. Keep default off; do not submit.
2026-07-14 mxfp4-small-m-n32-n128 rejected/unpromoted XPU kernels d553fd2ac0cfc86edbb4fe9c65d567318931fe91, /mnt/fast-ai/bench-results/deepseek-v4-flash-xpu/tp4-w8a8-woa-corrected-n32-20260714T1950Z, /mnt/fast-ai/bench-results/deepseek-v4-flash-xpu/tp4-w8a8-woa-corrected-n128-20260714T2000Z N32 passes low-level parity but fails exact changed-input replay. N128 reaches 30.5181/30.4817 tok/s, about 0.9% over N64, but changes early greedy output and does not justify a quality-changing promotion. Keep N64.
2026-07-14 tp4-collective-protocol-pivot custom-protocol-fail/control-path-pass ../notes/2026-07-14-tp4-collective-protocol-pivot.md, ../data/tp4-collective-protocol-gate-20260714.json, XPU kernels b84ac23, oneCCL 0277eab Local-mailbox polling and remote-atomic notification are unreliable on cross-card B70 mappings. The actual 8 KiB route is oneCCL’s in-band-sequenced Rt64_128_PCIE ring, not the peer-atomic small path. A direct replay hook is bitwise exact for 40 changing epochs but slower than normal XCCL (27.874 versus 21.666 us aligned max-rank median). Use it only to build the guarded ring-final-writeback plus MHC-post fusion.
2026-07-14 tp4-ring-writeback-mhc-post correctness-pass/performance-loss vLLM d7883b27a, XPU kernels 8e301dc, oneCCL edf0e17, ../data/tp4-ring-mhc-post-fullmodel-20260714.json, /mnt/fast-ai/bench-results/deepseek-v4-flash-xpu/tp4-ring-mhc-fusion-gmem94-20260714T221859Z All three ring final-writebacks now forward the raw reduced message and locally emit exact MHC-post residuals. Forty changing microgate epochs and real full-model graph replay pass; 1073 -> 437 -> 1073, copy, Paris, JSON, and all 12 cached-zero rows are exact. The strict median is 29.5955 tok/s versus the 30.2390 frontier (-2.13%). This partial boundary loses because it replaces the baseline fused MHC-post+pre kernel with ring+post followed by standalone MHC-pre. Do not submit; fold the following MHC-pre into ring completion before retesting.
2026-07-14 tp4-ring-mhc-post-pre-epilogue correctness-pass/performance-loss vLLM 73e39f50e, XPU kernels 80bc48d, oneCCL 8567899, ../data/mhc-post-pre-fusion-20260714.json, /mnt/fast-ai/bench-results/deepseek-v4-flash-xpu/tp4-ring-mhc-post-pre-gmem94-20260714T1852Z Forty changing cases and graph replay were bitwise exact, but putting six MHC-pre projection passes after the ring in the same 1024-thread kernel fell to 27.0187 tok/s. Graph replay already amortizes launch overhead; added register/work footprint backpressures the collective. Reject.
2026-07-14 tp4-ring-post-single-pre correctness-pass/performance-neutral vLLM 550117059, XPU kernels 45bd9ff, oneCCL 8567899, /mnt/fast-ai/bench-results/deepseek-v4-flash-xpu/tp4-ring-post-single-pre-gmem94-20260714T1906Z A dedicated one-launch M=1 pre kernel is exact and cuts isolated pre from 100.62 to 79.43 us. Moving it outside the ring recovered 29.6193 tok/s, but the custom ring remained slower than production oneCCL. Do not promote.
2026-07-14 mhc-post-pre-m1-single-kernel promoted vLLM f589f0f72, XPU kernels fa530f5fe, ../data/mhc-post-pre-fusion-20260714.json, /mnt/fast-ai/bench-results/deepseek-v4-flash-xpu/mhc-post-pre-m1-single-gmem95-20260714T1921Z, LocalMaxxing cmrl9xiwe06zzmj01cof0k38p A 256-thread Xe2 kernel preserves exact post accumulation and the BF16 boundary, then performs MHC-pre after a convergent barrier. It reduces the isolated boundary 119.262 -> 80.574 us. Forty changing cases and graph replay are bitwise exact; three strict suites reached 30.3404/30.2138/30.2405 tok/s, combined 36-prompt median 30.2707, all cold/cached-zero with exact canaries. This is the new trustworthy record.
2026-07-14 selective-block-fp8-w8a16-high4 promoted vLLM 726c1acfb, XPU kernels fa530f5fe, ../data/w8a16-shape-isolation-20260714.json, /mnt/fast-ai/bench-results/deepseek-v4-flash-xpu/w8a16-high4-no-shared-down-20260714T2346Z, LocalMaxxing cmrlb675r0705mj01k9psoub0 Shape isolation showed that all five projection families individually pass the frozen invariant; global W8A16 fails only after accumulated cross-family drift. Enable W8A16 for fused WQA/WKV, Q-B, O-B, and shared gate/up while retaining W8A8 for low-value, logit-sensitive shared-down. Two strict cold suites reached 33.4339/33.3632 tok/s, all cached-zero; sequential replay, exact canaries, 101! - 1, and executable quality gates pass. K160’s intermittent CJK corruption floor remains explicitly documented. This is the trustworthy single-session TP4 record.
2026-07-14 tp2-dp2-ep4-dpep-recovery correctness-pass/performance-closed ../notes/2026-07-14-tp2-dp2-dpep-recovery.md, ../data/tp2-dp2-dpep-recovery-20260714.json, vLLM 5e02991f3 Fixed equal/uneven XPU all-gatherv semantics and proved the original stall is a oneCCL fast-SYCL communicator-switch cycle between disjoint TP and crossed DP pairs. CPU/Gloo, global rendezvous, global generic oneCCL, and fast-TP/generic-DPEP all return exact 42; the best fresh 64-token screen is only 2.495917 tok/s. Close this topology until a communicator-scoped fast-SYCL or dedicated fused DPEP transport exists.
2026-07-15 shared-expert-clamped-swiglu-fp8-quant promoted vLLM 38260cda8, XPU kernels ae8151234, ../notes/2026-07-15-shared-expert-act-quant-record.md, ../data/shared-expert-fused-act-quant-20260715.json, /mnt/fast-ai/bench-results/deepseek-v4-flash-xpu/shared-expert-fused-act-quant-20260715T0140Z, LocalMaxxing cmrlf1hn609glmj019rsjdl4r A guarded Xe2 M=1 kernel preserves the clamp-at-10 BF16 tensor-op boundaries while fusing shared-expert SwiGLU with dynamic per-128 E4M3FN quantization. Eight bitwise device cases pass. The isolated boundary improves 77.326 -> 15.462 us (5.00x); paired strict cold suites reach 34.0671/34.0497 tok/s, all cached-zero, with replay, exact canaries, executable gates, and frozen 101! - 1 passing. This is the trustworthy TP4 single-session record; continue with register-resident MHC post/pre + exact RMSNorm. Dual FP8 output is deferred because selective W8A16 currently bypasses it.
2026-07-15 mhc-post-pre-m1-rms exactness/performance-closed vLLM 8bebe092f, XPU kernels efebdae, ../notes/2026-07-15-mhc-post-pre-rms-fusion-loss.md, ../data/mhc-post-pre-m1-rms-loss-20260715.json, guarded XPU operator and one-B70 probe The WG512/SG32 candidate preserves the old 256 logical MHC lanes, stages rounded BF16 in aligned SLM, and applies standalone-geometry RMSNorm in the same command. Forty changed states exposed 14 post-mix, 309 comb-mix, and 3 normalized-BF16 bit mismatches. It also regressed 20.326 -> 22.427 us, projecting a 0.179 ms/token loss over 85 boundaries. Graph replay itself passed. Close before any TP4 server run and return to the ordered collective critical path.
2026-07-15 pcie-aspm-performance-policy system-probe-failed/reboot-required ../notes/2026-07-15-aspm-device-recovery-blocker.md, ../data/aspm-device-recovery-blocker-20260715.json External links are healthy Gen4 x16, but changing global ASPM policy to performance stalled four-rank communicator setup and left three cards inaccessible in D3cold/runtime-error state after restoring default. FLR, xe unbind/rebind, and bus reset failed. No timing or performance claim; reboot before TP4 and do not repeat this mutation on the current driver/kernel.
2026-07-15 late-tp4-collective-and-placement-gates closed/rejected ../notes/2026-07-15-late-tp4-collective-and-placement-gates.md, ../data/late-tp4-gates-20260715.json, vLLM 62b8bed9, oneCCL c6aec66 and bef2321 LL workgroup geometry saves at most 0.360 ms/87 against mean controls and fails the 0.50 ms gate. Exact recursive doubling is 0.0656 ms/87 slower than paired ring controls. Interleaved expert ownership reaches the requested map but corrupts the first replay (1369 -> 361 -> 1369 versus 1073 -> 437 -> 1073), proving the packed MXFP4 path is not fully expert-map-clean. Preserve all three default-off; only the resident per-wire MHC consumer retains a positive measured upper bound.
2026-07-15 tp4-ring-readiness-markers prerequisite-pass ../notes/2026-07-15-ring-readiness-marker-gate.md, ../data/late-tp4-gates-20260715.json, oneCCL 1edec457 Release-ordered per-wire epochs in the unused LL scatter-buffer tail pass exact smoke and 24 changed epochs on every rank. Across 87 reductions the guarded candidate is only 0.038792 ms slower than the faster paired control, a conservative 0.446 us/boundary marker tax below the 1 us gate. This is synchronization infrastructure, not a decode record; proceed to the resident MHC consumer and require >=6 us/boundary saved before model integration.
2026-07-15 tp4-resident-mhc-consumer forward-progress-rejected ../notes/2026-07-15-resident-mhc-consumer-forward-progress-failure.md, ../data/tp4-resident-mhc-consumer-20260715.json, oneCCL 1e00a13, XPU kernels eac8ed6 All ranks derive epoch N while observing N-2, but the ring marker never advances while the normal-priority resident workgroup polls; bounded candidates take 294-314 ms versus 225-307 us and become exact only after timeout releases the ring. A low-priority queue makes no progress for >30 s. Finite two-stream overlap does not provide persistent cross-queue forward progress on this runtime. Reject before model integration; screen a compact 256-thread in-ring post/pre kernel instead.
2026-07-15 compact-tp4-ring-mhc-post-pre correctness-rejected ../notes/2026-07-15-compact-ring-mhc-post-pre-closure.md, ../data/compact-ring-mhc-post-pre-20260715.json, vLLM af8b8d776, XPU kernels 83ef7b6, oneCCL 88d90e0/8f415a1 The 256-thread microgate is exact and cuts the honest isolated boundary from 222.7-248.4 to 101.9-105.4 us, but full-model repeated changed inputs are nondeterministic. The exact 87-position graph, 42 alias pairs, four replays, rank skew, dependent producers, stable double buffers, 512 threads, and an explicit producer barrier do not repair it. Reject before speed; a future retry requires real-model intermediate-tensor capture.
2026-07-15 split-fp8-b4-qk16-geometry promoted ../notes/2026-07-15-split-fp8-geometry-record.md, ../data/split-fp8-geometry-record-20260715.json, vLLM fa3e27b461, /mnt/fast-ai/bench-results/deepseek-v4-flash-xpu/split-fp8-geometry-b4-qk16-recordidentity-20260715T0144Z, LocalMaxxing cmrlnp01l12q4mj01p58ynsyd Corrected profiling showed the 43-call BF16 kernel was prefill, then attributed 2.74 ms/token to split QK. Replacing four 16-head/8-warp QK programs with sixteen 4-head/16-warp programs cuts complete split attention 22-42% across short/128-token C4/C128 shapes. Focused output is bitwise exact; changed-input graph replay is 1073 -> 437 -> 1073. The strict cold suite reached 40.020972 tok/s median and 39.608039 p10, +17.48% over 34.067121, with all 12 cached-zero rows. This clears the 40 tok/s base gate.
2026-07-15 record-lane-noncollective-gates closed/rejected ../notes/2026-07-15-record-lane-noncollective-gates.md, ../data/record-lane-noncollective-gates-20260715.json, vLLM experiment/revert history through 44c5f959f, XPU scheduler fix/revert a7a0300/eeff530, /mnt/fast-ai/bench-results/deepseek-v4-flash-xpu/paired-control-20260715T0940Z The paired control reproduces 40.023086 tok/s. Correct profiling attributes 6.582/3.479/2.890/1.452 ms per token to dense/MXFP4/MHC/split-attention device kernels. Auxiliary streams regress, generic C4 fusion models skipped indexer work and reaches at most 39.6759, and approximate Triton compressor GEMV changes all 12 hashes at 39.7249. An in-kernel MXFP4 counter-reset race explains prior N32/N128 output changes; an ordered reset restores exact replay, but fixed N32 is immaterial and fixed N128 is slower than N64, so the fix was reverted. Record unchanged; require an exact >=0.50 ms/token real-model gate before another TP4 integration.
2026-07-15 real-mhc-capture-and-graph-fence correctness-root-cause-pass/performance-rejected ../notes/2026-07-15-real-mhc-capture-and-graph-fence-closure.md, ../data/real-mhc-capture-and-graph-fence-20260715.json, vLLM capture/fence/revert history through bba83e018, XPU kernels 748a59f, /mnt/fast-ai/bench-results/deepseek-v4-flash-xpu/real-mhc-boundary-capture-20260715T1200Z A validated real M=1 corpus preserves all 87 reductions, 85 MHC post/pre calls, final post, and 42 alias boundaries on every rank: 692 files, 571,072,236 bytes, aggregate SHA 6f8b7b9e...a5a59eb. The compact candidate is bitwise exact in eager mode and over eight graph replays on these real values. Full-model observers isolate the former corruption to a missing post-kernel graph-visible completion edge; a one-BF16 post read repairs six alternating requests. The repaired path is still rejected at 34.708355 tok/s, 13.28% below the 40.020972 record, with all 12 strict hashes changed. Close collective/MHC fusion; source and binaries are restored and no LocalMax submission is warranted.
2026-07-15 tp4-rank-arrival-trace measurement-invalid/closed ../notes/2026-07-15-tp4-rank-arrival-trace-closure.md, ../data/tp4-rank-arrival-trace-20260715.json, oneCCL b6b6481/14db31d, vLLM fc03ca89f/8721e07b4 A default-off same-device clock probe attempted to measure arrival skew without comparing raw cross-GPU timestamps. Device-USM snapshots and clock calibration worked, and completed all-reduce gates remained exact, but every LL256 marker sample timed out with mask 15, including self; calibration also exceeded 2% on some ranks. Reject the measurement before any full-model run. Preserve raw smoke1-8 evidence, restore production source, and make no skew, speed, or LocalMax claim.
2026-07-15 native-dual-qkv-rmsnorm exact-microgate-win/full-model-rejected ../notes/2026-07-15-native-dual-rmsnorm-graph-loss.md, ../data/native-dual-rmsnorm-20260715.json, vLLM d8d7cf198, XPU kernels ef307a8, candidate/control evidence under /mnt/fast-ai/bench-results/deepseek-v4-flash-xpu/native-dual-rmsnorm-* Matching Triton’s 128-thread/SG32 reduction order repaired a one-BF16 prototype mismatch. The final operator passes 160/160 changing eager cases and 32/32 changing graph replays across four B70s and projects 0.893-1.290 ms/token standalone savings. Two strict flag-on suites reach 39.9928/39.9174 tok/s versus a same-commit flag-off control at 40.0950, regressions of 0.255%/0.443%. Preserve default-off; standalone kernel timing did not survive reusable command-graph execution. Record and LocalMax remain unchanged.
2026-07-15 fused-qnorm-rope-fp8-kv-insert promoted ../notes/2026-07-15-fused-qnorm-rope-kv-insert-record.md, ../data/fused-qnorm-rope-kv-insert-record-20260715.json, vLLM 3a74a38a3, XPU kernels ef307a8, /mnt/fast-ai/bench-results/deepseek-v4-flash-xpu/fused-qnorm-rope-kv-insert-candidate-20260715T1040Z, LocalMaxxing cmrm601ig1hsmmj017npoivfd One Triton program preserves Q RMSNorm/RoPE and the KV BF16 rounding boundary while writing UE8M0 FP8 cache bytes directly, removing one graph node and the temporary KV row. All 160 changing eager cases and 32 graph replays are bitwise exact across four cards; the isolated boundary improves 2.02-2.08x. Two strict suites reach 40.135724/40.103728 tok/s, all cached-zero, versus the prior 40.020972 public record. Promote and continue into the WQ_B producer epilogue.
2026-07-15 wqb-m1-producer-fusion geometry-proven/integration-rejected ../notes/2026-07-15-wqb-m1-producer-fusion-closure.md, ../data/wqb-m1-producer-fusion-closure-20260715.json, XPU kernels de979b9, four-card raw SYCL results under /mnt/fast-ai/bench-results/deepseek-v4-flash-xpu/wqb-sycl-fp8-bf16-gemv-card*-20260715.json Padded-M16 Triton is 8.13x slower and inexact. A true-M1 subgroup kernel reaches 23.559-23.700 us, near oneDNN, but is still inexact and spreads each 512-wide head across workgroups. The head-contained geometry costs 53.330-53.644 us before the epilogue and cannot clear the 11.63 us/layer gate. Reject before model integration; preserve the geometry proof.
2026-07-17 TP4 M=2 producer/allreduce/MHC-consumer upper bound component-upper-bound-pass/not-integrated ../notes/2026-07-17-tp4-m2-producer-allreduce-consumer-upper-bound.md, ../data/tp4-m2-producer-allreduce-consumer-upper-bound-20260717.json, oneCCL 6fd2356, two raw confirmation directories under /mnt/fast-ai/bench-results/deepseek-v4-flash-xpu/ oneCCL’s Arc LL ring dropped producer dependencies and the native MHC operation required a one-BF16 graph-visible completion witness. After both repairs, two 40-epoch four-B70 gates are bitwise exact and save 0.953386/0.928339 ms/cycle at the slowest rank over 87 reductions and 85 consumers. A prior ReduceOp.MAX timing was invalid because the specialized ring is sum-only; the corrected harness gathers times. Proceed to the default-off finite event chain; no service or LocalMax claim yet.
2026-07-17 TP4 M=2 finite oneCCL event chain exact eager win/graph performance rejected ../notes/2026-07-17-tp4-m2-event-chain-closure.md, ../data/tp4-m2-event-chain-closure-20260717.json, ../scripts/tp4-m2-event-chain-gate.py, oneCCL 9636514, XPU kernels a609e1f The default-off fixed bridge and native M=2 consumer pass two 40-epoch eager gates bitwise, including rank skew, and bypassing c10d submissions saves 5.601/5.698 ms. Fixed-address graph replay is also exact, but captured ordinary XCCL already removes that host cost: the candidate saves only 0.109546 ms/cycle (4.265725 -> 4.156179), below the 0.50 ms gate. Close before 70 replays, service load, portfolio admission, or LocalMax; future communication work must remove device/collective work rather than submission overhead.
2026-07-17 content-addressed real M=2 cycle corpus and replay worker implemented/70-replay exact baseline ../notes/2026-07-17-m2-real-cycle-corpus-and-replay.md, ../data/m2-real-cycle-corpus-20260717.json, ../scripts/serve-k160-mtp1-m2-cycle-capture.sh, ../scripts/validate-m2-cycle-corpus.py, ../scripts/replay-m2-cycle-corpus.py, vLLM diagnostic 9fc754a One exact-record eager load captures 87 TP4 reductions and 85 native M=2 MHC boundaries per rank into 688 manifests/1,030 content-addressed blobs (150 MiB). All reduced results agree across ranks, all 85 MHC inputs link to collectives, and 84 storage aliases/rank are preserved. The no-model four-B70 worker passes 70/70 full fixed-address graph replays and establishes a 4.209382 ms slowest-rank component baseline. Promote as decoder-shell infrastructure, not model throughput or LocalMax evidence.
2026-07-17 fixed M=4/M=8 MHC over segmented exact collectives component-gate-pass/not-integrated ../notes/2026-07-17-m4-m8-fixed-mhc-component-gate.md, ../data/m4-m8-fixed-mhc-gate-20260717.json, ../scripts/benchmark-mwidth-cycle-corpus.py, ../scripts/run-mwidth-cycle-gate.sh, ../scripts/summarize-mwidth-cycle-gate.py, XPU kernels 50646a2 Row-tiled real M=2 tensors establish fixed-width economics without claiming sequential acceptance. Replacing repeated M=2 MHC calls with one fixed command saves 1.423781 ms/cycle at M=4 and 4.311293 ms/cycle at M=8; both controls and candidates pass 16 changed eager schedules and 70/70 graph replays on all four B70s. A single wide [4,4096] BF16 all-reduce is rejected after 427,072 mismatches/rank across 87 reductions in eager and graph modes. Retain segmented collectives; next require true sequential verifier tensors, held-out acceptance, and complete endpoint economics.
2026-07-18 DSpark exact-M7 sampler replay and context-WKV fusion exact profile complete/candidates rejected ../notes/2026-07-18-dspark-cycle-profile-and-fusion-closure.md, ../data/dspark-cycle-profile-and-fusion-closure-20260718.json, vLLM diagnostics e19c19f4c through ce70e1921 Named scopes replace the inferred cycle bucket: eager Markov sampling is approximately 10.50 ms/cycle. A combined model/sampler graph corrupts copy and JSON output. An isolated sampler graph passes 12/12 exact canaries but retains 83 kernels and 14/15 collective breaks and reaches only 62.460903 tok/s. Fusing the three context-WKV projections cuts that scope 1.914 -> 1.303 ms, passes 18/18 ordered exact canaries, and keeps every realistic row cache-zero, but two endpoint medians are only 64.269762/64.244449 tok/s, below the 64.661411 record. Preserve both default-off, make no LocalMax submission, and move the sampler/acceptance/commit transaction into the fixed Intel decoder shell.
2026-07-18 packed Markov max and generic M=8 DPAS MHC correctness/performance rejected ../notes/2026-07-18-packed-max-and-m8-dpas-closure.md, ../data/packed-max-and-m8-dpas-closure-20260718.json, ../scripts/bench-tp4-packed-max-allreduce.py, vLLM 217df8fc3, XPU kernels e6a5f6c / 827b779 Packed BF16-score/inverse-token int64 MAX is incorrect on all ranks and raises seven-step latency 1.557571 -> 2.351866 ms. Generic TF32 DPAS MHC saves 0.853521 ms/cycle in the row-tiled component but changes next mixes/layer input and returns 1053 instead of 1073 in both arithmetic canary passes. Both remain default-off; no performance suite or LocalMax submission. Continue with an exact pair-tiled vector MHC kernel that shares FN reads while preserving reduction order.
2026-07-18 exact fixed-M8 pair-tiled vector MHC exact graph pass/performance rejected ../notes/2026-07-18-exact-m8-pairtile-closure.md, ../data/exact-m8-pairtile-closure-20260718.json, XPU kernels 92b194a Pairing two rows/workgroup passes all outputs bit-for-bit over 16 changed eager schedules and 70 graph replays on every B70. Twelve-output blocks improve the first candidate from 9.143246 to 8.554636 ms/cycle, but the same-binary incumbent is 8.043785 ms; the existing exact staged pair-vector path is 8.999082 ms. Halving the M8 grid loses more occupancy than shared FN reads save. Keep default-off, skip endpoint/LocalMax, and return to the 10.50 ms DSpark sampler/acceptance/commit boundary.
2026-07-18 TP4 Level Zero IPC event pair transport exact primitive pass/proceed to sampler transaction ../notes/2026-07-18-level-zero-ipc-event-transport.md, ../data/tp4-level-zero-ipc-event-transport-20260718.json, ../scripts/bench-tp4-ipc-event-max-token.py, XPU kernels 88db339 Context-wide event-pool creation plus full opaque-handle SCM_RIGHTS brokering repairs peer opens. Across 12 warmups and 80 measured seven-step cycles, changing winners and forced ties pass with zero mismatches on all four B70s. One-shot device events cut seven tiny pair exchanges from 1,484.5065 to 184.7965 us at the slowest rank. A delayed-rank probe proves real peer waiting; two-slot reset/reuse hangs and is rejected. Production’s full-bias gather is already 371.347 us, so the honest transport-only ceiling is 186.5505 us/cycle, not 1.30 ms. Proceed only as part of a fixed M7 local-W2/sampler transaction; no endpoint or LocalMax claim.

Entry Template

## YYYY-MM-DD label

- Stage:
- Status: `failed|passed|rejected|promoted|inconclusive`
- Source/model/manifest revisions:
- vLLM/XPU-kernel commits and diffs:
- Command and environment:
- Hardware/topology:
- Result and profile paths:
- Correctness/quality artifacts:
- Memory/backend/graph trace:
- Decision and next gate: