Date: 2026-07-16
Status: active; easy bounded gates first, specialized Intel decoder strategic
The current target-verified record is now 80.820052 tok/s, with three
strict medians of 80.820052 / 76.900178 / 78.287226 tok/s and LocalMaxxing ID
cmrquta9905w3lg013m5vxoqx. The record stacks exact M=8 compressors, selective
M=8 W8A16, N128 routed MXFP4, native M=8 router normalization, and a guarded
sharded target-argmax/native rejection transaction. The latter removes the
full-vocabulary logits all-gather and FP32 sampler materialization for the
qualified greedy endpoint.
The original 63.851301 tok/s analysis below is retained as chronological plan context. At the current three-suite median of 78.287226 tok/s, reaching 100 tok/s still requires about 21.7% less time per emitted token or more useful accepted tokens per cycle. The immediate material boundary is no longer a generic sampler tweak: fold local top-1 into the sharded LM-head projection, then own the fixed verifier/accept/commit transaction in a reusable fixed-address SYCL/Level Zero command list. The fixed-geometry Intel decoder remains the strategic path if those bounded integrations do not close the gap.
The product objective is one active generation, never aggregate throughput:
The current record is the TP4+EP QNorm-M2 plus route-direct MTP1 portfolio at 63.851301 tok/s median, 59.718212 p10. With approximately 77.68% first-position acceptance, it emits about 1.7768 tokens per cycle and implies a roughly 27.83 ms speculative cycle.
At unchanged acceptance:
At the current cycle time:
Sub-millisecond wins remain useful but cannot complete the program alone. The final result needs target-cycle reduction and materially deeper useful speculation.
This is the immediate, lower-cost lane.
The first bounded experiment is an M=2 grouped-MXFP4 producer in which gate and up projections are owned by separate subgroups. The subgroups exchange the target-equivalent rounded BF16 fragments through SLM so no subgroup owns both B payloads and both FP32 accumulator sets. This directly addresses the occupancy failure that closed the bitwise-exact dual-accumulator producer.
Reject the candidate before model integration unless it:
Only a passing hardware gate earns a production selector and frozen same-binary B-A-B service run. Keep MXFP4 N64. N32, N128, four-lane scheduling, direct gather, routed activation, fused GEMM2 activation, paired dual-payload GEMM1, gather/shared-add, and unique-route emission remain preserved closures.
Expected role: produce nearer records and reusable Intel primitives. It is not by itself a credible explanation for the full 10 ms gap to 100 tok/s.
Normal-run evidence attributes roughly 8-9 ms per MTP1 cycle to TP communication. This is the only measured target-side scope large enough to close much of the 100 tok/s gap.
Do not reopen generic oneCCL flag sweeps, clone removal, LL geometry, recursive doubling, or the failed polling resident consumer. A new communication attempt must change the producer/consumer algebra or reduce collective count. Credible designs include:
Every proposal needs a real-tensor upper-bound proof of at least 0.50 ms/cycle before a TP4 service load. Cross-device profiler timestamps are not accepted as arrival-skew evidence.
Expected role: combine with Option 1 to make approximately 80-100 tok/s plausible. It is difficult and correctness-sensitive, but it attacks a large enough bucket.
The attached MTP2 reuse path is closed: the second position accepts only about 0.5-2.2% and the realistic service path deadlocks. Enabling more positions in that predictor is not an option.
The credible alternatives are:
Speculation evaluation must follow the frozen freeze-before-reveal contract:
TP2 plus two draft B70s is only a feasibility experiment. It proceeds only if the unchanged target fits TP2, target-only TP2 performance is competitive, and the two draft cards perform distinct useful work. Duplicating the same draft prediction is not a speedup. The recovered TP2+DP2+EP4 vLLM topology is only 2.495917 tok/s and cannot be used as evidence that this arrangement is ready.
Expected role: required for a credible 200 tok/s result. At the present cycle time, 200 tok/s needs roughly 5.6 emitted tokens per cycle; target fusion alone cannot supply that multiplier.
This is the strategic, highest-ceiling lane. vLLM can initially remain the model loader, API shell, and correctness oracle, while a fixed-geometry SYCL/Level Zero decoder owns the hot decode cycle.
The specialized decoder should provide:
Development artifacts should be cached so most kernel iterations avoid model reload:
Expected role: the most credible path to HIPfire-like efficiency, 100 tok/s, and eventually 200 tok/s when combined with Option 3. It is a substantial engineering program, not a parameter sweep.
The fresh post-portfolio attribution is complete at 17.8497 ms/cycle of noncollective device work. The subgroup-split producer is closed before build: against the now-promoted route-direct baseline its generous all-remote incremental ceiling is only about 0.123 ms/cycle.
The fixed-M2 producer/allreduce/consumer upper bound cleared Phase C twice, but the implemented finite event chain is now closed at the production-relevant graph gate. It is bitwise exact in two 40-epoch eager runs, under rank skew, and under fixed-address graph replay. Its apparent 5.60-5.70 ms eager saving is submission overhead that reusable graphs already remove: captured ordinary XCCL plus MHC takes 4.265725 ms versus 4.156179 ms for the direct chain, only a 0.109546 ms/cycle gain. Do not service-test or portfolio this candidate.
The later exact-M7 DSpark record moved the frontier to 64.661411 tok/s and its
named cycle attribution is now complete. The eager sequential Markov sampler
is approximately 10.50 ms/cycle. A separate replay graph is exact but retains
83 kernels and 14/15 collective breaks and regresses the endpoint to 62.460903
tok/s. Fusing all three context-WKV projections saves 0.611 ms in that local
scope but produces only 64.269762/64.244449 tok/s in two strict suites. Both
are closed default-off. Continue Phase C/D with a device-resident sampler,
acceptance, and commit pipeline inside the fixed-geometry decoder shell; do
not spend another load on generic sampler graph wrapping or context-only
fusion. See
../experiments/deepseek-v4-flash-reap-xpu-b70/notes/2026-07-18-dspark-cycle-profile-and-fusion-closure.md.
The easy bounded inventory is exhausted. Option 4’s first fixed-geometry shell artifact is operational: a 150 MiB content-addressed real M=2 corpus and a no-model four-B70 worker replay the 87-reduction/85-MHC cycle exactly 70/70 times at a 4.209382 ms slowest-rank median.
The first exact M-width extension now clears its component gate. Keeping the
proven segmented M=2 collectives and replacing repeated M=2 MHC calls with one
fixed command saves 1.423781 ms/cycle at M=4 and 4.311293 ms/cycle at M=8.
Both widths pass 16 changed eager schedules and 70/70 graph replays on all four
cards, including positions 28 and 58. A single wide [4,4096] BF16 collective
is not admissible: it corrupts 427,072 elements across 87 reductions on every
rank in eager and graph modes. Retain segmented collectives until that oneCCL
count/geometry defect is repaired.
The immediate gate is guarded fixed-M4/M8 integration with true sequential
verifier tensors and complete-cycle economics, followed by the frozen held-out
predictor evaluation. Attached MTP2 remains closed. Do not infer endpoint
throughput or acceptance from row-tiled component tensors, and do not revive
resident polling, rejected in-ring MHC implementations, or generic oneCCL flag
sweeps. Evidence is in
../experiments/deepseek-v4-flash-reap-xpu-b70/notes/2026-07-17-m4-m8-fixed-mhc-component-gate.md.