b70-optimization-lab

DeepSeek V4 Flash On Four B70s: 100/200 tok/s Roadmap

Date: 2026-07-16

Status: active; easy bounded gates first, specialized Intel decoder strategic

2026-07-18 Progress Update

The current target-verified record is now 80.820052 tok/s, with three strict medians of 80.820052 / 76.900178 / 78.287226 tok/s and LocalMaxxing ID cmrquta9905w3lg013m5vxoqx. The record stacks exact M=8 compressors, selective M=8 W8A16, N128 routed MXFP4, native M=8 router normalization, and a guarded sharded target-argmax/native rejection transaction. The latter removes the full-vocabulary logits all-gather and FP32 sampler materialization for the qualified greedy endpoint.

The original 63.851301 tok/s analysis below is retained as chronological plan context. At the current three-suite median of 78.287226 tok/s, reaching 100 tok/s still requires about 21.7% less time per emitted token or more useful accepted tokens per cycle. The immediate material boundary is no longer a generic sampler tweak: fold local top-1 into the sharded LM-head projection, then own the fixed verifier/accept/commit transaction in a reusable fixed-address SYCL/Level Zero command list. The fixed-geometry Intel decoder remains the strategic path if those bounded integrations do not close the gap.

Objective And Invariants

The product objective is one active generation, never aggregate throughput:

The current record is the TP4+EP QNorm-M2 plus route-direct MTP1 portfolio at 63.851301 tok/s median, 59.718212 p10. With approximately 77.68% first-position acceptance, it emits about 1.7768 tokens per cycle and implies a roughly 27.83 ms speculative cycle.

At unchanged acceptance:

At the current cycle time:

Sub-millisecond wins remain useful but cannot complete the program alone. The final result needs target-cycle reduction and materially deeper useful speculation.

Option 1: Finish High-Value Target-Kernel Fusion

This is the immediate, lower-cost lane.

The first bounded experiment is an M=2 grouped-MXFP4 producer in which gate and up projections are owned by separate subgroups. The subgroups exchange the target-equivalent rounded BF16 fragments through SLM so no subgroup owns both B payloads and both FP32 accumulator sets. This directly addresses the occupancy failure that closed the bitwise-exact dual-accumulator producer.

Reject the candidate before model integration unless it:

  1. matches the canonical remap -> GEMM1 -> clamp-at-10 SwiGLU -> GEMM2 -> gather chain bitwise on changed inputs;
  2. remains exact across fixed-address graph replays, duplicate routes, cross-row overlap, six-local, and valid all-remote EP routes;
  3. passes independently on all four physical B70s;
  4. saves at least 0.50 ms per 43-layer verifier cycle on the slowest card and worst valid route pattern.

Only a passing hardware gate earns a production selector and frozen same-binary B-A-B service run. Keep MXFP4 N64. N32, N128, four-lane scheduling, direct gather, routed activation, fused GEMM2 activation, paired dual-payload GEMM1, gather/shared-add, and unique-route emission remain preserved closures.

Expected role: produce nearer records and reusable Intel primitives. It is not by itself a credible explanation for the full 10 ms gap to 100 tok/s.

Option 2: Restructure TP4 Communication And Cycle Coordination

Normal-run evidence attributes roughly 8-9 ms per MTP1 cycle to TP communication. This is the only measured target-side scope large enough to close much of the 100 tok/s gap.

Do not reopen generic oneCCL flag sweeps, clone removal, LL geometry, recursive doubling, or the failed polling resident consumer. A new communication attempt must change the producer/consumer algebra or reduce collective count. Credible designs include:

Every proposal needs a real-tensor upper-bound proof of at least 0.50 ms/cycle before a TP4 service load. Cross-device profiler timestamps are not accepted as arrival-skew evidence.

Expected role: combine with Option 1 to make approximately 80-100 tok/s plausible. It is difficult and correctness-sensitive, but it attacks a large enough bucket.

Option 3: Develop Useful Deeper Speculation

The attached MTP2 reuse path is closed: the second position accepts only about 0.5-2.2% and the realistic service path deadlocks. Enabling more positions in that predictor is not an option.

The credible alternatives are:

Speculation evaluation must follow the frozen freeze-before-reveal contract:

TP2 plus two draft B70s is only a feasibility experiment. It proceeds only if the unchanged target fits TP2, target-only TP2 performance is competitive, and the two draft cards perform distinct useful work. Duplicating the same draft prediction is not a speedup. The recovered TP2+DP2+EP4 vLLM topology is only 2.495917 tok/s and cannot be used as evidence that this arrangement is ready.

Expected role: required for a credible 200 tok/s result. At the present cycle time, 200 tok/s needs roughly 5.6 emitted tokens per cycle; target fusion alone cannot supply that multiplier.

Option 4: Build The Intel Equivalent Of HIPfire

This is the strategic, highest-ceiling lane. vLLM can initially remain the model loader, API shell, and correctness oracle, while a fixed-geometry SYCL/Level Zero decoder owns the hot decode cycle.

The specialized decoder should provide:

Development artifacts should be cached so most kernel iterations avoid model reload:

Expected role: the most credible path to HIPfire-like efficiency, 100 tok/s, and eventually 200 tok/s when combined with Option 3. It is a substantial engineering program, not a parameter sweep.

Phase A: Protect And Re-attribute

  1. Keep the 63.851301 record binaries, launcher identity, result directory, exact outputs, and LocalMaxxing packet immutable.
  2. Capture a fresh exact-identity diagnostic twin of the record after the QNorm-M2 and route-direct portfolio. Reconcile noncollective device work, communication, queue gaps, host coordination, acceptance, and wall cycle.
  3. Do not interpret the older pre-portfolio eager profile as the current cycle.

Phase B: Easy Bounded Kernel Attempt

  1. Implement the subgroup-split/SLM-exchange M=2 MXFP4 producer in an isolated DeepSeek XPU-kernel worktree.
  2. Run the exact real-shape gate on all four B70s concurrently.
  3. If the worst-card/worst-route saving is below 0.50 ms/cycle, preserve the patch and negative result and stop this design before model loading.
  4. If it passes, add a guarded default-off production selector, rebuild once, run exact graph gates, and then run a frozen same-binary B-A-B suite.

Phase C: Large-Bucket Work

  1. Use the fresh profile to select a producer/reduction/consumer boundary with a conservative ceiling of at least 0.50 ms/cycle.
  2. Begin the fixed-buffer, device-resident cycle shell that can later become the specialized decoder.
  3. Promote individual kernels into that shell instead of accumulating more framework graph nodes.

Phase D: Speculation And Specialized Decoder

  1. Build the held-out speculation evaluator before training or selecting a deeper predictor.
  2. Establish exact M=4/M=8 target-verifier economics.
  3. Select MTP adaptation, DFlash/DEAGLE, or an external draft by emitted tokens per complete wall cycle, not acceptance percentage alone.
  4. Move the winning target and speculation pipeline into the fixed Intel decoder hot loop.

Decision Gates And Reporting

Current First Action

The fresh post-portfolio attribution is complete at 17.8497 ms/cycle of noncollective device work. The subgroup-split producer is closed before build: against the now-promoted route-direct baseline its generous all-remote incremental ceiling is only about 0.123 ms/cycle.

The fixed-M2 producer/allreduce/consumer upper bound cleared Phase C twice, but the implemented finite event chain is now closed at the production-relevant graph gate. It is bitwise exact in two 40-epoch eager runs, under rank skew, and under fixed-address graph replay. Its apparent 5.60-5.70 ms eager saving is submission overhead that reusable graphs already remove: captured ordinary XCCL plus MHC takes 4.265725 ms versus 4.156179 ms for the direct chain, only a 0.109546 ms/cycle gain. Do not service-test or portfolio this candidate.

The later exact-M7 DSpark record moved the frontier to 64.661411 tok/s and its named cycle attribution is now complete. The eager sequential Markov sampler is approximately 10.50 ms/cycle. A separate replay graph is exact but retains 83 kernels and 14/15 collective breaks and regresses the endpoint to 62.460903 tok/s. Fusing all three context-WKV projections saves 0.611 ms in that local scope but produces only 64.269762/64.244449 tok/s in two strict suites. Both are closed default-off. Continue Phase C/D with a device-resident sampler, acceptance, and commit pipeline inside the fixed-geometry decoder shell; do not spend another load on generic sampler graph wrapping or context-only fusion. See ../experiments/deepseek-v4-flash-reap-xpu-b70/notes/2026-07-18-dspark-cycle-profile-and-fusion-closure.md.

The easy bounded inventory is exhausted. Option 4’s first fixed-geometry shell artifact is operational: a 150 MiB content-addressed real M=2 corpus and a no-model four-B70 worker replay the 87-reduction/85-MHC cycle exactly 70/70 times at a 4.209382 ms slowest-rank median.

The first exact M-width extension now clears its component gate. Keeping the proven segmented M=2 collectives and replacing repeated M=2 MHC calls with one fixed command saves 1.423781 ms/cycle at M=4 and 4.311293 ms/cycle at M=8. Both widths pass 16 changed eager schedules and 70/70 graph replays on all four cards, including positions 28 and 58. A single wide [4,4096] BF16 collective is not admissible: it corrupts 427,072 elements across 87 reductions on every rank in eager and graph modes. Retain segmented collectives until that oneCCL count/geometry defect is repaired.

The immediate gate is guarded fixed-M4/M8 integration with true sequential verifier tensors and complete-cycle economics, followed by the frozen held-out predictor evaluation. Attached MTP2 remains closed. Do not infer endpoint throughput or acceptance from row-tiled component tensors, and do not revive resident polling, rejected in-ring MHC implementations, or generic oneCCL flag sweeps. Evidence is in ../experiments/deepseek-v4-flash-reap-xpu-b70/notes/2026-07-17-m4-m8-fixed-mhc-component-gate.md.