b70-optimization-lab

Model Effort Index

This page is the cross-model work queue and archive. It is meant to help the next agent switch models without rereading every historical note.

Hardware planning note: the active Intel lab has four B70 32 GB cards. That lets agents run four independent one-GPU screens or one TP4 service, but it does not leave spare VRAM for very large models or simultaneous production inference during multi-day optimization. Higher-VRAM Intel hardware would make larger future efforts, such as GLM 5.2 and DeepSeek Flash-class models, much more realistic to validate under the same quality rules.

How To Add A Model Effort

For the full optimization lifecycle, read model-optimization-guide.md before creating a new lane.

Create or update the smallest set of files that makes the lane understandable:

  1. results/<model>-<hardware>/README.md for promoted or closed-out outcomes.
  2. results/<model>-<hardware>/validity-gates.md for what counts as a record.
  3. results/<model>-<hardware>/reproduce.md for the best known commands.
  4. results/<model>-<hardware>/bugs-failed-paths.md for invalid fast lanes and failure signatures.
  5. notes/YYYY-MM-DD-<model>-...md for chronological experiment notes.
  6. patches/<model>-...patch for source or config deltas worth preserving.
  7. data/<model>-...json for compact structured result evidence.

Do not move old files just to make the tree look tidy. Add indexes and links unless a file is clearly misplaced and no one is likely to reference the old path.

Active / Recent Efforts

Laguna S 2.1 INT4 On Four B70s

Main entries:

Status: approved at 102.971435596 tok/s under the submitted legacy 100-event/99-interval convention and 101.941721240 tok/s under conventional interval accounting. It is 13/13 token-and-text exact against the canonical q1 teacher, cache-zero on all rows, and approved by LocalMaxxing as cms2ccv2d00lps201rej94pjy. The result uses exact width 12, DFlash depth 11, an audited 146/145 Breakable PIECEWISE topology, BF16 KV, and 31 runtime E4M3FN W8A16 DFlash projection conversions per rank.

This lane is sealed and closed; no benchmark or service is active. The conventional 102 objective remains short by 0.058278760 tok/s. Reopening it requires a new preregistration, not a continuation from the superseded 94.920 row.

Poolside’s quantized checkpoint officially ships calibrated FP8 KV; BF16 is a deliberate record-lane override. The earlier B70 A/B doubled cache capacity with FP8 but slowed short decode and changed output. Keep future official long-context FP8 service work separate from the BF16 bitwise-exact record.

Qwen3.6 27B INT4 AutoRound On B70

Main entries:

Status: active optimization target as of 2026-07-11. Current overall strict fresh-response best is TP2 on two B70s at 93.036242 tok/s, with exact + repeat128 + baseline + 1K quality pass and cached_tokens=0 throughout. Graph-safe FlashAttention enables one full four-row target graph; pair-swapped controls support the small headline gain. LocalMaxxing approved it as cmrgue7kl007pmj01yrkcyqmv. Start from ../results/qwen36-27b-autoround-int4-b70/tp2-fp16-graphsafe-flash-fullgraph-20260711.json. TP1 remains a separate active record class: 68.236 tok/s is the valid historical high (cmr9atqb800msqr01u760xh0t), while July 11 isolated reconfirmation produced a current 65.4-66.7 tok/s band with full quality on one row. Start TP1 from ../results/qwen36-27b-autoround-int4-b70/tp1-draftgraph-attribution-reconfirm-20260711.json. The older Intel-checkpoint promote-source row (53.522 tok/s, cmr4gokx90061nv01lhoe3ft8) remains a baseline/reference. Separate service/prompt-processing work is captured in ../experiments/qwen36-27b-autoround-int4-b70/notes/2026-07-04-long-context-ladder-baseline.md, including a 32K-capability exact-retrieval anchor through 17706 actual prompt tokens with cached_tokens=0.

Gemma 4 26B A4B Q8 / INT8 On B70

Main entries:

Status: production-servable one-B70 backend plus current frontier/reference, with diminishing returns unless the next change is a larger verifier/router/speculation or service-prefill design rather than another small flag sweep.

Best strict fresh-response result:

Important caveats:

Service/prefill status: UB2048 is the validated long-context candidate for the service lane. It passed fixed JSON-retrieval gates through 22730 actual prompt tokens and the corrected 30400 actual-token boundary case with cached_tokens=0, exact outputs, and no paired short-suite decode regression. It did not beat the short-decode record, so keep UB1024 for short-record reproduction.

Recent exhausted neighborhoods include adaptive MTP depth caps, tight p_min repeats, grouped reordered-Q8 duplicate-expert MoE, direct VDR2, top-8 reordered-Q8 slot blocking, Q4_K_M/Q5_K_M/Q6_K/Q8_0 draft swaps, fused verifier argmax, rowpack, non-direct top-k confidence gating, regular-Q8 top1 epilogue/partial reductions, direct BF16 routed gate/up+GEGLU, attention post-norm fusion, and per-layer post-norm fusion.

Gemma 4 12B IT INT4 AutoRound

Main entry:

Status: current model-slot production profile is c8. c10 is research-only; c12+ hit boundary failures.

MiniMax M2.7 INT4 AutoRound

Main entries:

Status: strong candidate to revisit when Gemma work stalls or when a cross-model collective/graph-boundary idea appears. Strict speed lane is 89.314195 output tok/s / 119.085594 total at p512/n1536; deployable 32K endpoint baseline is about 83-84 output tok/s.

Future speed work should target hidden-state collective and graph-boundary fusion, especially MoE-output allreduce plus epilogue or attention o_proj allreduce plus residual/RMSNorm. Do not spend much time on generic env flag sweeps.

Qwen3.6 35B A3B Quark W8A8 INT8

Main entries:

Status: closed reference packet for now, but preserve every lesson for a future return. No valid >150 tok/s path was found; best strict 4x baseline is 93.55 tok/s. The main carryover lesson is that graph/speculative speed paths must pass full-scale canaries, not smoke tests.

Qwen3.6 27B Q4_0 / FP8 Historical Lanes

Main entries:

Status: the intensive Q4_0/DFlash SYCL lane closed on 2026-07-13 at a strict one-B70 record of 47.818818 tok/s; the 100/200 tok/s single-session goals were not reached. Read the closure and transfer note before using its kernel, speculation, graph, or packing artifacts. Reopen only with one of the concrete scope changes listed there, not another flag sweep.

DeepSeek V4 Flash REAP/XPU On B70

Main entry:

Status: paused/closed on 2026-07-21 at a fully characterized frontier. The best verified one-session result is the experimental uniform-K160 target with target-verified DSpark7: 80.820052 tok/s strict high and 78.287226 tok/s three-suite median-of-medians on four B70s, with 36/36 cache-zero realistic rows and 24/24 exact canaries. LocalMaxxing approved cmrquta9905w3lg013m5vxoqx; no later verified endpoint exceeded it. Reopen only for the closeout’s 10-20M-token EAGLE/hybrid training condition or a new device-execution mechanism. The public checkpoint remains hash-pruned with unavailable calibration and must not be described as official true REAP.

Historical rejected fit/support evidence for the oversized Intel AutoRound artifact remains at the original AutoRound experiment packet.

Cross-Model Lessons

For evidence-linked strategies and their transfer boundaries, start with Cross-Model Patterns Worth Reusing.