b70-optimization-lab

Laguna exact shared-elementwise + QKNorm/RoPE stack record

Date: 2026-07-23 America/Toronto

Result

The preregistered A-B-B-A endpoint crossover passed every quality, causal, and record gate. The two candidate starts measured 34.55070137147406 and 33.89498511171744 tok/s for generated tokens 1-100 after TTFT. The lower candidate exceeds the prior approved 33.438926675602126 tok/s record cmrwot89400gqnz014oodtlbp by 0.45605843611531327 tok/s (1.363854888%).

All four starts matched the canonical q=1 greedy teacher 13/13 and matched each other 13/13. All 52 requests reported cached_tokens=0. Long-then-next passed 2/2 on every leg and the 863-input/512-output rollover row passed 1/1 on every leg. Every leg used a new service, a 60-second-or-longer device-idle gap separated adjacent legs, and the protocol allowed no warm-up generation, within-leg/service prompt repeat, warmed or cache-bearing repeat, retained history, or fifth run. The frozen 13-prompt suite was intentionally reused only across the four fresh ABBA services.

The conservative LocalMaxxing value is B2, the lower candidate: 33.89498511171744 tok/s. B1 is retained as supporting reproducibility evidence and is not submitted.

Frozen treatment

Both arms retained the approved exact eager depth-7 stack:

VLLM_XPU_LAGUNA_BATCHED_EXACT_MOE=1
VLLM_XPU_LAGUNA_M8_FUSED_W1_ROUTE_W2=1
VLLM_XPU_LAGUNA_M8_ROUTE_INTERLEAVE=1
VLLM_XPU_EXACT_SPEC_ATTN=1
LAGUNA_DFLASH_NUM_SPECULATIVE_TOKENS=7

The only A/B difference was:

A control:
  VLLM_XPU_LAGUNA_M8_SHARED_ELEMENTWISE=0
  VLLM_XPU_LAGUNA_M8_QKNORM_ROPE=0

B candidate:
  VLLM_XPU_LAGUNA_M8_SHARED_ELEMENTWISE=1
  VLLM_XPU_LAGUNA_M8_QKNORM_ROPE=1

VLLM_XPU_LAGUNA_M8_BF16_ROUTER_TOPK=0 was fixed off in both arms.

The shared-elementwise half replaces four literal BF16 operations per target layer with two native operations while preserving their incumbent rounding boundaries. Its four-card component gate was exhaustive over all 65,280 finite BF16 values for both operations, passed changing random inputs and post-timing replay, removed exactly 94 launches/cycle, and saved 0.699138-0.722866 ms/cycle.

The separately proven Q/K RMSNorm + RoPE half preserves the incumbent arithmetic, reduces its isolated launch count from 144 to 48 per target cycle, and saved 1.134039 ms/cycle in the component gate. The endpoint experiment measured the two exact launch-reduction bundles as one frozen treatment; component savings were not assumed to add.

Source and runtime identity

Each leg independently hashed all target and draft LFS files, checked the runtime packages and native libraries, required clean source trees, captured the full benchmark-sensitive environment, began with zero DFlash/request metrics, and ended with a bounded shutdown plus four-device idle proof.

Endpoint results

Leg Arm Headline tok/s p10 / mean Target-cycle ms Acceptance
A1 control 32.826917 25.939402 / 38.210676 92.438227 4642/12040
B1 candidate 34.550701 27.030435 / 39.694340 88.948299 4642/12040
B2 candidate 33.894985 27.232460 / 39.662490 88.886261 4642/12040
A2 control 33.273435 25.979233 / 37.828210 92.901753 4643/12033

B1 versus A1:

B2 versus A2:

Both adjacent comparisons independently clear every causal gate. The lower candidate also exceeds the lower control, so the result does not rely on choosing a favorable start.

Evidence

Canonical artifact root:

/media/steve/CorsairExternal/llm-optimization-artifacts/laguna-s-2.1/runs/shared-elementwise-qknorm-stack-abba-8936aac-b6076ce-052193e-20260723T063914Z

Key audit hashes:

59922f66bd27133d653843b9c9cdf7ca8c1b95519997354cae9fe71075a475fe  full-analysis.json
94461927f358d4809032b808c0effa2db6d9b4b56704c8a672e378cfd744c8ab  all-vs-canonical-teacher.json
788291299ec9416711d8867b92f6538265941b1fb92021b0357a12447102d144  cross-leg-exactness.json
456d7404fc04615e88409d35e0ca8a18eb646389cea5a836727e8aed3cf76808  03-B2-candidate/bench.json

The compact tracked packet is:

data/laguna-s-2.1-shared-elementwise-qknorm-stack-record-20260723.json

The LocalMaxxing queue is:

data/localmaxxing-laguna-s-2.1-int4-b70-dflash-shared-elementwise-qknorm-33.895tok-20260723.queue.json

The public API was rechecked immediately before submission: the matching approved 4x B70 Laguna INT4 record remained cmrwot89400gqnz014oodtlbp at 33.43892667560213 tok/s.

The preflighted queue was submitted exactly once. LocalMaxxing returned HTTP 201 and immediately approved the conservative B2 result as cmrx6p5dv001bo4017hb7sixz. A post-submission public API check found exactly one row with that ID and confirmed the intended model revision, four-B70 hardware identity, 112/512 median token counts, batch/concurrency 1, greedy temperature 0, target-verified speculation, and 33.89498511171744 tok/s. The sanitized response is:

data/localmaxxing-responses/laguna-s-2.1-int4-b70-dflash-shared-elementwise-qknorm-33.895tok-20260723.response.json

The public row reused LocalMaxxing hardware profile cmormmlvb0009ky04i6pvj96b, whose shared CPU/RAM fields render as null and 15 GB instead of the submitted host metadata. The four-B70 GPU identity and all benchmark fields are correct. Repair that shared hardware profile separately if the API gains an edit path; never duplicate-submit this result.

Disposition

This is the current approved strict 4x B70 Laguna S 2.1 INT4 record. Only the lower B2 value was submitted; B1 remains support evidence. Keep both selectors default-off globally and enable them together only in the pinned Laguna record launch command.

No held-out DeepSeek pack, cache/history acceleration, within-service prompt repeat, warmed repeat, or /mnt/fast-ai artifact was used. Postflight left no endpoint or worker running and all four B70s free.