b70-optimization-lab

Laguna routed-W1 N128 local-NVMe phase-one failed stop

Date: 2026-07-23 America/Toronto

Result

The recovered, frozen routed-W1 endpoint campaign stopped after A1/B1 exactly as preregistered. Both legs were honest and bitwise exact, and N128 removed real target-cycle work, but it lost the official endpoint comparison:

Metric A1 N64 control B1 N128 candidate B1 - A1
Headline median tok/s 34.969418822 34.029104704 -0.940314118 (-2.6890%)
Mean tok/s 38.223511425 39.176425110 +0.952913685
p10 tok/s 24.221875960 26.815770749 +2.593894789
Target-cycle time 94.093806528 ms 90.341118198 ms -3.752688330 ms (-3.9882%)
Acceptance 4,641/12,047 4,642/12,040 +0.030703 percentage point

B1 won only 3/13 paired prompt rows. The paired median was -0.940314118 tok/s (-3.057791%). It therefore failed the three frozen performance gates requiring a faster headline, at least 9/13 row wins, and a positive paired median. It passed the cycle-saving and bounded work-drift gates.

The classification is phase1_failed_stop. B2 and A2 were not run, no rescue or fifth run is permitted, and no LocalMaxxing payload was staged or submitted. The approved record remains 33.89498511171744 tok/s, cmrx6p5dv001bo4017hb7sixz.

Quality and honesty

Both fresh services passed every frozen source, model, runtime, freshness, accounting, exactness, and cleanup check:

The only treatment difference was literal VLLM_XPU_LAGUNA_M8_W1_N_TILE=64 versus 128. The complete approved shared-elementwise + QKNorm/RoPE + route-interleaved DFlash stack remained fixed.

What the result means

N128 is a real isolated kernel optimization, not an endpoint win. The prior four-card component gate measured an 8.7271% mean isolated W1 improvement, and this endpoint phase still saved 3.7527 ms per normalized target cycle. Nevertheless, ten prompt rows slowed and both robust endpoint summaries moved against the candidate. The larger tile is therefore closed as a promotion candidate for this exact stack.

The mean and p10 are reported as secondary observations only. A large cold A1 first-row slowdown raised B1’s mean comparison, but the frozen primary median, row-win count, and paired median all reject N128. No post-hoc metric replaces those gates.

Source and local-storage identity

All live model reads and evidence writes used internal NVMe/ext4:

/mnt/fast-ai/llm-models/laguna-s-2.1
/mnt/fast-ai/llm-optimization-artifacts/laguna-s-2.1

The external Corsair USB remained backup-only.

Evidence and seal

Canonical artifact root:

/mnt/fast-ai/llm-optimization-artifacts/laguna-s-2.1/runs/w1-n128-nvme-recovery-abba-8936aac-c59aaad-20260723T131632Z

Key hashes:

Both leg manifests verify. The seal reports valid=true, the campaign ledger and hash chain remain unchanged, and all publication-evidence checks pass. Postflight found no listener on port 8000, no model worker, kernel taint zero, and both leg cleanup records at status zero.

Compact tracked packet:

data/laguna-s-2.1-w1-n128-nvme-phase1-failed-stop-20260723.json

Disposition

Preserve N128 as negative evidence and keep N64 as the routed-W1 endpoint policy. The next clean lane should target shared-expert GEMM occupancy or a separately preregistered narrower routed-W1 geometry. Do not reinterpret the cycle saving as permission to rerun this campaign or stack the failed router candidate.