b70-optimization-lab

Laguna S 2.1 resume point

Last updated: 2026-07-26 America/Toronto

Status

The result is published, approved, sealed, and reproducible. A later metric audit found that the published helper used an inclusive-event numerator over an inter-event span, so the 102 tok/s objective is complete only under that historical convention, not under conventional interval accounting.

Read the accounting correction before making any speed claim. The approved receipt remains historical evidence and must not be duplicate-submitted.

Do not resume from the former 94.920 record, the obsolete recovery block, or the abandoned tree/selector ideas. They are historical evidence, not current work.

Record identity

Field Value
Target poolside/Laguna-S-2.1-INT4 at 4bbfc285f2f8b3b6b526274c133b7b17aae6c8cb
Draft poolside/Laguna-S-2.1-DFlash-INT4 at 5e07c246915c86dc6920fead03d019989224f2ba
vLLM e596ef1543466ae1a05e5bb8091f58872e2b18ba
XPU kernels 6f9dd3c3a7b1b677a992ca4f431a968408f9c816
Layout TP4+EP4, one active generation
Target verifier exact width 12
DFlash depth 11, greedy draft, standard rejection
Graph audited Breakable PIECEWISE capture size 12, 146 graphs / 145 eager breaks per rank
KV BF16
Treatment 31 runtime E4M3FN W8A16 DFlash dense-projection conversions per rank plus the exact auxiliary workspace
Selector VLLM_XPU_LAGUNA_DFLASH_FP8_W8A16=1

The intended separate FP8 draft-LM-head path exists in source, but its expected runtime preparation message is absent from the record log. Do not attribute the measured gain to that head. The evidence-backed treatment is limited to the 31 logged draft projections and auxiliary workspace.

Formal command and artifact

experiments/laguna-s-2.1-xpu-b70/tools/run_laguna_mwide_measurement_leg.sh \
  candidate B2 RUN_DIR 12 11 1 0 0 0 0 0 0 1 1

Sealed run:

/mnt/fast-ai/llm-optimization-artifacts/laguna-s-2.1/runs/
laguna-width12-dflash-fp8-e596ef154-20260726T214259Z

Promoted records:

Durable learning indexes:

The official quantized checkpoint declares calibrated per-tensor FP8 KV and vLLM auto resolves to it. This record explicitly overrides that with BF16 because the declared teacher is BF16, the benchmark used at most 0.5% of its allocated KV capacity, and the earlier controlled Laguna screen found FP8 exactly doubled capacity but was 4.132% slower and changed outputs. Do not confuse the record’s FP8 DFlash projection weights with its BF16 KV cache.

Gates that passed

What went wrong earlier

The old recovery wrapper resolved a nonexistent scratch-local probe path. Several summaries therefore reported 0/4 even though the probe never entered Python. Driver reload, FLR, and shared-memory cleanup conclusions made from those summaries were unfounded. Later corrected full-model runs proved the live TP4/XCCL path healthy, so the old recovery block is closed.

The submission helper also pointed speed records at the retired /api/benchmarks route and accepted an HTTP 200 HTML shell as success. It now uses /api/speed-tests, validates the exact projected request through the authenticated server dry-run, and accepts only HTTP 201 JSON containing a nonempty ID and status.

Rules for any future optimization

  1. Preregister the selector, treatment, control, primary metric, and stop conditions before running a score-bearing leg.
  2. Preserve the fixed suite and canonical teacher. Never change the target, omit slow prompts, cherry-pick starts, move setup outside the scored window, warm the service, or substitute a more favorable metric.
  3. Require one active generation, cached_tokens=0, one invocation per prompt, no prefix/history/response reuse, and target verification of every accepted draft token.
  4. Diff the complete run identity before interpreting speed. Require exact source commits, model revisions, binaries, flags, graph capture sizes, and 146/145 topology.
  5. Inspect the source file after editing it. Inspect per-rank logs and explicit execution markers before trusting summary counters.
  6. Treat a failed or missing probe as unknown. Never escalate privileged recovery unless the probe proves that it executed and classifies the failure boundary.
  7. Keep every negative patch/result in the experiment ledger, commit focused changes, and leave all experimental selectors default-off.
  8. Report the first valid result honestly, whether it wins or loses.
  9. For timestamp windows, record event and interval counts separately. N timestamped events span N-1 intervals; report the conventional interval field for new goals and qualify the historical compatibility field.

No recovery or hardware action is pending. Closing the conventional 0.05827875982973 tok/s gap, or any other benchmark work, requires a new preregistered experiment.