Last updated: 2026-07-26 America/Toronto
The result is published, approved, sealed, and reproducible. A later metric audit found that the published helper used an inclusive-event numerator over an inter-event span, so the 102 tok/s objective is complete only under that historical convention, not under conventional interval accounting.
102.97143559613157 tok/s;101.94172124017027 tok/s;102 tok/s;-0.05827875982973 tok/s;cms2ccv2d00lps201rej94pjy (APPROVED);Read the accounting correction before making any speed claim. The approved receipt remains historical evidence and must not be duplicate-submitted.
Do not resume from the former 94.920 record, the obsolete recovery block, or the abandoned tree/selector ideas. They are historical evidence, not current work.
| Field | Value |
|---|---|
| Target | poolside/Laguna-S-2.1-INT4 at 4bbfc285f2f8b3b6b526274c133b7b17aae6c8cb |
| Draft | poolside/Laguna-S-2.1-DFlash-INT4 at 5e07c246915c86dc6920fead03d019989224f2ba |
| vLLM | e596ef1543466ae1a05e5bb8091f58872e2b18ba |
| XPU kernels | 6f9dd3c3a7b1b677a992ca4f431a968408f9c816 |
| Layout | TP4+EP4, one active generation |
| Target verifier | exact width 12 |
| DFlash | depth 11, greedy draft, standard rejection |
| Graph | audited Breakable PIECEWISE capture size 12, 146 graphs / 145 eager breaks per rank |
| KV | BF16 |
| Treatment | 31 runtime E4M3FN W8A16 DFlash dense-projection conversions per rank plus the exact auxiliary workspace |
| Selector | VLLM_XPU_LAGUNA_DFLASH_FP8_W8A16=1 |
The intended separate FP8 draft-LM-head path exists in source, but its expected runtime preparation message is absent from the record log. Do not attribute the measured gain to that head. The evidence-backed treatment is limited to the 31 logged draft projections and auxiliary workspace.
experiments/laguna-s-2.1-xpu-b70/tools/run_laguna_mwide_measurement_leg.sh \
candidate B2 RUN_DIR 12 11 1 0 0 0 0 0 0 1 1
Sealed run:
/mnt/fast-ai/llm-optimization-artifacts/laguna-s-2.1/runs/
laguna-width12-dflash-fp8-e596ef154-20260726T214259Z
Promoted records:
patches/laguna-s-2.1-xpu-b70/;data/laguna-s-2.1-width12-dflash-fp8-record-20260726.json;experiments/laguna-s-2.1-xpu-b70/notes/2026-07-26-width12-dflash-fp8-w8a16-record.md;experiments/laguna-s-2.1-xpu-b70/notes/2026-07-26-throughput-window-accounting-correction.md;repro/laguna-s-2.1-int4-b70-102tps-20260726/;experiments/laguna-s-2.1-xpu-b70/notes/2026-07-26-reproducibility-provenance-audit.md;data/localmaxxing-laguna-s-2.1-int4-b70-width12-dflash-fp8-102.971tok-20260726.queue.json;data/localmaxxing-responses/laguna-s-2.1-int4-b70-width12-dflash-fp8-102.971tok-20260726.response.json.Durable learning indexes:
experiments/laguna-s-2.1-xpu-b70/notes/2026-07-26-campaign-transfer-ledger.md;experiments/laguna-s-2.1-xpu-b70/notes/2026-07-26-kv-cache-precision-decision.md;docs/research-workflow-playbook.md.The official quantized checkpoint declares calibrated per-tensor FP8 KV and
vLLM auto resolves to it. This record explicitly overrides that with BF16
because the declared teacher is BF16, the benchmark used at most 0.5% of its
allocated KV capacity, and the earlier controlled Laguna screen found FP8
exactly doubled capacity but was 4.132% slower and changed outputs. Do not
confuse the record’s FP8 DFlash projection weights with its BF16 KV cache.
cached_tokens=0 on all 13 requests;The old recovery wrapper resolved a nonexistent scratch-local probe path.
Several summaries therefore reported 0/4 even though the probe never entered
Python. Driver reload, FLR, and shared-memory cleanup conclusions made from
those summaries were unfounded. Later corrected full-model runs proved the live
TP4/XCCL path healthy, so the old recovery block is closed.
The submission helper also pointed speed records at the retired
/api/benchmarks route and accepted an HTTP 200 HTML shell as success. It now
uses /api/speed-tests, validates the exact projected request through the
authenticated server dry-run, and accepts only HTTP 201 JSON containing a
nonempty ID and status.
cached_tokens=0, one invocation per prompt,
no prefix/history/response reuse, and target verification of every accepted
draft token.N
timestamped events span N-1 intervals; report the conventional interval
field for new goals and qualify the historical compatibility field.No recovery or hardware action is pending. Closing the conventional
0.05827875982973 tok/s gap, or any other benchmark work, requires a new
preregistered experiment.