b70-optimization-lab

Gemma 4 26B A4B Validity Gates

This lane optimizes speed without accepting an INT4-quality downgrade. A result is not a record just because it is fast.

Required Identity

Every run summary must include:

Quality Rules

Initial Text Canaries

Use temperature 0, fixed seeds when supported, and repeat enough times to catch nondeterministic runtime bugs. The repo harness defaults to /v1/chat/completions and seed=1:

The Qwen lane demonstrated that smoke passes can hide 1-in-30 or 1-in-100 runtime corruptions. Full promotion needs repeat depth, not a single lucky run.

Speed Metrics

Report at least:

Promoted throughput requires non-null server usage.completion_tokens. Diagnostic runs may use --allow-missing-usage, but those are not record evidence until token counting is fixed.

For LocalMaxxing or cross-model comparison, prefer corrected/generated output throughput rather than total client throughput. Label total-token throughput separately.

Realistic Final Gate For Promotion

Diagnostic benchmarks may use synthetic or repetitive prompts while searching for optimization ideas. A result may be confirmed, promoted, or submitted only if it passes the realistic final gate:

Run it with the Gemma harness by setting REALISTIC_GATE=1; this writes realistic-suite.json and embeds realistic_final_gate in summary.json.

Thermal / Hardware Fairness

The current Gemma 26B Q8 record family has several-percent run-to-run variance. For promoted record repeats and close A/B decisions, capture XPU telemetry alongside the benchmark so temperature and throttling are not guessed after the fact.

Recommended telemetry on this host:

sudo -S -p '' xpu-smi dump -d <gpu> -m 0,1,2,3,4,5,17,18,35 -i 1 \
  < /home/steve/SUDOPASSWORD.txt

Never print or copy the sudo password. Record only the resulting telemetry file and summarize active core temperature start/max/end, memory temperature start/max/end, power mean/max, frequency mean/min/max, and throttle reasons. Flag real thermal-throttle samples. Burst Power Excursion should still be recorded, but do not treat it as equivalent to sustained thermal throttling without supporting frequency/power evidence.

The 2026-07-01 same-GPU thermal sweep did not find temperature to be the main driver of current final-postnorm variance: four exact GPU0 full512 repeats spanned 114.520-120.202 tok/s while active core max stayed 77-78 C, memory max stayed 86-90 C, frequency remained near max, and no thermal-throttle samples appeared. See experiments/gemma4-26b-a4b-q8-b70/sweeps/20260701-finalpostnorm-thermal-variance.md.

Do not compare a hot run against a cold historical outlier without telemetry, same-recipe repeats, or a same-window control lane. Small deltas near the frontier should be treated as variance until repeated under the fixed realistic gate with cached_tokens=0 and the exact same canary scale.

Micro-Improvement / A/B Reliability Gate

Near the current Gemma record, single-run medians are too noisy for micro-change decisions. The same-GPU thermal repeatability check measured 2.324% run-median CV and 4.409% p90 pairwise absolute run-median delta for the exact same recipe.

For source or config changes expected to move the short metric by only a few percent, use paired same-window A/B blocks and analyze per-prompt ratios with:

scripts/analyze-gemma-realistic-ab.py \
  --control data/<control-1>/summary.json \
  --control data/<control-2>/summary.json \
  --candidate data/<candidate-1>/summary.json \
  --candidate data/<candidate-2>/summary.json

The candidate is a micro-win only if the paired bootstrap 95% lower bound of the median candidate/control prompt ratio is above +1.0%, all candidate runs pass the realistic final gate, and p10/full512/wall/TTFT/canary/telemetry do not materially regress. Positive raw medians with a CI crossing zero are inconclusive_positive, not promotable. See results/gemma4-26b-a4b-q8-b70/reliability-protocol.md.

When the patch or config should only affect the target-side runtime, and not MTP/speculation behavior, add the no-spec calibration check before deciding a small delta:

The current no-spec calibration lane measured 0.330% run-median CV and 0.577% p90 pairwise same-recipe delta, much tighter than the 4.409% MTP noise band. Use it to resolve target-side changes that are otherwise lost in MTP variance. Do not use it for draft quality, verifier, acceptance-policy, handoff, or p-min/n-min/n-max changes.

Long-Context Service Gate

Prompt-processing and long-context service changes are useful only if they do not weaken the short decode record or hide behind cache reuse. Use the fixed long-context suite at repro/gemma4-26b-a4b-q8-b70/long-context-suite-v1.json and the harness scripts/bench-openai-long-context-suite.py.

A long-context/prefill result may be promoted as a service candidate only when:

Report long-context/prefill results separately from the short decode record. Use approximate prefill throughput (prompt_tokens / TTFT) only as a service diagnostic, alongside TTFT, generated-token throughput after TTFT, wall-clock throughput, prompt/output hashes, model identity, runtime commit, env vars, flags, and logs. Do not submit long-context service diagnostics to LocalMaxxing as the one-B70 short-decode headline unless the separate realistic final gate also beats the current record.

Run the paired service gate with repro/gemma4-26b-a4b-q8-b70/run-vdr2-long-context-service-gate.sh; run the paired short guard with repro/gemma4-26b-a4b-q8-b70/run-vdr2-short-decode-guard.sh.

Fresh-Response Vs Warmed/History Throughput

Headline throughput must apply to the realistic final gate above. A single synthetic first row may be useful for diagnosis, but it is not enough for promotion or LocalMaxxing submission.

Speculation is allowed, including MTP, draft-model speculation, n-gram speculation, and verifier-based multi-token acceptance. The validity question is whether the draft source could operate on a fresh request without already having seen the target response.

Separate every speculative result into:

If a benchmark mixes a cold first request with warmed repeated requests, report the first-request throughput separately and do not average warmed repeats into the fresh-response number. A draftless n-gram/history run that becomes fast only after the first identical output must be labeled history-accelerated and must not be submitted or promoted as fresh-response throughput.

The older benchmark harness supports BENCH_PROMPT_MODE=filled-long-unique and filled-fixed-line-unique for fresh-response aggregate checks. These modes generate a deterministic different prompt per repeat and store each row’s prompt_sha256. For those modes, the repeated-row mean may be treated as a fresh-response mean only when:

For historical filled-long and filled-fixed-line runs, keep the conservative policy: all rows are diagnostic unless the fixed realistic suite passes. Earlier “row0 fresh” labels in the repo should be read as pre-final-gate terminology, not as publishable real-world throughput.

Single-GPU records and four-replica aggregate capacity are different modes. Record and submit them separately: one full model on one B70 is the primary single-session decode record; four independent servers are an aggregate service capacity result only after each replica passes the same canary gate.

Promotion Thresholds

The first baseline can be promoted if it is quality-valid and reproducible, even if slow. Later “best” claims require: