b70-optimization-lab

20260623T0941 N7 No-Backend-Sampling Follow-Ups

Goal: take the 90.243 tok/s no-backend-sampling record and test the nearest confidence / draft-runtime knobs without changing model quality, prompt shape, or the 384-row chat canary gate.

Common identity:

Results

GPU Label Change Canary tok/s after TTFT tok/s wall Decision
0 gemma4-q8-gpu0-mtp-n7-aot-nmin2-pmin008-nobs-filled-long-deep-20260623T094131Z p-min=0.08 384/384 89.974 82.024 Valid, below record.
1 gemma4-q8-gpu1-mtp-n7-aot-nmin2-pmin012-nobs-filled-long-deep-20260623T094131Z p-min=0.12 384/384 90.301 82.334 Valid improvement over 90.243, but below GPU3.
2 gemma4-q8-gpu2-mtp-n7-aot-nmin3-pmin010-nobs-filled-long-deep-20260623T094131Z n-min=3, p-min=0.10 384/384 89.821 81.719 Valid, stricter minimum hurts.
3 gemma4-q8-gpu3-mtp-n7-aot-nmin2-pmin010-nobs-dthreads32-filled-long-deep-20260623T094131Z --spec-draft-threads 32 384/384 90.419 82.342 New valid record; submitted to LocalMaxxing.

Submitted record:

Takeaways

Follow-Up

Keep the promoted baseline:

MTP_N_MAX=7 MTP_N_MIN=2 MTP_P_MIN=0.10 MTP_BACKEND_SAMPLING=0 \
MTP_DRAFT_THREADS=32 BENCH_PROMPT_MODE=filled-long

Next four-at-a-time candidates: