This page is the cross-model work queue and archive. It is meant to help the next agent switch models without rereading every historical note.
Hardware planning note: the active Intel lab has four B70 32 GB cards. That lets agents run four independent one-GPU screens or one TP4 service, but it does not leave spare VRAM for very large models or simultaneous production inference during multi-day optimization. Higher-VRAM Intel hardware would make larger future efforts, such as GLM 5.2 and DeepSeek Flash-class models, much more realistic to validate under the same quality rules.
For the full optimization lifecycle, read
model-optimization-guide.md before creating a
new lane.
Create or update the smallest set of files that makes the lane understandable:
results/<model>-<hardware>/README.md for promoted or closed-out outcomes.results/<model>-<hardware>/validity-gates.md for what counts as a record.results/<model>-<hardware>/reproduce.md for the best known commands.results/<model>-<hardware>/bugs-failed-paths.md for invalid fast lanes and
failure signatures.notes/YYYY-MM-DD-<model>-...md for chronological experiment notes.patches/<model>-...patch for source or config deltas worth preserving.data/<model>-...json for compact structured result evidence.Do not move old files just to make the tree look tidy. Add indexes and links unless a file is clearly misplaced and no one is likely to reference the old path.
Main entries:
Status: approved at 102.971435596 tok/s under the submitted legacy
100-event/99-interval convention and 101.941721240 tok/s under conventional
interval accounting. It is 13/13 token-and-text exact against the canonical
q1 teacher, cache-zero on all rows, and approved by LocalMaxxing as
cms2ccv2d00lps201rej94pjy. The result uses exact width 12, DFlash depth 11,
an audited 146/145 Breakable PIECEWISE topology, BF16 KV, and 31 runtime
E4M3FN W8A16 DFlash projection conversions per rank.
This lane is sealed and closed; no benchmark or service is active. The
conventional 102 objective remains short by 0.058278760 tok/s. Reopening it
requires a new preregistration, not a continuation from the superseded
94.920 row.
Poolside’s quantized checkpoint officially ships calibrated FP8 KV; BF16 is a deliberate record-lane override. The earlier B70 A/B doubled cache capacity with FP8 but slowed short decode and changed output. Keep future official long-context FP8 service work separate from the BF16 bitwise-exact record.
Main entries:
Status: active optimization target as of 2026-07-11. Current overall strict
fresh-response best is TP2 on two B70s at 93.036242 tok/s, with exact +
repeat128 + baseline + 1K quality pass and cached_tokens=0 throughout.
Graph-safe FlashAttention enables one full four-row target graph; pair-swapped
controls support the small headline gain. LocalMaxxing approved it as
cmrgue7kl007pmj01yrkcyqmv.
Start from
../results/qwen36-27b-autoround-int4-b70/tp2-fp16-graphsafe-flash-fullgraph-20260711.json.
TP1 remains a separate active record class: 68.236 tok/s is the valid
historical high (cmr9atqb800msqr01u760xh0t), while July 11 isolated
reconfirmation produced a current 65.4-66.7 tok/s band with full quality on
one row. Start TP1 from
../results/qwen36-27b-autoround-int4-b70/tp1-draftgraph-attribution-reconfirm-20260711.json.
The older Intel-checkpoint promote-source row (53.522 tok/s,
cmr4gokx90061nv01lhoe3ft8) remains a baseline/reference. Separate
service/prompt-processing work is captured in
../experiments/qwen36-27b-autoround-int4-b70/notes/2026-07-04-long-context-ladder-baseline.md,
including a 32K-capability exact-retrieval anchor through 17706 actual prompt
tokens with cached_tokens=0.
Main entries:
Status: production-servable one-B70 backend plus current frontier/reference, with diminishing returns unless the next change is a larger verifier/router/speculation or service-prefill design rather than another small flag sweep.
Best strict fresh-response result:
c926ad098, one B70, UD-Q8_K_XL target/verifier, Q4_0 MTP draft
verified by the Q8 target;cached_tokens=0, no cache/history reuse;124.97714084813418 tok/s median generated-token throughput for tokens
1-100 after TTFT, p10 103.83610041293263, mean
122.47435471668817;cmr1u77na01k2ld01kalwzs1e.Important caveats:
2.324% run-median CV and 4.409% p90
pairwise absolute run-median delta; do not promote +1-4% single-run spikes;scripts/analyze-gemma-realistic-ab.py for close changes;104+ / 176+ tok/s rows are diagnostic/pre-final-gate
only;ngram-mod 245-280 tok/s rows are warmed/history artifacts, not
real fresh-response records.Service/prefill status: UB2048 is the validated long-context candidate for the
service lane. It passed fixed JSON-retrieval gates through 22730 actual
prompt tokens and the corrected 30400 actual-token boundary case with
cached_tokens=0, exact outputs, and no paired short-suite decode regression.
It did not beat the short-decode record, so keep UB1024 for short-record
reproduction.
Recent exhausted neighborhoods include adaptive MTP depth caps, tight p_min
repeats, grouped reordered-Q8 duplicate-expert MoE, direct VDR2, top-8
reordered-Q8 slot blocking, Q4_K_M/Q5_K_M/Q6_K/Q8_0 draft swaps, fused verifier
argmax, rowpack, non-direct top-k confidence gating, regular-Q8 top1
epilogue/partial reductions, direct BF16 routed gate/up+GEGLU, attention
post-norm fusion, and per-layer post-norm fusion.
Main entry:
Status: current model-slot production profile is c8. c10 is research-only; c12+ hit boundary failures.
Main entries:
Status: strong candidate to revisit when Gemma work stalls or when a
cross-model collective/graph-boundary idea appears. Strict speed lane is
89.314195 output tok/s / 119.085594 total at p512/n1536; deployable 32K
endpoint baseline is about 83-84 output tok/s.
Future speed work should target hidden-state collective and graph-boundary
fusion, especially MoE-output allreduce plus epilogue or attention o_proj
allreduce plus residual/RMSNorm. Do not spend much time on generic env flag
sweeps.
Main entries:
Status: closed reference packet for now, but preserve every lesson for a future
return. No valid >150 tok/s path was found; best strict 4x baseline is
93.55 tok/s. The main carryover lesson is that graph/speculative speed paths
must pass full-scale canaries, not smoke tests.
Main entries:
../notes/Status: the intensive Q4_0/DFlash SYCL lane closed on 2026-07-13 at a strict
one-B70 record of 47.818818 tok/s; the 100/200 tok/s single-session goals
were not reached. Read the
closure and transfer note
before using its kernel, speculation, graph, or packing artifacts. Reopen only
with one of the concrete scope changes listed there, not another flag sweep.
Main entry:
Status: paused/closed on 2026-07-21 at a fully characterized frontier. The best
verified one-session result is the experimental uniform-K160 target with
target-verified DSpark7: 80.820052 tok/s strict high and 78.287226 tok/s
three-suite median-of-medians on four B70s, with 36/36 cache-zero realistic
rows and 24/24 exact canaries. LocalMaxxing approved
cmrquta9905w3lg013m5vxoqx; no later verified endpoint exceeded it. Reopen
only for the closeout’s 10-20M-token EAGLE/hybrid training condition or a new
device-execution mechanism. The public checkpoint remains hash-pruned with
unavailable calibration and must not be described as official true REAP.
Historical rejected fit/support evidence for the oversized Intel AutoRound artifact remains at the original AutoRound experiment packet.
For evidence-linked strategies and their transfer boundaries, start with Cross-Model Patterns Worth Reusing.
data/, but keep API keys outside
Git as documented in localmaxxing.md.