This experiment lane tracks bring-up and optimization for:
Intel/Qwen3.6-27B-int4-AutoRoundQwen/Qwen3.6-27Bbits=4, group_size=128, symmetric,
packing_format=auto_round:auto_gptqThe initial TP1 single-B70 OpenAI-compatible endpoint works, and the lane now
has a strict fresh-response TP2 record after replacing the broken installed
oneCCL graph collective with pinned public oneCCL/libccl and capturing the
intrinsic-MTP draft through an opaque compiled all-gather boundary. Current
optimization work must beat the conservative 87.02911429766677 tok/s TP2
row toward the 100+ tok/s target, or improve service/max-context behavior
without using warmed/cache/history effects. TP1 remains a separate active
record class; it has not been declared exhausted.
Resume from these current entry points:
../../results/qwen36-27b-autoround-int4-b70/HANDOFF.md;../../results/qwen36-27b-autoround-int4-b70/tp2-public-oneccl-draftgraph-20260711.json;oneccl_ll256/README.md and
scripts/run-tp2-oneccl-public4ce-draftgraph-candidate.sh;scripts/run-tp1-current-candidate.sh;notes/2026-07-11-tp1-draftgraph-attribution-and-reconfirmation.md.Completed first milestone:
abc86de19eb1ebbf6a7df4582341325c22ddcb7d.max_model_len=2048.105/108 accepted draft tokens after
manual probes plus smoke).Current evidence:
/mnt/fast-ai/bench-results/qwen36-27b-autoround-int4-b70/servers/tp1-gpu0-port19410-20260703T012317Z.log;data/qwen36-27b-autoround-openai-smoke-20260703T013020Z.json.Current overall strict best:
VLLM_XPU_DDTREE_CAPTURE_GDN_CORE=1, reducing target graph pieces from 129
to 33;87.029114 tok/s, p10 79.941979, mean
87.913957; independent isolated high 87.815738;cached_tokens=0 all passed;../../results/qwen36-27b-autoround-int4-b70/tp2-capture-gdn-core-20260711.json;scripts/run-tp2-oneccl-public4ce-draftgraph-capturegdn-candidate.sh.Prior TP2 draft-graph milestone:
b52f40c / libccl 4ceafd1;82.893718 tok/s, p10 72.751868, mean
83.100685, all 12 strict prompts cached_tokens=0;85.393815 tok/s, exact cases + repeat128 + baseline
parity + 1K needle passed;82.894 as headline because the rows differ by 3.02%, inside the
established 4.4% variance band; a swapped four-GPU crossover measured a
same-direction +5.39% average over eager draft;../../results/qwen36-27b-autoround-int4-b70/tp2-public-oneccl-draftgraph-20260711.json;oneccl_ll256/README.md and
scripts/run-tp2-oneccl-public4ce-draftgraph-candidate.sh;cmrgjjw8n004qmj01cp91qxl0.Prior Intel-checkpoint TP1 strict best:
qwen3_next_mtp,
num_speculative_tokens=3, max_cudagraph_capture_size=8,
max_num_batched_tokens=1024;VLLM_XPU_GDN_PROMOTE_ACCEPTED_SPEC_STATE=1 and
VLLM_XPU_GDN_NONSPEC_POSTPROCESS_ACCEPTED_STATE=0;cached_tokens=0 every row, return_token_ids=true;53.522 tok/s for generated tokens 1-100 after
TTFT, p10 48.406, mean 53.986;data/qwen36-27b-autoround-int4-b70-baselines/intel-mtp3-xpugraph1-cg8-promotesource-noacceptedpost-repeat2-realistic128-chat-tokenids-qwensuite-20260703T044519Z.json;results/qwen36-27b-autoround-int4-b70/promote-source-noacceptedpost-20260703.json.TP1 historical high and current reproduced band:
webhie/Qwen3.6-27B-int4-AutoRound + runtime INT8 target
LM-head (BF16 scales) + runtime INT4 draft LM-head (BF16 scales);68.236 tok/s, p10 62.317, mean 67.830,
cached_tokens=0;65.359, 66.716, and
65.420 tok/s; all passed the strict cached-zero gate and the first passed
exact, repeat64, baseline parity, and the 1K check. Keep 68.236 as the
valid historical high and use 65.4-66.7 as the current reproduced band;-0.05%), so
the TP2 distributed all-gather graph fix is not a missing TP1 win. See
../../results/qwen36-27b-autoround-int4-b70/tp1-draftgraph-attribution-reconfirm-20260711.json;67.519 prior approved confirm, 68.397 same-recipe
quality-skipped control, 68.481 native-slot-copy smoke, 66.871
native-slot-copy confirm, and 67.300 PyTorch-slot-management same-window
control. Use 68.236 as the current quality-confirmed headline, not the
quality-skipped 68.397 or the one-off 68.481;data/qwen36-27b-autoround-int4-b70-baselines/qwen27-regressionfix-quality-confirm-20260706T102729Z-candidate-summary-20260706T102729Z.json
at 67.33805616805299 tok/s, strict fresh/cached-zero with repeat64
quality and baseline match all. The focused source guard is captured at
patches/qwen36-27b-autoround-int4-b70/vllm-qwen27-mixed-draft-kv-metadata-guard-20260706.patch;../../results/qwen36-27b-autoround-int4-b70/webhie-int8lmhead-bf16scale-draftint4-replayssm-current-confirm-20260706.json;notes/2026-07-06-current-confirm-68tok-and-textonlymtp-no-win.md;cmr9atqb800msqr01u760xh0t.Latest transaction diagnostic:
notes/2026-07-07-replayssm-state-digest-trace.md records a default-off
ReplaySSM state-digest trace patch and diagnostic endpoint run. It captured
commit/stage/spec-decode boundary digests for layer 0 at 67.453 tok/s
strict-fresh/cached-zero with quality skipped. Use it for future
graph-safe GDN/DeltaNet transaction/tape work, not as a promoted benchmark.notes/2026-07-07-replayssm-stage-decode-fusion-pregate.md records a
direct native-op microbench of gdn_replayssm_stage_conv +
gdn_replayssm_spec_decode at current MTP3/cache8 shape. Paired cost was
only 0.045 ms/layer (~2.18 ms over 48 GDN layers), so a fused op is
future ReplaySSM polish, not the main >100 tok/s path.notes/2026-07-07-draft-oracle-trace-branch-tail-screen.md records a
worker-side draft-token oracle trace on the current strict fresh recipe. It
paired 2143 verifier rows and found only 42 cases where the target-owned
bonus/replacement token appeared later in the unaccepted draft tail (1.96%
all rows, 3.13% partial rejects). Treat simple branch/tail rescue as closed
unless verifier-step cost drops or speculation depth changes.notes/2026-07-10-replayssm-vdim8-no-win.md closes a wider ReplaySSM
recurrent value bucket after a two-GPU reversed-order ABBA microbenchmark.
The card-balanced recurrent change was +0.024% (slower) and the complete
stage+decode pair changed by only -0.043%; the initial cross-card apparent
gain was a cold/card artifact. The source was restored to v_dim_per_sg=4
without spending an endpoint run.Latest stronger-drafter result:
Target-matched DFlash adaptation is reopened as a new architecture lane, not
as a retry of the weak public checkpoint. A corrected five-aux target corpus,
offline block-prefix evaluator, paper-style position-decay trainer, and
four-GPU scope/loss matrix are documented in
notes/2026-07-10-dflash-target-adaptation-lane.md. No offline acceptance
number is a speed claim; final-suite endpoint validation remains mandatory.
1,31,60 and no continuity breaks, and the new offline
evaluator successfully loads both Ex0bit compressed and full-vocab EAGLE3
checkpoints. Acceptance is far too low: compressed mean accepted
0.289908 over 14,784 starts (24.40% step-1 exact), while the full-vocab
spot check is the same class at 0.291016 over 512 starts. Do not spend
endpoint/kernel work on the off-the-shelf Ex0bit checkpoint as-is. The useful
artifact is the EAGLE3 aux collection/eval path, which can now support a
target-matched EAGLE3/DFlash training attempt. See
notes/2026-07-06-ex0bit-eagle3-aux-probe-no-win.md.scripts/train-qwen27-ex0bit-eagle3-adapter.py can adapt Ex0bit-format
checkpoints from qwen36_eagle_sequence_v2 data. The larger follow-up used
all 4 B70s to collect 384 target-owned prompts / 61,440 rows, trained
fc-lm-head on 288 prompts, and held out 96 prompts. It improved heldout
rollout from direct Ex0bit 0.289 and the first adaptation 0.539 to
0.6003787878787878 mean accepted, with 48.65% step-1 exact but only
20.10% step-2 conditional exact. This is still far below current MTP3
accepted depth and not endpoint-worthy. Next EAGLE/DFlash work needs a
multi-step rollout / accepted-prefix training objective, not endpoint
plumbing. A first rollout-objective implementation and 4-GPU screen improved
the best heldout mean to 0.6693046536796536 from original Ex0bit init, but
a narrowed original-init sweep improved it further to
0.973146645021645 (52.81% step-1 exact, 50.04% step-2 conditional,
49.40% step-3 conditional). This reopens the lane as training research but
is still below endpoint threshold. A continuation sweep reached only
1.0142045454545454 and widened the train/heldout gap. A larger v4 corpus
then collected 576 prompts / 92,160 rows and improved the best heldout mean
only modestly to 1.0592532467532467 (55.98% step-1 exact, 52.39%
step-2 conditional, 50.55% step-3 conditional). This is useful progress
but remains below the 1.5-2.0 offline threshold for endpoint/kernel
integration; the next move needs an objective or train-scope change, not
simple endpoint plumbing. See compact summaries
diagnostics/qwen27-eagle3-aux-v4-corpus-summary-20260706.json and
diagnostics/qwen27-ex0bit-eagle3-rollouttrain-v4-summary-20260706.json.
A bounded all-scope follow-up from the v4 best checkpoint moved heldout mean
only to 1.0707972582972582, while original all-scope low-LR runs underfit;
preserve it as no-endpoint evidence in
diagnostics/qwen27-ex0bit-eagle3-all-scope-v4-summary-20260706.json.
A later-step weighted continuation (decay=1.25, lr=2e-5) is the current
best diagnostic family. After rollout-3 continuation flattened at
1.1271645021645023, switching continuation training to rollout-5 produced
the current best diagnostic checkpoint at 1.1957070707070707 mean accepted
(56.68% step-1 exact, 53.61% step-2 conditional, 55.79% step-3
conditional, 1267 full-5 accepts). It remains below endpoint threshold and
repeated rollout-5 continuation is showing diminishing returns.
Summaries:
diagnostics/qwen27-ex0bit-eagle3-late-weight-v4-summary-20260706.json and
diagnostics/qwen27-ex0bit-eagle3-late-continuation-v4-summary-20260706.json
plus
diagnostics/qwen27-ex0bit-eagle3-late-continuation2-v4-summary-20260706.json
and
diagnostics/qwen27-ex0bit-eagle3-late-continuation3-v4-summary-20260706.json
plus
diagnostics/qwen27-ex0bit-eagle3-deep-continuation-v4-summary-20260706.json
and
diagnostics/qwen27-ex0bit-eagle3-deep-continuation2-v4-summary-20260706.json
and
diagnostics/qwen27-ex0bit-eagle3-deep-continuation3-v4-summary-20260706.json.
A broader v5 diagnostic corpus was collected (1152 prompts / 184,320
rows / zero continuity breaks) at
diagnostics/qwen27-eagle3-aux-v5-corpus-summary-20260706.json. The
previous v4-trained best scored 1.2001713564213565 mean accepted on the v5
heldout shard, and v5 rollout-5 training improved that to
1.2866838023088023 (59.02% step-1 exact, 55.32% step-2 conditional,
57.29% step-3 conditional, 3056 full-5 accepts). A disk-cleanup retry
plus accepted-prefix survival objective improved the current offline
diagnostic best to 1.340886544011544 mean accepted (59.92% step-1 exact,
56.26% step-2 conditional, 58.50% step-3 conditional, 3602 full-5
accepts). This is progress but still below endpoint threshold. Summaries:
diagnostics/qwen27-ex0bit-eagle3-v5-heldout-baseline-summary-20260706.json
and
diagnostics/qwen27-ex0bit-eagle3-v5-deep-continuation-summary-20260707.json,
plus
diagnostics/qwen27-ex0bit-eagle3-v5-continuation4-summary-20260707.json
and
diagnostics/qwen27-ex0bit-eagle3-v5-survival-objective-summary-20260707.json.
V6 broader chat-style aux-data collection is complete; compact summary:
diagnostics/qwen27-eagle3-aux-v6-corpus-summary-20260707.json, suite:
eagle-chat-corpus-v6-suite.json, raw root:
/mnt/fast-ai/bench-results/qwen36-27b-autoround-int4-b70/eagle-data/qwen27-eagle3-aux-v6-chat-4gpu-20260707T012928Z.
V5-survival-on-v6-heldout baseline:
diagnostics/qwen27-ex0bit-eagle3-v5-survival-on-v6-heldout-summary-20260707.json
at 0.8866846157479571 mean accepted. V6 survival-objective training
summary:
diagnostics/qwen27-ex0bit-eagle3-v6-survival-train-summary-20260707.json,
best 1.0069670776061594 mean accepted. This is a useful offline gain, but
still below endpoint threshold. V6 continuation summary:
diagnostics/qwen27-ex0bit-eagle3-v6-continuation-summary-20260707.json,
best 1.0401492607812575 mean accepted from rollout_loss_decay=0.5.
V6 step-focus summary:
diagnostics/qwen27-ex0bit-eagle3-v6-stepfocus-summary-20260707.json,
best 1.0493835907609466 mean accepted from
v6sf-r3-lr1e-5-decay0p25-rank0p1. This is a small diagnostic lift, not an
endpoint candidate. Do not keep sweeping the same v6 corpus/objective family
unless there is a new mechanism; next EAGLE move should improve data quality
with concrete embedded context or return to non-EAGLE speed work.
That data-quality follow-up is also complete: v6b concrete-context corpus
summary
diagnostics/qwen27-eagle3-aux-v6b-corpus-summary-20260707.json has
384 prompts / 61268 usable rows / zero continuity breaks; the best v6
draft scored 1.036561331974176 on v6b heldout; v6b training summary
diagnostics/qwen27-ex0bit-eagle3-v6b-stepfocus-summary-20260707.json
improved only to 1.0597349643221203 mean accepted. This is still far below
endpoint threshold. A four-GPU all-scope follow-up reached only
1.1014610941216445 mean accepted
(diagnostics/qwen27-ex0bit-eagle3-v6b-allscope-summary-20260707.json),
and a target-hidden trajectory distillation follow-up reached only
1.1023445463812436 mean accepted
(diagnostics/qwen27-eagle3-hidden-distill-screen-20260707.json;
notes/2026-07-07-eagle3-hidden-distill-no-endpoint.md). This closes small
EAGLE data/objective/all-scope/hidden-distill sweeps until there is a new
mechanism.
New five-aux mechanism screen:
notes/2026-07-07-eagle3-five-aux-tooling.md adds --aux-count 5 support
for aux layers [1,16,31,46,61], expanding old three-aux checkpoints into
slots [0,2,4]. The first v7 five-aux survival screen is closed in
notes/2026-07-07-eagle3-five-aux-survival-no-endpoint.md: clean corpus
(61,307 rows, zero aux bad files), training improved the expanded source
baseline 0.873 -> 1.082 mean accepted, but stayed below prior ~1.10
diagnostics and far below the 1.5-2.0 endpoint gate.
See
notes/2026-07-06-ex0bit-eagle3-target-adaptation-screen.md.notes/2026-07-07-targetbody-timing-and-mlp-workspace-no-win.md records a
graph-none/no-spec timing run, an enforce-eager layer split, and a closed
no-win screen for VLLM_XPU_SHARED_EXPERT_ACT_WORKSPACE=1. This Qwen27
checkpoint is dense qwen3_5_text, not MoE; do not route current work toward
MoE layerlets or the workspace flag.Previous fastest quality-gated variant:
webhie/Qwen3.6-27B-int4-AutoRound + runtime INT8 LM-head
(BF16 scales);VLLM_XPU_LM_HEAD_INT8=1 and
VLLM_XPU_LM_HEAD_INT8_SCALE_DTYPE=bf16;65.276 tok/s, p10 59.609, mean 65.077,
cached_tokens=0;65.005 and 64.864 tok/s;notes/2026-07-04-post-awq-record-repro-support.md at
66.12771533602819 tok/s, strict fresh/cached-zero gate passed, smoke
passed, no LocalMaxxing update because the recipe is unchanged and no fresh
quality rerun was required for a support check;64.234 and 64.090 tok/s;64.306 tok/s;results/qwen36-27b-autoround-int4-b70/webhie-int8-lmhead-bf16scale-20260703.json;notes/2026-07-03-int8-lmhead-bf16-scale-quality-pass.md.cmr5iu3gk00bfq901nidgcana.notes/2026-07-04-continuation-source-and-awq-state.md. This preserves the
active source snapshots, closes another no-repeat audit of cheap env/config
knobs, and records the cyankiwi/Qwen3.6-27B-AWQ-INT4 strict screen result:
compressed-tensors AWQ loaded and passed the fresh/cached-zero gate
mechanically, but only reached 56.565 tok/s and is closed no-win. See
notes/2026-07-04-cyankiwi-awq-int4-screen-no-win.md.Prior Intel-checkpoint fastest quality-gated variant:
AutoRound W4A16 + runtime INT8 LM-head;VLLM_XPU_LM_HEAD_INT8=1;62.628 tok/s, p10 58.104, mean 62.998,
cached_tokens=0;62.276 tok/s on GPU3;53.332 tok/s;results/qwen36-27b-autoround-int4-b70/int8-lmhead-20260703.json;patches/qwen36-27b-autoround-int4-b70/vllm-xpu-lm-head-int8-quality-pass-20260703.patch.Service-oriented INT8-head variant:
VLLM_XPU_LM_HEAD_INT8_SCOPE=target;61.898 tok/s, p10 57.494, mean 62.432;blue, green, red), so treat target-only as an
attribution/service idea that must be revalidated per checkpoint/revision and
scale dtype;notes/2026-07-03-int8-lmhead-scope-attribution.md;patches/qwen36-27b-autoround-int4-b70/vllm-xpu-lm-head-int8-scope-target-quality-pass-20260703.patch.Synthetic search reference:
vllm-random: 81.773 tok/s;data/qwen36-27b-autoround-int4-b70-baselines/intel-mtp5-xpugraph1-cg16-specmetrics-p512o512-r3-20260703T031846Z.json;Next milestone:
65.56930784255283 tok/s, and the timing refresh is captured in
notes/2026-07-04-phase0-phase1-baseline-and-timing.md.vllm-random metrics diagnostic-only.--return-token-ids before promoting
any change.
Use scripts/run-vllm-candidate.sh for single-replica strict candidate
screens so server logs, smoke output, strict fresh gate results, optional
quality checks, and compact summaries are captured consistently.notes/2026-07-05-replayssm-stage-profile-and-frontier.md.notes/2026-07-04-frontier-audit-onednn-graph-and-drafter.md,
notes/2026-07-04-compact-lmhead-top1-kernel-no-win.md, and
notes/2026-07-05-replayssm-stage-profile-and-frontier.md.
The latest synchronized MTP-forward diagnostic in
notes/2026-07-06-draft-proposer-timing-split.md closes the apparent
~11 ms recurrent MTP-next cost as async timing attribution: recurrent
dispatches are PIECEWISE graph mode and synchronized
model_forward_first/next are sub-millisecond. Do not chase MTP-next as an
eager-kernel bug; use accepted-token, target-forward, stronger-drafter, or
graph-safe state-transaction work for the next real speed attempt.
The later current-recipe subtiming check
notes/2026-07-07-current-mtp3-subtiming.md reproduced the current recipe
at 68.296 tok/s with quality skipped and showed the sampled decode bucket
is already fixed-shape MTP3 (4 unpadded, 4 padded, 3 scheduled spec
tokens, PIECEWISE graph). It reinforces that padding cleanup and noisy
async draft labels are not the next speed lever.
The target-body micro-screen
notes/2026-07-07-rmsnorm-gated-native-route-no-win.md is also closed:
existing _C.rms_norm plus SiLU multiply was faster in microbench, but not
bit-exact and not faster in same-window endpoint A/B. Do not repeat that
Python routing patch; only a true fused gated-RMSNorm kernel with better
numeric agreement would be a new idea.
The latest graph-safe transaction precheck is
notes/2026-07-06-replayssm-commit-pending-active-slot-guard.md: native
gdn_replayssm_commit_pending used to mutate metadata for null,
out-of-range, or inactive rows; the new guard script
../../scripts/check-gdn-replayssm-commit-pending.py now passes BF16/FP16/FP32
plus native prefix/recurrent checks after an active-slot guard. This is
infrastructure for partial-group / branch-regenerate work, not a throughput
win or LocalMaxxing row.
The follow-up branch-fork composition guard is
notes/2026-07-06-replayssm-branch-fork-composition-guard.md: native
copy_slots + compacted native commit_pending can fork a valid branch
slot and commit an accepted prefix exactly, but committing raw destination
rows after invalid-source copies corrupts unrelated pending slots. Any
branch/regenerate endpoint prototype must compact valid (src, dst) rows
before commit.
The 2026-07-06 replacement-suppression plumbing/margin follow-up is also
closed no-win: active scheduler recovery passed quality only at ~34-49
tok/s, and margin gating stayed below record. See
notes/2026-07-06-replacement-mask-plumbing-and-margin-no-win.md and
../../patches/qwen36-27b-autoround-int4-b70/vllm-qwen27-replacement-mask-plumbing-margin-no-win-20260706.patch.
The external Qwen/Qwen3.5-0.8B draft-model probe is also closed no-win:
compatibility work got explicit draft_model serving, text-only Qwen3.5
M-RoPE, mixed draft KV groups, and mixed block sizes past startup, but the
live k8 run accepted 0 draft tokens and fell to only ~2.3-2.6 tok/s
while rejecting every draft. See
notes/2026-07-06-qwen35-08b-external-draftmodel-zero-acceptance.md and
../../patches/qwen36-27b-autoround-int4-b70/vllm-qwen27-qwen35-08b-explicit-draftmodel-compat-zeroaccept-20260706.patch.65.25583870721442 tok/s, but it
was flat versus the 65.27648650325429 tok/s record because
get_top_tokens() still pays the dense LM-head. Treat
notes/2026-07-04-spec-greedy-topids-no-headline-win.md as integration
groundwork for a future true compact LM-head kernel, not a path to retest by
itself.notes/2026-07-04-compact-lmhead-top1-kernel-no-win.md. The native
int8_lm_head_top1_w8a8 prototype was exact but slower than dense oneDNN:
final 8x64 policy measured compact 2.66-2.68 ms versus dense
2.57-2.61 ms for rows 1-4, so do not wire it into vLLM.notes/2026-07-04-autoround-variant-screening-and-stepidx-audit.md.
Local webhie-Code and acyildirimer AutoRound variants passed the strict
gate but were slower than the same-window webhie control, and the possible
spec_step_idx MTP fix is a no-op for this lane because the checked
Qwen27 AutoRound configs all have mtp_num_hidden_layers=1. The focused
future-use patch is preserved at
../../patches/qwen36-27b-autoround-int4-b70/vllm-qwen-mtp-spec-step-idx-pass-through-future-20260706.patch
and documented in
notes/2026-07-06-qwen-mtp-spec-step-idx-pass-through.md; do not treat it
as a current Qwen27 speed candidate unless a checkpoint with multiple
mtp.layers.N tensors is introduced.notes/2026-07-04-webhie-depth-screen-no-win.md. On the fastest
webhie/BF16-scale INT8-LM-head recipe, strict same-window MTP4/cg8
(60.478 tok/s), MTP5/cg8 (59.257 tok/s), and MTP5/cg16
(59.817 tok/s) all lost to the MTP3/cg8 control (65.809 tok/s).
Treat the control as support/variance only, not a new LocalMaxxing row.notes/2026-07-04-webhie-mtp1-mtp2-depth-coverage-no-win.md.
The missing MTP1/MTP2 current-recipe coverage is closed: MTP1/cg8
51.246, MTP2/cg8 59.589, MTP3/cg8 control 64.730, MTP4/cg8
59.886, all cached_tokens=0 and gate-passing. Keep MTP3/cg8.notes/2026-07-04-webhie-bf16scale-capture-size-screen-no-win.md.
On the same webhie/BF16-scale MTP3 recipe, a four-GPU strict same-window
screen found cg8 remains best: cg4 64.507, cg8 control 65.153, cg16
63.500, cg32 64.071, all cached_tokens=0 and gate-passing. Keep
max_cudagraph_capture_size=8 unless a source change alters graph shapes,
row counts, or acceptance.notes/2026-07-04-int8-gemm-scratchpad-ring-screen-no-win.md.
The low-level ring-size screen found ring4 highs (65.708, 65.817) but
paired crossover deltas versus ring1 controls were only +0.42% and
+0.27%, below the practical variance band. Do not promote or submit;
keep default ring behavior unless a future trace shows scratchpad reuse as
a real issue.notes/2026-07-06-int4-gemm-scratchpad-ring-no-win.md.
A default-off VLLM_XPU_INT4_GEMM_SCRATCHPAD_RING_SIZE patch built and
endpoint-ran, but ring1 only measured +0.08% mean / +0.18%
median-of-runs over ring0 controls after crossover, while ring2/ring4 did
not help. The active source and live _C binary were restored; preserve
the patch as negative evidence only.notes/2026-07-05-gdn-qkvz-ba-quant-reuse-no-win.md.
A same-window four-GPU strict fresh screen of
VLLM_XPU_GDN_REUSE_QKVZ_BA_QUANT=clone, clone-ba, and clone-qkvz
found no credible win over control: control 64.398 tok/s, best
clone-qkvz 64.824 tok/s, inside variance. Keep the promoted recipe
unchanged.notes/2026-07-07-gdn-qkvz-ba-proj-pack-no-win.md.
Packing ba into one wider W4A16 qkvzba projection saved only
0.0034 ms/layer at rows=4 (~0.16 ms projected over 48 GDN layers),
far below the >=0.025 ms/layer gate. Do not implement endpoint/loader
packing from this signal.notes/2026-07-05-target-forward-low-risk-screens-and-backlog.md.
M-RoPE text-only fast path (65.797 vs control 65.960) and GDN fallback
prefill only (65.655 vs control 65.967) both passed the strict fresh
gate but did not beat controls. Continue with source/kernel work, not more
easy env knobs.notes/2026-07-05-qk-norm-rope-fused-spike-no-win.md.
A default-off Qwen3Next-specific XPU fusion for the gated
[q, gate, k, v] layout passed direct BF16 parity, but regressed the
strict fresh endpoint to 45.980 tok/s versus the 65.276 tok/s record.
The patch is preserved at
../../patches/qwen36-27b-autoround-int4-b70/vllm-qwen27-qk-norm-rope-fused-spike-20260705.patch.
Do not repeat this endpoint lane unless a new kernel first beats the
existing separate Q/K norm + RoPE primitives in a standalone microbench.notes/2026-07-05-gdn-output-norm-native-no-win.md.
A default-off _xpu_C.gdn_rms_norm_gated_xpu_out path was very fast in a
direct microbench and passed repeat32 quality, but same-window strict fresh
controls beat the endpoint candidate on average (65.299 control vs
64.569 native). The live vLLM/XPU source and local extension binary were
restored. Keep only the preserved no-win patches; do not repeat this lane
unless a future trace shows GDN output norm as a large standalone region.
17a. Latest GDN output norm + INT4 out-proj prototype:
notes/2026-07-07-gdn-fused-outproj-prototype-positive.md.
A default-off native _xpu_C.qwen_gdn_out_proj_int4_w4a16 prototype now
fuses the Qwen GDN gated RMSNorm workspace into the following INT4 W4A16
out_proj prologue. It built with oneAPI 2025.3, imported from a temporary
_xpu_C, passed FP16/BF16 synthetic parity, and measured ~5-6.7x faster
than the local PyTorch-workspace + existing INT4 GEMM subpath on Qwen27
shapes (~0.208 ms -> ~0.031-0.042 ms). This is diagnostic-only, not an
endpoint or LocalMaxxing result. Next step is TP1 INC W4A16 endpoint wiring
behind an env flag, then strict fresh same-window validation and repeat64
quality before promotion.notes/2026-07-05-native-promote-ssm-only-crash.md.
The Python-only VLLM_XPU_GDN_NATIVE_PROMOTE_CONV_STATE=0 switch passed
OpenAI smoke but hit UR_RESULT_ERROR_DEVICE_LOST during the strict run
before benchmark or quality artifacts. Treat it as crash/inconclusive, not
a speed or quality result. The next attempt must add a matching default-off
C++ gate around copy_conv_rows_to_indices in gdn_attention_spec_decode
so the packed native path and Python promotion agree on whether conv rows
are copied. That follow-up is now closed no-win too:
notes/2026-07-05-native-spec-conv-copy-gate-no-win.md disabled both
native conv promotion paths and made repeat64 quality worse (62/64
blue, green red yellow, plus one runaway repetition). Do not rerun blind
conv-copy disablement; future GDN state work needs a traced/taped exact
conv-window transaction.input_ids dispatch shortcut:
notes/2026-07-05-mtp-text-inputids-next-no-win.md.
Dispatch tracing showed recurrent Qwen3.5 MTP-next draft calls used
inputs_embeds=[1,5120] with input_ids=None, so a default-off source
spike tried to keep text-only embedding lookup inside the captured draft
forward by passing token IDs. Attempt 1 crashed before readiness because
torch compile tried to size inputs_embeds=None; the compile-shape
workaround got past that but stalled during decode PIECEWISE graph capture.
Active vLLM source was reverted, and the no-win patch is preserved at
../../patches/qwen36-27b-autoround-int4-b70/vllm-qwen27-mtp-text-inputids-next-no-win-20260705.patch.
Do not repeat this wrapper-level shortcut without a deeper compile/cudagraph
design change.notes/2026-07-05-gdn-packed-decode-with-source-no-win.md.
A default-off VLLM_XPU_GDN_PACKED_DECODE_WITH_SOURCE=1 patch allowed the
packed one-token GDN decode helper when accepted source rows were present
and promoted both conv and SSM before the packed update. Same-window strict
fresh/cached-zero screen passed mechanically, but the candidate lost to
control (65.077 vs 65.631 tok/s). Active vLLM source was reverted, and
the no-win patch is preserved at
../../patches/qwen36-27b-autoround-int4-b70/vllm-qwen27-gdn-packed-decode-with-source-no-win-20260705.patch.
Do not repeat Python-level accepted-source promotion shortcuts; future GDN
state work needs a traced exact accepted-prefix transaction or native
graph-safe tape/commit path.notes/2026-07-04-dynamic-drafter-depth-partial-group-crash.md.
A default-off prototype that actually shortened the MTP proposer loop
crashed with an XPU indexing assert when it created partial speculative
groups. A follow-up with upstream-style placeholder -1 rejection applied
also failed the same way:
notes/2026-07-04-dynamic-depth-placeholder-reject-retry-no-win.md.
Do not retry dynamic depth by changing only sampler placeholder handling;
partial groups need explicit support across proposer output, verifier
metadata, sampler rows, GDN state commit, and graph capture shapes.notes/2026-07-04-dflash-swa-revisit.md. Real mixed SWA support crashes
before readiness because the current DFlash/EAGLE proposer assumes all
draft layers belong to one KV-cache group. The target-verified
all-sliding single-group diagnostic passed the fresh gate but dropped to
20.630 tok/s, so DFlash remains closed no-win until multi-KV-group draft
metadata is implemented.notes/2026-07-04-dflash-multikv-mixed-swa-attempt.md. The experimental
patch
../../patches/qwen36-27b-autoround-int4-b70/vllm-dflash-multikv-mixed-swa-attempt-20260704.patch
gets mixed full/sliding DFlash through startup and graph capture with draft
KV groups [64, 65, 66, 67, 68], but endpoint testing is still no-win:
graph mode device-loses during the strict suite, and graph-off/no-async
shows only about 2-3% draft acceptance with low-teens or worse generation
throughput. Preserve the patch as upstream/plumbing research, not as a
record lane.notes/2026-07-04-language-model-only-no-win.md.
--language-model-only correctly enters text-only mode and reduced model
memory from 19.02 GiB to 18.15 GiB, but the MTP3/cg8 XPU graph path
hung before readiness at Capturing CUDA graphs (decode, PIECEWISE): 0/1.
Do not use it for the current strict decode recipe; only revisit for
non-MTP/non-graph service-memory work or after XPU graph capture changes.notes/2026-07-04-scheduler-mbt-and-chunked-prefill-screen.md.
MBT768 (64.131 tok/s) and MBT1280 (64.346 tok/s) passed the strict
fresh gate but lost to the current 65.276 tok/s record family; disabling
chunked prefill is invalid for MAX_MODEL_LEN=2048 / MBT1024. Keep
MAX_NUM_BATCHED_TOKENS=1024 and chunked prefill enabled.notes/2026-07-04-eagle1-endpoint-isolation-matrix.md.
The first local EAGLE1 draft looked promising offline (2.1016 mean
accepted tokens on held-out calibration starts) but did not survive endpoint
isolation. Current-state eager k3, default-state graph k3, and
current-state graph k1 all failed the calibration gate at only
19.828-22.410 tok/s; current-state graph k3 stalled before JSON output.
The compact summary is
../../data/qwen36-27b-autoround-int4-b70-baselines/qwen27-eagle1-endpoint-isolation-20260704T094450Z-summary.json.
Do not repeat endpoint config sweeps for this draft; future EAGLE work
starts with diverse chat-style corpus/eval v2 and stronger held-out
diagnostics.notes/2026-07-04-eagle-corpus-v2-tooling.md.
The collector now supports --suite, chat mode, request extra JSON, stable
request IDs, and prompt metadata; the dataset builder copies that metadata
into .pt samples; and offline eval reports acceptance by prompt family.
This is preparation only, not a speed result, but it is the correct restart
point if EAGLE is revisited.notes/2026-07-04-eagle-corpus-v2-chat-calib-smoke.md.
A one-GPU calibration-suite chat collection produced 3840 hidden rows,
24 samples, 0 continuity breaks, and metadata on 24/24 samples after
fixing suffix-tolerant request-ID matching. A tiny two-epoch draft reached
only 0.240 mean accepted offline, so it is not an endpoint candidate; it
only proves the metadata path works.notes/2026-07-04-eagle-corpus-v2-4gpu-heldout.md.
The four-GPU runner collected 96 chat prompts, 15360 hidden rows, 96
samples, metadata on 96/96 samples, and 0 continuity breaks. A compact
draft trained on shards 0-2 and evaluated on heldout shard 3 reached
only 0.489 mean accepted over 1024 starts, far below the prior 2.1016
offline draft that still failed endpoint quality. Do not endpoint-test this
draft; future EAGLE work needs materially stronger data/training/init first.notes/2026-07-04-eagle-corpus-v2-followups-closed.md.
The staged curriculum improved OOD-family heldout only to 0.616, a
balanced task-holdout split scored 0.601, the old stronger v1 draft
transferred poorly to v2 heldout (0.201), and all-96 training scored only
0.438 on the separate calibration suite. Current compact v2 EAGLE is
closed again; do not endpoint-test these drafts.notes/2026-07-04-eagle-v2-stronger-offline-screen-no-endpoint.md.
The stronger residual/two-layer screen improved heldout only to 0.695
mean accepted and separate calibration only to 0.441. It is
diagnostic-only and not an endpoint candidate. The reusable runner is
scripts/run-eagle-v2-stronger-offline-screen.sh.notes/2026-07-04-eagle-v3-target-loss-offline-no-endpoint.md.
Target-shaped one-layer drafts and token-heavy losses did not help on the
current v2 corpus. Best row was the compact frozen-base residual variant at
only 0.647 heldout mean accepted and 0.423 separate-calibration mean
accepted, below the prior v2 stronger screen and far below the offline
endpoint gate. Do not rerun larger/target-shaped EAGLE on this same corpus
without a materially new data or architecture idea. Reusable runner:
scripts/run-eagle-v3-target-loss-offline-screen.sh.notes/2026-07-06-eagle-v4-large-corpus-no-endpoint.md.
A larger four-GPU non-final chat corpus collected 384 prompts,
61,440 hidden rows, 384 samples, metadata on 384/384 samples, and
0 continuity breaks. The best larger compact draft reached only
0.718 heldout mean accepted and 0.512 separate-calibration mean
accepted, far below the endpoint gate (2.0 heldout / 1.5 calibration).
This closes “just more non-final data and hparams on the same compact
EAGLE architecture”; no endpoint run and no LocalMaxxing submission.notes/2026-07-04-draft-topk-calibration-diagnostic.md.
The target verifier token is in the built-in draft top-32 for 96-99% of
MTP positions and an oracle reranker would reach 3.910 target-verified
tokens/step on the 24-prompt trace; a larger 96-prompt non-final trace
confirmed base 2.595 vs oracle 3.864, but simple static token-bias and
margin rerankers were flat or worse on prompt-heldout split. A small
learned top-k MLP also moved only 2.7123 -> 2.7184 target tokens/step on
a separate calibration trace, too little for runtime overhead. Preserve the
trace/analyzers; do not ship a heuristic or tiny top-k reranker. Future
accepted-token work needs a materially stronger drafter/reranker on
isolated non-final data.notes/2026-07-04-draft-topk64-and-sequential-reranker-limit.md.
A full 96-prompt K64 diagnostic found target-in-top64 rates of
99.7%, 98.4%, and 96.8% by draft position, but held-out margin
reranking was flat and sparse-bias reranking regressed. The independent
oracle is an invalid post-hoc endpoint shortcut for sequential MTP, and
the final-slot upper bound still needs recomputing/branching the target
bonus row while adding only about +0.16 target tokens/step. Do not
reopen cheap top-k reranker endpoint patches.notes/2026-07-04-token-tree-mechanical-screen-no-win.md.
Existing vLLM speculative_token_tree support works mechanically for
Qwen27/XPU, but config-only tree shapes do not beat MTP3/cg8. A same-suite
24-prompt control reached 63.871 tok/s; binary depth-2 tree reached
60.526 tok/s; root top-3 reached 63.107 tok/s; all completed rows were
strict fresh/cached-zero diagnostics. Root top-2 stalled during drafter
checkpoint load, and root top-3 already closed the same root-alternative
idea. Do not reopen token-tree sweeps unless the branch design avoids the
current full-logits tree proposer cost or uses a stronger legal drafter.notes/2026-07-06-token-tree-current-recipe-no-win.md.
Same-window strict fresh diagnostics with the current target-INT8/draft-INT4
ReplaySSM recipe found no win: ordinary MTP3 control 67.797 tok/s,
root-3 67.691, root-2 59.159, and binary-depth-2 12.709. No quality
run was warranted. Do not repeat config-only token-tree sweeps on this
recipe.MAX_NUM_BATCHED_TOKENS screen:
notes/2026-07-04-short-decode-mbt-screen-no-win.md.
The current record recipe should stay at MBT1024 for short decode.
Same-window candidates MBT1536, MBT2048, and MBT4096 all passed the strict
fresh/cached-zero gate but landed at 63.829, 64.239, and 64.779 tok/s,
below the 65.276 record and inside recipe variance. The MBT1024 control
row is invalid because GPU0 device-lost during the first benchmark request
after smoke passed; existing support rows already cover the default recipe.notes/2026-07-04-frontier-closure-and-next-projects.md.
Independent audits found no unclosed non-cheating config/runtime lane and
no bounded atomic/single-pass/fused-quant LM-head kernel tweak likely to
beat dense oneDNN by >10%. Do not launch more Qwen27 endpoint/config
screens until the candidate is a real top-ID LM-head producer, a materially
stronger drafter/branch-regenerate architecture, or full partial-group
source support.notes/2026-07-06-branch-regenerate-feasibility-envelope.md.
The existing top-k64 trace was converted into a legal cost envelope. With
the then-current 67.519 tok/s record and 2.6243 target-verified tokens/step,
the inferred verifier step is 38.87 ms. A perfect MTP3 first-reject
branch/regenerate path with top-64 access reaches only 3.9565
tokens/step, or 101.8 tok/s if it adds zero step cost; it has only
0.697 ms/step budget for a 100 tok/s endpoint and cannot reach 125+
at the current step cost. Treat MTP3 branch work as a narrow ~100 tok/s
infrastructure lane, not the main 125+ route.
A 2026-07-07 refresh on the current 68.236 tok/s recipe and the fixed
strict Qwen suite closes MTP3 branch/regenerate even harder for >100:
notes/2026-07-07-current-recipe-strict-topk64-branch-envelope.md.
Current target tokens/step is 2.74695, inferred step cost is
40.2565 ms, and the perfect rank-64 suffix-regenerate envelope reaches
only 3.96813 tokens/step / 98.571 tok/s if it adds zero overhead.
At this step cost, 100 tok/s needs 4.02565 tokens/step, above the
MTP3 maximum of 4. Do not implement MTP3-only branch/regenerate as a
>100 tok/s lane until step cost is reduced or speculation depth changes.
A worker-side draft-token trace then closed the narrower bonus-tail rescue
variant: notes/2026-07-07-draft-oracle-trace-branch-tail-screen.md found
target bonus/replacement tokens later in the unaccepted draft tail only
1.96% of all rows / 3.13% of partial rejects on the strict fresh
current recipe.notes/2026-07-06-native-prefix-exact-state-rescreen-no-win.md.
The old July 5 exact-native/prefill replay flags were stale because
prefix-base later gained the needed extra state column. Rescreening them on
four GPUs closed the lane anyway: offset/writeout exact-native rows landed
at ~4.6-4.9 tok/s and failed quality, prefill-column replay collapsed
acceptance to zero, and replaypartial passed local repeat16 but only
reached 6.323 tok/s. The durable lesson is that same-forward exact
target-tail GDN state is impossible with current data flow because the
verifier-sampled replacement/bonus token has no projected GDN input row yet.
Do not continue native serial/prefill flag sweeps; move to a real
graph-safe state transaction/tape, target-tail projection/branch-regenerate
support, or a stronger drafter.configs/: reusable env files.scripts/: downloader, launcher, smoke, and future benchmark helpers.notes/: chronological run notes.patches/: patch snapshots for successful and failed source attempts.results/: compact experiment summaries.quality/: quality gates and prompt suites as they mature.localmaxxing/: queued payloads and response copies after valid records.The active headline family is now
webhie/Qwen3.6-27B-int4-AutoRound TP2 with FP16 target compute, runtime INT8
target LM-head BF16 scales, runtime INT4 draft LM-head BF16 scales, ReplaySSM,
public oneCCL, graph-safe FlashAttention, and one FULL four-row target graph.
The strict cold record is 93.036242 tok/s median for tokens 1-100 after TTFT;
every request reported cached_tokens=0, and exact cases, repeat128, baseline
parity, and the 1K needle passed. The current packet is
../../results/qwen36-27b-autoround-int4-b70/tp2-fp16-graphsafe-flash-fullgraph-20260711.json;
the implementation/replay record is in
../qwen27_graphsafe_flash_attention/README.md.
Older reference points remain useful for attribution: plain MTP3/cg8 was about
47.6-48.5 tok/s, and promote-source/no-accepted-postprocess lifted the lane
to 53.5-54.9 tok/s before the webhie variant and INT8 LM-head work. A fast
invalid flag, VLLM_XPU_GDN_NONSPEC_POSTPROCESS_FULL_ACCEPT=0, reached
51.273 tok/s on the strict suite and 74.877 tok/s synthetically, but failed
1024-token needle recall. Tracing explains the lift: the valid path copies
large GDN/Mamba state from the accepted speculative slot back to the running
slot after verification.
Current trace summary:
data/qwen36-27b-autoround-int4-b70-baselines/mamba-copy-trace-summary-mtp3-cg8-p512o128-20260703T042542Z.json.
Prompt-processing / long-context service work is tracked separately from the
short-decode record. The current service ladder lives in
notes/2026-07-04-long-context-ladder-baseline.md, with suite
../../repro/qwen36-27b-autoround-int4-b70/long-context-suite-v1.json and
runner scripts/run-long-context-ladder.sh. The latest 32K-capability anchor
uses the same webhie/BF16-scale INT8-LM-head MTP3/cg8 recipe at
MAX_MODEL_LEN=32768 and passes exact JSON retrieval through 17706 actual
prompt tokens with cached_tokens=0, TTFT median 22.443s, approximate
prefill median 224.67 tok/s, and after-TTFT short-output median
60.19 tok/s. This is a service-lane baseline, not a LocalMaxxing headline
decode row. For production-visible OpenAI content deltas, use the validated
no-parser service variant (QWEN36_27B_REASONING_PARSER=), which passed the
same 32K exact-retrieval gate with reasoning_delta_count=0.
The MBT follow-up is notes/2026-07-04-long-context-mbt-screen.md: same-window
screening kept MAX_NUM_BATCHED_TOKENS=4096 for the 32K no-parser service lane;
MBT2048 passed but was slower, and MBT8192 stalled without a complete gate
artifact.
The current valid env-only win appears to preserve the accepted-state transition
by reading from the accepted speculative slot as the running source, then
disabling the now-redundant accepted-state postprocess copy. Later timing first
moved attention toward full LM-head/logits, but the 2026-07-05 refresh corrected
that for the current record family: INT8 LM-head/local-argmax is already small,
and the active frontier is target forward plus recurrent MTP draft forward.
Bad candidates already closed: blind copy skips, skipping
full-accept state, changing the Triton memcpy block size, exact argmax plumbing
that still computes full logits, draft local-argmax plumbing that still
computes full logits, FP8 LM-head (quality fail), INT8 MTP k2/k4/k5,
INT8 cg4/cg16/cg32, draft-only INT8 LM-head, output-buffer reuse, bonus-token
argmax fast-path, chunked INT8 top-1 argmax-only verification, compressed/full
EAGLE3 (device-loss or too slow), DFlash (no-win locally), and simple draft
top-k reranking. Later target-matched EAGLE3 top-k oracle showed real
candidate-list headroom, but diagonal and small MLP rerankers were closed
no-win (1.1069 and 1.1193 accepted draft tokens versus 2.249 top-8
oracle); see notes/2026-07-07-eagle3-topk-oracle-and-diag-reranker.md.
The wider top-k oracle in
notes/2026-07-07-eagle3-wide-topk-oracle-extractor-gate.md shows top-64 and
top-128 contain enough signal to cross 100 tok/s only under an impossible
same-cost magic extractor (103.76 / 111.23 tok/s). Treat that as direction
for rank-promotion or selected-candidate extraction research, not as a
throughput claim and not as justification for naive full-tree verification.
The first direct rank-promotion screen is closed in
notes/2026-07-07-eagle3-v6b-rankpush-no-endpoint.md: listwise top-k rank
loss tooling worked, but best heldout accepted depth moved only
1.10146 -> 1.10506, so simple loss weighting around this checkpoint is not
the missing extractor.
The follow-up wide top-k MLP reranker screen is closed in
notes/2026-07-07-eagle3-wide-topk-reranker-no-endpoint.md: top-64/top-128
with hidden sizes 512/1024 peaked at 1.11539, below the prior small top-8
MLP reranker (1.11927). Cheap selected-candidate extraction from this frozen
draft is therefore closed.
The latest draft-INT4 fast-path screens are also closed:
keep-scheduled-spec-row routing, graph-off, graph-off/no-async, cg4, and normal
align/restore all kept the same repeat64 failure (55/64 expected
blue, green, red, yellow, 9/64 truncated blue, green, red) despite
strict fresh cached_tokens=0 speed rows at 68-72 tok/s; see
notes/2026-07-05-draft-int4-specrows-and-graph-bisect-no-win.md.
Serial GDN is closed as well: native-on SERIAL_SPEC_* rows were fast but
still repeat-invalid and likely bypassed the Python serial path, while
native-off serial/fallback actually exercised the path and fell to
~9.7-12.3 tok/s; see
notes/2026-07-05-draft-int4-serial-gdn-nativeoff-no-win.md.
The next GDN-state implementation lane should be a fixed-shape exact
accepted-prefix tape / GPU-side commit, not more serial source/offset sweeps.
The 2026-07-06 branch/regenerate feasibility model further narrows the MTP3
branch lane: even a perfect legal MTP3 first-reject correction and regenerated
suffix reaches only ~101.8 tok/s at zero overhead and cannot reach 125+
without reducing verifier-step cost or increasing speculative depth. See
notes/2026-07-06-branch-regenerate-feasibility-envelope.md.
The executable contract is ../../scripts/check-gdn-spec-recurrent-exact.py;
as of 2026-07-06 it validates exact recurrent prefix state,
accepted-prefix SSM/conv commit equality on XPU for k=3/4/5, and the
endpoint row-to-draft-prefix mapping for full reject, partial reject, full
accept with bonus, shifted full accept, draft-only, and suppressed
bonus/replacement tails. See notes/2026-07-05-accepted-prefix-tape-contract.md
and notes/2026-07-06-gdn-endpoint-row-contract-extension.md.
The native packed spec prefix contract is also now checked directly by
../../scripts/check-gdn-native-spec-prefix.py; see
notes/2026-07-05-native-spec-prefix-contract-check.md. It confirms the native
op publishes column j as the state after packed row j and selects source
column num_accepted_tokens - 1. That closes the simple off-by-one/source
column explanation for the invalid fast draft-INT4 rows and points future work
at an exact ReplaySSM/tape commit transaction.
A metadata-only attempt to feed a separate GDN accepted-prefix count buffer
into attention metadata is closed no-win:
notes/2026-07-06-gdn-accepted-prefix-counts-no-win.md and
../../patches/qwen36-27b-autoround-int4-b70/vllm-qwen27-gdn-accepted-prefix-counts-no-win-20260706.patch.
It passed GDN contracts and repeat64 quality, but the strict candidate was not
promotable and slowed to 37.451 tok/s; do not confuse this with the deeper
fixed-shape GDN transaction still on the backlog.
A first commit-overhead reduction for that direction was tested:
VLLM_XPU_GDN_REPLAYSSM_COMMIT_IN_FORWARD=1 plus skipping the redundant
post-verify commit when no restore correction is active. It passed strict fresh
and repeat64 quality at 63.854 tok/s, improving over prior ReplaySSM rows but
still below the 65.276 tok/s record. Preserve it as no-promote evidence:
notes/2026-07-05-replayssm-commit-in-forward-skippost-no-promote.md and
../../patches/qwen36-27b-autoround-int4-b70/vllm-qwen27-replayssm-commit-in-forward-skippost-no-promote-20260705.patch.
The corrected Ex0bit EAGLE3 nested-aux-layer patch is preserved as a
compatibility artifact, but the retest still showed prompt-dependent
acceptance collapse and unusable endpoint throughput, so EAGLE3 remains closed
for this local vLLM/XPU + webhie/Intel AutoRound target unless the drafter
runtime, accepted-token bookkeeping, or target/draft pairing changes
materially.
Draft-side mtp.fc runtime INT8 is closed too: the default-off patch quantized
only the BF16 Qwen3.5 MTP mtp.fc layer and kept target verification exact, but
same-window strict fresh screening showed the completed candidate slower
(66.777 tok/s) than controls (67.954 and 67.994 tok/s), while another
candidate hit TorchDynamo fake-tensor unsupported-op handling for the custom
INT8 GEMM inside compiled MTP. See
notes/2026-07-06-mtp-fc-int8-no-win.md and
../../patches/qwen36-27b-autoround-int4-b70/vllm-qwen27-mtp-fc-int8-no-win-20260706.patch.
GDN gated-RMSNorm rstd skip is another closed model-body micro-optimization:
the patch skipped an ignored Triton rstd allocation/writeback behind
VLLM_XPU_RMSNORM_SKIP_RSTD=1, but strict fresh candidates (66.329 and
66.595 tok/s) lost to controls (67.716 and 67.910 tok/s). See
notes/2026-07-06-rmsnorm-skip-rstd-no-win.md and
../../patches/qwen36-27b-autoround-int4-b70/vllm-qwen27-rmsnorm-skip-rstd-no-win-20260706.patch.
Current-recipe deeper MTP is closed. MTP4/MTP5 require
VLLM_XPU_GDN_REPLAYSSM_SPEC_CACHE_LEN>=16 because the ring length must
satisfy ring_len >= 2 * max_spec_len; with cache8 they fail readiness
(got 8 < 10 / got 8 < 12). A follow-up native cache16/spec6 dispatch patch
compiled and passed direct native-vs-fallback parity for BF16/FP16/FP32, but
endpoint screening still lost: same-window MTP3/cache8 control was
67.816 tok/s, MTP3/cache16 65.410, MTP4/cache16/cg16 61.637, and
MTP5/cache16/cg16 58.140. The wider ReplaySSM templates also emitted heavy
AOT spill warnings. See
notes/2026-07-06-draftint4-depth-cachelen-no-win.md,
notes/2026-07-06-replayssm-cache16-native-s6-no-win.md, and
../../patches/qwen36-27b-autoround-int4-b70/vllm-qwen27-replayssm-cache16-spec6-no-win-20260706.patch.
Do not repeat config-only or simple dispatch-widening MTP4/MTP5 sweeps on this
recipe.
Intrinsic-MTP adaptation is also closed for the currently tested corpus/scope
families. FC-only and FC+norms improved offline acceptance but did not transfer
to the endpoint; MTP5 FC+norms reached only 1.78198 accepted draft tokens
offline, not enough to justify cache16 endpoint overhead. A later deep-scope
diagnostic trained dequantized attention/MLP/all-dense MTP tensors on four
GPUs and reached only 1.416016 accepted draft tokens (2.416016 visible
tokens/step), with dense updates that are not compatible with the packed INT4
checkpoint. See
notes/2026-07-07-intrinsic-mtp-adaptation-screen.md,
notes/2026-07-07-intrinsic-mtp5-adaptation-no-endpoint.md,
notes/2026-07-07-intrinsic-mtp-deep-scope-no-endpoint.md, and
diagnostics/qwen27-intrinsic-mtp-deep-scope-4gpu-summary-20260707.json.
Do not repeat these intrinsic-MTP sweeps unless a new mechanism can prove
3+ accepted draft tokens before endpoint work.
The full-vocab five-aux EAGLE3/DFlash rank-push route is closed too. The
four-GPU screen using the full Ex0bit draft and v7 five-aux corpus was stopped
early because the best heldout step-1 exact rate reached only 0.343628 at
step 6000, giving an impossible upper bound of just 1.718 accepted draft
tokens (5 * step1_exact) versus the 3+ accepted-token endpoint gate. See
notes/2026-07-07-eagle3-fullvocab-5aux-rankpush-earlystop.md and
diagnostics/qwen27-eagle3-fullvocab-5aux-rankpush-earlystop-summary-20260707.json.
Do not continue this exact rank-push recipe or endpoint-wire this draft.
The latest DFlash revisit is also closed:
notes/2026-07-06-dflash-swa-pr40898-repair-no-record.md. After reviewing
upstream vLLM PR #40898, a local DFlash SWA/full-KV repair was implemented and
syntax-checked. It fixed the old catastrophic mixed-SWA acceptance symptom and
produced strict fresh diagnostic rows, but remained far below the 67.519 tok/s
record: k2 49.087, k4 54.836, k8 50.918 tok/s, all quality-skipped and
not promotable. Preserve the patch for future upstream/DFlash comparison, but
do not repeat k/capture-size DFlash sweeps for this draft.
The latest source-drift repair is
notes/2026-07-06-mixed-draft-kv-metadata-guard-and-draft-int4-group-screen.md.
The external Qwen/Qwen3.5-0.8B draft-model experiment accidentally enabled
mixed draft-KV metadata for normal intrinsic MTP and dropped the record recipe
to ~60-61 tok/s. Active source now keeps mixed draft-KV metadata DFlash-only
by default, with VLLM_XPU_SPEC_DECODE_MIXED_DRAFT_KV_METADATA=1 as an
explicit opt-in for future external-draft work. A quality-backed confirm
restored 67.338 tok/s. Same-window screens closed draft INT4 group64,
group256, and fp32-scale as no-win versus group128/BF16 scales.
cd /home/steve/llm-optimizations
experiments/qwen36-27b-autoround-int4-b70/scripts/download-model.sh
GPU_INDEX=0 PORT=19410 MAX_MODEL_LEN=2048 \
experiments/qwen36-27b-autoround-int4-b70/scripts/serve-vllm.sh
BASE_URL=http://127.0.0.1:19410/v1 MODEL=qwen36-27b-int4-autoround \
experiments/qwen36-27b-autoround-int4-b70/scripts/smoke-openai.sh
The pinned Intel snapshot currently lives on the internal NVMe HF cache:
/mnt/fast-ai/llm-cache/hf/hub/models--Intel--Qwen3.6-27B-int4-AutoRound/snapshots/abc86de19eb1ebbf6a7df4582341325c22ddcb7d
An external 4 TB USB drive is mounted at /mnt/usb-models for overflow model
variants and archived artifacts. Keep active hot-path benchmarks on the
internal NVMe when practical; use /mnt/usb-models/llm-cache/hf or
/mnt/usb-models/models for additional variants if internal space becomes a
constraint. Do not commit model weights or generated cache contents.
/home/steve/src/vllm