Model identity for all results in this file unless stated otherwise:
nameistoken/Qwen3.6-35B-A3B-Quark-W8A8-INT8;cced56592e8c8935f8220836b4baa04dfd389118;prefill-safe-int8-mixed-workspace-async-deep-gate, 20260615a13deep2.
| Metric | Value |
|---|---|
| Corrected output throughput | 93.55054235558917 tok/s |
| E2E output throughput | 90.62548580766561 tok/s |
| Total client token rate | 178.77293098777787 tok/s |
| Decode latency | 10.68988536269444 ms/token |
| TTFT client mean | 187.33663426246494 ms |
| JSON canary | 128/128, pass |
| Color canary | 256/256, pass |
| Quality suite | pass |
| Decision | accepted by requested gates |
| LocalMaxxing ID | cmqq4mw4c00yfqo01gb2ucgxj, APPROVED |
Primary artifacts:
deep-gate-summarydeep-gate-metricsdeep-gate-jsondeep-gate-colordeep-gate-qualityLocalMaxxing submission log2026-06-14-qwen36-recovery-implementation.mdKey identity fields:
COMPILATION_CONFIG='{"cudagraph_mode":"PIECEWISE"}'XPU_GRAPH=1VLLM_XPU_ENABLE_XPU_GRAPH=1VLLM_XPU_FORCE_GRAPH_WITH_COMM=1VLLM_XPU_GRAPH_NOOP_COMM_CAPTURE=1VLLM_XPU_GDN_NATIVE_FALLBACK=prefillVLLM_XPU_GDN_PREFILL_RECURRENT_FALLBACK=1VLLM_XPU_DISABLE_PREFILL_CUDAGRAPH_REPLAY=1VLLM_XPU_GREEDY_SAMPLE_TOPK_FALLBACK=1VLLM_XPU_INT8_MOE_MIXED_WORKSPACE=1GPU_MEMORY_UTILIZATION=0.90This is the current safe baseline for future 4x comparisons.
LocalMaxxing approved record:
localmaxxing-qwen36-35b-quark-int8-exacthf-20260612ak.json.
| Metric | Value |
|---|---|
| LocalMaxxing ID | cmq8yhxvo001ipb0149aoa79o |
| Status | APPROVED |
| Corrected output throughput | 99.42835812273452 tok/s |
| Total throughput | 196.3252731420561 tok/s |
| TTFT | 76.45406149094924 ms |
| Shape | p512/o512, streaming completions, temperature 0, 4 repeats after warmup |
Supporting artifacts:
Caveat: this was approved and should remain in the record ledger, but later work
added stricter repeat/canary discipline. Use the 93.55 tok/s deep gate as the
current strict-valid comparison point unless this older run is revalidated under
the newer gates.
prefill-safe-int8-mixed-workspace-async-smoke, 20260615a13.
| Metric | Value |
|---|---|
| Corrected output throughput | 95.01697182719025 tok/s |
| E2E output throughput | 92.004380418574 tok/s |
| Decode latency | 10.525270629841543 ms/token |
| JSON canary | 32/32, pass |
| Color canary | 32/32, pass |
| Quality suite | skipped |
Artifacts:
Use this only as a smoke reference, not a record claim.
These are valuable for direction-setting but are not valid records.
| Lane | Throughput | Status | Why not valid |
|---|---|---|---|
| ngram5 current-storeguard raw | 198.9479016380729 tok/s corrected |
raw artifact | ngram/spec artifact, not clean endpoint correctness |
| EAGLE2 tokenheavy synthetic accept | 181.9100662518911 tok/s corrected |
synthetic ceiling | canaries skipped; synthetic accept, not valid serving output |
| MTP k1 parity-fix-v2 | 107.76565033909118 tok/s corrected |
invalid | JSON and color failed on first repeat |
| MTP k1 ReplaySSM graph cap-nonuniform | 75.69792475939613 tok/s corrected |
invalid | JSON 96/96, color failed with first-decode garbage |
| MTP k1 ReplaySSM graph cap-small full | 76.244550687551 tok/s corrected |
invalid | JSON failed at 83, color failed at 23; full-accept double-processing signature |
| MTP k1 restore-sync | 73.45316484087655 tok/s corrected |
invalid | restore sync did not fix JSON/color; below the target direction |
| MTP k3 graph throughput probe | no throughput result | crash | engine-core cancellation after scheduling 4 tokens with 3 spec tokens |
Artifacts:
ngram5 raweagle2 synthetic ceilingmtp parity-fix-v2replayssm iter14 graph fullreplayssm iter19 graph cap-small fullreplayssm iter20 restore-syncmtp k3 crash logThe strict-valid non-spec graph baseline is near 93.55 tok/s. The >150 tok/s
target requires more accepted tokens per step or a fundamentally cheaper
proposer/verifier path. K1 MTP never got there: even when fast enough to appear
promising, it failed canaries, and the correct ReplaySSM path was either slow or
graph-racy.
The best use of this lane now is as a reference set for upstream/runtime bakeoffs and as a cautionary record for speculative decode state handling.