b70-optimization-lab

Qwen3.6 35B Quark INT8 on B70

This folder is the durable result packet for the Qwen3.6 35B A3B Quark W8A8 INT8 lane on Intel Arc Pro B70. It exists to make the final state easy to reference before moving effort to another model.

Bottom Line

The Qwen3.6 35B 4x B70 lane is exhausted for now. The best strict-valid TP4 result is 93.55 tok/s corrected output on the PIECEWISE forced-comm graph baseline. A legacy LocalMaxxing-approved run reached 99.43 tok/s, but newer deep gates are stricter and should be used for current claims.

No attempted speculative decode path produced a valid >150 tok/s result for this model. The fastest numbers above the baseline were invalid, synthetic, or crashed before validity gates.

Start Here

Best Results Summary

Scope Result Validity Primary artifact
4x strict-valid current base 93.55 tok/s JSON 128/128, color 256/256, quality pass; LocalMaxxing cmqq4mw4c00yfqo01gb2ucgxj deep-gate-summary
4x legacy public approved 99.43 tok/s LocalMaxxing approved, older gates localmaxxing snapshot
2x reference smoke 85.87 tok/s JSON 16/16, color 16/16, quality skipped; LocalMaxxing cmqq4mwgm00yiqo0133bj962q tp2-smoke-summary
2x older raw reference 91.59 tok/s smoke/reference only, not promoted tp2-latency-truth
Fastest raw artifact 198.95 tok/s invalid/reference only ngram5 raw
Fastest synthetic ceiling 181.91 tok/s canaries skipped, synthetic accept eagle2 ceiling summary
Fastest MTP-ish invalid 107.77 tok/s JSON and color failed immediately mtp parity fix v2

Status Recommendation

Stop spending ad hoc benchmark time on Qwen3.6 35B Quark INT8 unless the task is one of these controlled follow-ups:

For new optimization work, switch to another model. Prior successful lanes include Qwen 27B and MiniMax M2.7; Gemma 4 12B also has strong TP4 production material. See next model carryover.