b70-optimization-lab

Exact-Q4 Native DFlash Trace Closeout

Date: 2026-07-13

Status: native hook and ABI smoke complete; real capture and adaptation screen not run because the Qwen optimization lane closed before the protected runtime tree could be frozen.

Preserved implementation

The patch is deliberately labeled protected-stack: the shared speculative.cpp and server-context.cpp were already dirty with the active Q6 target-top1 experiment, so the snapshot contains that concurrent context as well as the trace plumbing. Do not blindly apply it to a different tree; use the isolated trace helper and named hook points from the blueprint when rebasing.

Verified

Remaining risks

Reopen procedure

  1. Rebase the isolated helper and lifecycle hooks onto a clean/frozen runtime.
  2. Compute the full dirty patch SHA-256, update capture-plan.json, and rebuild with LLAMA_DFLASH_TRACE_RUNTIME_COMMIT and LLAMA_DFLASH_TRACE_DIRTY_PATCH_SHA256 set to that exact identity.
  3. Launch TP1/parallel-one DFlash with draft K/V F16, n_max=n_min=0, p_min=0, checkpoints off, reasoning off, and the two deep model hashes.
  4. Run one collector prompt, parse it, and repeat the four corruption checks.
  5. Only then collect train/heldout traces and run the bounded layer-position-bias, draft-token-five screen. Stop unless it clears the 4 accepted / 5 emitted hard gate.