Date: 2026-07-13
Status: native hook and ABI smoke complete; real capture and adaptation screen not run because the Qwen optimization lane closed before the protected runtime tree could be frozen.
/home/steve/src/llama.cpp/common/dflash-target-trace.{h,cpp}.common/CMakeLists.txt, common/speculative.{h,cpp}, and
tools/server/server-context.cpp.experiments/qwen27-dflash-sycl-b70/tests/dflash-q4-target-trace-native-smoke.cpp.2026-07-13-dflash-q4-target-trace-hook-blueprint.md.patches/qwen27-dflash-exact-q4-native-trace-protected-stack-20260713.patch,
SHA-256 ef865ff8397bba95a98c33e7e92bfe537ac05c84134a0031795891be9c4b8970.The patch is deliberately labeled protected-stack: the shared
speculative.cpp and server-context.cpp were already dirty with the active
Q6 target-top1 experiment, so the snapshot contains that concurrent context as
well as the trace plumbing. Do not blindly apply it to a different tree; use
the isolated trace helper and named hook points from the blueprint when
rebasing.
llama-server build completed with both trace compile-time
identities set; speculative.cpp, the helper, and server-context.cpp all
compiled and linked.[3,5,5120] features, contiguous positions, and delayed labels.qwen36_eagle_sequence_v2 with two training
anchors. Native identity, payload-byte, position, and label corruptions were
rejected.4.0 accepted drafts and 5.0 emitted tokens per favorable
width-six cycle.capture-plan.json, and rebuild
with LLAMA_DFLASH_TRACE_RUNTIME_COMMIT and
LLAMA_DFLASH_TRACE_DIRTY_PATCH_SHA256 set to that exact identity.n_max=n_min=0,
p_min=0, checkpoints off, reasoning off, and the two deep model hashes.layer-position-bias, draft-token-five screen. Stop unless it clears the
4 accepted / 5 emitted hard gate.