llama.cpp b11307: Improved Batch Order Preservation and Layer Input Handling
llama.cpp b11307 introduces accurate preservation of the original batch order for speculative decoding layer inputs, better handling of token order for NextN embeddings, compatibility improvements with tensor split, and regression coverage for various decoding scenarios. Some backend and test updates are deferred to future releases.
What Changed
- Original batch order is now preserved for speculative decoding layer inputs.
- Layer input reordering is fully compatible with tensor splitting.
- Token order for unmasked NextN embeddings is restored using the original token mapping, even when layer-input capture is disabled.
- Support for masking/restoring original row order after synchronization.
- Extended regression tests cover tensor split, repeated reads/decodes, CUDA devices, and NextN embedding scenarios.
- WebGPU reservation across batch sizes and OpenVINO hidden-state capture have improved coverage and allocation checks.
- Variable names standardized: batch indices to batch_idxs and embedding indices to embd_batch_idxs.
- n_embd variable declaration position preserved to avoid side effects.
Implications for Developers
These updates make batch processing and speculative decoding more reliable, particularly in multi-device or high-concurrency scenarios. Developers using custom KV layouts, tensor split, or complex decoding strategies benefit from improved correctness and debugging. Regression test coverage is stronger, though some additional tests and backend-specific fixes are scheduled for future releases.
