llama.cpp b11345: Adds q2_k and q3_k Quantization Support for Hexagon Backend
The latest llama.cpp release (b11345) brings support for q2_k and q3_k quantization types to Hexagon backend, enabling more efficient LLM inference on devices with Qualcomm Snapdragon NPUs.
What changed?
llama.cpp release b11345 introduces support for the q2_k and q3_k quantization types on the Hexagon backend, which targets Qualcomm Snapdragon NPUs (neural processing units). The release also improves internal memory allocation consistency related to these quantized kernels. No breaking API changes or backward-incompatible modifications were reported.
Why does it matter to an everyday developer?
With quantization support for q2_k and q3_k types on Hexagon, developers targeting Snapdragon-powered devices—for example, smartphones or edge AI devices—can use smaller and more memory-efficient LLM models. This makes running large language models locally on-device more practical, improving speed and conserving battery or compute resources. If you are developing mobile, IoT, or embedded AI applications, this improvement unlocks more deployment options and potentially better end-user experiences.
What can the developer do now?
Developers building or deploying LLM solutions on Snapdragon-based platforms can now take advantage of q2_k and q3_k quantized models for the Hexagon backend using llama.cpp. Download the latest binaries or update your build to b11345, convert compatible models to the new quantization types, and deploy or test them on Hexagon-compatible devices. Official binaries are available for a wide range of platforms, including an Android and Linux ARM64 package specifically supporting the Snapdragon ecosystem.
