llama.cpp b11335: CUDA Volta (sm70) Now Uses Turing MMVQ Tuning for Improved Performance
llama.cpp's latest release reroutes Volta (sm70) CUDA kernels to use Turing's MMVQ parameter table. This improves K-quant batch-1 decode performance on Tesla V100 GPUs by 3.17% when running quantized models.
What changed?
llama.cpp b11335 changes how CUDA kernels are tuned for NVIDIA Volta (sm_70, e.g. Tesla V100) GPUs. Previously, Volta used a generic MMVQ parameter table, which launched K-quant batch-1 decode at nwarps=4. Now, Volta is routed to use the Turing MMVQ parameters, specifically launching the K-quant vec_dot kernel at nwarps=2. This small, targeted tuning change improves throughput for quantized model decoding tasks on relevant GPUs.
Why does it matter to an everyday developer?
If you run quantized LLM inference (K-quant models) on NVIDIA Tesla V100 GPUs, this update gives you about 3% faster batch-1 decode performance—with the same model accuracy (perplexity remains identical). No code or model changes are required: just use the latest llama.cpp release. Developers deploying inference backends on V100s, especially in cost-sensitive or latency-sensitive environments, directly benefit.
What can the developer do now?
- 1
Upgrade to llama.cpp b11335
Download the latest release binaries for your platform (macOS, Windows, Linux, etc.) from the official llama.cpp release page.
- 2
Test on NVIDIA V100 (sm_70)
Deploy or run your workload on a V100 GPU. Compare inference throughput using K-quantized models before and after the upgrade if you wish to confirm the improvement.
