llama.cpp b11390 Fixes CUDA MMQ Memory Fault
The b11390 release of llama.cpp addresses a CUDA memory fault affecting MMQ when the number of experts greatly exceeds the number of microbatches. The update improves stability for developers running CUDA workloads and is already available across major platforms.
What changed?
llama.cpp b11390 introduces a fix for a CUDA memory fault that could occur in the Mixture of Experts (MMQ) configuration when the number of experts (n_expert) is much larger than the number of microbatches (n_ubatch). This update prevents specific crashes or memory faults during CUDA execution in such scenarios. Pre-built binaries are available for macOS, iOS, Ubuntu, Linux, Android, and Windows with support for various hardware and backends.
Why does it matter to an everyday developer?
If you use CUDA acceleration with llama.cpp, particularly for models configured with a high number of experts compared to microbatches, this fix prevents memory faults that could otherwise terminate your inference jobs. This is relevant for developers experimenting with or deploying large language models using CUDA and MMQ features. The fix increases reliability and reduces debugging time due to intermittent GPU memory errors.
What can the developer do now?
Developers using CUDA in their llama.cpp pipelines should upgrade to b11390 to avoid possible memory faults when n_expert is much greater than n_ubatch. Updated binaries are available for all supported platforms, including macOS (Apple Silicon and Intel), Ubuntu, Windows, Android, and more. Download the latest release and update your deployment or testing environments as appropriate.
