llama.cpp b11388: Activation-Based Layer Statistics and New CLI Options
llama.cpp b11388 adds support for calculating activation-based statistics on GGUF models, offering new insights into model behavior and quality directly from the CLI, and introducing several new reporting metrics and API extensions.
What changed?
Release b11388 of llama.cpp adds the ability to calculate and report a range of activation-based statistics for GGUF-format language models. Newly introduced metrics include entropy, cosine similarity, L2 norm (both raw and per-layer aggregated), sum of squared activations, and the new Euclidean–Cosine Score (ECS). These can be computed for model 'imatrices' (internal matrices) directly via a new --activation-statistics CLI option. In addition, the API gains a new compute_layer_statistics() function, output formatting is improved, and reporting of ZD Scores is now two-tailed. The release is broadly available across desktop, server, and mobile platforms for most popular architectures; openEuler builds are disabled.
Why does it matter to an everyday developer?
These activation-based statistics give developers and researchers a way to better understand model internals, monitor quantization effects, and detect outliers or problematic layers before or during deployment. This is valuable when evaluating model quality, comparing model versions, or debugging unexpected inference behavior. The new CLI and API options make it much easier to automate statistics collection and integrate these checks into model pipelines or CI workflows.
What can the developer do now?
- Use the --activation-statistics flag in the CLI to generate detailed statistics on GGUF-format models without custom code.
- Leverage the compute_layer_statistics() function in your own code to automate or customize reporting.
- Monitor entropy, cosine similarity, aggregated L2 norm, ECS, and other metrics to catch quantization errors, outliers, or unexpected distributional changes.
