Calibrated Statistical Methods for LLM Judge Scores: evalstats Python Package Released
A new paper presents evalstats, an open-source package addressing inflated false positives in AI evaluation based on LLM judge scores. Developers gain access to robust, calibrated statistical methods and actionable guidelines for small-sample studies.
Researchers often use LLMs as judges in small-sample AI evaluation studies, but recent findings indicate that standard statistical analyses over these scores can lead to unreliable significance claims and inflated false positives—especially when human-LLM agreement appears high. The new evalstats Python package addresses these risks with calibrated statistical inference methods tailored for mixed human-LLM judge setups.

Key Changes and Capabilities
- Implements nine hypothesis tests using prediction-powered inference (PPI), including the first known PPI corrections for Wilcoxon signed-rank, Mann-Whitney U, and related omnibus rank-based tests.
- Introduces bootstrap-adaptive power tuning to stabilize inference in small human-labeled calibration sets by adaptive shrinkage and variance adjustment.
- Automatically selects the most calibrated statistical methods for confidence intervals and p-value corrections in small-sample AI evaluations (N<100).
Recommended Practices for Developers
- Avoid running simple statistics over raw LLM judge scores without calibration, as this can yield inflated false positive rates—even where human-LLM agreement is "almost perfect."
- For small test sets (N<100), do not use bootstrap confidence intervals; instead, use the methods provided in evalstats.
