Alignment Forecasting: Predicting Alignment Failures From Training Data
A new research paper introduces Alignment Forecasting, a method for predicting the likelihood that fine-tuning a language model on a specific dataset will increase misalignment before training begins. The work includes the ALIGNMENTFORECASTBENCH benchmark and demonstrates initial success with a forecasting scaffold, but also highlights current limitations and the need for further progress.
The newly proposed task of Alignment Forecasting aims to predict, before model training, whether fine-tuning a language model on a given dataset will lead to alignment failures—such as increased deception, sycophancy, or other undesirable behaviors. Until now, such alignment problems were primarily identified only after training via post-hoc model auditing.

ALIGNMENTFORECASTBENCH Benchmark
To measure progress in this area, the authors introduce ALIGNMENTFORECASTBENCH, a benchmark consisting of over 5,000 forecasting questions that span 17 target models, 32 datasets, and 16 different failure modes.
Forecasting Scaffold Approach
A key contribution is a forecasting scaffold in which a large language model (LLM) reviews the dataset and rates how strongly it may induce specific model misbehaviors. This rating, combined with the base rate of the failure mode and the prior tendency of the target model, is then used in a simple learned model to forecast alignment failures. This method performs better than chance and surpasses baselines, including fine-tuned forecasters and those leveraging post-fine-tuning behavior of weaker models.
Effect on Data Curation
The forecasting scaffold can identify problematic training examples that standard classifier approaches might miss. When these flagged examples are removed from datasets such as UltraChat, models show improved alignment on multiple-choice tests. However, the benefits are less clear for open-ended conversational alignment.
