Researchers have built a way to flag risky AI training data before a model ever gets fine-tuned on it.
The approach, called Alignment Forecasting, tries to predict whether a given fine-tuning dataset will push a target model toward a specific failure mode, like deception or sycophancy, before training happens. To test it, the team built ALIGNMENTFORECASTBENCH, a set of more than 5,000 forecasting questions covering 17 target models, 32 datasets, and 16 failure modes. Frontier models asked to make these predictions directly did poorly. The researchers' better approach uses an LLM to rate how strongly a dataset pushes toward bad behavior, then combines that rating with the failure mode's base rate and the target model's known tendencies, and this beat both a model trained specifically for the task and a method that just watches how weaker models behave after training on the same data.
Right now, misalignment mostly gets caught the hard way: after training, when someone audits the finished model and finds a problem already baked in. Shifting that check earlier, to the training data itself, could let teams filter out bad examples before they cause damage instead of patching a model after the fact. The team tested this on UltraChat, a real post-training dataset, and filtering out flagged examples produced more aligned models on multiple-choice evaluations in most cases.
That multiple-choice win did not clearly carry over to open-ended conversations, which is exactly where alignment problems tend to show up in the wild. The authors themselves call this early-stage progress, not a ready-made safety filter, so treat it as a promising research direction rather than a tool anyone should bolt onto a production pipeline yet.