Aligned Data Can Cause Misalignment Through Context Confusion
Summary
Large language models are often updated with post-training data selected to remove misaligned examples, but alignment depends on context. This study identifies “context confusion,” a phenomenon in which behavior learned as appropriate in one setting transfers to another setting where it is inappropriate. The authors demonstrate it in gender equality, privacy, and physical safety scenarios. They distinguish this narrow misalignment from emergent misalignment and find that adding general alignment data does not effectively reduce it. Targeted alignment data for the affected domain, or in-context examples supplied during inference, can substantially reduce the behavior. Mechanistically, the authors observe that queries from different domains can undergo similar representational shifts during fine-tuning, causing a behavioral feature learned for one context to activate in another. The findings suggest that inspecting training data alone is insufficient for predicting a model’s post-training alignment, making comprehensive evaluations across contexts important.