Fine-Tuning Changes in LLM Representations Do Not Predict Causal Importance
Summary
This study examines how fine-tuning changes internal mechanisms in large language models, focusing on attention patterns and layer-wise activations. It compares those changes with task-relevant components identified by EAP, including attention heads and logit-level activations that contribute to performance. EAP components cluster in particular layers, suggesting that task-specific behavior is functionally localized. However, the layers with the largest representational changes during fine-tuning are largely different from the layers containing those causally important components. The study also finds that shared EAP components across tasks do not guarantee transfer when the tasks differ, such as classification and generation. In some cases, fine-tuning for one task reduces performance on another, even when the tasks have substantial overlap in their identified components.