Representational Simplicity Does Not Consistently Predict Smaller Causal Circuits
Summary
This study tests whether cleaner internal representations and more concentrated feature attributions actually imply smaller causal circuits that are easier to reverse-engineer. The authors use matched standard and adversarial continual training from the same pretrained GPT-2 Small checkpoint, with both models required to retain competence on indirect object identification and pass independent robustness verification. They compare sparse-autoencoder decomposability, the number of SAE features engaged in task attribution, and the size of faithful circuits recovered from the raw computational graph. The robustly trained model is more SAE-decomposable and uses fewer SAE features for task attribution. Circuit size, however, depends on the required faithfulness level: on the competence-matched IOI task, the standard model leads or ties below 85% faithfulness, while the robust model requires substantially fewer edges at 90% and 95% faithfulness. The threshold-dependent pattern is established on the primary model pair. Representational trends also generalize across a seven-point sweep and a second corpus. The results therefore separate representational or attributional simplicity from causal simplicity rather than treating the former as a reliable proxy for the latter.