Sparse Autoencoders (SAEs) decompose model activations into sparse combinations of dictionary atoms that can be interpreted as concepts. The paper argues that standard SAEs introduce an independence prior across image patches through their objective, even though natural images and the activations they produce contain spatial dependencies. To address this mismatch, the authors specialize the Linear Representation Hypothesis for vision as the Markov-Field Linear Representation Hypothesis (MFLRH), which explicitly includes those dependencies. They then propose Spatial-SAE as an amortized maximum-a-posteriori estimator under MFLRH. Across four variants, Spatial-SAE has a 96% average win rate over standard SAEs on synthetic concept-recovery tests and also improves interpretability on DINOv2 activations. The improvement comes with a reconstruction trade-off concentrated in high spatial frequencies, indicating that better concept structure is obtained at some cost to fine-detail reconstruction.
AI News
The latest AI releases, research, products, and industry updates.
Loading...