Tokenized SAEs: Disentangling SAE Reconstructions (2502.17332v1)

Published 24 Feb 2025 in cs.LG

Abstract: Sparse auto-encoders (SAEs) have become a prevalent tool for interpreting LLMs' inner workings. However, it is unknown how tightly SAE features correspond to computationally important directions in the model. This work empirically shows that many RES-JB SAE features predominantly correspond to simple input statistics. We hypothesize this is caused by a large class imbalance in training data combined with a lack of complex error signals. To reduce this behavior, we propose a method that disentangles token reconstruction from feature reconstruction. This improvement is achieved by introducing a per-token bias, which provides an enhanced baseline for interesting reconstruction. As a result, significantly more interesting features and improved reconstruction in sparse regimes are learned.