Determine the source of the discrepancy between trained SAEs and theoretical minimizers

Determine whether the discrepancy between trained many-feature amortized sparse autoencoders and the globally optimal dictionaries of the exact-coding objective is caused by amortization, the many-feature data distribution, finite-sample effects, or the optimization path.

Background

The paper contrasts its trained amortized ReLU sparse autoencoders with a two-feature theoretical benchmark in which the globally optimal dictionary merges nested features. The experiments do not exhibit the threshold-defined pairwise merging predicted by that benchmark.

The authors explicitly leave unresolved which aspect of the experimental setting explains the difference: amortized inference, the many-feature distribution, finite data, or optimization dynamics. They relate the amortization possibility to open problem MAIS-O39.

References

Whether that difference is due to amortization (the concern of open problem MAIS-O39), the many-feature distribution, finite-sample effects, or the optimization path itself remains open.

— A Dominant Diffuse Phase in the Sparse Autoencoder Phase Diagram  (2609.10299 - Plascencia, 9 Sep 2026) in Section 7, Discussion, paragraph “Trained SAEs versus minimizers”

Whether it persists across additional dictionary and data draws, SAE variants, and global minimizers of $G_\lambda$ remains an open question.

— A Dominant Diffuse Phase in the Sparse Autoencoder Phase Diagram  (2609.10299 - Plascencia, 9 Sep 2026) in Section 8, Conclusions, final paragraph