Determine whether the causal gap is dissociated from reconstruction improvement

Determine whether improving sparse-autoencoder reconstruction quality without a corresponding increase in the trained-versus-soft-frozen causal-effect ratio represents a genuine dissociation between reconstruction and causal performance.

Background

Across four training budgets, the explained-variance gap between trained and soft-frozen autoencoders increases significantly, whereas the causal-effect ratio is non-monotone and its fitted trend is not statistically significant. This pattern is consistent with reconstruction quality and causal performance changing independently.

The authors explicitly state that the available four-budget experiment does not establish a dissociation. Additional training budgets, larger samples, or a more adequately powered comparison would be needed to determine whether the apparent separation is a reproducible phenomenon.

References

The reconstruction gap grows significantly and the causal gap does not. With four budgets, $120$ latents each, and non-monotone point estimates, this is consistent with a dissociation but does not establish one.

Where You Measure Decides What You Measure: Position Selection in Ablation-Based SAE Evaluation  (2608.13337 - Noël, 13 Aug 2026) in Appendix, Section “Reconstruction improves without a matching causal gain”