Determine the appropriate estimand for position-varying latent effects

Determine whether ablation-based causal evaluation of a sparse-autoencoder latent should report a distribution over firing tokens, an activation-weighted expectation, or a context-conditioned estimand rather than a single scalar, and identify which alternative is most stable across dictionaries at equal computational cost.

Background

The paper shows that a latent’s measured causal effect can vary substantially across the tokens at which it fires, with measurement position accounting for a large share of the observed variance. Because the conventional top-activating-token protocol allows the dictionary to select the evaluation position, a scalar causal score may conflate the latent’s effect with the particular token chosen.

The authors explicitly leave unresolved what quantity should replace or supplement that scalar. They propose three empirically testable possibilities—an across-token distribution, an activation-weighted expectation, and a context-conditioned estimand—and state that the alternatives should be compared for cross-dictionary stability using the paper’s crossed design.

References

What we cannot say is why one firing position differs from another, and that names the next question: if a latent's effect varies this much across the tokens where it fires, the quantity to report may be a distribution over those tokens, an activation-weighted expectation, or a context-conditioned estimand rather than a scalar, decidable empirically, by whichever is more stable across dictionaries at equal cost on the crossed design used here.

Where You Measure Decides What You Measure: Position Selection in Ablation-Based SAE Evaluation  (2608.13337 - Noël, 13 Aug 2026) in Section 6, Limitations and conclusion