Determine the appropriate estimand for position-varying latent effects
Determine whether ablation-based causal evaluation of a sparse-autoencoder latent should report a distribution over firing tokens, an activation-weighted expectation, or a context-conditioned estimand rather than a single scalar, and identify which alternative is most stable across dictionaries at equal computational cost.
References
What we cannot say is why one firing position differs from another, and that names the next question: if a latent's effect varies this much across the tokens where it fires, the quantity to report may be a distribution over those tokens, an activation-weighted expectation, or a context-conditioned estimand rather than a scalar, decidable empirically, by whichever is more stable across dictionaries at equal cost on the crossed design used here.