Explanation for U-shaped attribution curves

Explain why token-level attribution methods, particularly EK-FAC, can produce U-shaped decile-training curves in which tokens assigned negative influence have a positive effect on the subliminal response rate comparable to that of positively attributed tokens.

Background

The appendix reports that token-level filtering for some target animals in the Llama and OLMo experiments yields U-shaped relationships between attribution rank and subliminal response rate. In these cases, tokens with intermediate attribution appear to have little effect, whereas tokens assigned negative scores can produce subliminal learning at rates similar to tokens assigned positive scores.

The authors explicitly state that they do not know why this pattern occurs. Resolving its cause could clarify whether the behavior reflects a limitation of the attribution approximations, an artifact of ranking or filtering, or a more complex relationship between token-level parameter updates and subliminal learning.

References

We see that for token-level filtering of some animals in Llama and OLMo, some of the attribution methods (especially EKFAC) yield a U-shaped curve, which means that these negative score tokens have positive influence on subliminal learning rate. We are unsure as to why this is.

Can Data Attribution Filter Out Subliminal Learning? Not Reliably  (2609.20027 - Weckbecker et al., 17 Sep 2026) in Appendix, Section “Full results under decile training” (Appendix A)