Identify the mechanism of relevance inference in feature-attention pathways

Determine where relevance is inferred in the alternating-axis models and whether reduced influence from inactive feature sources reflects suppression within the attention mechanism itself.

Background

The paper studies whether feature-axis attention in alternating-axis tabular transformers enables task-dependent routing of computation under sparse priors. Interventions that attenuate messages associated with individual features show that active sources tend to exert more specific and stronger control over their own predictive coefficients than inactive sources, particularly in final-layer query messages. These findings provide evidence for feature-aligned computational routing but do not reveal how the model determines feature relevance.

The unresolved issue is whether relevance is inferred before the intervened feature messages are formed, within the attention operation itself, or through other upstream or downstream computations. Establishing this would require tracing the formation and propagation of the relevant internal representations rather than only measuring the effects of message attenuation on fitted coefficients. The question is included because the paper explicitly states that its intervention experiments do not resolve it.

References

Selective computational routing is a candidate mechanism underlying the alternating-axis advantage, but these experiments do not establish where relevance is inferred or whether reduced inactive-source influence reflects suppression within attention itself.

— Architecture Alignment With Sparse Priors in Tabular Foundation Models  (2609.36883 - Zhao et al., 29 Sep 2026) in Section 4, subsection “Evaluating and interpreting selective routing” (Section~\ref{sec:cross-scale-gating})