Determine the causes of U-shaped attentional bias in transformer models

Determine the causes of the U-shaped attentional bias exhibited by transformer-based language models, in which material at the beginnings and ends of long texts receives greater weight than material in the middle, in order to clarify why this bias arises and contributes to degraded recall and hallucination.

Background

The paper explains that transformer-based LLMs exhibit a U-shaped attentional bias: they tend to weight information at the beginnings and ends of long texts more heavily than information in the middle. This behavior can produce recall degradation and increase the risk of hallucination when models analyze long archival documents or document sets.

The authors note that several possible explanations have been proposed, including the distribution of valuable framing material in the beginning and end of documents and the autoregressive structure of language-model prediction. However, they explicitly state that the causes of this attentional pattern remain uncertain, leaving the underlying mechanism unresolved.

References

There are a few probable reasons for this, though the causes are uncertain.

INDRA: A New AI Tool for Exploring Tobacco, Fossil Fuel, and Chemical Industry Archives  (2609.11261 - Akselrad et al., 10 Sep 2026) in Endnote discussing the "lost in the middle" effect and attention sinks in the Introduction–Background section