Mechanism of in‑context learning in Transformers
Characterize the mechanism by which Transformer-based Large Language Models perform in-context learning without parameter updates.
References
Despite the good performance of the ICL capabilities, the mechanism of ICL still remains an open question.
— Beyond the Black Box: Theory and Mechanism of Large Language Models
(2601.02907 - Gan et al., 6 Jan 2026) in Subsubsection In-Context Learning, Section 6: Inference Stage (Core Theories and Methods)
We were unable to definitively identify or exclude additional mechanisms beyond learned retrieval that may underlie out-of-distribution generalization. Further research is therefore required to achieve a more comprehensive understanding of how in-context learning operates in TFMs.
— Understanding the Surprising Generalization Properties of Tabular Foundation Models
(2608.17957 - Shaheen et al., 18 Aug 2026) in Section 6, “Conclusion and limitations”