Empirical validity of within-task exchangeability for real pretraining corpora

Determine whether within-task exchangeability is a good approximation for real pretraining corpora, thereby assessing the empirical validity of the hierarchical task-prior model used to represent in-context learning.

Background

The paper models in-context learning using a hierarchical task-prior representation in which examples generated within a task are exchangeable. Appendix A.4 argues that, under infinite within-task exchangeability, this representation follows canonically from the de Finetti–Hewitt–Savage theorem rather than constituting an additional modeling assumption.

The unresolved issue is whether this exchangeability assumption is empirically appropriate for real pretraining corpora. Natural-language data may contain document order, structural dependencies, and mixtures of tasks that violate or only approximately satisfy exchangeability; resolving this question would test the applicability of the paper’s Bayesian in-context-learning framework to real-world data.

References

Whether within-task exchangeability is a good approximation for real pretraining corpora is ultimately an empirical question.

— Unifying ICL, SFT, KL-Regularized RL Through a Bayesian Lens  (2609.05111 - Fan, 4 Sep 2026) in Appendix A.4.4, Remark (4) “Empirical plausibility”