When and how synthetic data improve generalization and transfer
Characterize the regimes in which synthetic data augmentation improves out‑of‑distribution generalization and transfer learning performance, including identifying beneficial types of distributional shift, quantifying the impact of generative‑model estimation error, and developing diagnostics for harmful extrapolation.
References
A central open problem is therefore to characterize when and how synthetic data improve generalization ability and transferability. This includes, but is not limited to, identifying the types of distributional shifts for which synthetic augmentation is beneficial, understanding the role of the estimation error of the generative model, and developing diagnostics to detect harmful extrapolation.
This question becomes more practical when considering data scalability: curating large-scale real videos with explicit reasoning goals is expensive, and synthetic data offers a more controllable and scalable alternative, but it remains unclear whether such scaling can systematically improve video reasoning in realistic scenarios.
Relatedly, mean-pooled representations from frozen transformer encoders like RoBERTa are known to exhibit anisotropy \citep{ethayarajh-2019-contextual}, which can distort cosine-based distance measures; it remains an open question whether our distance-sensitive findings would hold under embeddings specifically designed for cosine geometry, such as sentence-transformers \citep{reimers-gurevych-2019-sentence}.