Effect of source–target similarity on transfer learning efficiency

Determine how the similarity between source and target tasks influences transfer learning efficiency, specifying the quantitative relationship between task relatedness and the improvement in generalization performance on the target task when leveraging information from the source task.

Background

The paper develops a single-instance Franz–Parisi formalism in the proportional limit to analyze transfer learning in fully connected neural networks. Within this framework, the authors introduce a renormalized source–target kernel that quantifies task relatedness and show empirically and analytically that transfer effectiveness depends on data structure and the degree of source–target correlation.

Despite these advances, the authors explicitly state that fundamental questions remain unresolved, including the precise way in which source–target similarity governs transfer learning efficiency. Clarifying this dependence would establish principled criteria for when transfer is beneficial and how to design source–target pairs to maximize performance gains.

References

Despite being among the dominating paradigms in deep learning applications, TL remains poorly understood from a theoretical perspective, with several fundamental questions still open. For instance, (i) how does the source-target similarity affect TL efficiency?

— Statistical mechanics of transfer learning in fully-connected networks in the proportional limit  (2407.07168 - Ingrosso et al., 2024) in Introduction (Section 1)

We do not yet have a strong explanation for the poor performance of transfer with the embedding model. However, the non-trivial performance of model selection indicates that the issue may be with the optimization approach for the target embedding.

— Multi-Task Learning for Sparsely-Labeled Time Series: A Case Study on Cold-Hardiness Modeling  (2609.09062 - Saxena et al., 8 Sep 2026) in Section 6.2, subsection “Transfer Learning”

A possible explanation for the origin of the observed asymmetry for NoPE and ALiBi could be the higher structural complexity of tab_red. If correct, this explanation would entail that transfer learning from a more complex to a less complex task might be favored. In our view, this question deserves further attention.

— Distance generalization in transformers: why bother with positional encoding?  (2609.11913 - Nevermann et al., 10 Sep 2026) in Section 3, subsection “Complexity of transfer learning”

Both benchmarks preserve the evaluation rule; transfer to changed reward criteria or outcomes requiring unobserved information remains untested.

— JET: Judge-Guided Evolution at Test Time for Agent Programs  (2609.34126 - Teng et al., 28 Sep 2026) in Section 6, Conclusion and Limitations