Mechanism of knowledge generalization from task–insight training

Determine how knowledge learned from task–insight training is generalized through backpropagation to semantically similar tasks and different output formats, including full rollout generation.

Background

The proposed method trains the model to predict compact, self-generated insights from the task description alone, while evaluating the model without explicitly providing those insights. The authors observe that this training can improve performance on full task rollouts and hypothesize that parameter smoothness and semantic similarity enable transfer across tasks and output formats.

The paper does not establish the mechanism responsible for this transfer. Understanding it is important for determining when task–insight training will reliably generalize beyond the literal task and insight representation.

References

The generalization dynamics that drive this transfer are currently unknown.

— RLTL;DR: Self-improvement by Internalizing Self-generated Feedback  (2609.37633 - Kirchhof et al., 29 Sep 2026) in Section 6.1, “The surprising learning only from (task, insight) tuples”; Conclusion and outlook

The generalization dynamics that drive this transfer are currently unknown. We attribute the effectiveness of (task, insight) training to smoothness during the backpropagation.

— RLTL;DR: Self-improvement by Internalizing Self-generated Feedback  (2609.37633 - Kirchhof et al., 29 Sep 2026) in Section 6.2, “Limits of (task, insight) learning”; Conclusion and outlook

This gives rise to multiple next questions: First, how exactly is the knowledge generalized through the backpropagation? We hypothesize this has to do with the smoothness of the parameters of a (sufficiently pretrained) model, implicitly routing the knowledge not just naively to the literal task and literal format of task \rightarrow insight, but to any semantically similar task and output format, including generating a full rollout. Second, where are the limits of this paradigm? Tool-calling might be special in its hard to find but easy to apply insights. We expect that training only on compressed insights is not feasible in all domains, especially in domains where the policy possesses too little pretraining capabilities for the smoothness to emerge, or domains where tasks are so specific that strong enough insights are not applicable to similar problems. Third, which other forms of training become possible if we remove the need for full rollouts?

— RLTL;DR: Self-improvement by Internalizing Self-generated Feedback  (2609.37633 - Kirchhof et al., 29 Sep 2026) in Conclusion and outlook