Mechanism of knowledge generalization from task–insight training
Determine how knowledge learned from task–insight training is generalized through backpropagation to semantically similar tasks and different output formats, including full rollout generation.
References
The generalization dynamics that drive this transfer are currently unknown.
The generalization dynamics that drive this transfer are currently unknown. We attribute the effectiveness of (task, insight) training to smoothness during the backpropagation.
This gives rise to multiple next questions: First, how exactly is the knowledge generalized through the backpropagation? We hypothesize this has to do with the smoothness of the parameters of a (sufficiently pretrained) model, implicitly routing the knowledge not just naively to the literal task and literal format of task \rightarrow insight, but to any semantically similar task and output format, including generating a full rollout. Second, where are the limits of this paradigm? Tool-calling might be special in its hard to find but easy to apply insights. We expect that training only on compressed insights is not feasible in all domains, especially in domains where the policy possesses too little pretraining capabilities for the smoothness to emerge, or domains where tasks are so specific that strong enough insights are not applicable to similar problems. Third, which other forms of training become possible if we remove the need for full rollouts?