Robust Multi-Source Behavior Aggregation and Budget-Aware On-Policy Fusion

Develop robust methods for aggregating behaviors from multiple source models, reducing the query cost of on-policy supervision, and jointly incorporating process-level and outcome-level feedback in behavior-level model fusion.

Background

Behavior-level fusion transfers source-model capabilities through output distributions, demonstrations, preferences, critiques, rewards, or verifier feedback. On-policy variants obtain supervision on states generated by the target model, which can improve coverage but increases source-query and computation costs. The survey identifies robust multi-source aggregation, budget-aware on-policy querying, and the combination of process- and outcome-level feedback as open problems.

References

Open problems include robust multi-source aggregation, budget-aware on-policy queries, and joint process- and outcome-level feedback.

From Parameters to Behaviors: A Survey of Model Fusion for Large Language Models  (2609.19553 - Cai et al., 17 Sep 2026) in Section 3.3, Behavior-Level Fusion