Generalization of ACL findings across models and curricula

Determine whether the forgetting patterns and trade-offs observed in ACLArena generalize across different model families and training orders.

Background

The experiments use Qwen3-8B-Base and a fixed four-stage curriculum consisting of mathematical reasoning, search, e-commerce, and instruction following. Although this controlled design enables systematic comparisons among sequential training, MMOPD, SDFT, model merging, and MLE, it leaves unresolved whether the observed capability forgetting and method trade-offs are specific to this model and curriculum or recur more broadly.

The authors identify extension to additional agent environments, longer training sequences, and alternative curricula as necessary for distinguishing general ACL phenomena from effects caused by the present experimental setup.

References

This controlled setting supports systematic comparisons, but the extent to which the observed forgetting patterns and method trade-offs generalize across model families and training orders remains unclear.

— ACLArena: Agent Continue Learning in Multi-stage Post-training  (2609.23989 - Wang et al., 21 Sep 2026) in Section “Limitations & Future Work”