Generalization beyond a single model and isolated harness training
Evaluate whether LEGO-RL generalizes to model architectures other than Qwen3.5-35B-A3B and to training a single policy across multiple coding-agent harnesses.
References
First, all experiments use Qwen3.5-35B-A3B, and each coding-agent harness is trained separately, so generalization to other model architectures and mixed-harness training remains to be evaluated.
— LEGO-RL: Harness-Native Reinforcement Learning for Coding Agents
(2608.17393 - Du et al., 18 Aug 2026) in Limitations and Future Work, p. 17