Generalization beyond a single model and isolated harness training

Evaluate whether LEGO-RL generalizes to model architectures other than Qwen3.5-35B-A3B and to training a single policy across multiple coding-agent harnesses.

Background

The evaluation trains Qwen3.5-35B-A3B separately with each of three coding-agent harnesses—OpenHands SDK, Claude Code, and OpenCode. Consequently, the reported gains do not establish whether LEGO-RL transfers across model architectures or whether one policy can be trained effectively under multiple harness control flows. The authors explicitly identify both forms of generalization as unevaluated, and later describe mixed-harness training as ongoing work.

References

First, all experiments use Qwen3.5-35B-A3B, and each coding-agent harness is trained separately, so generalization to other model architectures and mixed-harness training remains to be evaluated.

LEGO-RL: Harness-Native Reinforcement Learning for Coding Agents  (2608.17393 - Du et al., 18 Aug 2026) in Limitations and Future Work, p. 17