Persistence and scaling of SMITH’s performance gains

Determine whether the performance gains produced by SMITH persist, saturate, or invert when the policy and judge are scaled to models with 70B or more parameters.

Background

SMITH’s experiments train policies with at most 8B parameters and use a quality judge with 30B activated parameters. The authors therefore do not establish how the framework behaves at substantially larger scales, particularly when both the trained policy and the judging model become much larger.

The paper’s Self-Judge ablation provides only a partial investigation of judge-size dependence: removing the external judge improves held-out Reasoning-Gym performance but degrades out-of-domain GQA performance. Consequently, the scaling behavior of SMITH’s gains at the 70B-plus scale remains unresolved, including whether larger-scale training will preserve the gains, lead to saturation, or cause performance degradation.

References

All trained policies are at most 8B parameters and the quality judge has 30B activated parameters, a regime chosen for compute feasibility; whether the gains persist, saturate, or invert at the 70B+ scale is an open question, and our Self-Judge ablation (Table~\ref{tab:ablation_8b}) is only a partial probe of judge-size dependence, since dropping the external judge improves held-out RG but degrades OOD GQA.

Joint Optimization of Tool Creation and Use for Large Language Model Agents  (2608.24571 - Tam et al., 25 Aug 2026) in Section 5, Limitations and discussions, paragraph “Scale, judge dependence, and base priors”