Persistence and scaling of SMITH’s performance gains
Determine whether the performance gains produced by SMITH persist, saturate, or invert when the policy and judge are scaled to models with 70B or more parameters.
References
All trained policies are at most 8B parameters and the quality judge has 30B activated parameters, a regime chosen for compute feasibility; whether the gains persist, saturate, or invert at the 70B+ scale is an open question, and our Self-Judge ablation (Table~\ref{tab:ablation_8b}) is only a partial probe of judge-size dependence, since dropping the external judge improves held-out RG but degrades OOD GQA.
— Joint Optimization of Tool Creation and Use for Large Language Model Agents
(2608.24571 - Tam et al., 25 Aug 2026) in Section 5, Limitations and discussions, paragraph “Scale, judge dependence, and base priors”