Determine the reliability of JET’s diagnostic-feedback gains
Determine whether the performance gains attributed to diagnostic feedback in Judge-Guided Evolution at Test Time (JET) persist under broader campaign replication, given that limited replication leaves the observed diagnostic-feedback differences uncertain.
References
Adapted programs attain higher observed mean rewards, including on unseen target tasks, but limited campaign replication leaves smaller differences, including diagnostic-feedback gains, uncertain.
— JET: Judge-Guided Evolution at Test Time for Agent Programs
(2609.34126 - Teng et al., 28 Sep 2026) in Section 6, Conclusion and Limitations