Determine the reliability of JET’s diagnostic-feedback gains

Determine whether the performance gains attributed to diagnostic feedback in Judge-Guided Evolution at Test Time (JET) persist under broader campaign replication, given that limited replication leaves the observed diagnostic-feedback differences uncertain.

Background

The primary experiments use five independent campaigns per condition. The paper reports gains from the evolved judge and analyzes textual diagnostics as an additional source of guidance for target-side program rewrites.

The conclusion cautions that limited campaign replication makes smaller differences, including diagnostic-feedback gains, uncertain. Additional independent campaigns are therefore needed to establish whether the diagnostic-feedback effect is robust rather than an artifact of limited replication.

References

Adapted programs attain higher observed mean rewards, including on unseen target tasks, but limited campaign replication leaves smaller differences, including diagnostic-feedback gains, uncertain.

— JET: Judge-Guided Evolution at Test Time for Agent Programs  (2609.34126 - Teng et al., 28 Sep 2026) in Section 6, Conclusion and Limitations