Training-time effects of the EvoResearcher reward decomposition
Determine whether training the EvoResearcher reward decomposition through the proposed GRPO objective causes the training-time version to inherit, amplify, or invert the effects observed for the inference-time protocol.
References
Whether the training-time version inherits, amplifies, or inverts the inference-time effects is an open empirical question that the protocol makes substantially cheaper to investigate, since the same reward decomposition drives both.
— Training-Free Inference-Time Self-Reflection and Cost-Bounded Early Stopping for Large Language Models
(2608.18884 - Yu et al., 19 Aug 2026) in Section 6, Discussion, subsection “Implications”