Verifier hacking under extended training in Trade-R1
Determine whether extending the training duration of the Trade-R1 policy model enables the learned policy to discover strategies that bypass the Retrieval-Augmented Verification protocol with the Triangular Consistency metric (i.e., verifier hacking).
References
Whether longer training might enable the model to discover subtle strategies to bypass the verification protocol (i.e., "verifier hacking") remains an open question.
Whether a strong verifier inverts at pressures beyond N{=}4096 is open; what is already established is that its compute is increasingly wasted relative to a sound verifier---which is the economic argument for settlement either way.
A shared benchmark that holds the candidates fixed and scores the exploitability of each verifier has yet to appear.