Verifier hacking under extended training in Trade-R1

Determine whether extending the training duration of the Trade-R1 policy model enables the learned policy to discover strategies that bypass the Retrieval-Augmented Verification protocol with the Triangular Consistency metric (i.e., verifier hacking).

Background

Trade-R1 introduces a Retrieval-Augmented Verification protocol with a Triangular Consistency metric to gate stochastic market rewards by checking pairwise alignment among retrieved evidence, the model’s reasoning chain, and the final decision. The training in reported experiments was stopped at a predefined step due to computational budget constraints rather than full convergence.

The authors caution that longer training might allow the model to learn to circumvent the verification protocol—a potential failure mode they term "verifier hacking." Establishing whether this occurs under extended training is thus an explicit unresolved question.

References

Whether longer training might enable the model to discover subtle strategies to bypass the verification protocol (i.e., "verifier hacking") remains an open question.

Trade-R1: Bridging Verifiable Rewards to Stochastic Environments via Process-Level Reasoning Verification  (2601.03948 - Sun et al., 7 Jan 2026) in Limitations, Section 5 (Conclusion)

Whether a strong verifier inverts at pressures beyond N{=}4096 is open; what is already established is that its compute is increasingly wasted relative to a sound verifier---which is the economic argument for settlement either way.

Proof-Carrying Cognition: Closing the Verification Gap with Reality-Settled Reward  (2609.09776 - M et al., 9 Sep 2026) in Section 4, “What failed—and how we repair it”

A shared benchmark that holds the candidates fixed and scores the exploitability of each verifier has yet to appear.

No Free Checker: A Survey of Verifiers for Robot Policies  (2609.09250 - Wan et al., 8 Sep 2026) in Section 7.3, subsection "Measuring a Verifier Under Optimization"