Establish whether BTS reward ordering decreases sycophancy
Establish whether the Group Relative Policy Optimization update driven by the Bayesian Truth Serum reward decreases the model’s sycophancy rate \(\lambda\), rather than merely assigning honest responses a higher expected reward under fixed response profiles.
References
The theorem delivers an ordering of expected rewards under fixed profiles. It does not follow that the GRPO update decreases $\lambda$, and we do not prove that it does. The step from a reward ordering to a claim about where the optimization converges belongs to the optimization rather than to the mechanism.
Rewards that pay for agreement inside the sampled group, such as the majority-vote reward of test-time reinforcement learning, have been used for reasoning accuracy but not for sycophancy; whether they reduce sycophancy can be tested, but it lies outside the scope of this paper.
We conjecture that the premium a mechanism pays for a rarer answer drives the reduction in sycophancy, not the focality of the truthful equilibrium.
It also remains to test whether values of $G$ smaller than $64$ reduce sycophancy as much as $G = 64$.