Validate the human-frequency proxy for group frequency

Validate whether the percentage of people predicted by the Bayesian Truth Serum prompt is a reliable proxy for the realized frequency of answers within the sampled GRPO response group.

Background

The implementation asks the LLM to predict what percentage of people on Earth would answer “true,” while the theoretical Bayesian Truth Serum mechanism scores predictions against the answer frequency within the sampled model group. The paper treats human frequency as a proxy for group frequency because the model has no natural peer population to name, but this substitution is not empirically verified.

References

The prediction report asks for a human frequency and is scored against the group's, which the results suggest is a workable proxy but which we do not verify.

Mitigating LLM sycophancy with RL-based fine-tuning: Bayesian Truth Serum approach  (2608.25267 - Mytsyk et al., 26 Aug 2026) in Section 3.1, paragraph “TF Dataset: Fine-tuning and Evaluation”; Section 6, paragraph “Directions”