On-policy rollout consistency for numerical reasoning

Determine whether on-policy rollouts produce more consistent numerical reasoning patterns within rollout groups and thereby benefit tasks requiring fine-grained numerical precision.

Background

The paper reports that GRPO can be competitive with or superior to Echo-GRPO on VSI-Bench, a benchmark involving exact numerical estimations such as object size and distance. The authors attribute this result to a hypothesized advantage of on-policy rollouts: reasoning patterns within each rollout group may be more consistent, which could support the fine-grained numerical precision required by the task.

This explanation is explicitly presented as a conjecture rather than an established result. The unresolved issue is therefore whether the consistency of on-policy numerical reasoning patterns is in fact responsible for the observed performance advantage on precision-sensitive tasks.

References

We conjecture that on-policy rollouts produce more consistent numerical reasoning patterns within rollout groups, which is beneficial for tasks requiring fine-grained numerical precision.

Reason in the Words You Speak: Idiolectal Paraphrasing Off-Policy Traces for Reasoning Distillation in VideoLLMs  (2608.26684 - Lee et al., 27 Aug 2026) in Section 4, subsection “Main results”