Papers
Topics
Authors
Recent
Search
2000 character limit reached

Self-rewarding correction for mathematical reasoning

Published 26 Feb 2025 in cs.AI and cs.LG | (2502.19613v1)

Abstract: We study self-rewarding reasoning LLMs, which can simultaneously generate step-by-step reasoning and evaluate the correctness of their outputs during the inference time-without external feedback. This integrated approach allows a single model to independently guide its reasoning process, offering computational advantages for model deployment. We particularly focus on the representative task of self-correction, where models autonomously detect errors in their responses, revise outputs, and decide when to terminate iterative refinement loops. To enable this, we propose a two-staged algorithmic framework for constructing self-rewarding reasoning models using only self-generated data. In the first stage, we employ sequential rejection sampling to synthesize long chain-of-thought trajectories that incorporate both self-rewarding and self-correction mechanisms. Fine-tuning models on these curated data allows them to learn the patterns of self-rewarding and self-correction. In the second stage, we further enhance the models' ability to assess response accuracy and refine outputs through reinforcement learning with rule-based signals. Experiments with Llama-3 and Qwen-2.5 demonstrate that our approach surpasses intrinsic self-correction capabilities and achieves performance comparable to systems that rely on external reward models.

Summary

  • The paper proposes a self-rewarding reasoning framework integrating LLM generation and evaluation for autonomous mathematical self-correction.
  • A two-stage algorithm combines self-rewarding instruction-following fine-tuning and reinforcement learning, using only self-generated data.
  • Empirical results demonstrate the framework significantly outperforms intrinsic self-correction, improving accuracy on mathematical reasoning benchmarks.

The paper introduces a novel self-rewarding reasoning framework for LLMs that integrates generation and evaluation into a single model. This framework aims to enhance the self-correction capabilities of LLMs in mathematical reasoning tasks, reducing computational overhead compared to approaches relying on external reward models.

The key contributions of the paper are:

  • A self-rewarding reasoning framework integrating the generator and reward model into a single LLM, enabling autonomous reasoning, evaluation, and correction.
  • A two-stage algorithmic framework for self-correction in mathematical reasoning, relying only on self-generated data. The first stage uses sequential rejection sampling to construct long chain-of-thought (CoT) trajectories encoding self-rewarding and self-correction behaviors. The second stage enhances these behaviors through reinforcement learning with rule-based signals.
  • Empirical validation demonstrating that self-rewarding correction significantly outperforms intrinsic self-correction.

The self-rewarding reasoning process is formulated as a multi-turn Markov Decision Process (MDP). An LLM generates an initial reasoning attempt a1∼π1(⋅∣s1)a_1 \sim \pi_1(\cdot | s_1) given a prompt s1=x∈Xs_1 = x \in \mathcal{X} from a distribution D0\mathcal{D}_0, where π\pi is the LLM. It then self-rewards its response by generating an evaluation y1∼π1(⋅∣s1,a1)y_1 \sim \pi_1(\cdot | s_1, a_1). If the model assesses its answer as correct (y1y_1 = [VERIFY] correct), the generation stops. Otherwise, the LLM generates a refined response and evaluation (a2,y2)∼π2(⋅∣s2)(a_2, y_2) \sim \pi_2(\cdot | s_2), conditioned on the updated state s2=(s1,a1,y1)s_2 = (s_1, a_1, y_1). The self-refinement continues until the model produces a self-evaluation yhy_h assessing the answer as correct.

y1∼π1(⋅∣s1,a1)y_1 \sim \pi_1(\cdot | s_1, a_1)

s1=x∈Xs_1 = x \in \mathcal{X}0

where:

  • s1=x∈Xs_1 = x \in \mathcal{X}1 is the initial prompt.
  • s1=x∈Xs_1 = x \in \mathcal{X}2 is the initial reasoning attempt.
  • s1=x∈Xs_1 = x \in \mathcal{X}3 is the LLM.
  • s1=x∈Xs_1 = x \in \mathcal{X}4 is the self-rewarding evaluation.

The two-stage training framework consists of:

  1. Self-rewarding instruction-following fine-tuning (IFT): An initial LLM s1=x∈Xs_1 = x \in \mathcal{X}5 is fine-tuned using demonstration data collected via sequential rejection sampling, resulting in an improved model s1=x∈Xs_1 = x \in \mathcal{X}6 integrating self-rewarding reasoning abilities.
  2. Reinforcement learning (RL) optimization: s1=x∈Xs_1 = x \in \mathcal{X}7 is further refined using RL, leveraging it as the reference model. This stage enhances the model's ability to assess correctness and refine responses.

The self-rewarding signal is trained by token prediction, where models include reasoning in their evaluations and output specific tokens to indicate their evaluation results, such as "[VERIFY] correct" and "[VERIFY] wrong". Data collection uses a rejection sampling approach, generating self-correction trajectories and preserving desired ones. The process includes generating initial reasoning responses, sampling self-rewarding signals, and correction sampling. The LLMs are fine-tuned using a standard SFT pipeline to maximize:

s1=x∈Xs_1 = x \in \mathcal{X}8

where:

  • s1=x∈Xs_1 = x \in \mathcal{X}9 is the initial prompt.
  • D0\mathcal{D}_00 is the initial reasoning attempt.
  • D0\mathcal{D}_01 is the self-rewarding evaluation of the first turn.
  • D0\mathcal{D}_02 is the revised reasoning attempt.

For the RL stage, the paper considers both deep RL methods and direct alignment algorithms. A trajectory-wise reward function D0\mathcal{D}_03 is used for trajectory D0\mathcal{D}_04, where D0\mathcal{D}_05 is the horizon. The oracle reward D0\mathcal{D}_06 is used, where D0\mathcal{D}_07 is the ground-truth verifier. The KL-regularized objective is:

D0\mathcal{D}_08

where:

  • D0\mathcal{D}_09 is the policy being optimized.
  • Ï€\pi0 is the initial LLM.
  • Ï€\pi1 is the reference model.
  • Ï€\pi2 is the trajectory-wise reward.
  • Ï€\pi3 is the Kullback-Leibler divergence.
  • Ï€\pi4 is a regularization coefficient.

The paper also adopts Direct Preference Optimization (DPO) to solve the equation, using the multi-turn DPO (M-DPO) framework. The loss function π\pi5 is:

Ï€\pi6

where:

  • Ï€\pi7 is the winning trajectory.
  • Ï€\pi8 is the losing trajectory.
  • Ï€\pi9 is the policy being optimized.
  • y1∼π1(⋅∣s1,a1)y_1 \sim \pi_1(\cdot | s_1, a_1)0 is the reference policy.
  • y1∼π1(⋅∣s1,a1)y_1 \sim \pi_1(\cdot | s_1, a_1)1 is the sigmoid function.
  • y1∼π1(⋅∣s1,a1)y_1 \sim \pi_1(\cdot | s_1, a_1)2 is a regularization coefficient.

The models are evaluated on mathematical reasoning abilities using benchmarks including MATH500, OlympiadBench, and Minerva Math. Evaluation metrics include turn 1 accuracy, final accuracy, improvement in accuracy from the first attempt to the final answer (y1∼π1(⋅∣s1,a1)y_1 \sim \pi_1(\cdot | s_1, a_1)3), fraction of problems changed from incorrect to correct (y1∼π1(⋅∣s1,a1)y_1 \sim \pi_1(\cdot | s_1, a_1)4), and fraction of problems changed from correct to incorrect (y1∼π1(⋅∣s1,a1)y_1 \sim \pi_1(\cdot | s_1, a_1)5).

The main results demonstrate that intrinsic self-correction with prompting generally fails, while self-rewarding reasoning models significantly outperform existing baselines. For instance, on MATH500, self-rewarding IFT achieves y1∼π1(⋅∣s1,a1)y_1 \sim \pi_1(\cdot | s_1, a_1)6 = 5.0% and y1∼π1(⋅∣s1,a1)y_1 \sim \pi_1(\cdot | s_1, a_1)7 = 0.4%. Self-rewarding reasoning models also improve final accuracy compared to single-turn baselines. The paper finds that deep RL algorithms outperform direct alignment algorithms.

Further experiments were performed using a simplified two-turn conversation framework and Llama models, confirming the generality of the proposed framework. Ablation studies on data distribution show that the data composition in self-rewarding IFT influences the outcome supervised reward model (ORM) accuracy. The paper also investigates additional rule designs in RL training.

The paper concludes by highlighting the effectiveness of the self-rewarding reasoning framework in enhancing self-correction capabilities and computational efficiency. Future research directions include addressing the lower reward model accuracy compared to external ORMs, incorporating multi-turn RL methods, and extending the framework to step-wise correction.

Paper to Video (Beta)

No one has generated a video about this paper yet.

Whiteboard

No one has generated a whiteboard explanation for this paper yet.

Open Problems

We haven't generated a list of open problems mentioned in this paper yet.

Tweets

Sign up for free to view the 6 tweets with 3 likes about this paper.