Annealed Expert-Bonus Reward Shaping
- The paper introduces a mechanism that adds a linearly decaying bonus to expert-conditioned trajectories, stabilizing early policy optimization.
- It employs Implicit Expert Forcing (IEF) with ERRS gating to ensure only flawless expert trajectories receive the bonus, facilitating efficient exploration.
- Empirical evaluations show improved accuracy, faster convergence, and greater training stability compared to standard GRPO methods.
Annealed Expert-Bonus Reward Shaping is a temporal reward augmentation mechanism utilized within the In-Context Steered Policy Optimization (ICPO) framework for Reinforcement Learning from Verifiable Rewards (RLVR) in large reasoning models. This method targets the stabilization of policy optimization by judiciously leveraging expert-generated trajectories early in training through an additional, progressively diminishing reward bonus that fosters efficient exploration before yielding autonomy to the agent’s own reward dynamics (Huang et al., 30 Oct 2025).
1. Shaped Reward Definition and Application Criteria
Within ICPO, the principal reward signal, , is binary: a value of $1$ is assigned if the final answer in trajectory is correct, and $0$ otherwise. The Annealed Expert-Bonus Reward Shaping introduces an additive bonus on top of under specific gating conditions. The bonus is exclusively conferred when a trajectory:
- Is generated via expert-conditioning with Implicit Expert Forcing (IEF),
- Passes the Expert-Region Reject Sampling (ERRS) filter, and
- Achieves perfect verifiable correctness, i.e., .
Formally, for training step , given decay factor , fixed bonus weight , and the indicator function over eligible expert regions $1$0:
$1$1
This directly enforces that only expert-accepted, high-reward trajectories contribute to the bonus component.
2. Linear Annealing Schedule
To ensure the expert bonus does not dominate learning indefinitely, the contribution is linearly annealed over the total training steps. Letting the annealing schedule be
$1$2
where $1$3 is the current optimization step count and $1$4 is the total number of RL update steps (typically $1$5 in the reported experiments). With $1$6 as default, the effective bonus weight at step $1$7 is $1$8. The annealed shaping mechanism ensures maximal expert steering at initialization, diminishing linearly to zero at the conclusion of training. This provides a principled transition from heavy reliance on imitation to autonomous reward optimization.
3. Functional Rationale
Reward shaping in this form biases the Q-values toward expert-preferred actions early, promoting coverage of high-quality regions in the state-action space. When $1$9 is large (initial stages), the gradient of the policy objective is dominated by the expert-shaping term, rapidly directing the policy towards modes where expert rollouts succeed. As annealing progresses and 0, the influence of the expert term wanes, asymptotically recovering pure RL over 1. Thus, the approach:
- Reduces policy variance and accelerates the discovery of high-reward strategies during the noisy initial regime,
- Prevents over-imitation in later phases by ceding optimization to the agent’s verifiable reward signal.
This delineation into an "imitation-primed" phase followed by a "fine-tuning-autonomous" phase is essential for both efficient convergence and generalization.
4. Algorithmic Integration
The implementation of Annealed Expert-Bonus Reward Shaping within the ICPO update cycle is minimal yet precise. After generating both on-policy and IEF (expert-conditioned) rollouts for each prompt:
- When an IEF trajectory passes the reward and ERRS gate (2 and 3, with default 4), a randomly chosen on-policy trajectory is replaced by the successful IEF trace,
- Its reward is overwritten as 5,
- Mixed-policy advantages and the surrogate loss are computed using all (possibly shaped) rewards,
- Policy parameters are updated accordingly.
The only deviation from vanilla mixed-policy GRPO is this explicit overwrite of the reward for expert-accepted trajectories:
$0$9
This guarantees that shaping is local, targeted, and decaying, rather than a pervasive modification across all trajectories.
5. Empirical Evaluation and Ablation Findings
Ablation studies and benchmarks across Qwen3-1.7B and Qwen3-8B models substantiate that annealed expert-bonus shaping confers tangible gains:
- On a six-task in-distribution benchmark:
- GRPO baseline: 48.37
- ICPO without reward shaping: 52.54 (+4.17)
- ICPO†(with annealed bonus): 51.35 (+2.98)
- For the MATH-500 expert domain (Qwen3-1.7B):
- GRPO: 83.60 → ICPO†: 87.20 (+3.60)
- On Qwen3-8B (Table 3) average performance:
- Full ICPO†: 66.51
- –Reward Shaping: 65.78
- –ERRS: 64.99
- No IEF (vanilla GRPO): 63.76
Learning-curve analyses (Figure 1) illustrate that ICPO†(annealed shaping) achieves earlier, higher, and smoother validation rewards with reduced variance between steps. This suggests improved policy stability and more reliable early convergence when integrating Annealed Expert-Bonus Reward Shaping.
6. Practical Parameterization
Guidelines for operationalizing Annealed Expert-Bonus Reward Shaping include:
- Initial expert-bonus weight 6: 1.0 is standard; set 7 lower (e.g., 0.5 or 0.2) for noisier expert data.
- Total annealing horizon 8: Align with the expected number of RL update steps (e.g., 300–500). For extended training, consider slower or non-linear decay schedules (e.g., 9).
- Expert-region threshold $0$0: With strictly binary reward, $0$1 ensures only flawless expert trajectories receive the bonus; in settings with graded rewards, set $0$2 higher (e.g., 90th percentile).
- Off-policy ratio: A value of $0$3 is typical, but increase $0$4 for trusted experts, monitoring KL divergence.
- Few-shot demonstration count $0$5 (for IEF): $0$6 provides a minimal expert prior; higher $0$7 (2 or 3) offers stronger prior at increased computational expense.
The interdependence between shaping hyperparameters and optimization trajectory underscores the need for experimental calibration in novel domains.
7. Summary and Significance
Annealed Expert-Bonus Reward Shaping in ICPO constitutes a minimal additive reward, $0$8, restricted to expert-conditioned, high-reward trajectories and linearly annealed throughout training. This mechanism amplifies the effect of expert demonstration early while phasing out reliance as the model matures, systematically mediating between imitation learning and autonomous reinforcement learning. Empirical results confirm its benefit for improved accuracy and training stability in mathematical reasoning tasks, supporting its adoption as a scalable, general-purpose tool in RLVR for large language and reasoning models (Huang et al., 30 Oct 2025).