---
title: Contrastive Evidence Policy Optimization (CEPO)
url: https://www.emergentmind.com/topics/contrastive-evidence-policy-optimization-cepo
type: topic
---

# Contrastive Evidence Policy Optimization (CEPO)

Searching arXiv for the CEPO paper and closely related work to ground citations.
Contrastive Evidence Policy Optimization (CEPO) is a reinforcement-learning-with-verifiable-rewards (RLVR) self-distillation method introduced in “CEPO: RLVR Self-Distillation using Contrastive Evidence Policy Optimization” [2605.19436]. It addresses the token-level credit-assignment problem that arises when a correct rollout receives a single sequence-level reward and every token in that rollout is updated identically, regardless of whether the token was a decisive reasoning step or merely grammatical filler. CEPO conditions the model on both a correct answer and a wrong answer, then asks at each token whether the correct answer favors that token while the wrong answer disfavors it. The wrong-answer teacher is constructed from rejected rollouts already present in the batch, so no additional sampling is required. The method is presented as inheriting the structural safety guarantees of prior RLVR self-distillation while sharpening credit at decisive tokens and leaving the improvement to vanish at filler positions; empirically, it reports higher average accuracy than GRPO on five multimodal mathematical reasoning benchmarks at both 2B and 4B scale [2605.19436].

## 1. RLVR setting and the credit-assignment bottleneck

In the RLVR formulation used by CEPO, a policy $\pi_\theta$ receives a problem prompt $x$ and samples a sequence, or chain of thought, $y=(y_1,\dots,y_T)$. A deterministic verifier assigns a binary outcome,
$$
R(x,y)\in\{0,1\},
$$
labeling the rollout correct or wrong [2605.19436].

A practical instantiation discussed in the CEPO paper is Group Relative Policy Optimization (GRPO), which samples $G$ rollouts per prompt and computes sequence-level advantages
$$
A^{(i)}=\frac{R(x,y^{(i)})-\mu_G}{\sigma_G},
$$
where $\mu_G$ and $\sigma_G$ are the mean and standard deviation of the rewards across the $G$ samples. GRPO then applies that same scalar advantage uniformly to every token in a rollout: all tokens in a correct rollout receive $+A^{(i)}$, and all tokens in a wrong rollout receive $-A^{(i)}$ [2605.19436].

The CEPO paper identifies the resulting limitation as uniform credit assignment over long chains. On this account, grammar, framing, and other filler tokens receive the same reinforcement signal as key arithmetic steps or logical inferences. As chains grow, the signal allocated to decisive reasoning steps becomes diluted. The paper characterizes the consequence as reduced sample efficiency and impaired convergence in long-horizon reasoning settings [2605.19436].

## 2. Self-distillation in RLVR and the leakage problem

The CEPO formulation introduces three next-token distributions that share the same parameters $\theta$ but differ in conditioning context:
$$
P_S(y_t)\coloneqq \pi_\theta(y_t\mid x,y_{<t}),
$$
$$
P_T^+(y_t)\coloneqq \pi_\theta(y_t\mid x,r^+,y_{<t}),
$$
$$
P_T^-(y_t)\coloneqq \pi_\theta(y_t\mid x,r^-,y_{<t}),
$$
where $P_S$ is the student, $P_T^+$ is a teacher conditioned on the correct answer $r^+$, and $P_T^-$ is a teacher conditioned on a wrong answer $r^-$ [2605.19436].

The paper contrasts CEPO with on-policy self-distillation methods such as OPSD and SDPO, which minimize
$$
L_{KL}=\mathbb{E}_{\tau\sim \pi_\theta}\left[\sum_t D_{KL}\!\left(P_T^+(\cdot\mid\cdot)\,\|\,P_S(\cdot\mid\cdot)\right)\right].
$$
As described in the paper, the gradient of this objective contains a vocabulary-wide term,
$$
\sum_{v\in V} P_T^+(v)\nabla_\theta \log P_S(v),
$$
and prior work is cited there as showing that this term inevitably leaks the answer into the gradient and degrades generalization. The CEPO paper further states that distribution-matching self-distillation methods, specifically OPSD and SDPO, empirically fall below the untrained baseline, which it interprets as confirmation of the predicted information leakage [2605.19436].

The immediate precursor presented in the CEPO paper is RLSD. Its remedy is to evaluate only the sampled token under stop-gradient, yielding the evidence ratio
$$
w_t^{RLSD}=\exp\!\left(\operatorname{sign}(A)\cdot \operatorname{sg}\!\big(\log P_T^+(y_t)-\log P_S(y_t)\big)\right),
$$
and then reweight the sequence-level advantage according to
$$
A\rightarrow \bar A_t = A\cdot \operatorname{clip}(w_t,1-\epsilon_w,1+\epsilon_w).
$$
In the presentation of the method, this removes the vocabulary-wide sum and anchors the update direction to $\operatorname{sign}(A)$ [2605.19436].

## 3. Contrastive evidence formulation

CEPO sharpens the token-level signal by replacing the one-sided question “does the correct answer favor this token?” with the contrastive question “does the correct answer favor it while the wrong answer disfavors it?” [2605.19436].

Its token-level contrastive quantity is the raw stop-gradiented delta
$$
\Delta_t^{CE}\coloneqq \operatorname{sg}\!\left[\log P_T^+(y_t)-\log P_T^-(y_t)\right].
$$
The paper also gives an equivalent probability-difference form,
$$
\delta_t\coloneqq P_T^+(y_t)-P_T^-(y_t).
$$
These quantities distinguish tokens that are positively supported by the correct-answer teacher and negatively supported by the wrong-answer teacher from tokens for which the two teachers behave similarly [2605.19436].

Given a normalized sequence-level advantage $A$, CEPO defines the contrastive weight
$$
w_t^{CE}=\exp\!\big(\operatorname{sign}(A)\cdot \Delta_t^{CE}\big)
       =\left(\frac{P_T^+(y_t)}{P_T^-(y_t)}\right)^{\operatorname{sign}(A)},
$$
and constructs token-level advantages
$$
\bar A_t
= A\cdot\Big[(1-\lambda)+\lambda\cdot \operatorname{clip}(w_t^{CE},1-\epsilon_w,1+\epsilon_w)\Big],
$$
where $\lambda$ is annealed from $\lambda_0\to 0$ over $T_{\text{warm}}$ steps and $\epsilon_w\in(0,1)$ is a clipping bound [2605.19436].

Because $\Delta_t^{CE}$ is under stop-gradient, the policy-gradient loss is written as
$$
L_{CEPO}(\theta)
=
-\mathbb{E}_{\tau\sim \pi_\theta}
\left[
\sum_{t=1}^T
\bar A_t \nabla_\theta \log \pi_\theta(y_t\mid x,y_{<t})
\right].
$$
The paper’s interpretation is asymmetric in the sign of $A$. For correct rollouts, tokens with large $P_T^+/P_T^-$ receive $\bar A_t\gg A$, whereas filler tokens with $\Delta_t\approx 0$ remain near the original sequence-level weight. For wrong rollouts, the same mechanism sharpens blame on tokens favored by the wrong-answer teacher [2605.19436].

## 4. Wrong-answer teachers, computational overhead, and algorithmic realization

A central operational feature of CEPO is that the wrong-answer teacher is obtained from rejected rollouts already computed within each GRPO group. If $G^-$ denotes the rejected subset, the method chooses $r^-$ as the final answer of the lowest-reward rollout. When all rollouts are correct, the pseudocode specifies the fallback $P_T^-\leftarrow P_S$ [2605.19436].

The paper emphasizes that no extra sampling is needed. One forward pass of $\pi_\theta$ conditioned on $(x,r^-)$ yields $P_T^-(\cdot)$, and together with $P_T^+(\cdot)$ conditioned on the ground-truth answer, this amounts to one extra teacher pass per trajectory, which the paper states is identical overhead to RLSD [2605.19436].

The training loop is correspondingly simple. For each $(x,r^+)$ in the batch, the policy samples $G$ rollouts, computes $A^{(i)}$, selects $r^-$ from the lowest-reward rollout, evaluates $P_T^+$ and $P_T^-$, forms per-token contrastive deltas $\Delta_t$, computes weights $w_t$, converts sequence-level advantages into token-level advantages $\bar A_t$, and finally applies a PPO-clipped surrogate update using those token-level advantages [2605.19436].

This suggests that CEPO is designed as a drop-in modification to existing RLVR pipelines built around grouped rollouts and PPO-style policy optimization, rather than as a separate sampling or verification framework.

## 5. Structural guarantees and discriminative sharpness

The CEPO paper states a theorem labeled “CEPO Safety” for $\lambda\in[0,1]$ and $\epsilon_w\in(0,1)$ with three properties [2605.19436].

First, **direction anchoring** states that
$$
\operatorname{sign}(\bar A_t)=\operatorname{sign}(A)\quad \text{for all } t,
$$
so privileged information cannot flip the update direction. Second, **leakage-free gradients** state that $\nabla_\theta L_{CEPO}$ contains no sum over $v\in V$; the correct and wrong answers enter only as scalars at the sampled token under stop-gradient. Third, **RLSD containment** states that if
$$
P_T^-(y_t)=P_S(y_t)\quad \forall t,
$$
then
$$
\Delta_t^{CE}=\log(P_T^+/P_S),
$$
and CEPO is exactly RLSD [2605.19436].

The paper also gives a proposition labeled “Discriminative Sharpness.” On a correct trajectory with $A>0$,
$$
w_t^{CE}>w_t^{RLSD}\iff P_T^-(y_t)<P_S(y_t).
$$
Accordingly, the paper concludes that CEPO strictly amplifies tokens that the wrong-answer teacher disfavors, identifying these as genuine reasoning steps, while CEPO and RLSD coincide near filler positions where $P_T^-\approx P_S$ [2605.19436].

Within the internal logic of the method, this proposition is the main justification for the term “contrastive evidence.” The additional teacher is not used to introduce a second target distribution for matching; it is used to sharpen token-level discrimination relative to the student baseline while preserving the stop-gradient safety structure.

## 6. Empirical profile and nomenclature

The empirical evaluation in the CEPO paper uses Qwen3-VL-2B-Instruct and Qwen3-VL-4B-Instruct, LoRA-tuned for 50 steps on Geo3k, a set of 3 K geometry problems. Evaluation is conducted on held-out multimodal reasoning benchmarks: DynaMath, LogicVista, MathVision$_{mini}$, MMMU, and WeMath. Baselines under identical budgets are GRPO sequence-level RL, OPSD, SDPO, and RLSD [2605.19436].

| Method | 2B average accuracy | 4B average accuracy |
|---|---:|---:|
| Base | 39.73% | 58.36% |
| GRPO | 41.17% | 57.43% |
| RLSD | 40.05% | 58.51% |
| OPSD | 34.96% | 56.23% |
| SDPO | 35.70% | 55.75% |
| CEPO | 43.43% | 60.56% |

On the averages reported over the five benchmarks, CEPO improves over GRPO by $+2.26$ percentage points at 2B and $+3.13$ percentage points at 4B. The paper further states that OPSD and SDPO collapse below the base model, whereas CEPO yields the best sample efficiency across scales and tasks while sharply localizing credit on decisive reasoning tokens and preserving the safety guarantees of prior RLVR self-distillation. The code release is listed at `https://github.com/ahmedheakl/CEPO` [2605.19436].

The acronym “CEPO” is not unique in the literature. An earlier, unrelated paper, “Policy Optimization with Sparse Global Contrastive Explanations” [2207.06269], uses CEPO to denote a constrained-MDP framework in which a target policy $\pi'$ maximizes expected return subject to a budget on the expected number of deviations from a behavior policy $\pi$. In that setting, explanation length is the number of diverging states or regions included in a global contrastive explanation, and the optimization is posed with a hard cap $K$ on
$$
\mathbb{E}_{\pi'}\!\left[\sum_{t=0}^\infty C(s_t,a_t)\right].
$$
That 2022 usage concerns sparse policy modification and interpretable contrastive explanations in discrete and continuous MDPs, whereas the 2026 usage concerns RLVR self-distillation and token-level contrastive credit assignment [2207.06269]. A common misconception is therefore to treat the acronym as denoting a single method family; in the arXiv record, it refers to two distinct formulations that share only the abbreviation.

Source: https://www.emergentmind.com/topics/contrastive-evidence-policy-optimization-cepo