---
title: Curiosity-Driven RLHF Advances
url: https://www.emergentmind.com/topics/curiosity-driven-rlhf-cd-rlhf
type: topic
---

# Curiosity-Driven RLHF Advances

Searching arXiv for the specified CD-RLHF papers and closely related work to ground the article.
Curiosity-Driven Reinforcement Learning from Human Feedback (CD-RLHF) denotes RLHF variants that supplement preference-derived extrinsic rewards with intrinsic rewards that favor exploration. Current arXiv usage suggests that the label covers at least two distinct 2025 formulations: a token-level curiosity mechanism for improving output diversity in summarization and instruction-following, introduced in "Curiosity-Driven Reinforcement Learning from Human Feedback" [2501.11463], and a belief-based curiosity mechanism for eliciting latent user traits in personalized multi-turn dialogue, developed in "Enhancing Personalized Multi-Turn Dialogue with Curiosity Reward" [2504.03206]. In both cases, the central idea is to preserve the optimization structure of RLHF while adding a shaped signal that rewards informative or novel trajectories.

## 1. Scope and conceptual variants

CD-RLHF emerged from two different problem settings. In one setting, standard RLHF is described as improving alignment to human preferences while reducing output diversity relative to supervised fine-tuning, producing repetitive or homogenized outputs; CD-RLHF addresses this by adding token-level curiosity for novel states during generation [2501.11463]. In the other setting, multi-turn RLHF is described as prioritizing helpfulness and safety but falling short in empathetic, adaptive, and personalized interaction; CD-RLHF addresses this by rewarding reductions in uncertainty about latent user traits during dialogue [2504.03206].

The two formulations differ in what constitutes a “state novelty” signal and in what exploration is intended to achieve. The diversity-oriented formulation uses prediction error in a learned latent state space and applies it to tokens outside the top-$k$ by probability. The personalized-dialogue formulation uses a belief over hidden user types and rewards actions that improve the agent’s user model [2501.11463][2504.03206].

| Formulation | Intrinsic signal | Primary aim |
|---|---|---|
| Diversity-oriented CD-RLHF | Prediction error in an Intrinsic Curiosity Module | Optimize both output diversity and alignment quality |
| Personalized dialogue CD-RLHF | Changes in belief over latent user type | Reveal user preferences and adapt to them |

This suggests that CD-RLHF is better understood as a family of RLHF augmentations rather than a single algorithm. What is common across variants is the addition of intrinsic motivation to a standard RLHF backbone; what differs is the latent variable being explored, the reward granularity, and the failure modes.

## 2. Integration with RLHF objectives

Both formulations preserve the standard separation between an extrinsic alignment signal and an auxiliary intrinsic signal. In the personalized multi-turn setting, the combined objective is reported as

$$
J(\theta)=\mathbb{E}_{\tau\sim\pi_\theta}\Bigg[\sum_{t=0}^{T}\big(r_{\mathrm{RLHF}}(s_t,a_t)+\lambda\, r_{\mathrm{cur}}(s_t,a_t)\big)\Bigg],
$$

with turn-level curiosity rewards added to the RLHF reward and optimized using the multi-turn RLHF training recipe of Shani et al. (2024); for Education Dialogue, the paper reports $\lambda=9.0$ [2504.03206].

In the diversity-oriented formulation, the combined reward is written at token level as

$$
r_t = r^{(e)}_t + \eta\, r^{(i)}_t,
$$

where the extrinsic reward combines a sequence-level reward model score with a KL penalty to the SFT reference model,

$$
r^{(e)} = R - \beta D_{\mathrm{KL}}\big(\pi_{\mathrm{policy}}(\cdot|s_t)\,\|\,\pi_{\mathrm{ref}}(\cdot|s_t)\big),
$$

and the policy is updated with PPO using generalized advantage estimation [2501.11463].

The personalized-dialogue paper gives a stronger theoretical account through a POMDP over observable dialogue histories $s_t$ and latent user type $u\in\mathcal{U}$. With extended state $s_t'=\langle s_t,u\rangle$, transition and reward are

$$
\mathcal{T}'(s'_{t+1}\mid s'_t,a_t)=\mathcal{T}(s_{t+1}\mid s_t,a_t,u), \quad
\mathcal{R}'(s'_t,a_t)=\mathcal{R}(s_t,a_t\mid u),
$$

and the belief update is

$$
b_{t+1}(u)\propto \mathcal{T}(s_{t+1}\mid s_t,a_t,u)\, b_t(u).
$$

Using a belief-based expected reward,

$$
\mathcal{R}^b(s_t,b_t,a_t)=\sum_u b_t(u)\,\mathcal{R}(s_t,a_t\mid u),
$$

the paper invokes potential-based reward shaping over beliefs,

$$
r^b(s_t,b_t,a_t)=\mathcal{R}^b(s_t,b_t,a_t)+\gamma\,\phi(b_{t+1})-\phi(b_t),
$$

and states, via Eck et al. (2013), that such shaping leaves the optimal policy invariant [2504.03206].

By contrast, the diversity-oriented paper is framed operationally rather than through PBRS. It retains the conventional SFT $\rightarrow$ RM $\rightarrow$ PPO pipeline, initializes policy and reference models via SFT on the same preference dataset, trains the reward model on paired preferences, and inserts an Intrinsic Curiosity Module during PPO rollouts [2501.11463].

## 3. Belief-based CD-RLHF for personalized multi-turn dialogue

The personalized formulation targets latent user inference during online interaction. The agent does not assume pre-existing user profiles or histories; instead, it must infer the hidden user type from observed dialogue. Two environments are reported. In Education Dialogue, the latent trait is the student’s preferred learning style, described in prompts and analyses as “hands-on” versus “story-telling,” and earlier as “lecture-based” versus “hands-on.” In Exercise Recommendation, the user profile has 20 attributes, of which five are relevant to the ideal strategy: age/injury status, extroversion, motivation, outdoor preference, and socioeconomic status [2504.03206].

In practice, beliefs are not updated by an explicit Bayesian transition model. They are set by a fixed oracle user model:

$$
p_t(u)=M_\phi(u\mid s_t), \qquad b_t \equiv p_t(\cdot).
$$

For Education Dialogue, the oracle classifier is a Gemma 7B model that predicts preferred learning style from the conversation. For Exercise Recommendation, the oracle is a decision-tree-like classifier, with an LLM-based extractor when relevant attributes are implicit [2504.03206].

The intrinsic rewards are defined directly over changes in these beliefs. The paper reports potential functions requiring the ground-truth user type $u^*$ in simulated settings,

$$
\phi_{\mathrm{acc}}(b)=b(u^*)-\frac{1}{|\mathcal{U}|},
$$

$$
\phi_{\mathrm{log\text{-}acc}}(b)=\log b(u^*)+\log|\mathcal{U}|,
$$

$$
\phi_{\mathrm{neg\text{-}ent}}(b)=\sum_u b(u)\log b(u)+\log|\mathcal{U}|.
$$

From these, it defines PBRS-consistent differential rewards:

$$
r_{\mathrm{cur}}^{\mathrm{DiffAcc}}(t)=p_{t+1}(u^*)-p_t(u^*),
$$

$$
r_{\mathrm{cur}}^{\mathrm{DiffLogAcc}}(t)=\log p_{t+1}(u^*)-\log p_t(u^*),
$$

$$
r_{\mathrm{cur}}^{\mathrm{DiffEnt}}(t)=\sum_u p_t(u)\log p_t(u)-\sum_u p_{t+1}(u)\log p_{t+1}(u).
$$

It also reports non-differential rewards,

$$
r_{\mathrm{cur}}^{\mathrm{Acc}}(t)=p_{t+1}(u^*)-\frac{1}{|\mathcal{U}|},
$$

$$
r_{\mathrm{cur}}^{\mathrm{Ent}}(t)=\sum_u p_{t+1}(u)\log p_{t+1}(u)+\log|\mathcal{U}|,
$$

and an information-gain surrogate,

$$
r_{\mathrm{cur}}^{\mathrm{InfoGain}}(t)=D_{\mathrm{KL}}\big[p_{t+1}(u)\,\|\,p_t(u)\big].
$$

The policy action space is the full natural-language response at turn $t$. There is no hard-coded controller. The policy may progress the task or elicit user traits by asking clarifying questions or proposing preference-sensitive options. Differential curiosity rewards are reported to encourage early elicitation without artificially prolonging conversations, whereas non-differential rewards can incentivize longer dialogues [2504.03206].

The environments are fully simulated. Education Dialogue uses a Gemma 2B student simulator fine-tuned via supervised learning and a Gemma 2B reward model from Shani et al. (2024). Exercise Recommendation uses Gemini 1.5 Pro for detailed backstory generation consistent with sampled attributes; users are simulated by an LLM environment that stays consistent with those backstories. The same shaping mechanism is used in both domains, which the paper presents as domain-agnostic personalization [2504.03206].

## 4. Token-level CD-RLHF for diversity-preserving alignment

The diversity-oriented formulation adapts curiosity-driven exploration from RL to autoregressive text generation. The state at time $t$ is the prompt plus the partial generation, with $s_0=x$ and $s_t=\{s_0,a_{<t}\}$, the action is token $a_t\sim\pi_\theta(a_t\mid s_t)$, and the transition is $s_{t+1}=\{s_t,a_t\}$ [2501.11463].

Its intrinsic reward is produced by an Intrinsic Curiosity Module consisting of a feature encoder $\phi$ and a forward model $f$, both two-layer MLPs. The forward model predicts the next-state feature,

$$
\hat{\phi}(s_{t+1}) = f(\phi(s_t), a_t),
$$

and the ICM loss is

$$
\mathcal{L}_{\mathrm{ICM}}=\frac{1}{2}\,\big\|\hat{\phi}(s_{t+1})-\phi(s_{t+1})\big\|_2^2.
$$

Curiosity is gated by token probability. If $V^{(k)}$ denotes the top-$k$ tokens by probability, then

$$
r^{(i)}_t =
\begin{cases}
0, & a_t \in V^{(k)},\\[4pt]
\frac{1}{2}\big\|\hat{\phi}(s_{t+1})-\phi(s_{t+1})\big\|_2, & \text{otherwise}.
\end{cases}
$$

The intrinsic reward is then whitened within each trajectory as

$$
r^{(i)} \leftarrow (r^{(i)}-\mu)/\sigma^2,
$$

where $\mu$ and $\sigma$ are the mean and standard deviation of $r^{(i)}$ over the trajectory [2501.11463].

The representation design is specific to LLMs. The method uses the last-layer hidden states of the reference model to represent $s_t$ and $s_{t+1}$, and the actor model’s token embedding to represent $a_t$. Novelty is local rather than global: there is no memory bank, and the curiosity reward depends on the prediction error of the immediate next-state feature given the current state and action [2501.11463].

The paper keeps the curiosity coefficient $\eta$ small and constant to avoid intrinsic rewards overwhelming extrinsic rewards. Reported values are $0.04$ for Gemma-2B/7B, $0.06$ for Llama-3.2-1B TL;DR, $0.04$ for Llama-3.2-1B UltraFeedback, and $0.08$ for Llama-3.2-3B. Top-$k$ gating uses $k=1$ by default, giving intrinsic rewards at approximately $20\%$ of positions, with ablations for $k\in\{3,10\}$. PPO is implemented with DeepSpeed-Chat; key hyperparameters include PPO batch size $256$, PPO epochs $1$, rollout per step $1$, clip ratio $\epsilon=0.2$, GAE $\lambda=0.95$, and $\gamma=1.0$ [2501.11463].

The reported training datasets are TL;DR summarization and UltraFeedback instruction following. Data are split across SFT, reward-model training, and PPO in a $20\%/40\%/40\%$ partition. Evaluation further includes out-of-distribution testing on MT-Bench and extended story generation with ROC Stories prompts [2501.11463].

## 5. Empirical findings

In personalized multi-turn dialogue, the main reported metric is pairwise win rate judged by Gemini 1.5 Pro. Against the multi-turn RLHF baseline, DiffAcc achieves $75.25\%$ wins, Acc $63.00\%$, DiffLogAcc $74.00\%$, and InfoGain $74.00\%$; the Entropy variant averages $48.25\%$ and shows type-specific asymmetry [2504.03206]. For overall conversation quality, DiffLogAcc achieves $57.50\%$ wins against baseline, while Baseline versus DiffAcc is $59.75\%$, indicating that DiffAcc slightly hurts overall quality relative to the baseline. The paper also reports that the differential-accuracy model learns to ask about user type at turn 1 early, from approximately $10$k steps, producing higher $p_1(u^*)$ than the baseline, whose accuracy improves only when the student explicitly states preferences [2504.03206].

A central qualitative result in the dialogue setting concerns the distinction between grounded and ungrounded curiosity. Entropy reward performs best on the “hands-on” type but worst on “story-telling,” and this is attributed to “controlling behavior,” in which the policy drives the classifier toward one type rather than grounding in user statements. Accuracy-based rewards are described as grounded and as avoiding this pathology. Non-differential rewards such as Acc are also reported to lengthen conversations, harming quality, whereas differential rewards such as DiffLogAcc avoid prolongation [2504.03206].

In the diversity-oriented setting, CD-RLHF is reported to improve diversity metrics over RLHF while preserving RM-based alignment. On TL;DR, representative results include Diversity $+33.16\%$, EAD $+6.07\%$, Self-BLEU $-23.08\%$, and SentBERT $-4.33\%$ for Gemma-2B versus RLHF, with RM score $0.95$ versus $0.90$; for Llama-3.2-1B, Diversity is $+40.26\%$ and RM score is $1.17$ versus $1.14$; for Llama-3.2-3B, Diversity is $+28.01\%$ and RM score improves from $3.33$ to $3.49$ [2501.11463]. On UltraFeedback, representative gains include Diversity $+12.63\%$ and EAD $+14.06\%$ for Gemma-2B, with RM score $-0.90$ versus $-1.01$, and Diversity $+23.16\%$ for Llama-3.2-3B, with RM score $1.43$ versus $1.35$ [2501.11463].

External evaluation is also reported. TL;DR GPT-4 win rates in favor of CD-RLHF are $58\%$ for Gemma-2B, $46\%$ for Gemma-7B, $58\%$ for Llama-3.2-1B, and $56\%$ for Llama-3.2-3B; UltraFeedback averages approximately $62\%$ GPT-4 win rates across models, and human evaluations consistently prefer CD-RLHF on diversity [2501.11463]. On MT-Bench, overall GPT-4 scores improve from $5.35$ to $5.68$ for Gemma-2B, from $5.75$ to $5.96$ for Gemma-7B, from $3.71$ to $4.18$ for Llama-3.2-1B, and from $5.98$ to $6.08$ for Llama-3.2-3B [2501.11463].

Ablations in the diversity-oriented paper emphasize that reward frequency and KL strength remain consequential. Expanding intrinsic rewards from approximately $20\%$ of positions to $60$–$100\%$ yields small additional diversity gains but degrades alignment; for TL;DR Gemma-2B, RM score drops from $0.95$ to $0.88$ at $100\%$. For top-$k$ gating on TL;DR Gemma-2B, the paper reports: $k=1$, Diversity $0.2839$ and RM $0.95$; $k=3$, Diversity $0.2690$ and RM $0.94$; $k=10$, Diversity $0.2613$ and RM $0.95$ [2501.11463].

## 6. Limitations, safety, and unresolved questions

The two CD-RLHF lines expose different limitations. In personalized dialogue, the user traits are simplified, with only two learning styles in education and fixed attribute logic in exercise. The user model is assumed to be an optimal fixed oracle, whereas real deployment would require learned or adaptive user models together with robust uncertainty quantification. The experiments rely on simulated users, and distribution shift to real-world diversity or non-cooperative behavior is explicitly identified as a limitation. The paper further notes that no special safe-RLHF components are implemented, and that privacy, transparency, and bias mitigation remain future work [2504.03206].

In diversity-oriented CD-RLHF, the principal limitation is scalar mismatch between intrinsic and extrinsic rewards. The paper states that this is mitigated by whitening and a small constant $\eta$, but not eliminated. It also notes that the diversity–alignment trade-off remains and that balancing still depends on $\beta$, $\eta$, and gating. Degenerate novelty signals, reward hacking, and unanticipated behaviors from increased diversity are identified as concerns; the proposed safeguards are a strong KL constraint, GPT-4 and human evaluation, top-$k$ gating, and keeping curiosity confined to controllable contexts [2501.11463].

A recurring misconception is that curiosity automatically improves the desired downstream behavior. The dialogue paper shows that entropy-based shaping can induce “controlling behavior” when the external reward is user-agnostic, while the diversity paper shows that applying curiosity too frequently can degrade alignment [2504.03206][2501.11463]. Another misconception is that curiosity necessarily implies longer or more exploratory interactions in an unconstrained sense. In fact, the reported results distinguish between differential and non-differential shaping: differential curiosity in dialogue is explicitly presented as a safeguard against gratuitous prolongation, and top-$k$ gating in token-level generation is presented as a safeguard against unnecessary exploration [2504.03206][2501.11463].

These results place CD-RLHF at the intersection of RLHF, intrinsic motivation, active learning, and reward shaping. The dialogue formulation explicitly links its method to Bayesian experimental design and active preference elicitation; the text-generation formulation links its token-level curiosity to count-based exploration and prediction-error bonuses in RL [2504.03206][2501.11463]. A plausible implication is that future CD-RLHF research will continue to diversify along the latent variable being explored—user traits, stylistic modes, or under-covered semantic regions—while retaining the common structure of RLHF augmented with carefully controlled intrinsic rewards.

Source: https://www.emergentmind.com/topics/curiosity-driven-rlhf-cd-rlhf