Papers
Topics
Authors
Recent
Search
2000 character limit reached

Curiosity-Driven RLHF Advances

Updated 15 July 2026
  • CD-RLHF is a family of methods that augments RLHF with intrinsic rewards to balance alignment with exploration.
  • The diversity-oriented approach leverages token-level curiosity to enhance output variety while maintaining model alignment.
  • The personalized dialogue variant uses belief-based curiosity to infer latent user traits and drive adaptive multi-turn interactions.

Searching arXiv for the specified CD-RLHF papers and closely related work to ground the article. Curiosity-Driven Reinforcement Learning from Human Feedback (CD-RLHF) denotes RLHF variants that supplement preference-derived extrinsic rewards with intrinsic rewards that favor exploration. Current arXiv usage suggests that the label covers at least two distinct 2025 formulations: a token-level curiosity mechanism for improving output diversity in summarization and instruction-following, introduced in "Curiosity-Driven Reinforcement Learning from Human Feedback" (Sun et al., 20 Jan 2025), and a belief-based curiosity mechanism for eliciting latent user traits in personalized multi-turn dialogue, developed in "Enhancing Personalized Multi-Turn Dialogue with Curiosity Reward" (Wan et al., 4 Apr 2025). In both cases, the central idea is to preserve the optimization structure of RLHF while adding a shaped signal that rewards informative or novel trajectories.

1. Scope and conceptual variants

CD-RLHF emerged from two different problem settings. In one setting, standard RLHF is described as improving alignment to human preferences while reducing output diversity relative to supervised fine-tuning, producing repetitive or homogenized outputs; CD-RLHF addresses this by adding token-level curiosity for novel states during generation (Sun et al., 20 Jan 2025). In the other setting, multi-turn RLHF is described as prioritizing helpfulness and safety but falling short in empathetic, adaptive, and personalized interaction; CD-RLHF addresses this by rewarding reductions in uncertainty about latent user traits during dialogue (Wan et al., 4 Apr 2025).

The two formulations differ in what constitutes a “state novelty” signal and in what exploration is intended to achieve. The diversity-oriented formulation uses prediction error in a learned latent state space and applies it to tokens outside the top-kk by probability. The personalized-dialogue formulation uses a belief over hidden user types and rewards actions that improve the agent’s user model (Sun et al., 20 Jan 2025, Wan et al., 4 Apr 2025).

Formulation Intrinsic signal Primary aim
Diversity-oriented CD-RLHF Prediction error in an Intrinsic Curiosity Module Optimize both output diversity and alignment quality
Personalized dialogue CD-RLHF Changes in belief over latent user type Reveal user preferences and adapt to them

This suggests that CD-RLHF is better understood as a family of RLHF augmentations rather than a single algorithm. What is common across variants is the addition of intrinsic motivation to a standard RLHF backbone; what differs is the latent variable being explored, the reward granularity, and the failure modes.

2. Integration with RLHF objectives

Both formulations preserve the standard separation between an extrinsic alignment signal and an auxiliary intrinsic signal. In the personalized multi-turn setting, the combined objective is reported as

J(θ)=Eτ∼πθ[∑t=0T(rRLHF(st,at)+λ rcur(st,at))],J(\theta)=\mathbb{E}_{\tau\sim\pi_\theta}\Bigg[\sum_{t=0}^{T}\big(r_{\mathrm{RLHF}}(s_t,a_t)+\lambda\, r_{\mathrm{cur}}(s_t,a_t)\big)\Bigg],

with turn-level curiosity rewards added to the RLHF reward and optimized using the multi-turn RLHF training recipe of Shani et al. (2024); for Education Dialogue, the paper reports λ=9.0\lambda=9.0 (Wan et al., 4 Apr 2025).

In the diversity-oriented formulation, the combined reward is written at token level as

rt=rt(e)+η rt(i),r_t = r^{(e)}_t + \eta\, r^{(i)}_t,

where the extrinsic reward combines a sequence-level reward model score with a KL penalty to the SFT reference model,

r(e)=R−βDKL(πpolicy(⋅∣st) ∥ πref(⋅∣st)),r^{(e)} = R - \beta D_{\mathrm{KL}}\big(\pi_{\mathrm{policy}}(\cdot|s_t)\,\|\,\pi_{\mathrm{ref}}(\cdot|s_t)\big),

and the policy is updated with PPO using generalized advantage estimation (Sun et al., 20 Jan 2025).

The personalized-dialogue paper gives a stronger theoretical account through a POMDP over observable dialogue histories sts_t and latent user type u∈Uu\in\mathcal{U}. With extended state st′=⟨st,u⟩s_t'=\langle s_t,u\rangle, transition and reward are

T′(st+1′∣st′,at)=T(st+1∣st,at,u),R′(st′,at)=R(st,at∣u),\mathcal{T}'(s'_{t+1}\mid s'_t,a_t)=\mathcal{T}(s_{t+1}\mid s_t,a_t,u), \quad \mathcal{R}'(s'_t,a_t)=\mathcal{R}(s_t,a_t\mid u),

and the belief update is

bt+1(u)∝T(st+1∣st,at,u) bt(u).b_{t+1}(u)\propto \mathcal{T}(s_{t+1}\mid s_t,a_t,u)\, b_t(u).

Using a belief-based expected reward,

J(θ)=Eτ∼πθ[∑t=0T(rRLHF(st,at)+λ rcur(st,at))],J(\theta)=\mathbb{E}_{\tau\sim\pi_\theta}\Bigg[\sum_{t=0}^{T}\big(r_{\mathrm{RLHF}}(s_t,a_t)+\lambda\, r_{\mathrm{cur}}(s_t,a_t)\big)\Bigg],0

the paper invokes potential-based reward shaping over beliefs,

J(θ)=Eτ∼πθ[∑t=0T(rRLHF(st,at)+λ rcur(st,at))],J(\theta)=\mathbb{E}_{\tau\sim\pi_\theta}\Bigg[\sum_{t=0}^{T}\big(r_{\mathrm{RLHF}}(s_t,a_t)+\lambda\, r_{\mathrm{cur}}(s_t,a_t)\big)\Bigg],1

and states, via Eck et al. (2013), that such shaping leaves the optimal policy invariant (Wan et al., 4 Apr 2025).

By contrast, the diversity-oriented paper is framed operationally rather than through PBRS. It retains the conventional SFT J(θ)=Eτ∼πθ[∑t=0T(rRLHF(st,at)+λ rcur(st,at))],J(\theta)=\mathbb{E}_{\tau\sim\pi_\theta}\Bigg[\sum_{t=0}^{T}\big(r_{\mathrm{RLHF}}(s_t,a_t)+\lambda\, r_{\mathrm{cur}}(s_t,a_t)\big)\Bigg],2 RM J(θ)=Eτ∼πθ[∑t=0T(rRLHF(st,at)+λ rcur(st,at))],J(\theta)=\mathbb{E}_{\tau\sim\pi_\theta}\Bigg[\sum_{t=0}^{T}\big(r_{\mathrm{RLHF}}(s_t,a_t)+\lambda\, r_{\mathrm{cur}}(s_t,a_t)\big)\Bigg],3 PPO pipeline, initializes policy and reference models via SFT on the same preference dataset, trains the reward model on paired preferences, and inserts an Intrinsic Curiosity Module during PPO rollouts (Sun et al., 20 Jan 2025).

3. Belief-based CD-RLHF for personalized multi-turn dialogue

The personalized formulation targets latent user inference during online interaction. The agent does not assume pre-existing user profiles or histories; instead, it must infer the hidden user type from observed dialogue. Two environments are reported. In Education Dialogue, the latent trait is the student’s preferred learning style, described in prompts and analyses as “hands-on” versus “story-telling,” and earlier as “lecture-based” versus “hands-on.” In Exercise Recommendation, the user profile has 20 attributes, of which five are relevant to the ideal strategy: age/injury status, extroversion, motivation, outdoor preference, and socioeconomic status (Wan et al., 4 Apr 2025).

In practice, beliefs are not updated by an explicit Bayesian transition model. They are set by a fixed oracle user model:

J(θ)=Eτ∼πθ[∑t=0T(rRLHF(st,at)+λ rcur(st,at))],J(\theta)=\mathbb{E}_{\tau\sim\pi_\theta}\Bigg[\sum_{t=0}^{T}\big(r_{\mathrm{RLHF}}(s_t,a_t)+\lambda\, r_{\mathrm{cur}}(s_t,a_t)\big)\Bigg],4

For Education Dialogue, the oracle classifier is a Gemma 7B model that predicts preferred learning style from the conversation. For Exercise Recommendation, the oracle is a decision-tree-like classifier, with an LLM-based extractor when relevant attributes are implicit (Wan et al., 4 Apr 2025).

The intrinsic rewards are defined directly over changes in these beliefs. The paper reports potential functions requiring the ground-truth user type J(θ)=Eτ∼πθ[∑t=0T(rRLHF(st,at)+λ rcur(st,at))],J(\theta)=\mathbb{E}_{\tau\sim\pi_\theta}\Bigg[\sum_{t=0}^{T}\big(r_{\mathrm{RLHF}}(s_t,a_t)+\lambda\, r_{\mathrm{cur}}(s_t,a_t)\big)\Bigg],5 in simulated settings,

J(θ)=Eτ∼πθ[∑t=0T(rRLHF(st,at)+λ rcur(st,at))],J(\theta)=\mathbb{E}_{\tau\sim\pi_\theta}\Bigg[\sum_{t=0}^{T}\big(r_{\mathrm{RLHF}}(s_t,a_t)+\lambda\, r_{\mathrm{cur}}(s_t,a_t)\big)\Bigg],6

J(θ)=Eτ∼πθ[∑t=0T(rRLHF(st,at)+λ rcur(st,at))],J(\theta)=\mathbb{E}_{\tau\sim\pi_\theta}\Bigg[\sum_{t=0}^{T}\big(r_{\mathrm{RLHF}}(s_t,a_t)+\lambda\, r_{\mathrm{cur}}(s_t,a_t)\big)\Bigg],7

J(θ)=Eτ∼πθ[∑t=0T(rRLHF(st,at)+λ rcur(st,at))],J(\theta)=\mathbb{E}_{\tau\sim\pi_\theta}\Bigg[\sum_{t=0}^{T}\big(r_{\mathrm{RLHF}}(s_t,a_t)+\lambda\, r_{\mathrm{cur}}(s_t,a_t)\big)\Bigg],8

From these, it defines PBRS-consistent differential rewards:

J(θ)=Eτ∼πθ[∑t=0T(rRLHF(st,at)+λ rcur(st,at))],J(\theta)=\mathbb{E}_{\tau\sim\pi_\theta}\Bigg[\sum_{t=0}^{T}\big(r_{\mathrm{RLHF}}(s_t,a_t)+\lambda\, r_{\mathrm{cur}}(s_t,a_t)\big)\Bigg],9

λ=9.0\lambda=9.00

λ=9.0\lambda=9.01

It also reports non-differential rewards,

λ=9.0\lambda=9.02

λ=9.0\lambda=9.03

and an information-gain surrogate,

λ=9.0\lambda=9.04

The policy action space is the full natural-language response at turn λ=9.0\lambda=9.05. There is no hard-coded controller. The policy may progress the task or elicit user traits by asking clarifying questions or proposing preference-sensitive options. Differential curiosity rewards are reported to encourage early elicitation without artificially prolonging conversations, whereas non-differential rewards can incentivize longer dialogues (Wan et al., 4 Apr 2025).

The environments are fully simulated. Education Dialogue uses a Gemma 2B student simulator fine-tuned via supervised learning and a Gemma 2B reward model from Shani et al. (2024). Exercise Recommendation uses Gemini 1.5 Pro for detailed backstory generation consistent with sampled attributes; users are simulated by an LLM environment that stays consistent with those backstories. The same shaping mechanism is used in both domains, which the paper presents as domain-agnostic personalization (Wan et al., 4 Apr 2025).

4. Token-level CD-RLHF for diversity-preserving alignment

The diversity-oriented formulation adapts curiosity-driven exploration from RL to autoregressive text generation. The state at time λ=9.0\lambda=9.06 is the prompt plus the partial generation, with λ=9.0\lambda=9.07 and λ=9.0\lambda=9.08, the action is token λ=9.0\lambda=9.09, and the transition is rt=rt(e)+η rt(i),r_t = r^{(e)}_t + \eta\, r^{(i)}_t,0 (Sun et al., 20 Jan 2025).

Its intrinsic reward is produced by an Intrinsic Curiosity Module consisting of a feature encoder rt=rt(e)+η rt(i),r_t = r^{(e)}_t + \eta\, r^{(i)}_t,1 and a forward model rt=rt(e)+η rt(i),r_t = r^{(e)}_t + \eta\, r^{(i)}_t,2, both two-layer MLPs. The forward model predicts the next-state feature,

rt=rt(e)+η rt(i),r_t = r^{(e)}_t + \eta\, r^{(i)}_t,3

and the ICM loss is

rt=rt(e)+η rt(i),r_t = r^{(e)}_t + \eta\, r^{(i)}_t,4

Curiosity is gated by token probability. If rt=rt(e)+η rt(i),r_t = r^{(e)}_t + \eta\, r^{(i)}_t,5 denotes the top-rt=rt(e)+η rt(i),r_t = r^{(e)}_t + \eta\, r^{(i)}_t,6 tokens by probability, then

rt=rt(e)+η rt(i),r_t = r^{(e)}_t + \eta\, r^{(i)}_t,7

The intrinsic reward is then whitened within each trajectory as

rt=rt(e)+η rt(i),r_t = r^{(e)}_t + \eta\, r^{(i)}_t,8

where rt=rt(e)+η rt(i),r_t = r^{(e)}_t + \eta\, r^{(i)}_t,9 and r(e)=R−βDKL(πpolicy(⋅∣st) ∥ πref(⋅∣st)),r^{(e)} = R - \beta D_{\mathrm{KL}}\big(\pi_{\mathrm{policy}}(\cdot|s_t)\,\|\,\pi_{\mathrm{ref}}(\cdot|s_t)\big),0 are the mean and standard deviation of r(e)=R−βDKL(πpolicy(⋅∣st) ∥ πref(⋅∣st)),r^{(e)} = R - \beta D_{\mathrm{KL}}\big(\pi_{\mathrm{policy}}(\cdot|s_t)\,\|\,\pi_{\mathrm{ref}}(\cdot|s_t)\big),1 over the trajectory (Sun et al., 20 Jan 2025).

The representation design is specific to LLMs. The method uses the last-layer hidden states of the reference model to represent r(e)=R−βDKL(πpolicy(⋅∣st) ∥ πref(⋅∣st)),r^{(e)} = R - \beta D_{\mathrm{KL}}\big(\pi_{\mathrm{policy}}(\cdot|s_t)\,\|\,\pi_{\mathrm{ref}}(\cdot|s_t)\big),2 and r(e)=R−βDKL(πpolicy(⋅∣st) ∥ πref(⋅∣st)),r^{(e)} = R - \beta D_{\mathrm{KL}}\big(\pi_{\mathrm{policy}}(\cdot|s_t)\,\|\,\pi_{\mathrm{ref}}(\cdot|s_t)\big),3, and the actor model’s token embedding to represent r(e)=R−βDKL(πpolicy(⋅∣st) ∥ πref(⋅∣st)),r^{(e)} = R - \beta D_{\mathrm{KL}}\big(\pi_{\mathrm{policy}}(\cdot|s_t)\,\|\,\pi_{\mathrm{ref}}(\cdot|s_t)\big),4. Novelty is local rather than global: there is no memory bank, and the curiosity reward depends on the prediction error of the immediate next-state feature given the current state and action (Sun et al., 20 Jan 2025).

The paper keeps the curiosity coefficient r(e)=R−βDKL(πpolicy(⋅∣st) ∥ πref(⋅∣st)),r^{(e)} = R - \beta D_{\mathrm{KL}}\big(\pi_{\mathrm{policy}}(\cdot|s_t)\,\|\,\pi_{\mathrm{ref}}(\cdot|s_t)\big),5 small and constant to avoid intrinsic rewards overwhelming extrinsic rewards. Reported values are r(e)=R−βDKL(πpolicy(⋅∣st) ∥ πref(⋅∣st)),r^{(e)} = R - \beta D_{\mathrm{KL}}\big(\pi_{\mathrm{policy}}(\cdot|s_t)\,\|\,\pi_{\mathrm{ref}}(\cdot|s_t)\big),6 for Gemma-2B/7B, r(e)=R−βDKL(πpolicy(⋅∣st) ∥ πref(⋅∣st)),r^{(e)} = R - \beta D_{\mathrm{KL}}\big(\pi_{\mathrm{policy}}(\cdot|s_t)\,\|\,\pi_{\mathrm{ref}}(\cdot|s_t)\big),7 for Llama-3.2-1B TL;DR, r(e)=R−βDKL(πpolicy(⋅∣st) ∥ πref(⋅∣st)),r^{(e)} = R - \beta D_{\mathrm{KL}}\big(\pi_{\mathrm{policy}}(\cdot|s_t)\,\|\,\pi_{\mathrm{ref}}(\cdot|s_t)\big),8 for Llama-3.2-1B UltraFeedback, and r(e)=R−βDKL(πpolicy(⋅∣st) ∥ πref(⋅∣st)),r^{(e)} = R - \beta D_{\mathrm{KL}}\big(\pi_{\mathrm{policy}}(\cdot|s_t)\,\|\,\pi_{\mathrm{ref}}(\cdot|s_t)\big),9 for Llama-3.2-3B. Top-sts_t0 gating uses sts_t1 by default, giving intrinsic rewards at approximately sts_t2 of positions, with ablations for sts_t3. PPO is implemented with DeepSpeed-Chat; key hyperparameters include PPO batch size sts_t4, PPO epochs sts_t5, rollout per step sts_t6, clip ratio sts_t7, GAE sts_t8, and sts_t9 (Sun et al., 20 Jan 2025).

The reported training datasets are TL;DR summarization and UltraFeedback instruction following. Data are split across SFT, reward-model training, and PPO in a u∈Uu\in\mathcal{U}0 partition. Evaluation further includes out-of-distribution testing on MT-Bench and extended story generation with ROC Stories prompts (Sun et al., 20 Jan 2025).

5. Empirical findings

In personalized multi-turn dialogue, the main reported metric is pairwise win rate judged by Gemini 1.5 Pro. Against the multi-turn RLHF baseline, DiffAcc achieves u∈Uu\in\mathcal{U}1 wins, Acc u∈Uu\in\mathcal{U}2, DiffLogAcc u∈Uu\in\mathcal{U}3, and InfoGain u∈Uu\in\mathcal{U}4; the Entropy variant averages u∈Uu\in\mathcal{U}5 and shows type-specific asymmetry (Wan et al., 4 Apr 2025). For overall conversation quality, DiffLogAcc achieves u∈Uu\in\mathcal{U}6 wins against baseline, while Baseline versus DiffAcc is u∈Uu\in\mathcal{U}7, indicating that DiffAcc slightly hurts overall quality relative to the baseline. The paper also reports that the differential-accuracy model learns to ask about user type at turn 1 early, from approximately u∈Uu\in\mathcal{U}8k steps, producing higher u∈Uu\in\mathcal{U}9 than the baseline, whose accuracy improves only when the student explicitly states preferences (Wan et al., 4 Apr 2025).

A central qualitative result in the dialogue setting concerns the distinction between grounded and ungrounded curiosity. Entropy reward performs best on the “hands-on” type but worst on “story-telling,” and this is attributed to “controlling behavior,” in which the policy drives the classifier toward one type rather than grounding in user statements. Accuracy-based rewards are described as grounded and as avoiding this pathology. Non-differential rewards such as Acc are also reported to lengthen conversations, harming quality, whereas differential rewards such as DiffLogAcc avoid prolongation (Wan et al., 4 Apr 2025).

In the diversity-oriented setting, CD-RLHF is reported to improve diversity metrics over RLHF while preserving RM-based alignment. On TL;DR, representative results include Diversity st′=⟨st,u⟩s_t'=\langle s_t,u\rangle0, EAD st′=⟨st,u⟩s_t'=\langle s_t,u\rangle1, Self-BLEU st′=⟨st,u⟩s_t'=\langle s_t,u\rangle2, and SentBERT st′=⟨st,u⟩s_t'=\langle s_t,u\rangle3 for Gemma-2B versus RLHF, with RM score st′=⟨st,u⟩s_t'=\langle s_t,u\rangle4 versus st′=⟨st,u⟩s_t'=\langle s_t,u\rangle5; for Llama-3.2-1B, Diversity is st′=⟨st,u⟩s_t'=\langle s_t,u\rangle6 and RM score is st′=⟨st,u⟩s_t'=\langle s_t,u\rangle7 versus st′=⟨st,u⟩s_t'=\langle s_t,u\rangle8; for Llama-3.2-3B, Diversity is st′=⟨st,u⟩s_t'=\langle s_t,u\rangle9 and RM score improves from T′(st+1′∣st′,at)=T(st+1∣st,at,u),R′(st′,at)=R(st,at∣u),\mathcal{T}'(s'_{t+1}\mid s'_t,a_t)=\mathcal{T}(s_{t+1}\mid s_t,a_t,u), \quad \mathcal{R}'(s'_t,a_t)=\mathcal{R}(s_t,a_t\mid u),0 to T′(st+1′∣st′,at)=T(st+1∣st,at,u),R′(st′,at)=R(st,at∣u),\mathcal{T}'(s'_{t+1}\mid s'_t,a_t)=\mathcal{T}(s_{t+1}\mid s_t,a_t,u), \quad \mathcal{R}'(s'_t,a_t)=\mathcal{R}(s_t,a_t\mid u),1 (Sun et al., 20 Jan 2025). On UltraFeedback, representative gains include Diversity T′(st+1′∣st′,at)=T(st+1∣st,at,u),R′(st′,at)=R(st,at∣u),\mathcal{T}'(s'_{t+1}\mid s'_t,a_t)=\mathcal{T}(s_{t+1}\mid s_t,a_t,u), \quad \mathcal{R}'(s'_t,a_t)=\mathcal{R}(s_t,a_t\mid u),2 and EAD T′(st+1′∣st′,at)=T(st+1∣st,at,u),R′(st′,at)=R(st,at∣u),\mathcal{T}'(s'_{t+1}\mid s'_t,a_t)=\mathcal{T}(s_{t+1}\mid s_t,a_t,u), \quad \mathcal{R}'(s'_t,a_t)=\mathcal{R}(s_t,a_t\mid u),3 for Gemma-2B, with RM score T′(st+1′∣st′,at)=T(st+1∣st,at,u),R′(st′,at)=R(st,at∣u),\mathcal{T}'(s'_{t+1}\mid s'_t,a_t)=\mathcal{T}(s_{t+1}\mid s_t,a_t,u), \quad \mathcal{R}'(s'_t,a_t)=\mathcal{R}(s_t,a_t\mid u),4 versus T′(st+1′∣st′,at)=T(st+1∣st,at,u),R′(st′,at)=R(st,at∣u),\mathcal{T}'(s'_{t+1}\mid s'_t,a_t)=\mathcal{T}(s_{t+1}\mid s_t,a_t,u), \quad \mathcal{R}'(s'_t,a_t)=\mathcal{R}(s_t,a_t\mid u),5, and Diversity T′(st+1′∣st′,at)=T(st+1∣st,at,u),R′(st′,at)=R(st,at∣u),\mathcal{T}'(s'_{t+1}\mid s'_t,a_t)=\mathcal{T}(s_{t+1}\mid s_t,a_t,u), \quad \mathcal{R}'(s'_t,a_t)=\mathcal{R}(s_t,a_t\mid u),6 for Llama-3.2-3B, with RM score T′(st+1′∣st′,at)=T(st+1∣st,at,u),R′(st′,at)=R(st,at∣u),\mathcal{T}'(s'_{t+1}\mid s'_t,a_t)=\mathcal{T}(s_{t+1}\mid s_t,a_t,u), \quad \mathcal{R}'(s'_t,a_t)=\mathcal{R}(s_t,a_t\mid u),7 versus T′(st+1′∣st′,at)=T(st+1∣st,at,u),R′(st′,at)=R(st,at∣u),\mathcal{T}'(s'_{t+1}\mid s'_t,a_t)=\mathcal{T}(s_{t+1}\mid s_t,a_t,u), \quad \mathcal{R}'(s'_t,a_t)=\mathcal{R}(s_t,a_t\mid u),8 (Sun et al., 20 Jan 2025).

External evaluation is also reported. TL;DR GPT-4 win rates in favor of CD-RLHF are T′(st+1′∣st′,at)=T(st+1∣st,at,u),R′(st′,at)=R(st,at∣u),\mathcal{T}'(s'_{t+1}\mid s'_t,a_t)=\mathcal{T}(s_{t+1}\mid s_t,a_t,u), \quad \mathcal{R}'(s'_t,a_t)=\mathcal{R}(s_t,a_t\mid u),9 for Gemma-2B, bt+1(u)∝T(st+1∣st,at,u) bt(u).b_{t+1}(u)\propto \mathcal{T}(s_{t+1}\mid s_t,a_t,u)\, b_t(u).0 for Gemma-7B, bt+1(u)∝T(st+1∣st,at,u) bt(u).b_{t+1}(u)\propto \mathcal{T}(s_{t+1}\mid s_t,a_t,u)\, b_t(u).1 for Llama-3.2-1B, and bt+1(u)∝T(st+1∣st,at,u) bt(u).b_{t+1}(u)\propto \mathcal{T}(s_{t+1}\mid s_t,a_t,u)\, b_t(u).2 for Llama-3.2-3B; UltraFeedback averages approximately bt+1(u)∝T(st+1∣st,at,u) bt(u).b_{t+1}(u)\propto \mathcal{T}(s_{t+1}\mid s_t,a_t,u)\, b_t(u).3 GPT-4 win rates across models, and human evaluations consistently prefer CD-RLHF on diversity (Sun et al., 20 Jan 2025). On MT-Bench, overall GPT-4 scores improve from bt+1(u)∝T(st+1∣st,at,u) bt(u).b_{t+1}(u)\propto \mathcal{T}(s_{t+1}\mid s_t,a_t,u)\, b_t(u).4 to bt+1(u)∝T(st+1∣st,at,u) bt(u).b_{t+1}(u)\propto \mathcal{T}(s_{t+1}\mid s_t,a_t,u)\, b_t(u).5 for Gemma-2B, from bt+1(u)∝T(st+1∣st,at,u) bt(u).b_{t+1}(u)\propto \mathcal{T}(s_{t+1}\mid s_t,a_t,u)\, b_t(u).6 to bt+1(u)∝T(st+1∣st,at,u) bt(u).b_{t+1}(u)\propto \mathcal{T}(s_{t+1}\mid s_t,a_t,u)\, b_t(u).7 for Gemma-7B, from bt+1(u)∝T(st+1∣st,at,u) bt(u).b_{t+1}(u)\propto \mathcal{T}(s_{t+1}\mid s_t,a_t,u)\, b_t(u).8 to bt+1(u)∝T(st+1∣st,at,u) bt(u).b_{t+1}(u)\propto \mathcal{T}(s_{t+1}\mid s_t,a_t,u)\, b_t(u).9 for Llama-3.2-1B, and from J(θ)=Eτ∼πθ[∑t=0T(rRLHF(st,at)+λ rcur(st,at))],J(\theta)=\mathbb{E}_{\tau\sim\pi_\theta}\Bigg[\sum_{t=0}^{T}\big(r_{\mathrm{RLHF}}(s_t,a_t)+\lambda\, r_{\mathrm{cur}}(s_t,a_t)\big)\Bigg],00 to J(θ)=Eτ∼πθ[∑t=0T(rRLHF(st,at)+λ rcur(st,at))],J(\theta)=\mathbb{E}_{\tau\sim\pi_\theta}\Bigg[\sum_{t=0}^{T}\big(r_{\mathrm{RLHF}}(s_t,a_t)+\lambda\, r_{\mathrm{cur}}(s_t,a_t)\big)\Bigg],01 for Llama-3.2-3B (Sun et al., 20 Jan 2025).

Ablations in the diversity-oriented paper emphasize that reward frequency and KL strength remain consequential. Expanding intrinsic rewards from approximately J(θ)=Eτ∼πθ[∑t=0T(rRLHF(st,at)+λ rcur(st,at))],J(\theta)=\mathbb{E}_{\tau\sim\pi_\theta}\Bigg[\sum_{t=0}^{T}\big(r_{\mathrm{RLHF}}(s_t,a_t)+\lambda\, r_{\mathrm{cur}}(s_t,a_t)\big)\Bigg],02 of positions to J(θ)=Eτ∼πθ[∑t=0T(rRLHF(st,at)+λ rcur(st,at))],J(\theta)=\mathbb{E}_{\tau\sim\pi_\theta}\Bigg[\sum_{t=0}^{T}\big(r_{\mathrm{RLHF}}(s_t,a_t)+\lambda\, r_{\mathrm{cur}}(s_t,a_t)\big)\Bigg],03–J(θ)=Eτ∼πθ[∑t=0T(rRLHF(st,at)+λ rcur(st,at))],J(\theta)=\mathbb{E}_{\tau\sim\pi_\theta}\Bigg[\sum_{t=0}^{T}\big(r_{\mathrm{RLHF}}(s_t,a_t)+\lambda\, r_{\mathrm{cur}}(s_t,a_t)\big)\Bigg],04 yields small additional diversity gains but degrades alignment; for TL;DR Gemma-2B, RM score drops from J(θ)=Eτ∼πθ[∑t=0T(rRLHF(st,at)+λ rcur(st,at))],J(\theta)=\mathbb{E}_{\tau\sim\pi_\theta}\Bigg[\sum_{t=0}^{T}\big(r_{\mathrm{RLHF}}(s_t,a_t)+\lambda\, r_{\mathrm{cur}}(s_t,a_t)\big)\Bigg],05 to J(θ)=Eτ∼πθ[∑t=0T(rRLHF(st,at)+λ rcur(st,at))],J(\theta)=\mathbb{E}_{\tau\sim\pi_\theta}\Bigg[\sum_{t=0}^{T}\big(r_{\mathrm{RLHF}}(s_t,a_t)+\lambda\, r_{\mathrm{cur}}(s_t,a_t)\big)\Bigg],06 at J(θ)=Eτ∼πθ[∑t=0T(rRLHF(st,at)+λ rcur(st,at))],J(\theta)=\mathbb{E}_{\tau\sim\pi_\theta}\Bigg[\sum_{t=0}^{T}\big(r_{\mathrm{RLHF}}(s_t,a_t)+\lambda\, r_{\mathrm{cur}}(s_t,a_t)\big)\Bigg],07. For top-J(θ)=Eτ∼πθ[∑t=0T(rRLHF(st,at)+λ rcur(st,at))],J(\theta)=\mathbb{E}_{\tau\sim\pi_\theta}\Bigg[\sum_{t=0}^{T}\big(r_{\mathrm{RLHF}}(s_t,a_t)+\lambda\, r_{\mathrm{cur}}(s_t,a_t)\big)\Bigg],08 gating on TL;DR Gemma-2B, the paper reports: J(θ)=Eτ∼πθ[∑t=0T(rRLHF(st,at)+λ rcur(st,at))],J(\theta)=\mathbb{E}_{\tau\sim\pi_\theta}\Bigg[\sum_{t=0}^{T}\big(r_{\mathrm{RLHF}}(s_t,a_t)+\lambda\, r_{\mathrm{cur}}(s_t,a_t)\big)\Bigg],09, Diversity J(θ)=Eτ∼πθ[∑t=0T(rRLHF(st,at)+λ rcur(st,at))],J(\theta)=\mathbb{E}_{\tau\sim\pi_\theta}\Bigg[\sum_{t=0}^{T}\big(r_{\mathrm{RLHF}}(s_t,a_t)+\lambda\, r_{\mathrm{cur}}(s_t,a_t)\big)\Bigg],10 and RM J(θ)=Eτ∼πθ[∑t=0T(rRLHF(st,at)+λ rcur(st,at))],J(\theta)=\mathbb{E}_{\tau\sim\pi_\theta}\Bigg[\sum_{t=0}^{T}\big(r_{\mathrm{RLHF}}(s_t,a_t)+\lambda\, r_{\mathrm{cur}}(s_t,a_t)\big)\Bigg],11; J(θ)=Eτ∼πθ[∑t=0T(rRLHF(st,at)+λ rcur(st,at))],J(\theta)=\mathbb{E}_{\tau\sim\pi_\theta}\Bigg[\sum_{t=0}^{T}\big(r_{\mathrm{RLHF}}(s_t,a_t)+\lambda\, r_{\mathrm{cur}}(s_t,a_t)\big)\Bigg],12, Diversity J(θ)=Eτ∼πθ[∑t=0T(rRLHF(st,at)+λ rcur(st,at))],J(\theta)=\mathbb{E}_{\tau\sim\pi_\theta}\Bigg[\sum_{t=0}^{T}\big(r_{\mathrm{RLHF}}(s_t,a_t)+\lambda\, r_{\mathrm{cur}}(s_t,a_t)\big)\Bigg],13 and RM J(θ)=Eτ∼πθ[∑t=0T(rRLHF(st,at)+λ rcur(st,at))],J(\theta)=\mathbb{E}_{\tau\sim\pi_\theta}\Bigg[\sum_{t=0}^{T}\big(r_{\mathrm{RLHF}}(s_t,a_t)+\lambda\, r_{\mathrm{cur}}(s_t,a_t)\big)\Bigg],14; J(θ)=Eτ∼πθ[∑t=0T(rRLHF(st,at)+λ rcur(st,at))],J(\theta)=\mathbb{E}_{\tau\sim\pi_\theta}\Bigg[\sum_{t=0}^{T}\big(r_{\mathrm{RLHF}}(s_t,a_t)+\lambda\, r_{\mathrm{cur}}(s_t,a_t)\big)\Bigg],15, Diversity J(θ)=Eτ∼πθ[∑t=0T(rRLHF(st,at)+λ rcur(st,at))],J(\theta)=\mathbb{E}_{\tau\sim\pi_\theta}\Bigg[\sum_{t=0}^{T}\big(r_{\mathrm{RLHF}}(s_t,a_t)+\lambda\, r_{\mathrm{cur}}(s_t,a_t)\big)\Bigg],16 and RM J(θ)=Eτ∼πθ[∑t=0T(rRLHF(st,at)+λ rcur(st,at))],J(\theta)=\mathbb{E}_{\tau\sim\pi_\theta}\Bigg[\sum_{t=0}^{T}\big(r_{\mathrm{RLHF}}(s_t,a_t)+\lambda\, r_{\mathrm{cur}}(s_t,a_t)\big)\Bigg],17 (Sun et al., 20 Jan 2025).

6. Limitations, safety, and unresolved questions

The two CD-RLHF lines expose different limitations. In personalized dialogue, the user traits are simplified, with only two learning styles in education and fixed attribute logic in exercise. The user model is assumed to be an optimal fixed oracle, whereas real deployment would require learned or adaptive user models together with robust uncertainty quantification. The experiments rely on simulated users, and distribution shift to real-world diversity or non-cooperative behavior is explicitly identified as a limitation. The paper further notes that no special safe-RLHF components are implemented, and that privacy, transparency, and bias mitigation remain future work (Wan et al., 4 Apr 2025).

In diversity-oriented CD-RLHF, the principal limitation is scalar mismatch between intrinsic and extrinsic rewards. The paper states that this is mitigated by whitening and a small constant J(θ)=Eτ∼πθ[∑t=0T(rRLHF(st,at)+λ rcur(st,at))],J(\theta)=\mathbb{E}_{\tau\sim\pi_\theta}\Bigg[\sum_{t=0}^{T}\big(r_{\mathrm{RLHF}}(s_t,a_t)+\lambda\, r_{\mathrm{cur}}(s_t,a_t)\big)\Bigg],18, but not eliminated. It also notes that the diversity–alignment trade-off remains and that balancing still depends on J(θ)=Eτ∼πθ[∑t=0T(rRLHF(st,at)+λ rcur(st,at))],J(\theta)=\mathbb{E}_{\tau\sim\pi_\theta}\Bigg[\sum_{t=0}^{T}\big(r_{\mathrm{RLHF}}(s_t,a_t)+\lambda\, r_{\mathrm{cur}}(s_t,a_t)\big)\Bigg],19, J(θ)=Eτ∼πθ[∑t=0T(rRLHF(st,at)+λ rcur(st,at))],J(\theta)=\mathbb{E}_{\tau\sim\pi_\theta}\Bigg[\sum_{t=0}^{T}\big(r_{\mathrm{RLHF}}(s_t,a_t)+\lambda\, r_{\mathrm{cur}}(s_t,a_t)\big)\Bigg],20, and gating. Degenerate novelty signals, reward hacking, and unanticipated behaviors from increased diversity are identified as concerns; the proposed safeguards are a strong KL constraint, GPT-4 and human evaluation, top-J(θ)=Eτ∼πθ[∑t=0T(rRLHF(st,at)+λ rcur(st,at))],J(\theta)=\mathbb{E}_{\tau\sim\pi_\theta}\Bigg[\sum_{t=0}^{T}\big(r_{\mathrm{RLHF}}(s_t,a_t)+\lambda\, r_{\mathrm{cur}}(s_t,a_t)\big)\Bigg],21 gating, and keeping curiosity confined to controllable contexts (Sun et al., 20 Jan 2025).

A recurring misconception is that curiosity automatically improves the desired downstream behavior. The dialogue paper shows that entropy-based shaping can induce “controlling behavior” when the external reward is user-agnostic, while the diversity paper shows that applying curiosity too frequently can degrade alignment (Wan et al., 4 Apr 2025, Sun et al., 20 Jan 2025). Another misconception is that curiosity necessarily implies longer or more exploratory interactions in an unconstrained sense. In fact, the reported results distinguish between differential and non-differential shaping: differential curiosity in dialogue is explicitly presented as a safeguard against gratuitous prolongation, and top-J(θ)=Eτ∼πθ[∑t=0T(rRLHF(st,at)+λ rcur(st,at))],J(\theta)=\mathbb{E}_{\tau\sim\pi_\theta}\Bigg[\sum_{t=0}^{T}\big(r_{\mathrm{RLHF}}(s_t,a_t)+\lambda\, r_{\mathrm{cur}}(s_t,a_t)\big)\Bigg],22 gating in token-level generation is presented as a safeguard against unnecessary exploration (Wan et al., 4 Apr 2025, Sun et al., 20 Jan 2025).

These results place CD-RLHF at the intersection of RLHF, intrinsic motivation, active learning, and reward shaping. The dialogue formulation explicitly links its method to Bayesian experimental design and active preference elicitation; the text-generation formulation links its token-level curiosity to count-based exploration and prediction-error bonuses in RL (Wan et al., 4 Apr 2025, Sun et al., 20 Jan 2025). A plausible implication is that future CD-RLHF research will continue to diversify along the latent variable being explored—user traits, stylistic modes, or under-covered semantic regions—while retaining the common structure of RLHF augmented with carefully controlled intrinsic rewards.

Definition Search Book Streamline Icon: https://streamlinehq.com
References (2)

Topic to Video (Beta)

No one has generated a video about this topic yet.

Whiteboard

No one has generated a whiteboard explanation for this topic yet.

Follow Topic

Get notified by email when new papers are published related to Curiosity-Driven RLHF (CD-RLHF).