CD-RLHF is a family of methods that augments RLHF with intrinsic rewards to balance alignment with exploration.
The diversity-oriented approach leverages token-level curiosity to enhance output variety while maintaining model alignment.
The personalized dialogue variant uses belief-based curiosity to infer latent user traits and drive adaptive multi-turn interactions.
Searching arXiv for the specified CD-RLHF papers and closely related work to ground the article.
Curiosity-Driven Reinforcement Learning from Human Feedback (CD-RLHF) denotes RLHF variants that supplement preference-derived extrinsic rewards with intrinsic rewards that favor exploration. Current arXiv usage suggests that the label covers at least two distinct 2025 formulations: a token-level curiosity mechanism for improving output diversity in summarization and instruction-following, introduced in "Curiosity-Driven Reinforcement Learning from Human Feedback" (Sun et al., 20 Jan 2025), and a belief-based curiosity mechanism for eliciting latent user traits in personalized multi-turn dialogue, developed in "Enhancing Personalized Multi-Turn Dialogue with Curiosity Reward" (Wan et al., 4 Apr 2025). In both cases, the central idea is to preserve the optimization structure of RLHF while adding a shaped signal that rewards informative or novel trajectories.
1. Scope and conceptual variants
CD-RLHF emerged from two different problem settings. In one setting, standard RLHF is described as improving alignment to human preferences while reducing output diversity relative to supervised fine-tuning, producing repetitive or homogenized outputs; CD-RLHF addresses this by adding token-level curiosity for novel states during generation (Sun et al., 20 Jan 2025). In the other setting, multi-turn RLHF is described as prioritizing helpfulness and safety but falling short in empathetic, adaptive, and personalized interaction; CD-RLHF addresses this by rewarding reductions in uncertainty about latent user traits during dialogue (Wan et al., 4 Apr 2025).
The two formulations differ in what constitutes a “state novelty” signal and in what exploration is intended to achieve. The diversity-oriented formulation uses prediction error in a learned latent state space and applies it to tokens outside the top-k by probability. The personalized-dialogue formulation uses a belief over hidden user types and rewards actions that improve the agent’s user model (Sun et al., 20 Jan 2025, Wan et al., 4 Apr 2025).
Formulation
Intrinsic signal
Primary aim
Diversity-oriented CD-RLHF
Prediction error in an Intrinsic Curiosity Module
Optimize both output diversity and alignment quality
Personalized dialogue CD-RLHF
Changes in belief over latent user type
Reveal user preferences and adapt to them
This suggests that CD-RLHF is better understood as a family of RLHF augmentations rather than a single algorithm. What is common across variants is the addition of intrinsic motivation to a standard RLHF backbone; what differs is the latent variable being explored, the reward granularity, and the failure modes.
2. Integration with RLHF objectives
Both formulations preserve the standard separation between an extrinsic alignment signal and an auxiliary intrinsic signal. In the personalized multi-turn setting, the combined objective is reported as
with turn-level curiosity rewards added to the RLHF reward and optimized using the multi-turn RLHF training recipe of Shani et al. (2024); for Education Dialogue, the paper reports λ=9.0 (Wan et al., 4 Apr 2025).
In the diversity-oriented formulation, the combined reward is written at token level as
rt=rt(e)+ηrt(i),
where the extrinsic reward combines a sequence-level reward model score with a KL penalty to the SFT reference model,
The personalized-dialogue paper gives a stronger theoretical account through a POMDP over observable dialogue histories st and latent user type u∈U. With extended state st′=⟨st,u⟩, transition and reward are
and states, via Eck et al. (2013), that such shaping leaves the optimal policy invariant (Wan et al., 4 Apr 2025).
By contrast, the diversity-oriented paper is framed operationally rather than through PBRS. It retains the conventional SFT J(θ)=Eτ∼πθ[t=0∑T(rRLHF(st,at)+λrcur(st,at))],2 RM J(θ)=Eτ∼πθ[t=0∑T(rRLHF(st,at)+λrcur(st,at))],3 PPO pipeline, initializes policy and reference models via SFT on the same preference dataset, trains the reward model on paired preferences, and inserts an Intrinsic Curiosity Module during PPO rollouts (Sun et al., 20 Jan 2025).
3. Belief-based CD-RLHF for personalized multi-turn dialogue
The personalized formulation targets latent user inference during online interaction. The agent does not assume pre-existing user profiles or histories; instead, it must infer the hidden user type from observed dialogue. Two environments are reported. In Education Dialogue, the latent trait is the student’s preferred learning style, described in prompts and analyses as “hands-on” versus “story-telling,” and earlier as “lecture-based” versus “hands-on.” In Exercise Recommendation, the user profile has 20 attributes, of which five are relevant to the ideal strategy: age/injury status, extroversion, motivation, outdoor preference, and socioeconomic status (Wan et al., 4 Apr 2025).
In practice, beliefs are not updated by an explicit Bayesian transition model. They are set by a fixed oracle user model:
For Education Dialogue, the oracle classifier is a Gemma 7B model that predicts preferred learning style from the conversation. For Exercise Recommendation, the oracle is a decision-tree-like classifier, with an LLM-based extractor when relevant attributes are implicit (Wan et al., 4 Apr 2025).
The intrinsic rewards are defined directly over changes in these beliefs. The paper reports potential functions requiring the ground-truth user type J(θ)=Eτ∼πθ[t=0∑T(rRLHF(st,at)+λrcur(st,at))],5 in simulated settings,
The policy action space is the full natural-language response at turn λ=9.05. There is no hard-coded controller. The policy may progress the task or elicit user traits by asking clarifying questions or proposing preference-sensitive options. Differential curiosity rewards are reported to encourage early elicitation without artificially prolonging conversations, whereas non-differential rewards can incentivize longer dialogues (Wan et al., 4 Apr 2025).
The environments are fully simulated. Education Dialogue uses a Gemma 2B student simulator fine-tuned via supervised learning and a Gemma 2B reward model from Shani et al. (2024). Exercise Recommendation uses Gemini 1.5 Pro for detailed backstory generation consistent with sampled attributes; users are simulated by an LLM environment that stays consistent with those backstories. The same shaping mechanism is used in both domains, which the paper presents as domain-agnostic personalization (Wan et al., 4 Apr 2025).
4. Token-level CD-RLHF for diversity-preserving alignment
The diversity-oriented formulation adapts curiosity-driven exploration from RL to autoregressive text generation. The state at time λ=9.06 is the prompt plus the partial generation, with λ=9.07 and λ=9.08, the action is token λ=9.09, and the transition is rt=rt(e)+ηrt(i),0 (Sun et al., 20 Jan 2025).
Its intrinsic reward is produced by an Intrinsic Curiosity Module consisting of a feature encoder rt=rt(e)+ηrt(i),1 and a forward model rt=rt(e)+ηrt(i),2, both two-layer MLPs. The forward model predicts the next-state feature,
Curiosity is gated by token probability. If rt=rt(e)+ηrt(i),5 denotes the top-rt=rt(e)+ηrt(i),6 tokens by probability, then
rt=rt(e)+ηrt(i),7
The intrinsic reward is then whitened within each trajectory as
rt=rt(e)+ηrt(i),8
where rt=rt(e)+ηrt(i),9 and r(e)=R−βDKL(πpolicy(⋅∣st)∥πref(⋅∣st)),0 are the mean and standard deviation of r(e)=R−βDKL(πpolicy(⋅∣st)∥πref(⋅∣st)),1 over the trajectory (Sun et al., 20 Jan 2025).
The representation design is specific to LLMs. The method uses the last-layer hidden states of the reference model to represent r(e)=R−βDKL(πpolicy(⋅∣st)∥πref(⋅∣st)),2 and r(e)=R−βDKL(πpolicy(⋅∣st)∥πref(⋅∣st)),3, and the actor model’s token embedding to represent r(e)=R−βDKL(πpolicy(⋅∣st)∥πref(⋅∣st)),4. Novelty is local rather than global: there is no memory bank, and the curiosity reward depends on the prediction error of the immediate next-state feature given the current state and action (Sun et al., 20 Jan 2025).
The paper keeps the curiosity coefficient r(e)=R−βDKL(πpolicy(⋅∣st)∥πref(⋅∣st)),5 small and constant to avoid intrinsic rewards overwhelming extrinsic rewards. Reported values are r(e)=R−βDKL(πpolicy(⋅∣st)∥πref(⋅∣st)),6 for Gemma-2B/7B, r(e)=R−βDKL(πpolicy(⋅∣st)∥πref(⋅∣st)),7 for Llama-3.2-1B TL;DR, r(e)=R−βDKL(πpolicy(⋅∣st)∥πref(⋅∣st)),8 for Llama-3.2-1B UltraFeedback, and r(e)=R−βDKL(πpolicy(⋅∣st)∥πref(⋅∣st)),9 for Llama-3.2-3B. Top-st0 gating uses st1 by default, giving intrinsic rewards at approximately st2 of positions, with ablations for st3. PPO is implemented with DeepSpeed-Chat; key hyperparameters include PPO batch size st4, PPO epochs st5, rollout per step st6, clip ratio st7, GAEst8, and st9 (Sun et al., 20 Jan 2025).
The reported training datasets are TL;DR summarization and UltraFeedback instruction following. Data are split across SFT, reward-model training, and PPO in a u∈U0 partition. Evaluation further includes out-of-distribution testing on MT-Bench and extended story generation with ROC Stories prompts (Sun et al., 20 Jan 2025).
5. Empirical findings
In personalized multi-turn dialogue, the main reported metric is pairwise win rate judged by Gemini 1.5 Pro. Against the multi-turn RLHF baseline, DiffAcc achieves u∈U1 wins, Acc u∈U2, DiffLogAcc u∈U3, and InfoGainu∈U4; the Entropy variant averages u∈U5 and shows type-specific asymmetry (Wan et al., 4 Apr 2025). For overall conversation quality, DiffLogAcc achieves u∈U6 wins against baseline, while Baseline versus DiffAcc is u∈U7, indicating that DiffAcc slightly hurts overall quality relative to the baseline. The paper also reports that the differential-accuracy model learns to ask about user type at turn 1 early, from approximately u∈U8k steps, producing higher u∈U9 than the baseline, whose accuracy improves only when the student explicitly states preferences (Wan et al., 4 Apr 2025).
A central qualitative result in the dialogue setting concerns the distinction between grounded and ungrounded curiosity. Entropy reward performs best on the “hands-on” type but worst on “story-telling,” and this is attributed to “controlling behavior,” in which the policy drives the classifier toward one type rather than grounding in user statements. Accuracy-based rewards are described as grounded and as avoiding this pathology. Non-differential rewards such as Acc are also reported to lengthen conversations, harming quality, whereas differential rewards such as DiffLogAcc avoid prolongation (Wan et al., 4 Apr 2025).
In the diversity-oriented setting, CD-RLHF is reported to improve diversity metrics over RLHF while preserving RM-based alignment. On TL;DR, representative results include Diversity st′=⟨st,u⟩0, EADst′=⟨st,u⟩1, Self-BLEU st′=⟨st,u⟩2, and SentBERT st′=⟨st,u⟩3 for Gemma-2B versus RLHF, with RM score st′=⟨st,u⟩4 versus st′=⟨st,u⟩5; for Llama-3.2-1B, Diversity is st′=⟨st,u⟩6 and RM score is st′=⟨st,u⟩7 versus st′=⟨st,u⟩8; for Llama-3.2-3B, Diversity is st′=⟨st,u⟩9 and RM score improves from T′(st+1′∣st′,at)=T(st+1∣st,at,u),R′(st′,at)=R(st,at∣u),0 to T′(st+1′∣st′,at)=T(st+1∣st,at,u),R′(st′,at)=R(st,at∣u),1 (Sun et al., 20 Jan 2025). On UltraFeedback, representative gains include Diversity T′(st+1′∣st′,at)=T(st+1∣st,at,u),R′(st′,at)=R(st,at∣u),2 and EAD T′(st+1′∣st′,at)=T(st+1∣st,at,u),R′(st′,at)=R(st,at∣u),3 for Gemma-2B, with RM score T′(st+1′∣st′,at)=T(st+1∣st,at,u),R′(st′,at)=R(st,at∣u),4 versus T′(st+1′∣st′,at)=T(st+1∣st,at,u),R′(st′,at)=R(st,at∣u),5, and Diversity T′(st+1′∣st′,at)=T(st+1∣st,at,u),R′(st′,at)=R(st,at∣u),6 for Llama-3.2-3B, with RM score T′(st+1′∣st′,at)=T(st+1∣st,at,u),R′(st′,at)=R(st,at∣u),7 versus T′(st+1′∣st′,at)=T(st+1∣st,at,u),R′(st′,at)=R(st,at∣u),8 (Sun et al., 20 Jan 2025).
External evaluation is also reported. TL;DR GPT-4 win rates in favor of CD-RLHF are T′(st+1′∣st′,at)=T(st+1∣st,at,u),R′(st′,at)=R(st,at∣u),9 for Gemma-2B, bt+1(u)∝T(st+1∣st,at,u)bt(u).0 for Gemma-7B, bt+1(u)∝T(st+1∣st,at,u)bt(u).1 for Llama-3.2-1B, and bt+1(u)∝T(st+1∣st,at,u)bt(u).2 for Llama-3.2-3B; UltraFeedback averages approximately bt+1(u)∝T(st+1∣st,at,u)bt(u).3 GPT-4 win rates across models, and human evaluations consistently prefer CD-RLHF on diversity (Sun et al., 20 Jan 2025). On MT-Bench, overall GPT-4 scores improve from bt+1(u)∝T(st+1∣st,at,u)bt(u).4 to bt+1(u)∝T(st+1∣st,at,u)bt(u).5 for Gemma-2B, from bt+1(u)∝T(st+1∣st,at,u)bt(u).6 to bt+1(u)∝T(st+1∣st,at,u)bt(u).7 for Gemma-7B, from bt+1(u)∝T(st+1∣st,at,u)bt(u).8 to bt+1(u)∝T(st+1∣st,at,u)bt(u).9 for Llama-3.2-1B, and from J(θ)=Eτ∼πθ[t=0∑T(rRLHF(st,at)+λrcur(st,at))],00 to J(θ)=Eτ∼πθ[t=0∑T(rRLHF(st,at)+λrcur(st,at))],01 for Llama-3.2-3B (Sun et al., 20 Jan 2025).
Ablations in the diversity-oriented paper emphasize that reward frequency and KL strength remain consequential. Expanding intrinsic rewards from approximately J(θ)=Eτ∼πθ[t=0∑T(rRLHF(st,at)+λrcur(st,at))],02 of positions to J(θ)=Eτ∼πθ[t=0∑T(rRLHF(st,at)+λrcur(st,at))],03–J(θ)=Eτ∼πθ[t=0∑T(rRLHF(st,at)+λrcur(st,at))],04 yields small additional diversity gains but degrades alignment; for TL;DR Gemma-2B, RM score drops from J(θ)=Eτ∼πθ[t=0∑T(rRLHF(st,at)+λrcur(st,at))],05 to J(θ)=Eτ∼πθ[t=0∑T(rRLHF(st,at)+λrcur(st,at))],06 at J(θ)=Eτ∼πθ[t=0∑T(rRLHF(st,at)+λrcur(st,at))],07. For top-J(θ)=Eτ∼πθ[t=0∑T(rRLHF(st,at)+λrcur(st,at))],08 gating on TL;DR Gemma-2B, the paper reports: J(θ)=Eτ∼πθ[t=0∑T(rRLHF(st,at)+λrcur(st,at))],09, Diversity J(θ)=Eτ∼πθ[t=0∑T(rRLHF(st,at)+λrcur(st,at))],10 and RM J(θ)=Eτ∼πθ[t=0∑T(rRLHF(st,at)+λrcur(st,at))],11; J(θ)=Eτ∼πθ[t=0∑T(rRLHF(st,at)+λrcur(st,at))],12, Diversity J(θ)=Eτ∼πθ[t=0∑T(rRLHF(st,at)+λrcur(st,at))],13 and RM J(θ)=Eτ∼πθ[t=0∑T(rRLHF(st,at)+λrcur(st,at))],14; J(θ)=Eτ∼πθ[t=0∑T(rRLHF(st,at)+λrcur(st,at))],15, Diversity J(θ)=Eτ∼πθ[t=0∑T(rRLHF(st,at)+λrcur(st,at))],16 and RM J(θ)=Eτ∼πθ[t=0∑T(rRLHF(st,at)+λrcur(st,at))],17 (Sun et al., 20 Jan 2025).
6. Limitations, safety, and unresolved questions
The two CD-RLHF lines expose different limitations. In personalized dialogue, the user traits are simplified, with only two learning styles in education and fixed attribute logic in exercise. The user model is assumed to be an optimal fixed oracle, whereas real deployment would require learned or adaptive user models together with robust uncertainty quantification. The experiments rely on simulated users, and distribution shift to real-world diversity or non-cooperative behavior is explicitly identified as a limitation. The paper further notes that no special safe-RLHF components are implemented, and that privacy, transparency, and bias mitigation remain future work (Wan et al., 4 Apr 2025).
In diversity-oriented CD-RLHF, the principal limitation is scalar mismatch between intrinsic and extrinsic rewards. The paper states that this is mitigated by whitening and a small constant J(θ)=Eτ∼πθ[t=0∑T(rRLHF(st,at)+λrcur(st,at))],18, but not eliminated. It also notes that the diversity–alignment trade-off remains and that balancing still depends on J(θ)=Eτ∼πθ[t=0∑T(rRLHF(st,at)+λrcur(st,at))],19, J(θ)=Eτ∼πθ[t=0∑T(rRLHF(st,at)+λrcur(st,at))],20, and gating. Degenerate novelty signals, reward hacking, and unanticipated behaviors from increased diversity are identified as concerns; the proposed safeguards are a strong KL constraint, GPT-4 and human evaluation, top-J(θ)=Eτ∼πθ[t=0∑T(rRLHF(st,at)+λrcur(st,at))],21 gating, and keeping curiosity confined to controllable contexts (Sun et al., 20 Jan 2025).
A recurring misconception is that curiosity automatically improves the desired downstream behavior. The dialogue paper shows that entropy-based shaping can induce “controlling behavior” when the external reward is user-agnostic, while the diversity paper shows that applying curiosity too frequently can degrade alignment (Wan et al., 4 Apr 2025, Sun et al., 20 Jan 2025). Another misconception is that curiosity necessarily implies longer or more exploratory interactions in an unconstrained sense. In fact, the reported results distinguish between differential and non-differential shaping: differential curiosity in dialogue is explicitly presented as a safeguard against gratuitous prolongation, and top-J(θ)=Eτ∼πθ[t=0∑T(rRLHF(st,at)+λrcur(st,at))],22 gating in token-level generation is presented as a safeguard against unnecessary exploration (Wan et al., 4 Apr 2025, Sun et al., 20 Jan 2025).
These results place CD-RLHF at the intersection of RLHF, intrinsic motivation, active learning, and reward shaping. The dialogue formulation explicitly links its method to Bayesian experimental design and active preference elicitation; the text-generation formulation links its token-level curiosity to count-based exploration and prediction-error bonuses in RL (Wan et al., 4 Apr 2025, Sun et al., 20 Jan 2025). A plausible implication is that future CD-RLHF research will continue to diversify along the latent variable being explored—user traits, stylistic modes, or under-covered semantic regions—while retaining the common structure of RLHF augmented with carefully controlled intrinsic rewards.