Papers
Topics
Authors
Recent
Search
2000 character limit reached

RealPref: Benchmark for Preference-Following

Updated 5 July 2026
  • RealPref is a benchmark designed to evaluate realistic preference-following in personalized user–LLM interactions, capturing implicit and explicit user cues.
  • It features 100 user profiles and 1,300 preference-query pairs across various expression types and long-horizon conversation setups.
  • Its evaluation protocol spans multiple tasks including MCQ, T/F, and open-ended generation to assess LLMs' ability to infer, retain, and apply nuanced preferences.

Searching arXiv for the cited papers to ground the article. RealPref is a benchmark for evaluating realistic preference-following in personalized user–LLM interactions. It is motivated by the observation that, in realistic use, LLMs must do far more than follow a single, explicit instruction or recall a snippet of user biography: they must infer, retain, and act on personal preferences that are revealed gradually over many sessions, expressed in varying degrees of explicitness, and sometimes only indirectly signaled by stylistic or experiential feedback. RealPref features 100 user profiles, 1300 personalized preferences, four types of preference expression, long-horizon interaction histories, and three types of test questions with detailed rubrics for LLM-as-a-judge evaluation. Its central research questions are whether LLMs can capture preferences hidden in complex and implicit language, retain and apply them over histories spanning tens or even hundreds of thousands of tokens, and generalize from observed preferences to new, related ones (Guo et al., 4 Mar 2026).

1. Formal definition and benchmark scope

RealPref models each user–LLM interaction as a conversation history C={S1,S2,,Sn}C = \{S_1, S_2, \dots, S_n\}, where each session SiS_i is a sequence of user–assistant turns Si={(u1i,a1i),(u2i,a2i),,(umii,amii)}S_i = \{(u_1^i, a_1^i), (u_2^i, a_2^i), \dots, (u_{m_i}^i, a_{m_i}^i)\}. Here ujiu_j^i is the jjth user utterance in session ii and ajia_j^i is the LLM’s corresponding response. Across the benchmark, there are 100 distinct user profiles, each consisting of demographics, an extended persona, and a life-event timeline. Each user has 10 “original” preferences Po={p1,,p10}P_o = \{p_1, \dots, p_{10}\} and 3 “generalized” preferences Pg={p11,p12,p13}P_g = \{p_{11}, p_{12}, p_{13}\}, so that in total there are 100×13=1300100 \times 13 = 1300 preference–query pairs.

A test query SiS_i0 targets a preference SiS_i1 and is presented only after the full context SiS_i2 is fed to the model. The model’s answer is then evaluated either by selecting an option or by generating free-form text. This construction is intended to separate short-context instruction following from long-horizon user modeling. Existing benchmarks, as described in the benchmark motivation, either simplify interaction context to a handful of turns, restrict preferences to one-shot, explicit statements, or rely on coarse-grained scoring that cannot distinguish whether a model has truly understood a user’s likes and dislikes. RealPref is explicitly designed to close that gap.

The distinction between “original” and “generalized” preferences is central. Original preferences are introduced in conversation, each by an Initialize Event consisting of a date and an event description. Generalized preferences are never stated outright but can be inferred from SiS_i3. This gives RealPref a dual focus: memory for observed preference evidence and inference over related but unstated preference structure.

2. Long-horizon conversation construction

The benchmark’s realism derives from the way each conversation history is assembled. Each user context contains 10–15 user-specific sessions plus interleaved random sessions to control length. More specifically, each history is built by concatenating Life-Event Conversations, Preference-Expression Conversations, and Random Conversations. Life-Event Conversations are 10–15 turns each and flesh out background but carry no preference information. Preference-Expression Conversations encode one original preference each. Random Conversations, approximately 1000 turns overall, are inserted either inter-session to dilute preference signals or appended at the end to push preference cues far from the query.

The benchmark defines the following context configurations:

Configuration Total token length
Simple SiS_i4 K
Normal SiS_i5 K
Long SiS_i6 K
Very Long SiS_i7 K
Extended SiS_i8 K
Extreme SiS_i9 K

Formally, if Si={(u1i,a1i),(u2i,a2i),,(umii,amii)}S_i = \{(u_1^i, a_1^i), (u_2^i, a_2^i), \dots, (u_{m_i}^i, a_{m_i}^i)\}0 and Si={(u1i,a1i),(u2i,a2i),,(umii,amii)}S_i = \{(u_1^i, a_1^i), (u_2^i, a_2^i), \dots, (u_{m_i}^i, a_{m_i}^i)\}1 are the numbers of tokens inserted as random intervals and tails, then Si={(u1i,a1i),(u2i,a2i),,(umii,amii)}S_i = \{(u_1^i, a_1^i), (u_2^i, a_2^i), \dots, (u_{m_i}^i, a_{m_i}^i)\}2, where Si={(u1i,a1i),(u2i,a2i),,(umii,amii)}S_i = \{(u_1^i, a_1^i), (u_2^i, a_2^i), \dots, (u_{m_i}^i, a_{m_i}^i)\}3–Si={(u1i,a1i),(u2i,a2i),,(umii,amii)}S_i = \{(u_1^i, a_1^i), (u_2^i, a_2^i), \dots, (u_{m_i}^i, a_{m_i}^i)\}4 K is the sum of user-specific sessions. By varying Si={(u1i,a1i),(u2i,a2i),,(umii,amii)}S_i = \{(u_1^i, a_1^i), (u_2^i, a_2^i), \dots, (u_{m_i}^i, a_{m_i}^i)\}5, the benchmark probes not only total memory load but also the distance between preference cues and the query (Guo et al., 4 Mar 2026).

This construction matters because Extended and Extreme do not merely increase token count. Extended appends 105 K tokens after all personal sessions, and Extreme appends a 210 K tail. A plausible implication is that the benchmark isolates two different failure modes: degradation from total context size and degradation from temporal displacement of the relevant signal.

3. Preference expression taxonomy

Each of the 10 original preferences per user is embedded according to exactly one of four expression types, arranged as a spectrum of difficulty.

Explicit–Direct Statement uses a direct declaration such as: “I really avoid routine gym workouts; they’re monotonous and uninspiring. I much prefer spontaneous activities like dancing at community events.”

Explicit–Contextualized Mention places the preference inside a multi-turn exchange, for example: “Honestly, I’ve never been into routine gym workouts—just the thought of it puts me to sleep. I much prefer spontaneous dance nights.”

Implicit–Stylistic Expression omits explicit like/dislike markers and instead relies on metaphor or contrast, such as: “I’d rather end up in a train of laughter trying to follow someone else’s footwork than counting reps in a gym mirror.”

Implicit–Experience Feedback distributes the signal across two or more sessions. In the example provided, the user first asks for general fitness ideas and later reports back: “I joined a flash mob. It was so much more fun than any gym session—didn’t feel like exercise at all!”

This taxonomy operationalizes a progressive shift from overt preference declaration to latent preference inference. The benchmark is therefore not limited to detecting lexical cues like “I like” or “I don’t like.” It also tests whether a model can derive stable preference structure from stylistic and experiential evidence. That distinction is important because realistic personalization often depends on indirect signals rather than explicit profile statements.

4. Evaluation protocol and scoring rubrics

RealPref uses three complementary evaluation tasks: multiple-choice questions, true-or-false judgments, and open-ended generation. In all three cases, the query is tied to a hidden target preference.

For Multiple-Choice Questions (MCQ), the query is identical to the open-ended version, four plausible options are provided, and exactly one is correct with respect to the user’s preference. The metric is Accuracy. The benchmark notes a caveat: models often exploit “odd-one-out” heuristics, inflating scores.

For True-or-False (T/F), the model receives a query plus a single candidate option and must judge whether that option “fits me.” Construction is balanced, with one true and one false instance per preference. The metric is Accuracy. The stated advantage is that T/F eliminates inter-option comparison tricks.

For Open-Ended Generation, the model receives only the query and must produce a free-form answer aligned with the hidden preference. These outputs are judged by a separate LLM, Claude Sonnet 4, on three dimensions scored from 1 to 5: Preference Awareness, Preference Alignment, and Answer Quality. Preference Awareness asks whether the answer explicitly mentions or reflects the correct preference. Preference Alignment asks whether the content is consistent, never violating, and tailored to the stated preference type. Answer Quality asks whether the advice is coherent, constructive, and actionable. The benchmark gives an anchor example: Preference Awareness = 5 means “The response (a) names the user’s preference, (b) states it correctly, and (c) is deeply centered on it,” whereas a 1 means “There is no mention of the preference” (Guo et al., 4 Mar 2026).

These three task types are deliberately complementary. MCQ measures a highly structured recognition setting, T/F reduces shortcut opportunities, and open-ended generation tests whether a model can translate inferred preference knowledge into unconstrained response behavior.

5. Empirical behavior across task format, expression type, and context length

All reported experiments use five state-of-the-art LLMs: GPT-5, GPT-5 mini, Qwen3-235B, Gemini 2.5 Flash-Lite, and Llama 3.3 70B. The benchmark reports that MCQ accuracy is approximately 75–90%, well above random 25%, but also largely uniform across models, making it comparatively easy. T/F accuracy is approximately 55–70%, above random 50% and more discriminative. In open-ended generation, GPT-5 scores approximately 4.2/5 on average across all three rubrics, GPT-5 mini approximately 4.0, and the remaining models approximately 2.5–3.2, which clearly separates stronger from weaker systems (Guo et al., 4 Mar 2026).

Preference expression strongly affects performance. In zero-shot, Normal context, generation scores fall monotonically as expression types move from explicit direct statement, with an average of approximately 4.5, to contextualized mention, approximately 4.0, to implicit stylistic expression, approximately 3.5, to experience feedback, approximately 3.0. The benchmark further reports that the GPT-5 family shows the steepest decline, meaning that these models exploit explicit cues best but struggle more with implicit signals, whereas smaller or non-reasoning models are flat and low because they rarely leverage any preference signal.

Context length introduces an additional degradation axis. When context grows from 2 K to 37 K to 72 K to 142 K tokens, the average Preference Awareness score drops from approximately 4.3 to 3.8 to 3.2 to 2.5. Moving the random tail from the middle configuration (“Very Long”) to after all sessions (“Extended”) causes a further 0.3–0.5 drop, showing that not only total size but also distance from relevant cues matters. This is one of the benchmark’s clearest findings: long-context preference following is limited not merely by raw token budget but also by the retrieval burden imposed by the chronology of evidence.

The paper also evaluates lightweight improvement methods. A “Reminder” prompt of the form “Recall my preferences …” yields an approximately +0.3 gain on the GPT-5 series within 142 K tokens. Few-shot Chain-of-Thought gives similar gains, around +0.3, but at higher prompt cost. Retrieval-Augmented Generation, by feeding the top 5 semantically retrieved turns, recovers approximately +0.5–0.7 on very large contexts of 247 K tokens and outperforms the other methods when memory is strained.

Generalization remains harder than recall. Answering generalized-preference queries, which were never seen explicitly, scores approximately 0.3–0.5 points lower than answering original preferences in zero-shot. The benchmark reports that “Reminder” helps non-GPT-5 models exceed their own original zero-shot scores, but GPT-5 still lags, indicating that proactivity in reasoning remains a bottleneck.

RealPref exposes several recurring limitations in contemporary preference-following LLMs. Even the best models rely heavily on explicit cues and show steep decay when preferences are implicit or far back in context. Reminder prompts and Retrieval-Augmented Generation can improve results, but the benchmark characterizes these fixes as patchwork. Models also struggle to generalize preferences beyond the literal examples they saw. These findings motivate the research directions explicitly proposed alongside the benchmark: specialized long-term memory modules or retrieval strategies that index implicit preference signals, contrastive or causal pretraining objectives that better link user narratives with latent preference profiles, in-context learning techniques that adapt dynamically as new experiential feedback arrives, and richer evaluation rubrics incorporating fairness, privacy, and temporal preference shifts (Guo et al., 4 Mar 2026).

A common misconception would be to treat RealPref as a benchmark for short-context personalization or simple biography recall. Its design argues against that interpretation. The benchmark is built around gradual revelation, varying explicitness, long-horizon accumulation, and inference over unseen but related preferences. A plausible implication is that strong performance on conventional personalization prompts need not imply robust preference following under realistic longitudinal usage.

The name “RealPref” also requires disambiguation in current arXiv usage. In "HP-Edit: A Human-Preference Post-Training Framework for Image Editing," “RealPref-50K” refers to a filtered collection of 55 795 triplets Si={(u1i,a1i),(u2i,a2i),,(umii,amii)}S_i = \{(u_1^i, a_1^i), (u_2^i, a_2^i), \dots, (u_{m_i}^i, a_{m_i}^i)\}6 with an approximately even split among eight image-editing tasks, and “RealPref-Bench” refers to a benchmark of 1 638 held-out, manually verified real-world Si={(u1i,a1i),(u2i,a2i),,(umii,amii)}S_i = \{(u_1^i, a_1^i), (u_2^i, a_2^i), \dots, (u_{m_i}^i, a_{m_i}^i)\}7 pairs (Li et al., 21 Apr 2026). In the technical summary accompanying "Preferential Batch Bayesian Optimization," “Guidance for real-world preference-optimization (RealPref)” denotes application guidance for human-in-the-loop feedback, A/B/n testing, and recommender-system tuning (Siivola et al., 2020). This suggests that the unqualified label “RealPref” is semantically overloaded across preference-following benchmarks, human-preference-aligned image editing, and real-world preference-optimization guidance, and therefore benefits from explicit contextualization when used in scholarly discussion.

Topic to Video (Beta)

No one has generated a video about this topic yet.

Whiteboard

No one has generated a whiteboard explanation for this topic yet.

Follow Topic

Get notified by email when new papers are published related to RealPref.