Generalization of preference-update failures to naturalistic interactions

Determine whether the preference-update failure observed in HorizonBench persists with the same severity in naturalistic human–AI interaction.

Background

HorizonBench evaluates whether conversational systems update user preferences as those preferences evolve, but it uses simulated users and constructed interaction histories. The paper notes that the reported tendency to select pre-evolution values may not transfer directly to conversations arising from ordinary, longitudinal human–AI interaction.

Resolving this problem would establish whether failures identified by synthetic longitudinal benchmarks reflect a genuine property of deployed systems or an artifact of benchmark construction.

References

Its authors state that whether the failure holds at the same severity in naturalistic interaction is unresolved, and position the benchmark as preparation ``in advance of longitudinal human data.''

— RealCompanion: Benchmarking Human Understanding from Reasoning over Longitudinal Real-World Conversations  (2610.01780 - Behnam et al., 1 Oct 2026) in Appendix, Section “Related Works,” subsection “Three Substitutions,” paragraph on authored questions and HorizonBench