Generalization of preference-update failures to naturalistic interactions
Determine whether the preference-update failure observed in HorizonBench persists with the same severity in naturalistic human–AI interaction.
References
Its authors state that whether the failure holds at the same severity in naturalistic interaction is unresolved, and position the benchmark as preparation ``in advance of longitudinal human data.''
— RealCompanion: Benchmarking Human Understanding from Reasoning over Longitudinal Real-World Conversations
(2610.01780 - Behnam et al., 1 Oct 2026) in Appendix, Section “Related Works,” subsection “Three Substitutions,” paragraph on authored questions and HorizonBench