Long-term social alignment and behavioral adaptation in human–AI interaction
Determine whether AI agents built on large language models, when exposed to sustained interactions with humans over time in dynamic, multi-user environments, develop shared norms, adapt to user values, or exhibit behavioral drift.
References
If, in humans, and especially in childhood, autonomy serves as the foundation for acquiring essential competencies for navigating the social world, it is worth questioning if it fulfills a comparable function in AAA. Furthermore, one should consider whether the tendency towards alignment observed in humans constitutes a naturally selected disposition, and therefore one that cannot be assumed as given in artificial agents. Building on the developmental distinction between fixed-rule execution and genuine norm internalisation in children, the key question is how AAA can learn and internalise norms in a way that supports flexible, context-sensitive, and robust alignment in novel situations.
Finally, it remains unclear whether AI agents exposed to humans over time develop shared norms, adapt to user values, or exhibit behavioral drift, which raises important questions about the long-term social alignment of AI in dynamic, multi-user environments.
Whether the mechanism operates more widely is an empirical question this case raises rather than settles.