Pluralistic alignment: satisfying diverse moral preferences
Develop methods for building artificial agents that satisfy the moral preferences of a wide range of individuals, operationalizing pluralistic alignment, for example by integrating multiple moral value signals within a single learning agent.
References
This may provide a promising direction for building agents that are able to satisfy the moral preferences of a wide range of individuals, which currently remains an open problem in alignment \citep{anwar2024foundational,ji2024aialignmentcomprehensivesurvey}.
While we acknowledge pluralistic alignment as a critical open problem of RL from human feedback (Casper et al., 2023; Sorensen et al., 2024), we view our work as addressing the complementary and largely orthogonal challenge of improving sample efficiency of preference-based reward learning.
Data visualization is comparably more open-ended, thus stressing models on divergent human expertise, whose effective integration remains an open challenge \citep{park2024rlhf,chakraborty2024maxminrlhf}.