Pluralistic alignment: satisfying diverse moral preferences

Develop methods for building artificial agents that satisfy the moral preferences of a wide range of individuals, operationalizing pluralistic alignment, for example by integrating multiple moral value signals within a single learning agent.

Background

The authors propose intrinsic moral rewards for aligning LLM agents and note that multi-objective formulations could encode multiple moral values within a single agent.

They explicitly state that building agents capable of satisfying a broad diversity of moral preferences remains an open problem in alignment, suggesting their approach as a possible direction.

References

This may provide a promising direction for building agents that are able to satisfy the moral preferences of a wide range of individuals, which currently remains an open problem in alignment \citep{anwar2024foundational,ji2024aialignmentcomprehensivesurvey}.

Moral Alignment for LLM Agents  (2410.01639 - Tennant et al., 2024) in Section 6 (Discussion)

While we acknowledge pluralistic alignment as a critical open problem of RL from human feedback (Casper et al., 2023; Sorensen et al., 2024), we view our work as addressing the complementary and largely orthogonal challenge of improving sample efficiency of preference-based reward learning.

Subspace Inference Enables Efficient Active Reward Learning from Preferences  (2609.04066 - Zhou et al., 3 Sep 2026) in Section 5.1, subsection “On the unimodality of EKF”

Data visualization is comparably more open-ended, thus stressing models on divergent human expertise, whose effective integration remains an open challenge \citep{park2024rlhf,chakraborty2024maxminrlhf}.

Efficient Test-Time Adaptation through Human-AI Interaction  (2609.04141 - Wang et al., 3 Sep 2026) in Section 4.4, “Exploration: Deriving Community Expertise from Individual Adaptations”