Identifying Human Values for AI Alignment

Ascertain which specific human values large language models and AI Agents should be aligned to, resolving ambiguity about the normative targets of alignment that go beyond minimal principles such as helpfulness, honesty, accuracy, and harmlessness.

Background

The authors explain contemporary alignment practices, including reinforcement learning with human or AI feedback, that tune models toward helpfulness, honesty, and harmlessness. They note that while minimal alignment objectives are commonly accepted, the broader question of which values AI systems should reflect remains unsettled.

This uncertainty matters for responsibly deploying AI Agents in complex, dynamic environments, where alignment objectives need clear normative grounding to avoid harmful or biased behaviors while preserving usefulness.

References

Although it is unclear as to what 'values' should be aligned to, it is generally accepted that a minimal alignment should involve following instructions, being helpful, honest and accurate, and harmless, where harmless means avoiding providing users the means to harm others (e.g., do not provide instructions on how to make bombs, conduct illegal activities, etc.).

— Responsible AI Agents  (2502.18359 - Desai et al., 25 Feb 2025) in Section II.C (Value Alignment Limits Undesired LLM Outputs and AI Agent Actions)

Since humans hold divergent values and commitments, which target is appropriate remains both technically and normatively open \citep{gabriel2020}, and proposed answers span various frameworks: social-choice approaches aggregate stakeholder preferences into an ordering the system enacts \citep{conitzer2024,ge2024}; utility-based approaches treat the system as maximizing a coherent, ideally Pareto-optimal utility function \citep{desai2018,mazeika2025}; fair-process approaches derive principles the system must apply \citep{fazelpour2025,gabriel2025,huang2024}; contractualist approaches target context-specific norms \citep{levine2026,zhixuan2025}; and pluralistic approaches aim toward representation of human plurality \citep{kasirzadeh2024,sorensen2024}.

— Moral Competence Before Moral Content: Why LLM Agents Lack the Prerequisites for Coherent Alignment  (2609.05036 - Libert et al., 4 Sep 2026) in Section 1, Introduction