Identifying Human Values for AI Alignment
Ascertain which specific human values large language models and AI Agents should be aligned to, resolving ambiguity about the normative targets of alignment that go beyond minimal principles such as helpfulness, honesty, accuracy, and harmlessness.
References
Although it is unclear as to what 'values' should be aligned to, it is generally accepted that a minimal alignment should involve following instructions, being helpful, honest and accurate, and harmless, where harmless means avoiding providing users the means to harm others (e.g., do not provide instructions on how to make bombs, conduct illegal activities, etc.).
Since humans hold divergent values and commitments, which target is appropriate remains both technically and normatively open \citep{gabriel2020}, and proposed answers span various frameworks: social-choice approaches aggregate stakeholder preferences into an ordering the system enacts \citep{conitzer2024,ge2024}; utility-based approaches treat the system as maximizing a coherent, ideally Pareto-optimal utility function \citep{desai2018,mazeika2025}; fair-process approaches derive principles the system must apply \citep{fazelpour2025,gabriel2025,huang2024}; contractualist approaches target context-specific norms \citep{levine2026,zhixuan2025}; and pluralistic approaches aim toward representation of human plurality \citep{kasirzadeh2024,sorensen2024}.