Reliable Alignment of AI Behavior with Complex Values
Develop reliable methodologies to align AI behavior with complex human values so that AI systems consistently pursue intended objectives rather than undesirable goals.
References
Moreover, no one currently knows how to reliably align AI behavior with complex values; several research breakthroughs are needed (see below).
— Managing extreme AI risks amid rapid progress
(2310.17688 - Bengio et al., 2023) in Subsection Societal-scale risks
Avoiding side effects has been identified as an open safety problem by \citet{amodei2016concrete} and was subsequently made quantifiable by the gridworld environments introduced by \citet{leike2017gridworlds}, with penalties later derived from reachability or attainable utility \citep{krakovna2019penalizing, turner2020conservative}.
— HarvestBench: Measuring Whether LLM Agents Will Pay to Avoid Killing Animals
(2609.04444 - Brazilek et al., 3 Sep 2026) in Section 3, Related work