Value-based preferences beyond GPT-3.5/4
Determine whether large language model species other than GPT-3.5 and GPT-4 exhibit value-based strategic preferences when strategies are associated with explicit numerical payoffs.
References
Although past work has shown that GPT-3.5 and GPT-4 have preferences for higher valued strategies in a dictator game , it is not clear whether other model species possess similar preferences.
— Do Large Language Models Learn Human-Like Strategic Preferences?
(2404.08710 - Roberts et al., 2024) in Section 3, Do LLMs Prefer Strategies Based on Value?, opening paragraph
The represented payoff is the one factor these experiments could not settle: its representation checks failed in the conditions that would license the reading, so it is reported without that interpretation.
— From Deceptive Outputs to Deceptive Mechanisms: A Causal Framework for Language-Model Deception Research
(2609.04166 - Shkolnikov, 3 Sep 2026) in Section 3, subsection “What do these interventions establish?”