Extent and Modulation of LLM Behavioral Consistency with Human Decision-Making

Determine the extent to which large language models exhibit behavior consistent with human decision-making, and ascertain whether their behavior can be modulated through targeted interventions.

Background

The paper examines whether LLMs can replicate core properties of human decision-making, particularly the stochastic variability and adaptive behavior observed in dynamic tasks. Although LLMs often match or exceed human performance on standard reasoning benchmarks, their ability to reproduce human-like noise and process-level dynamics is uncertain.

To address this, the authors propose a process-oriented evaluation framework with progressive interventions (Intrinsicality, Instruction, Imitation) and validate it on two classic economic tasks (second-price auction and newsvendor problem). The open question frames the central goal: assessing behavioral fidelity and the potential to modulate LLM behavior through targeted interventions.

References

While LLMs now match or surpass human accuracy on standard reasoning benchmarks , their ability to reproduce these stochastic patterns remains an open question:

To what extent do LLMs exhibit behavior consistent with human decision-making, and can this behavior be modulated through targeted interventions?

Noise, Adaptation, and Strategy: Assessing LLM Fidelity in Decision-Making  (2508.15926 - Feng et al., 21 Aug 2025) in Introduction (Section 1), immediately preceding and including the boxed question

Whether summarization bias survives Stage 3 is genuinely open.

The findings are also strikingly mixed: some studies report that models reproduce a wide range of human capabilities, biases, and classic experimental effects, whereas others find that they diverge substantially from human behavior, leaving the overall degree of alignment unresolved and the field polarized.

CogGym: Towards Large-Scale Comparative Evaluation of Human and Machine Cognition  (2609.21259 - Ying et al., 18 Sep 2026) in Introduction, paragraph beginning “To address these limitations”; Section 1, paragraph beginning “A natural pathway to evaluate intelligence in machines”