Robust LLM reasoning under internally discordant evidence
Determine whether contemporary large language model systems can reason robustly in scenarios where the available evidence is internally discordant, i.e., where different evidence sources conflict with each other.
References
As a result, it remains unclear whether current LLM systems can reason robustly when the available evidence is internally discordant.
Correct cues have high valid-response accuracy, but the study does not settle when deference is rational.
What constitutes a decision-relevant consequence? Safety-critical decisions involve physical harm, social interaction, comfort, rule compliance, task completion, and recoverability. These dimensions cannot always be reduced to a single scalar risk score. A key challenge is to determine which consequential differences are sufficient to change an action and how heterogeneous forms of evidence and constraints should be combined.