Systematic Exploration of Admissible Hypothesis Sets by LLMs under Underdetermination
Determine whether large language models can systematically explore the full set of admissible explanatory hypotheses consistent with a fixed set of observations in controlled underdetermination settings, by generating multiple non-redundant, valid hypotheses rather than converging on a small subset or a single answer.
References
Yet contemporary LLM evaluations largely reward one-shot correctness~\citep{shojaee2025llm,koblischke2025gravity,shojaee2024llm,wang2024mmlu,coignion2024performance,hendrycks2020measuring}, leaving open whether models can systematically explore sets of valid explanations under controlled underdetermination.
\citet{salimi2026wiring}, synthesizing more than sixty studies across eleven model families, report a similar pattern and identify ``limited mechanistic understanding of abductive processes'' as an open problem.