Usefulness of LLM-generated hypotheses and ideas for scientific discovery
Determine whether hypotheses and research ideas generated by large language models are useful in practice and lead to new scientific discoveries, given their typically theoretical nature and the costly validation required to justify them.
References
Additionally, given that hypotheses and ideas are typically theoretical and cannot be validated without costly justification, it is unclear whether generated hypotheses and ideas are truly useful and lead to new scientific discoveries.
— Transforming Science with Large Language Models: A Survey on AI-assisted Scientific Discovery, Experimentation, Content Generation, and Evaluation
(2502.05151 - Eger et al., 7 Feb 2025) in Subsection "Limitations and future directions", Designing and conducting experiments; AI-based discovery (Section \ref{sec:experiments})
Directly evaluating such claims is challenging since determining whether a research idea is genuinely useful remains an open problem and lacks reliable evaluation metrics (Si et al., 2025; 2026).
— FormalTCS: Benchmarking End-to-End Frontier Formal Theoretical Computer Science Research of Large Language Models
(2608.20153 - Wang et al., 20 Aug 2026) in Section 5, opening paragraph, page 5