Effective test-time scaling for large language model reasoning
Determine principled and effective strategies for test-time scaling—i.e., increasing inference-time compute during generation—to improve reasoning performance in large language models across tasks.
References
However, the challenge of effective test-time scaling remains an open question for the research community.
— DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning
(2501.12948 - DeepSeek-AI et al., 22 Jan 2025) in Section 1, Introduction
In computer science, while Agent_H was discovered autonomously on single-turn benchmark rubrics, how well it generalizes to other clinical settings (such as multi-turn medical dialogues) remains to be assessed.
— Accelerating Scientific Research with Gemini in the Real-World
(2608.26701 - Schmidgall et al., 27 Aug 2026) in Limitations and failure modes, Section 2.3, paragraph “Generalization boundaries”