Generalization of deep research agent results to broader grounded reasoning tasks
Determine whether the results reported for state-of-the-art deep research agents—systems that conduct multi-step research on the public internet using web search tools—generalize across other grounded reasoning tasks that operate over different data regimes and requirements beyond internet-based deep research.
References
However, deep research relies on publicly available, non-proprietary, knowledge, and black-box web search tools. Thus, it is not entirely clear whether the reported state-of-the art deep research results indeed generalize across other grounded reasoning tasks.
DeepResearch Bench \citep{du-etal-2025-drb} spans 22 domains and includes questions in both English and Chinese. In our experiments, we used only English examples and leave multilingual evaluation to future work.