Control the effect of model capability across agent roles

Isolate the effect of model capability on the dominant error type in deep research systems by running the same system with the same language model across different agent roles.

Background

The paper observes that the dominant error type appears to track the capability of the LLM assigned to an agent, but this conclusion is observational rather than based on a controlled comparison. The evaluated systems use different models at different stages, making it difficult to distinguish the effect of an agent’s position in the multi-agent architecture from the effect of model capability.

A controlled study using the same deep research system and the same LLM across all agent roles would address this confounding factor and establish whether model capability, architectural depth, or their interaction determines the observed error distribution.

References

Our observation that the dominant error type tracks the capability of the model an agent runs (\cref{sec:agent_analysis,sec:tracing_algo_results}) is observational rather than controlled. Isolating the factor would require running the same system with the same model across roles, which we leave to future work.

Who is the Agent to Blame? Localizing Faithfulness and Citation Mistakes in Agentic Deep Research  (2608.24306 - Hirsch et al., 25 Aug 2026) in Limitations, paragraph “Attributing error types to model capability”