Reliability of LLM-based agents for end-to-end reproduction of scientific papers
Determine whether AI agents powered by large language models can reliably perform end-to-end reproduction of computational results from published scientific papers when provided only with the paper content, by autonomously extracting the methodology, implementing the described algorithms from scratch, executing the full pipeline, and generating quantitative outputs that match the original publication.
References
However, whether these agents can reliably perform end-to-end reproduction from real scientific papers remains an open question.
Importantly, the design of our study does not allow us to draw any conclusions as to whether humans prefer Faraday over Claude and Codex on average. Gathering conclusive evidence for this preference would require a larger-scale study across a randomly selected set of tasks, and is an important direction for future work.