Reliability of language agents in creating real economic value in professional settings
Determine whether contemporary language-model-based agents can reliably create economic value in real-world professional environments, beyond performance on exam-style benchmarks, by consistently delivering expert-level work products that meet domain constraints and professional standards.
References
As traditional benchmarks reach saturation, it remains fundamentally unclear whether today’s agents can reliably create value in economically valued, professional environments.
— \$OneMillion-Bench: How Far are Language Agents from Human Experts?
(2603.07980 - Yang et al., 9 Mar 2026) in Section 1, Introduction
The open question is whether such a component adds enough value over a deterministic script or a human-written procedure to outweigh its nondeterminism and the extra verification it requires.
— SoK: Cross-Chain Transaction Identification and Matching
(2608.17532 - Zheng et al., 18 Aug 2026) in Section 5, “Open Challenges,” subsection “Verifiable use of models and agents”