Assessing Meta-Harness performance across different proposer agents
Determine how the performance and effectiveness of Meta-Harness depend on the choice of the coding-agent proposer by conducting a broader, systematic study across multiple proposer agents (beyond Claude Code) and across diverse task domains to quantify variation in outcomes.
References
While we evaluate on three diverse domains, our experiments demonstrate that harness search can work with one particularly strong coding-agent proposer (Claude Code); a broader study of how the effect varies across proposer agents remains for future work.
— Meta-Harness: End-to-End Optimization of Model Harnesses
(2603.28052 - Lee et al., 30 Mar 2026) in Discussion (Section 5)
We have not tested other model families, and treat comparable-capability portability as a reasonable conjecture rather than a result.
— An Agentic Workflow for Legacy HPC Modernization: Converting the Two-Electron-Integral Core of GAMESS
(2608.12249 - Shen et al., 12 Aug 2026) in Section III.A.1, “Defining the Agent” (also discussed in Section V.C, “Limitations”), p. 4