Generalizability of harness evolution versus repeated sampling
Determine whether automatic harness evolution for large language model agents yields generalizable improvements in harness design, or whether the observed gains are primarily due to repeated sampling during the search process under matched feedback and inference budgets.
References
The existing evaluation protocol therefore leaves a fundamental question unresolved: Does harness evolution yield generalizable improvements in harness design, or are its gains primarily due to repeated sampling?
The artifacts do not identify whether the cause was optimizer choice, instruction interaction, subject-model compliance, or sampling variation.
We hypothesize that scaling both executable environments and the diversity of research ideas can yield sustained gains; we call this long-term research direction pretraining the harness.
These results show that models can make useful local improvements, while robust evolution across unseen tasks and runtime models remains an open challenge.
Our results show that a learned anchored rule is effective, but not that learning is necessary.
The scaling result stops at $B=100$: held-out success was still rising at the largest batch we could afford, so whether it keeps improving at larger batch sizes or saturates remains open.
Future work could reduce this training cost through more sample-efficient optimization or reuse of existing rollout data, and investigate whether editors and playbooks can transfer across tasks, domains, or harness-search runs.
Whether such automation improves on simpler test-time scaling is under active investigation~\citep{wang2026rethinkingevaluationharnessevolution}.