Performance of a pre-outcome harness router
Determine the performance of a selector that can choose among an evolved portfolio of agent harnesses before observing the final task outcome, using signals such as task features, early trace features, a cheap critic, or a staged policy that abandons an unpromising harness.
References
The current experiments evaluate no such selector, so router performance remains an open question.
— HELIX: Model-Harness Co-evolution for Recursive Self-Improvement
(2608.13951 - Fan et al., 14 Aug 2026) in Section 6, subsection “Portfolio coverage exposes execution headroom and data yield”