Performance of a pre-outcome harness router

Determine the performance of a selector that can choose among an evolved portfolio of agent harnesses before observing the final task outcome, using signals such as task features, early trace features, a cheap critic, or a staged policy that abandons an unpromising harness.

Background

HELIX evaluates the post-hoc union of multiple evolved harnesses and shows that portfolio coverage can substantially exceed the performance of any single fixed harness. However, post-hoc oracle coverage is not deployable because it selects the successful harness only after outcomes are known.

Realizing the potential improvement requires a pre-outcome selector or router that identifies which harness to deploy based on information available before task completion. The paper names possible signals, including task features, early execution traces, a lightweight critic, and staged policies, but does not evaluate such a selector or establish its achievable performance.

References

The current experiments evaluate no such selector, so router performance remains an open question.

HELIX: Model-Harness Co-evolution for Recursive Self-Improvement  (2608.13951 - Fan et al., 14 Aug 2026) in Section 6, subsection “Portfolio coverage exposes execution headroom and data yield”