- The paper introduces harness policy learning, using cost-aware GRPO to teach agents when to track beliefs, commit progress, recall experience, and record notes during long-horizon tasks.
- EvoHarness-RL raises Qwen3-8B performance on ALFWorld from 47.9% with ReAct to 96.9% on seen tasks and 86.6% on unseen tasks, matching or exceeding larger-model baselines.
- The results show that reinforcement learning selectively reduces routine harness calls while evolving a compact, revisable experience store, though generalization beyond ALFWorld remains untested.
EvoHarness-RL addresses a specific gap in long-horizon LLM agent design: while external harness components (memories, state trackers, skill libraries, verifiers) are increasingly common, the runtime policy governing when an agent reads, writes, or consolidates that external state is typically fixed by prompts or heuristics. The paper proposes treating this coordination as a learnable policy problem, and demonstrates that doing so yields large gains on ALFWorld with a modest 8B backbone (2608.05446).
The authors define harness policy learning as learning, offline, a policy that constructs and updates external harness state online during task execution. They abstract heterogeneous harness components into three policy-facing functional roles:
- Belief: persistent estimates of environment state (object states, locations, relations), so the policy need not rely solely on context-window memory.
- Progress: subgoal-status records (gi,σi) externalizing what has been attempted, what remains open, and where execution is blocked.
- Experience: cross-episode knowledge — skills, failure modes, search priors, strategies.
The agent interacts with this workspace through four compact meta-actions: track (read belief), commit (write progress), recall (retrieve experience), and note (record new insights). Crucially, harness actions consume the same interaction budget as environment actions, so the value of external-state access must be weighed against its step cost. This cost coupling is what motivates RL rather than prompt-time conventions.
An environment adapter grounds BPE in domain-specific implementations. In ALFWorld, Belief is a rule-based world-state tracker queried selectively via track[object] or track[world]; Progress is a bounded committed-subgoal list; Experience is a skill bank organized into general skills, task-specific skills, common mistakes, and object-location priors, retrieved by keyword overlap with LFU eviction. A consolidation model merges buffered notes and trajectory summaries into the store at epoch boundaries via add/update/remove operations.
Two-stage training recipe
Training proceeds in two stages on Qwen3-8B. First, supervised harness fine-tuning uses successful teacher trajectories collected with the same BPE interface (Claude Opus teacher; 87 trajectories yielding 1,153 next-action pairs from 500 training games) to teach both task-solving behavior and the semantics of when to track, commit, recall, or note. Second, cost-aware GRPO optimizes the SFT-initialized policy with a reward combining a dominant sparse success signal (Rsucc=10⋅1[solved]), a success-gated efficiency bonus, a time-annealed action-diversity bonus that encourages broad harness exploration early and decays toward specialization, and fixed penalties for degenerate repetition ("spam") and malformed actions.
Main results
On the ALFWorld seen split (140 tasks across six families), EvoHarness-RL reaches 96.9% average success, a +49.0 absolute improvement over the ReAct baseline for the same Qwen3-8B backbone, decisively outperforming frozen memory methods (ExpeL, ReasoningBank, MemP, Dynamic Cheatsheet, ACE) and trainable baselines including SkillOS (80.2%) and SkillRL (89.9%). The staged progression — 56.4% (prompt-time harness) → 68.6% (SFT) → 96.9% (GRPO) — supports the central claim that optimization, not merely scaffolding, converts the harness into an effective decision interface. Notably, the trained 8B model matches Claude Opus 4.5's ReAct performance (96.4%).
| Variant |
Avg SR (%) |
Δ over ReAct |
| ReAct (Qwen3-8B) |
47.9 |
— |
| EvoHarness-Base (prompt-time) |
56.4 |
+8.5 |
| EvoHarness-SFT |
68.6 |
+20.7 |
| EvoHarness-RL |
96.9 |
+49.0 |
| SkillOS (Qwen3-8B) |
80.2 |
+32.3 |
| SkillRL (Qwen2.5-7B) |
89.9 |
+42.0 |
The prompt-time BPE harness also transfers across model scales without any training: it lifts GPT-4.1 by +22.1 points (47.9% → 70.0%) and GPT-5 by +25.7 points (60.7% → 85.0%), and pushes already-strong Claude Opus 4.5 to 98.5%. This indicates that explicit belief/progress/experience externalization benefits even frontier policies, independent of the learned-coordination contribution.
Ablations confirm all three BPE components contribute synergistically at inference time: removing Experience yields the lowest average (48.6%), removing Belief disproportionately harms Clean and Cool tasks requiring localization and state verification, and removing Progress degrades Pick2 tasks with dependent subgoals.
On the unseen split, the full RL policy reaches 86.6% versus 50.0% for ReAct. An instructive negative result appears here: EvoHarness-SFT drops to 69.4%, below even the prompt-time harness (77.6%), because imitation of teacher harness-use patterns does not calibrate access costs in novel environments. Cost-aware GRPO corrects this, suggesting the RL stage learns a transferable access strategy rather than memorizing seen environments — though the paper does not test generalization beyond ALFWorld's distribution shift.
Co-evolutionary dynamics
Two empirical dynamics characterize training. Harness annealing: harness-call frequency drops sharply during GRPO and stabilizes near one call per episode. Action-level decomposition shows this pruning is selective rather than uniform — recall persists longest (cross-episode priors remain valuable), while commit and note decay toward zero once stable strategies are internalized, and track follows an intermediate pattern. Training-reward curves show EvoHarness-RL outpacing standard GRPO throughout, indicating the call reduction reflects internalization of scaffolded behavior into the policy weights, not collapse. The authors explicitly caution that the resulting action distribution is environment-dependent: ALFWorld favors Experience due to reusable household procedures, whereas software-engineering settings would likely favor Progress.
Harness evolution: the experience store expands rapidly early in training, then growth becomes selective as consolidation merges redundant entries, evicts rarely used skills via LFU, and preserves frequently recalled knowledge. The final bank remains compact yet diverse, functioning as a curated, revisable substrate rather than append-only memory. A qualitative case study illustrates the loop: the agent recalls a stale "kettle on countertop" prior, detects contradiction through failed searches, finds the kettle on a stoveburner, and writes corrective evidence back — treating recalled experience as revisable rather than oracle-like.
Limitations and open questions
Several constraints bound these findings. Evaluation is confined to ALFWorld, a text-based embodied benchmark with mostly sequential tasks; the paper itself notes the Progress instantiation (a bounded sequential subgoal list) may be insufficient for branching workflows, and that the learned harness-action mix should not be read as universal. The Belief tracker is rule-based and deterministic, sidestepping noisy state estimation — a setting the paper identifies but does not solve. Harness actions consume interaction budget, making results sensitive to Tmax and reward shaping choices that are only validated here. Finally, experience consolidation depends on an external frontier model (Claude Opus) at epoch boundaries, so the system is not fully self-contained at the 8B scale; whether policy-model consolidation suffices is left unexamined. Whether the annealing dynamic (near-total internalization of routine harness use) holds in environments with genuinely non-stationary state is an open question the paper raises implicitly but does not answer.
Conclusion
EvoHarness-RL reframes harness use as a trainable runtime policy over a compact BPE workspace, combining supervised bootstrapping with cost-aware GRPO. On ALFWorld it achieves 96.9% seen and 86.6% unseen success with Qwen3-8B, matching frontier-model performance, and shows that training shifts agents from frequent scaffold calls toward selective, cost-aware external-state access while the experience store co-evolves through consolidation and forgetting. The result argues that gains in long-horizon reliability can come from learning coordination with external workspaces, not only from stronger tools or larger memories.