- The paper introduces CurateEvo, a failure-driven framework that evolves executable curation code to jointly produce SFT data, RL data, and inference-time memory from fixed corpora.
- CurateEvo achieves the best results across ACEBench-Agent, BFCL-V4, and τ²-Bench, improving average scores by 3.2 points on labeled data and 2.7 points on wild data over prior methods.
- The framework improves efficiency by using 48% fewer tokens and 50% less wall-clock time per 1,000 retained training turns, while failure diagnosis and adaptive filtering preserve rare, high-value agent behaviors.
Motivation and problem statement
Agentic post-training pipelines for LLM agents—SFT warm-up followed by multi-turn RL such as GRPO—are highly sensitive to the quality of the interaction data they consume. Existing curation approaches, however, exhibit two systematic weaknesses. First, curation neglect: most methods concentrate on augmenting additional trajectories (via simulated users, synthetic environments, or online synthesis) while underusing filtering and refinement, which prior work shows can improve performance at lower training cost. Second, adaptation rigidity: curation pipelines are typically fixed, so they cannot be re-targeted when downstream failure modes shift or when the raw corpus distribution changes.
CurateEvo addresses both weaknesses by making the curation strategy itself the optimization target (2607.06140). The strategy is represented as executable code that transforms a fixed raw corpus into three resources: a supervised fine-tuning dataset Dsft, an RL dataset Drl, and an inference-time memory bank Mmem. An LLM-based code-evolution agent iteratively rewrites this code using failed trajectories from a held-out development set as feedback, while the base model, training recipe (SFT+GRPO), environment, and inference budget remain fixed across epochs.
Method
The framework formalizes curation as a cost-aware optimization over strategies:
ρ⋆=argρmaxJ(ρ)=P(ρ)−λC(ρ)
where P(ρ) is dev-set performance of the policy trained on curated data, and C(ρ) is a log-scaled, tanh-normalized measure of total retained training turns. Because exhaustively evaluating candidate code revisions is infeasible, each epoch applies two ordered revisions that approximate the objective:
- Effectiveness revision: failed dev trajectories are converted into actionable "failed points" (capability weakness, trajectory evidence, curation hint). The evolution agent diagnoses recurring failure modes—incorrect tool selection, invalid argument grounding, weak multi-turn planning, poor error recovery, memory mismatch, noisy supervision—and rewrites the curation code to augment uncovered modes, filter misleading examples, refine weak trajectories, and extract memory entries.
- Efficiency revision: conditioned on the diagnosed failure modes and a curation profile (Σe: scale, turn distributions, redundancy), the code is further revised to remove duplicates, prune low-utility samples, compress repeated patterns, and truncate long trajectories, while preserving samples covering rare failures and critical tool-use steps.
A rollback mechanism prevents error cascading: any epoch whose revision does not improve J reverts to the previous strategy. The test set is never touched during evolution.
Experimental setup
Experiments use Qwen3-4B with LoRA (rank 16) and a fixed SFT+GRPO recipe, evaluated on ACEBench-Agent, BFCL-V4, and τ2-Bench. Two raw-corpus settings are studied: labeled (human-annotated trajectories from SWE-chat, AgentRewardBench, OpenHands-Feedback) and wild (real user interaction logs from open sources). The code-evolution agent uses GPT-5.4 within mini-SWE-agent over 3 epochs, with λ=0.3 and Drl0. Baselines include GRPO without curation, MUA-RL, EnvScaler, AWM, RODS, and FunReason-MT.
Main results
CurateEvo achieves the best score on all three benchmarks in both settings, improving average scores by 3.2 points over the strongest prior method on labeled data and 2.7 points on wild data. Notably, CurateEvo with Qwen3-4B outperforms several baselines built on Qwen3-8B, supporting the claim that adaptive curation can substitute for model-scale increases. Gains on Drl1-Bench—which requires long-horizon state tracking and user coordination—indicate the improvements extend beyond tool-call accuracy. Wild-data results are particularly relevant: the framework extracts competitive supervision from noisy real interaction logs without human annotation.
Curation overhead is also reduced: CurateEvo consumes 0.51M tokens and 405 seconds per 1K retained training turns, roughly 48% fewer tokens and 50% less wall-clock time than the average prior baseline, because it evolves code over a fixed corpus rather than generating new environments or simulated interactions.
Ablations and analyses
Removing the effectiveness-oriented revision causes the largest degradation (~7 points on average), confirming failure-driven diagnosis as the primary driver; removing the efficiency revision costs about 1 point, showing pruning improves quality rather than merely reducing scale. All three resources are complementary, with SFT and RL data contributing most and memory providing inference-time task knowledge; ablation drops are larger on wild data, consistent with noisy corpora depending more heavily on adaptive filtering.
Three further analyses sharpen the picture:
| Analysis |
Finding |
| Recipe compatibility |
Adding CurateEvo's curated resources to GRPO, AgentGym-RL, and ProRL-Agent yields an average gain of 21.3 points, strongest with ProRL-Agent |
| Evolution epochs |
Most gains occur in the first three epochs (24.2→50.3 labeled; 21.8→47.7 wild); later epochs mainly shrink curated scale while performance saturates |
| Cost weight Drl2 |
Performance peaks at Drl3; large Drl4 over-prunes rare-failure and long-horizon trajectories |
The epoch analysis implies curated scale need not grow monotonically with performance—the goal is a more useful training distribution, not a larger one. The recipe analysis positions CurateEvo strictly as a data-side module orthogonal to policy-optimization advances.
The appendix characterization of the final evolved strategy is informative: it decomposes trajectories into decision-centric segments rather than treating whole trajectories as selection units, separates transferable policy behaviors (kept for SFT/RL) from task-specific knowledge (compressed into retrievable memory), constructs GRPO groups around high-frequency failure families to provide positive-negative contrasts, and treats discarding as an active design choice to preserve LoRA capacity and avoid diluting preference learning.
Failure mode analysis
Across evolution, failures shift from shallow format and routing errors toward three structural bottlenecks: state binding drift (stale identifiers, lost cross-turn parameters), adaptive recovery failure (repeating failed calls, premature termination, write-before-confirmation), and the grounded execution gap (actions not supported by evidence, schema constraints, or executability). The authors attribute these to supervision that rewards locally plausible actions and clean success paths, and they propose—but do not implement—remedies such as state-centric contrastive curation, step-level feedback for invalid repetitions, and execution-aware filtering.
Limitations and open questions
Several caveats bear directly on the reported results. Evaluation uses a single 4B base model and LoRA fine-tuning; whether the gains transfer to full-parameter tuning or larger models is untested. Evolution relies on a proprietary frontier LLM (GPT-5.4) as the code-evolution agent, so cost and reproducibility of the evolution loop itself depend on that component, and only 3 epochs are run despite evidence of near-saturation beyond them. The dev set serves both as the feedback signal and the strategy-selection criterion, raising a mild risk of dev-set overfitting that the held-out test evaluation mitigates but does not eliminate. Finally, the identified failure modes (state binding drift, adaptive recovery, grounded execution) are diagnosed but not resolved by curation alone; the paper leaves open how verifier modules or execution monitors could be integrated into the curation loop, and how the cost weight Drl5 should be set adaptively rather than fixed at 0.3.
Conclusion
CurateEvo reframes agentic post-training data curation as iterative, failure-driven evolution of executable curation code, jointly producing SFT data, RL data, and inference-time memory under a cost-aware objective. Consistent gains of 3.2 and 2.7 average points over prior curation methods on labeled and wild corpora, a 21.3-point average gain when combined with alternative post-training recipes, and roughly halved curation overhead collectively indicate that feedback-driven adaptation of the data pipeline is a practical complement to algorithmic advances in agent RL.