- The paper introduces ToolVerse, combining automated MCP environment synthesis, graph-guided task generation, and Turn-Aware Relative Advantage (TARA) for long-horizon agentic RL.
- The framework creates 422 executable environments with 4,438 tools and the 2,987-item GUST dataset, enabling stateful tasks with 3–7 sequential subtasks and turn-level golden-trace supervision.
- TARA consistently outperforms naive GRPO, raising Qwen3-8B performance from 46.51% to 61.66% on ACEBench-Agent and improving results across BFCL-v3, τ²-Bench, and ACEBench-Agent.
ToolVerse is a framework for agentic reinforcement learning (RL) that addresses three coupled bottlenecks in training tool-using LLM agents: limited environment diversity, difficulty of constructing long-horizon multi-turn tasks, and coarse-grained credit assignment under sparse rewards (2607.15660). The paper makes three concrete contributions: an automated pipeline that converts raw tool definitions into executable MCP environments, a graph-based task synthesis method that yields the GUST dataset, and a Turn-Aware Relative Advantage (TARA) algorithm for turn-level credit assignment. Training on ToolVerse produces consistent gains across BFCL-v3 Multi-Turn, τ2-Bench, and ACEBench-Agent on models ranging from Qwen3-4B to Qwen2.5-14B-Instruct.
Scaling executable agentic environments
The environment construction pipeline starts from raw JSON tool schemas, drawn primarily from the multi-turn subset of the Toucan dataset, and transforms them into executable MCP toolsets through a closed-loop synthesis process. Schema refactoring first normalizes function signatures and removes textual noise; a scenario-driven mock database (Python dictionaries modeling entities such as users, orders, and inventory, with 3–5 coherent records per scenario) provides stateful interaction; LLM-generated Python functions then implement each tool against this database. Only toolsets passing all syntax checks and unit tests are retained; roughly 20% of initial toolsets are filtered out, mostly because their APIs depend on external services or side effects that cannot be faithfully mocked.
The resulting corpus comprises 422 executable MCP environments containing 4,438 tools, spanning eight macro-domains, with most environments holding 7–13 tools. These environments are integrated into Verl as the action space for RL training. A notable design constraint is that only toolsets with 5–20 tools are kept, which ensures sufficient intra-environment dependency structure for long-horizon task generation but also restricts coverage of very large toolsets.
Graph-guided long-horizon task synthesis
Task generation is structured around a tool dependency graph (TDG), whose nodes are tools and whose edges encode parameter dependencies (one tool's output feeds another's input) and semantic ordering constraints inferred by an LLM judge. Because LLM-inferred edges can be noisy, the TDG is sanitized into a DAG via topological traversal, prioritizing parameter-dependency edges over semantic ones when breaking cycles. The authors argue that residual edge noise does not corrupt supervision because the task instruction, golden trace, and rule-based validator are all constructed against the same executable dependency structure—an assumption that holds only insofar as execution-based verification removes unexecutable trajectories.
The Dynamic Unlocking Sampling (DUS) algorithm performs constrained sampling over this DAG: starting from zero-in-degree root nodes, it samples batches of tools per trajectory stage, decrements successor in-degrees, and unlocks nodes once all prerequisites are satisfied. This topological progression induces an implicit curriculum in which simple queries necessarily precede complex operations. Sampled skeletons are grounded through inverse context reconstruction—arguments instantiated in topological order from prior tool outputs and database state—to produce executable Golden Traces. An LLM converts each trace into a natural-language user query, with dependency edges supplied explicitly so the generated instruction preserves causal structure. Verification proceeds in two stages: LangGraph replay against the stateful mock database, followed by teacher-agent Pass@8 filtering using GPT-4.1, DeepSeek-V3.2, and Qwen3-32B; tasks unsolved in any attempt are discarded.
The final GUST dataset contains 2,987 data items averaging 3.8 tasks per item across 3–7 sequential subtasks, with at most ten retained tasks per environment to limit overfitting. Because golden traces provide ground truth for every turn, they enable the turn-level reward signals used in training—a coupling between dataset design and algorithm design that is central to the framework.
Turn-aware relative advantage estimation
Standard GRPO assigns a single trajectory-level advantage to all tokens, which cannot distinguish correct intermediate steps from fatal later errors in long-horizon rollouts. TARA decomposes the advantage at each turn t into two group-normalized components. The local advantage standardizes a binary turn reward ri,t—set to 1 when the rollout's tool-call JSON sequence covers the golden set under dictionary-level matching (tool names and arguments), allowing any TDG-compatible order within a turn—against the same-turn group distribution. The gated future advantage computes a discounted sum of future turn rewards multiplied by a consistency gate δi,t=ri,t, so future credit flows only through currently valid steps; this value is likewise normalized within the group. The total advantage fuses the two terms as Atotal=Alocal+λAfuture with λ=0.5, and is propagated uniformly to all tokens in the turn. Zero-variance turns receive zero advantage, preventing degenerate gradients.
The appendix supplies two theoretical results. Under the assumption of deterministic turn rewards, the variance of the TARA value estimate is λ2σfuture2, strictly below the return-to-go variance whenever λ<1. Second, for "distractor" actions—incorrect turns that nonetheless precede positive outcomes—the gate forces the future term to zero while the local term is strictly negative given a positive group mean, guaranteeing non-positive total advantage and thus suppression of spurious credit. Both results depend on the binary, golden-trace-derived reward being deterministic and well-defined at every turn, which is precisely what the GUST construction guarantees but which would not transfer to settings lacking per-turn ground truth.
Empirical results
Training uses GRPO on Qwen3-4B, Qwen3-8B, and Qwen2.5-14B-Instruct with DeepSeek-V3.2 as user simulator, on 32 A100 GPUs with global batch size 128, KL loss disabled, and four-run averaged evaluations.
| Model |
BFCL-v3 Overall |
τ2-Bench Overall |
ACEBench-Agent Overall |
| Qwen3-4B base → +TARA |
25.38 → 28.25 |
21.06 → 26.83 |
41.28 → 55.00 |
| Qwen3-8B base → +TARA |
28.88 → 37.50 |
27.87 → 32.37 |
46.51 → 61.66 |
| Qwen2.5-14B base → +TARA |
17.88 → 24.12 |
24.38 → 32.40 |
47.78 → 61.66 |
Gains are largest on ACEBench-Agent (+15.15 for Qwen3-8B, +13.88 for Qwen2.5-14B), and TARA consistently outperforms naive GRPO trained on the same data—for example, 61.66% versus 56.66% on ACEBench-Agent for Qwen3-8B—with the gap widest in multi-turn and multi-step subtasks where credit assignment is most acute. Against reproducible public baselines on Qwen2.5-7B-Instruct, TARA reaches 20.00% on BFCL-v3 and 29.87% on τ2-Bench, exceeding ToolRL, AgentFlow, and SimpleTIR; head-to-head comparison with SALT and FTRL was not feasible because their artifacts are not publicly available, a concession the authors acknowledge.
Ablations isolate the algorithm's components: turn-local credit alone underperforms even naive GRPO on t0-Bench (27.33% vs. 30.10%), and ungated future credit also falls short of full TARA, confirming that both future propagation and the consistency gate are necessary. Performance is best at t1 and t2, degrading at extremes. Environment scaling shows a monotonic trend: growing from 100 to 422 training environments raises BFCL-v3 from 35.00% to 37.50% and t3-Bench from 27.33% to 32.37%, indicating that environment diversity—not merely additional rollouts—drives generalization. Training curves show TARA converging faster and more stably than naive GRPO in both trace score and validation score.
Limitations and open questions
The authors identify two principal limitations. First, task diversity is bounded by the predefined dependencies in available protocols; DUS cannot discover novel tool interactions autonomously, so emergent behaviors beyond the LLM-inferred TDG structure are out of scope. Second, TARA presupposes reliable per-turn reward signals derived from golden traces; in highly sparse or ambiguous reward settings without such supervision, the mechanism may fail to provide useful feedback. Two further caveats bear on interpretation: the theoretical guarantees assume deterministic binary turn rewards, and the environment pipeline discards roughly 20% of toolsets whose APIs resist dictionary-based mocking, so the 422-environment corpus skews toward mockable, stateless-backend scenarios. Whether turn-level golden-trace supervision can be obtained at scale outside synthetic environments remains an open question the paper does not resolve.
Conclusion
ToolVerse couples three mechanisms—automated MCP environment synthesis, DAG-constrained long-horizon task generation, and gated turn-level advantage estimation—into a single training framework, and demonstrates that the combination yields broad, model-agnostic improvements on multi-turn tool-use benchmarks, with TARA's gains attributable specifically to finer-grained credit assignment. The framework's dependence on golden-trace-derived turn rewards and predefined dependency graphs delineates the boundary of its applicability and defines the most direct questions for subsequent work.