---
title: GTA-Workflow Benchmark Suite
url: https://www.emergentmind.com/topics/gta-workflow
type: topic
---

# GTA-Workflow Benchmark Suite

GTA-Workflow denotes the long-horizon, open-ended workflow evaluation setting introduced as part of the GTA-2 benchmark suite. GTA-Workflow is intended for benchmarking general tool-using agents under realistic productivity task distributions, moving beyond short atomic tool chains to complex, composite deliverables that more closely reflect real-world applications. In this framework, a task is the production of a heterogeneous artifact (e.g., a multi-section PDF, a slide presentation, a multimedia report) from a raw set of files and a natural user request, utilizing a specified suite of API-accessible tools. GTA-Workflow is architected to evaluate not only model-level tool-use precision but also the orchestration and execution capabilities of agent frameworks—such as planning, state management, error recovery, and adaptive sub-goal decomposition—across extended task horizons [2604.15715].

## 1. Conceptual Foundations and Motivation

GTA-Workflow arises from recognition that benchmarks for general tool agents must reflect the open-ended, multi-modal, multi-step nature of real productivity workflows. Existing tool-use evaluations (cf. GTA-Atomic [2407.08713]) focus on closed-ended, “atomic” tasks—few-step tool chains with single reference outputs. However, practical AI assistants must execute dynamic plans, integrate across multimodal inputs, and synthesize diverse deliverables without fixed scripts or templates. GTA-Workflow addresses these requirements by defining a benchmark in which each task specifies a workflow objective (e.g., “prepare a ten-slide PDF synthesizing findings from five research papers,” or “extract, denoise, and summarize a 3-minute audio section from a podcast”), along with an input context, a toolset, and a recursive checkpoint tree for outcome-driven evaluation [2604.15715].

## 2. Task Definition and Tool Universe

Each GTA-Workflow task is formalized as a quadruple $(F, Q, T, \mathit{Cpt})$, where:

- $F$: a set of input files (potentially including images, videos, audio, DOCX/XLSX/PDF/PPTX).
- $Q$: a real user-authored natural language query describing the composite deliverable.
- $T$: a set of 37 deployed APIs covering perception (e.g., image/audio/video analysis), document and structured data operations (e.g., ReadPDF, CsvFileGenerator), reasoning (e.g., Prover via Z3), and creative generation (e.g., TextToImage, TextToVideoTool).
- $\mathit{Cpt}$: a hierarchical tree of verifiable, outcome-oriented checkpoints, assigning weights to each sub-goal.

Unlike atomic tasks, for which there exists a fixed ground-truth tool chain of length $m$, GTA-Workflow tasks are open-ended: the deliverable $D$ may be achieved via many valid execution paths, and success is defined extensionally by the satisfaction of the sub-goals in $\mathit{Cpt}$, not by strict stepwise matching [2604.15715].

## 3. Recursive Checkpoint-Based Evaluation

GTA-Workflow employs a recursive checkpoint scoring mechanism to assess deliverables. Each node in the checkpoint tree $\mathit{Cpt}$ is either a composite sub-goal (internal) or an atomic outcome (leaf). The evaluation proceeds as follows:

1. Traverse the tree, for each node $n$:
   - If $n$ is a leaf, invoke a powerful LLM judge $M$ (e.g., GPT-5.2) with $(D, \text{Requirements}(n), \text{Rubric}(n))$; assign a score $s_n \in [0, 10]$.
   - For internal nodes, recursively aggregate child scores: $S(n) = \sum_{i=1}^k w_i \cdot S(c_i)$, with $w_i$ the normalized weights and $c_i$ the children.
2. The root score $S_\mathrm{root}$ is the workflow deliverable score. Success is binary: the workflow “passes” if $S_\mathrm{root} > k$, where $k = 7$ by default.

Leaf Success Rate (fraction of checkpoints with $s_n > 7$) and Tool Success Rate (fraction of tool invocations running without execution errors) are also reported [2604.15715].

## 4. Dataset Construction and Task Diversity

The GTA-Workflow corpus includes 132 workflow tasks. Source diversity is ensured by:

- Sourcing ~50% of workflows from production agent platforms (Manus, Minimax Agent, Kortix, Flowith, CrewAI), thus capturing authentic user queries and actual business or productivity workflows.
- Curating the remainder from high-engagement Stack Exchange and Reddit queries, which are then further refined by LLMs under human supervision.
- Each workflow specifies 3–19 checkpoints, elaborating requirements on structure, correctness, and presentation of the final artifact.

Examples illustrate the spectrum: scientific reviews, data analysis with tabular visualization, multimedia extraction, document synthesis, chaining of perception, reasoning, and generation tools. Table 2 in [2604.15715] details the filtration and selection process, starting from 154 candidates and resulting in 132 finalized workflows.

## 5. Experimental Analysis and Performance Results

Empirical evaluation demonstrates a pronounced “capability cliff” in contemporary tool agents:

- While leading LLMs and frameworks (e.g., Gemini-2.5-Pro, GPT-5, Claude-Sonnet-4.5) maintain high Tool Success Rates (e.g., 91.2% for Gemini-2.5-Pro), Root Success Rates and average $S_\mathrm{root}$ drop sharply in GTA-Workflow: maximum observed Root SR is 14.39%, with $S_\mathrm{root} = 3.64/10$ for Gemini-2.5-Pro and similarly low for others.
- In contrast, the same models achieve >40% end-to-end solution rates on GTA-Atomic (short, closed tasks).
- This outcome demonstrates that correct atomic tool use is necessary but insufficient for robust workflow completion, with failures driven by inadequate state tracking, error recovery, and sub-goal decomposition [2604.15715].

Advancements in execution harnesses (e.g., OpenClaw, Manus, Kortix) offer substantial improvement. OpenClaw increases Root SR from 0% to 50% (S_root from 2.49 to 6.82) for a fixed base LLM. Manus and Kortix achieve over 53% in subset analysis, indicating that harness design (system-level memory, persistent state, advanced error handling) is a key determinant of actual workflow success, surpassing differences in base LLM reasoning when Tool SR is saturated.

## 6. Feedback, Diagnostics, and Future Directions

The impact of feedback granularity is quantified:

- Generic “coarse” feedback during re-tries yields only a 4% relative gain in $S_\mathrm{root}$, while detailed checkpoint diagnostics raise $S_\mathrm{root}$ by 12% (from 2.83 to 3.15).
- This demonstrates that fine-grained, task-decomposed feedback is critical to guiding agent improvement on workflow-scale problems [2604.15715].

Recommended future directions include:

- Extending evaluation beyond deliverable quality to encompass safety and governance.
- Rigorous system-level ablations across model types and harnesses to distinguish execution and reasoning contributions.
- Refinement of checkpoint taxonomy for causal modeling of error pathways and more robust diagnostic signals.
- Releasing both raw and reformulated task pairs to control for bias in benchmark construction.

## 7. Significance and Implications

GTA-Workflow establishes a benchmark for the next generation of agent evaluation: its open-ended, multimodal, verifiably-scored workflows expose bottlenecks not only in LLM reasoning but, more critically, in the orchestration, planning, and execution machinery underpinning modern tool agents. The empirical gap revealed by GTA-Workflow argues that future progress in personal and professional AI assistants depends as much on the architecture of execution frameworks as on scaling model capacity per se. As such, GTA-Workflow quantifies the steep drop from atomic precision to reliable end-to-end workflows, and provides a framework for systematically closing this gap via improved harnesses, feedback granularity, and workflow planning methodologies [2604.15715].

Source: https://www.emergentmind.com/topics/gta-workflow