---
title: Long-Horizon ASR in Agentic Systems
url: https://www.emergentmind.com/topics/long-horizon-action-success-rate-asr
type: topic
---

# Long-Horizon ASR in Agentic Systems

Long-horizon Action Success Rate (ASR) quantifies the fraction of trials in which an agent successfully completes all required steps or subgoals in complex, temporally extended tasks. ASR serves as the principal metric for evaluating the ability of agentic systems—spanning LLM-based planners, Vision-Language-Action (VLA) stacks, and hierarchical controllers—to execute multi-step protocols, solve sequential tasks, and maintain performance as the task horizon increases. Unlike stepwise or local metrics, long-horizon ASR measures strict end-to-end success, making it sensitive to compounded errors, long-range memory limitations, partial observability, and failure recovery. The metric is now widely adopted in robotics, manipulation, navigation, and general agentic AI research as an anchor for benchmarking and diagnosis of system reliability under horizon scaling.

## 1. Formal Definition and Computation

The canonical definition of long-horizon Action Success Rate (ASR) is the episode-level completion metric:
\[
\mathrm{ASR} = \frac{1}{N} \sum_{i=1}^N \mathbb{I}[\text{all } M_i \text{ subgoals succeeded in episode } i]
\]
where $N$ is the total number of evaluation episodes (trials), $M_i$ is the number of subgoals or atomic actions in episode $i$, and $\mathbb{I}[\cdot]$ is the indicator function returning 1 if all subgoals are achieved and 0 otherwise. In many domains, a subgoal is completed when a precise geometric, semantic, or kinematic predicate is met (e.g., object placed within target region, logical state transition reached). Some works also report step-wise ASR (fraction of successful subgoals across all episodes) and average subtask success for finer-grained diagnosis [2604.20721][2601.01618][2604.18791].

The metric generalizes naturally to more complex compositional horizons. For variable task depths:
\[
\mathrm{ASR}(H^*(s)) = \frac{N_{\mathrm{succ}}(s)}{N_{\mathrm{att}}(s)}
\]
where $H^*(s)$ is the intrinsic horizon at extension level $s$ in nested task families [2604.11978].

Across domains such as manipulation, navigation, and agentic reasoning, long-horizon ASR consistently refers to the fraction of episodes in which all steps succeed, no matter the local reward structure or auxiliary sub-metrics.

## 2. Benchmark Domains, Task Construction, and Success Criteria

ASR underpins evaluation protocols in diverse high-complexity agentic scenarios:

- **Symbolic and embodied planning:** AgentBoard games (Blocksworld, Tyreworld, Jericho), where full success requires satisfying all atomic environment goal conditions [2408.09559].
- **Robotic manipulation:** Multi-object interaction, protocol-based lab automation, or 3D rearrangement, with geometric and semantic constraints on all final placements [2509.08820][2603.23676][2601.01618][2604.18791].
- **Navigation:** Episodic success is determined by reaching the goal state and halting within physical and timing thresholds [2602.12351].
- **Human-scene interaction:** Ordered composite skills (Follow, Carry, Climb, Sit) chained together with stringent per-step criteria and global episode timeouts [2604.20721].
- **Multimodal and database/Web domains:** Cross-domain composition with depth/breadth extension, where ASR measures robustness as horizon grows [2604.11978].

Each setting enforces strict full-sequence requirements: an episode is only counted toward ASR if no subgoal (according to task-specific binary or graded predicates) fails. Protocol-level ASR often drops precipitously as horizon or scene complexity increases, exposing compounding bottlenecks absent from single-step metrics.

## 3. Representative Empirical Results and Baseline Comparisons

Long-horizon ASR reveals sharp delineations in the capabilities of various agentic architectures, particularly under scaling of the horizon and complexity:

| Domain/Baseline                  | SR/ASR (%) | Context                                                          |
|-----------------------------------|-----------:|------------------------------------------------------------------|
| PCArena (HiAgent) [2408.09559]   | 42         | 2x improvement over Standard (21), 5 long-horizon AgentBoard tasks|
| RAMP-3D [2603.23676]              | 79.5       | 3D box rearrangement, 11 variants, 1–30 objects                  |
| Goal2Skill [2604.13942]           | 32.4       | RMBench, adaptive subtask composition, vs. 9.8% best baseline    |
| SD-VLA [2602.03983]               | 76.4       | Memory-dependent LIBERO-Memory, +39.8 pp over best prior         |
| LiLo-VLA [2602.21531]             | 69         | LIBERO-Long++, Ultra-Long, best baseline: 28                     |
| HELM [2604.18791]                 | 81.5       | LIBERO-LONG, +23.1 pp over OpenVLA-H=8                           |
| Action-Sketcher [2601.01618]      | 96.0       | LIBERO Long-Horizon (8–16 subtasks), best on complex manipulation|
| ALAS [2604.20721]                 | 72         | HSI-LH1, vs. TokenHSI 55, CML 30                                 |
| RoboChemist [2509.08820]          | 72         | Protocol SR, full chemical sequence, vs. highest baseline 38      |
| RoboClaw [2603.11558]             | 75         | Real-world manipulation, +25 pp over VLA open-loop (50)           |

These results consistently show that naive or non-hierarchical models suffer exponential-like decay in ASR as tasks lengthen. In contrast, architectures with explicit memory, subgoal chunking, hierarchical working memory, and closed-loop recovery (e.g. HELM, PCArena, Goal2Skill, Action-Sketcher, ALAS) dramatically improve episode completion rates.

## 4. Mechanisms That Affect Long-Horizon ASR

Three primary system-level factors systematically impact ASR:

- **Temporal Memory:** Hierarchical, episodic, or layer-wise KV memory mitigates “memory gap” failures typical in purely Markovian or short-context models [2408.09559][2603.07647][2604.18791]. Failure ablations show loss of 8–23 pp ASR when disabling key episodic or chunked working memory modules.
- **Verification and Recovery:** Learned state verifiers and recovery controllers (rollback, reflection, multi-policy orchestration) address cumulative subgoal errors and allow for local correction without global episode reset [2604.13942][2604.18791][2603.11558].
- **Task/Plan Decomposition:** Explicit subgoal chunking, plan summarization, and trajectory retrieval localize memory, reduce context overload, and enable targeted reasoning [2408.09559][2601.01618].

Ablation studies consistently show double-digit drops in ASR on disabling these components. For example, PCArena’s observation summarization and retrieval module yields 30–50% ASR swings; HELM’s episodic memory and state verifier contribute the majority of its 23.1-point gain.

## 5. Statistical Analysis and Horizon-Scaling Behavior

Empirical studies document nonlinear ASR collapse as horizon or step count grows. In HORIZON [2604.11978], both LLM agents and RL/VLA agents display sharp transition regions: high ASR is sustained up to a domain/model-specific critical extension, after which episode success rapidly collapses to zero. Domain compositional strategies (depth vs. breadth) reveal the susceptibility of each architecture to structurally distinct error accumulation.

ASR is typically reported as mean ± standard deviation across random seeds or trial replicates. Extensive bootstrapping or runwise aggregation yields reliable uncertainty quantification, especially at deep compositional levels.

## 6. Diagnosing and Mitigating ASR Degradation

Comprehensive failure attribution, as exemplified in HORIZON [2604.11978], employs LLM-as-a-Judge pipelines and error taxonomy labeling to analyze breakdowns. Failures transition from process-level (local subplan issues) to design-level (memory and catastrophic forgetting), aligning with ASR collapse.

Mitigation strategies empirically shown to elevate ASR include:

- Hierarchical subplanning and plan repair [2408.09559][2604.11978]
- Enhanced working and episodic memory modules [2604.18791][2603.07647][2408.09559]
- Closed-loop, reflection-based recovery [2604.13942][2603.11558]
- Task- and memory-aware verification logic [2604.18791]
- Explicit decomposition and subgoal summarization [2408.09559][2601.01618]

Simply scaling context length or model size without architectural improvements yields limited (often <6 pp) ASR gains.

## 7. Relation to Partial-Completion, Average Progress, and Other Metrics

While ASR strictly tracks full-sequence completion, several works also report:

- **Average Progress (AP):** Mean fraction of ordered subgoals completed before failure [2602.21531].
- **Average subtask success:** Mean over all subtask hits in all episodes [2604.20721][2601.01618].
- **Progress Rate (PR):** Fraction of goal conditions satisfied at episode end [2408.09559].

ASR remains the most stringent and informative measure for true end-to-end reliability, with significant implications for system deployment in robotics, automated lab environments, and decision-critical agentic domains.

---

**References**

- [2408.09559] "HiAgent: Hierarchical Working Memory Management for Solving Long-Horizon Agent Tasks with Large Language Model"
- [2604.13942] "Goal2Skill: Long-Horizon Manipulation with Adaptive Planning and Reflection"
- [2603.23676] "Grounding Vision and Language to 3D Masks for Long-Horizon Box Rearrangement"
- [2602.12351] "LongNav-R1: Horizon-Adaptive Multi-Turn RL for Long-Horizon VLA Navigation"
- [2601.01618] "Action-Sketcher: From Reasoning to Action via Visual Sketches for Long-Horizon Robotic Manipulation"
- [2509.08820] "RoboChemist: Long-Horizon and Safety-Compliant Robotic Chemical Experimentation"
- [2603.07647] "TempoFit: Plug-and-Play Layer-Wise Temporal KV Memory for Long-Horizon Vision-Language-Action Manipulation"
- [2604.20721] "ALAS: Adaptive Long-Horizon Action Synthesis via Async-pathway Stream Disentanglement"
- [2604.11978] "The Long-Horizon Task Mirage? Diagnosing Where and Why Agentic Systems Break"
- [2602.21531] "LiLo-VLA: Compositional Long-Horizon Manipulation via Linked Object-Centric Policies"
- [2603.11558] "RoboClaw: An Agentic Framework for Scalable Long-Horizon Robotic Tasks"
- [2602.03983] "Efficient Long-Horizon Vision-Language-Action Models via Static-Dynamic Disentanglement"
- [2604.18791] "HELM: Harness-Enhanced Long-horizon Memory for Vision-Language-Action Manipulation"

Source: https://www.emergentmind.com/topics/long-horizon-action-success-rate-asr