Papers
Topics
Authors
Recent
Search
2000 character limit reached

How Fast Do Agents Rot? An Empirical Study of Long-Horizon Degradation in LLM Agents for Production Decision-Making

Published 31 Aug 2026 in physics.soc-ph and cs.AI | (2609.01660v1)

Abstract: Production deployments of LLM agents remain unreliable on long, multi-step workflows even as benchmark success rates climb steadily. We argue this gap is largely an artifact of task horizon: benchmarks are dominated by short-to-medium horizons where success remains high, while production workloads demand an order of magnitude more dependent steps. We measure the effect directly, characterizing the shape of agent degradation and disentangling its cause across a large controlled study spanning nine models, six open models from 1.2B to 671B parameters, and three deployed proprietary systems; four task families, including a genuinely agentic tool-use loop; five horizons; and three context regimes. Task success follows a geometric law governed by a single per-step reliability parameter, which rises with model scale but saturates well below 1 even for the strongest models, guaranteeing eventual collapse at sufficiently long horizons. The effect is sharpest on the agentic task, where every model tested, including widely deployed systems, falls from near-perfect success to near zero within sixteen steps of (n=10,664 analyzed trajectories. Degradation is driven by step count rather than context length: bounding the context window steepens decay rather than easing it (logit slope -0.69 vs. -0.44), p=3x10-6), contradicting a lost-in-the-middle explanation and warning against a common production shortcut. Projecting measured reliability onto representative benchmark horizons quantifies a substantial gap between benchmark and production conditions, from 0.42 at GAIA-length horizons to 0.24 at hundred-step production horizons. For teams responsible for agent orchestration and reliability at scale, these results argue for horizon-aware evaluation and reliability budgeting in place of aggregate pass-rate metrics. Code, prompts, seeds, and raw trajectories are released.

Authors (1)

Summary

  • The paper identifies geometric decay as the dominant degradation pattern in 28 of 36 model–task cells, highlighting that long-horizon reliability is severely impacted by response chain error accumulation.
  • Longer and more dependent steps directly relate to model performance drop with the guiding hypothesis that as per-step reliability $(r)$ increases performance is typically defined as $r^H$.
  • Per-step reliability increases with model scale, but even large models (eg proprietary models) cannot eliminate compounding error weill below their initial performance on longer tasks.

Research question and central thesis

“How Fast Do Agents Rot?” (2609.01660) investigates why LLM agents that perform well on conventional benchmarks remain unreliable on production workflows requiring many dependent actions. The paper’s central hypothesis is that task success is governed by compounding per-step reliability. If an agent must execute HH dependent steps and succeeds at each step with probability rr, end-to-end success is approximately rHr^H. Even a high per-step reliability therefore produces severe degradation as the horizon increases.

The paper argues that standard benchmark scores obscure this phenomenon because many benchmarks evaluate relatively short horizons. Production workflows, by contrast, often require substantially more tool calls, state transitions, and dependent decisions. The resulting benchmark–production gap is presented not as an anomalous deployment failure but as a predictable consequence of evaluating the flat upper portion of a geometric decay curve.

The empirical study examines nine instruction-tuned models from six vendor families: six open models ranging from 1.2B to 671B parameters and three proprietary systems accessed through hosted interfaces. It analyzes 10,664 trajectories across four task families, five horizons, and three context regimes. The design emphasizes exact oracle verification, paired comparisons across context conditions, and model selection among geometric, threshold, and linear decay functions.

Experimental design

The four task families are synthetic but deliberately structured to isolate different forms of sequential dependence. Ledger requires maintaining account balances through a sequence of operations. Refchain tests retrieval and state tracking by requiring the model to follow assignments amid distractors. Cipher requires ordered string transformations, making early execution errors consequential for all later states. ToolQA is the most agentic condition: the model must traverse a hidden chain by issuing inspection calls, with each subsequent target revealed only after the preceding tool interaction. Consequently, the full path cannot be precomputed in advance.

This distinction is important. The first three tasks could be criticized as structured-output or symbolic-execution probes rather than representative agent workloads. ToolQA addresses that concern by using an unmodified ReAct-style loop in which the model selects tool calls while operating under partial information. The paper therefore evaluates both controlled sequential computation and self-directed tool use.

The context manipulation is designed to separate step count from context length. In the natural regime, the model receives one instruction per turn with the complete multi-turn history. In the compressed regime, the same number of operations and turns is retained while prior context is shortened using carried state and windowed history. In the padded regime, all operations for a horizon are presented in a single turn. Reusing the same underlying task instances across regimes permits paired rather than independent-sample comparisons.

The statistical procedure is appropriate to the study’s goals. Wilson intervals are used for success rates near zero or one, and AIC determines whether geometric, threshold, or linear functional forms better describe each model–task cell. Temperature is fixed, seeds are recorded, and the complete code, prompts, task generators, raw trajectories, and analysis scripts are released.

Geometric degradation across horizons

The primary result is that geometric decay is the best-fitting functional form in 28 of 36 model–task cells. The remaining cells are better described by threshold or cliff-like behavior, in which performance remains high until a critical horizon and then declines sharply. Thus, the paper does not claim that all agents exhibit a single universal curve; rather, it identifies geometric compounding as the dominant pattern under the tested conditions.

The task-specific ordering is also informative. Refchain produces the shallowest degradation for stronger models, Ledger shows intermediate degradation, and Cipher produces the steepest decline. Cipher’s behavior follows directly from its dependency structure: one incorrectly applied edit can invalidate the final answer even if subsequent reasoning is locally correct. This result implies that horizon alone is insufficient as a reliability descriptor. The dependency topology and error recoverability of the task materially affect the location and slope of the degradation curve.

The paper makes a strong claim that deserves careful qualification: because no model reaches per-step reliability of exactly one on a non-trivial task, eventual collapse is mathematically unavoidable under the independent-step approximation. This conclusion is valid as a statement about the fitted compounding model, but its empirical force depends on how well the approximation represents actual trajectories. The later finding that hazard increases over time suggests that the constant-rr geometric model is, if anything, optimistic at long horizons.

Collapse on the agentic tool-use task

ToolQA yields the paper’s most operationally consequential result. Eight of nine models perform at or near ceiling at the shortest horizon, indicating that the task is not intrinsically difficult when only a few steps are required. Nevertheless, by horizon 16 every model has fallen to a small fraction of its initial performance.

The Qwen2.5-72B trajectory illustrates the inadequacy of short-horizon pass rates. It achieves perfect performance at horizons 2 and 4 but succeeds in only approximately one in eight trials at horizon 16. A benchmark restricted to the shorter conditions would therefore classify the model as highly reliable while providing no evidence of the subsequent failure regime. The paper’s broader claim is that a high aggregate score at short horizons provides little information about long-horizon reliability unless the decay curve itself is measured.

The result is particularly significant because it holds for proprietary systems as well as open models. It also occurs in a self-directed tool-use loop rather than only in tasks where the sequence of operations is externally specified. However, the authors note that the ToolQA API budget was exhausted before longer conditions could be completed. The reported collapse at horizon 16 is therefore a lower bound on the severity of degradation, not an estimate of the eventual asymptote.

Scaling and per-step reliability

Per-step reliability increases with model scale across the open-model ladder, with a reported Pearson correlation of r=0.36r=0.36 against the logarithm of parameter count. This is a positive but modest association, and it does not establish that parameter count is the primary determinant of reliability. The study also reports diminishing returns and observes that proprietary models cluster near the strongest open models rather than substantially exceeding them.

The operational interpretation is more important than the correlation itself. Scaling improves the per-step parameter that controls compounding, but it does not eliminate residual error. Consequently, even a model that appears highly capable on short tasks can remain unsuitable for sufficiently long dependent workflows. The result supports reliability budgeting based on measured step-level performance rather than model class, parameter count, or aggregate benchmark rank.

The paper’s phrasing that per-step reliability “saturates well below 1” should be understood as an empirical observation within the evaluated models, tasks, and decoding configuration. The study does not establish a universal scaling law or an irreducible upper bound on per-step reliability.

Increasing hazard within trajectories

The data provide evidence against a strictly constant per-step hazard. Mean per-step accuracy decreases from 0.58 in the first third of long trajectories to 0.44 in the final third. The first error occurs approximately one-third of the way through a trajectory, and recovery after an error is uncommon.

This pattern suggests that geometric decay is a useful aggregate approximation but not a complete mechanistic model. If later steps are less reliable than earlier steps, then a stationary-rr projection will overestimate long-horizon success. Errors appear to behave as absorbing states: once an invalid state, incorrect interpretation, or malformed tool action enters the trajectory, later steps generally propagate rather than repair it.

The paper also identifies format and tool-call drift in 21% of trajectories. This is an important decomposition because it shows that long-horizon degradation is not exclusively a reasoning failure. Interface-conformance errors, parsing failures, and invalid actions compound with state-tracking and decision-making errors. In production systems, step-level validation could therefore detect both semantic and interface failures before they become end-to-end task failures.

Step count versus context length

The context-regime experiment rejects the paper’s principal alternative explanation: that degradation is primarily caused by long contexts or lost-in-the-middle effects. The natural regime has a logit slope of −0.44-0.44 per horizon doubling. The compressed regime, which limits retained context, performs worse and has a steeper slope of −0.69-0.69; this difference is reported as statistically significant with p=3×10−6p=3 \times 10^{-6}. The padded regime has a slope of −0.40-0.40, statistically indistinguishable from the natural regime with rr0.

These results imply that the number of dependent operations, rather than the amount of context presented to the model, is the dominant driver under the experimental conditions. Collapsing all operations into one prompt does not remove the degradation, although it begins from a lower baseline. Similarly, truncating history does not improve reliability and appears to worsen it, presumably because the agent loses the scaffold of its prior state and reasoning.

The practical conclusion is direct: naive context truncation is not a general reliability intervention for long-horizon agents. It may reduce latency or token cost while simultaneously removing information required for state maintenance. The result does not show that context length is irrelevant in all agentic systems; it shows that, in this paired design, context reduction fails to explain the dominant degradation pattern.

The claimed benchmark–production gap

The paper projects measured per-step reliability onto representative benchmark horizons and reports success estimates of 0.42 at eight steps, 0.36 at 15 steps, 0.33 at 20 steps, 0.30 at 30 steps, and 0.24 at 100 steps. These figures are used to argue that common benchmarks sample horizons before the major collapse becomes visible.

There is, however, an internal quantitative inconsistency in the supplied manuscript. The paper states that the projection uses mean per-step reliability rr1, but direct geometric compounding would yield rr2, not 0.42, and would produce substantially smaller values at the longer horizons. The reported sequence is also not consistent with a single constant per-step reliability: the implied reliability varies considerably across the listed horizons. The benchmark-gap argument may still be valid, but these particular numerical projections require clarification, such as a definition of rr3 different from per-step success, a nonstandard horizon mapping, or a transcription error.

This inconsistency is consequential because the benchmark-gap quantification is one of the paper’s headline results. The qualitative conclusion—that short-horizon evaluation can substantially overstate production reliability—is supported by the trajectory results and the observed decay curves. The exact projected values, however, should not be treated as reproducible until the underlying calculation is reconciled with the stated geometric model.

Limitations and open questions

The strongest limitation is ecological validity. The task families are synthetic, enabling exact verification and controlled horizon manipulation, but they do not fully represent open-ended software engineering, web navigation, organizational workflows, or multi-agent production systems. ToolQA improves the study’s agentic relevance, yet it remains a constrained hidden-chain environment.

The model sample is broad but limited to nine systems. Three additional small models were excluded because their hosted routes failed to follow the structured response protocol, which may introduce selection effects. The fixed, moderate decoding temperature also leaves sensitivity to temperature, sampling strategy, retries, verifier design, and inference-time search unresolved. These factors could alter both per-step reliability and recovery after errors.

The geometric form is selected in most, but not all, model–task cells. The increasing hazard observed within trajectories further indicates that a single stationary parameter is an approximation. A key open question is whether production agents with explicit state validation, rollback, branching, retries, or external memory exhibit the same decay law, or whether orchestration can convert absorbing errors into recoverable failures. Another unresolved issue is whether the compressed-regime penalty results from loss of state, loss of intermediate reasoning traces, or a particular implementation of context summarization.

Conclusion

The paper establishes a clear empirical relationship between task horizon and LLM-agent reliability. Across controlled tasks and nine models, success usually declines geometrically, while the agentic ToolQA condition shows collapse from near-ceiling performance to near failure within 16 steps. Larger models improve per-step reliability but do not eliminate compounding error, and later trajectory steps are less reliable than earlier ones.

Its most defensible practical conclusion is that agent evaluation should report success as a function of horizon and should measure step-level reliability directly. Context truncation alone is not supported as a remedy and may worsen performance. The qualitative evidence for horizon-driven degradation is strong, although the stated benchmark-gap projections contain a numerical inconsistency that must be corrected before those figures can support precise operational planning.

Paper to Video (Beta)

No one has generated a video about this paper yet.

Whiteboard

Explain it Like I'm 14

1. What is this paper about?

This paper studies why AI agents powered by LLMs often work well in short tests but fail during long, complicated tasks.

An LLM agent is a program that can:

  • think about what to do,
  • use tools,
  • observe the result,
  • and then choose its next action.

For example, an agent might search a website, read information, fill out a form, and check whether the task worked. Each action depends on the actions before it.

The paper’s main idea is that agents become less reliable as the number of required steps increases. The authors describe this as agents “rotting” over long tasks.

2. What questions did the researchers ask?

The researchers wanted to answer four main questions:

  1. How does success change as a task gets longer? Does success decrease slowly, suddenly, or according to a predictable pattern?
  2. Is the problem caused by the number of steps or by the amount of text the agent must remember? In other words, does an agent fail because it has too many actions to perform, or because its conversation history becomes too long?
  3. Do larger and more advanced models remain reliable for longer?
  4. Do ordinary benchmarks give an overly optimistic picture of how agents perform in real-world work?

3. How did the researchers conduct the study?

The tasks

The study tested nine different LLMs. Six were openly available models, ranging from relatively small models with 1.2 billion parameters to very large models with 671 billion parameters. The other three were commercial systems accessed through their normal online services.

The researchers created four types of tasks:

  • Ledger: The agent had to keep track of changing bank-account-like balances.
  • Refchain: The agent had to follow a variable through many changes and report its final value.
  • Cipher: The agent had to perform a series of edits on a short piece of text.
  • ToolQA: The agent had to explore a hidden chain by using tools. It could not plan the entire route at the beginning because it only learned the next location after inspecting the current one.

ToolQA was especially important because it acted more like a real AI agent. The agent had to repeatedly decide what tool to use and what to do next.

In total, the researchers analyzed 10,664 agent attempts, called trajectories. A trajectory is simply the complete sequence of actions an agent takes during one task.

Different context conditions

The researchers compared three ways of presenting information:

  • Natural: The agent received one instruction at a time and kept the full conversation history.
  • Compressed: Some earlier information was removed or shortened.
  • Padded: All the instructions were given together in one large message.

This comparison helped separate two possible causes of failure:

  • the agent having to perform many dependent steps, and
  • the agent having to process a large amount of text.

Measuring success

The tasks had exact, computer-checked answers. This meant the researchers did not need another AI system—or a human—to judge whether an answer was correct.

They also used statistical tests to compare different possible patterns of decline, such as:

  • a straight-line decrease,
  • a sudden cliff after a certain point,
  • or a geometric decline.

A geometric decline is similar to what happens when a small loss is repeated again and again. For example, if an agent has a 90% chance of getting each step right, then completing 10 steps correctly requires success at every step:

0.910≈0.350.9^{10} \approx 0.35

So a 90% reliable step-by-step system would complete the whole 10-step task only about 35% of the time. This is because small mistakes multiply together.

4. What did the researchers find?

Success decreases rapidly as tasks get longer

The study found that task success usually followed a geometric pattern. This pattern appeared in 28 out of 36 model-and-task combinations.

The key idea is called per-step reliability. This means the chance that the agent performs one step correctly. Even when this chance is fairly high, it is almost never exactly 100%.

If every step has a small chance of failure, then a long task becomes much less likely to succeed completely.

For example, an agent that is 95% reliable at each step might seem very good. But over 50 dependent steps:

0.9550≈0.080.95^{50} \approx 0.08

That means the complete task would succeed only about 8% of the time under this simple model.

Larger models helped, but did not solve the problem

Larger models generally had better per-step reliability. However, none of the models reached perfect reliability on the more difficult tasks.

This means that bigger models can postpone failure, but they cannot remove the basic problem. If a task is long enough, even a very strong model will eventually have a high chance of making a mistake.

The researchers also found that improvements became smaller as models got larger. This suggests that simply increasing model size may not be enough to make agents dependable on very long tasks.

The agentic tool-use task showed the sharpest collapse

The ToolQA task produced the most worrying result.

At very short horizons, most models performed almost perfectly. However, by 16 steps, every model had fallen to a small fraction of its original success rate.

One example was the Qwen2.5-72B model. It completed the task perfectly at horizons of 2 and 4 steps, but by 16 steps it succeeded in only about one out of eight attempts.

This shows why short benchmarks can be misleading. A model can look completely reliable in a short test while being very unreliable in a longer workflow.

Errors often became permanent

The agents did not simply make occasional mistakes and then recover. Instead, an early mistake often caused later actions to go wrong as well.

The researchers found that average step accuracy fell from about 58% near the beginning of long tasks to about 44% near the end. Once an agent took a wrong turn, it rarely found its way back.

This is similar to following a map:

  1. You make a small wrong turn.
  2. You use your new location as if it were correct.
  3. Every later decision is based on the wrong location.
  4. You end up farther and farther from the destination.

The paper also found that problems with tool-call format—such as producing an action that the system could not understand—affected about 21% of trajectories. These technical mistakes added to reasoning mistakes.

The number of steps mattered more than the amount of context

The researchers expected that shortening the conversation history might help the agent by reducing the amount of information it had to process. Instead, shortening the context made performance worse.

Giving all the instructions at once produced a decline similar to the natural multi-turn condition, although it started from a lower level of success.

These results suggest that the main problem is not simply that the agent is “lost in the middle” of a long conversation. The deeper problem is that it must carry out many dependent actions correctly.

The paper therefore warns that a common production shortcut—cutting off old conversation history to save money or reduce delay—may actually reduce reliability.

Common benchmarks may overestimate real-world reliability

Using the measured average per-step reliability of 0.61, the researchers estimated success at different task lengths:

Approximate task length Estimated complete-task success
8 steps 42%
15 steps 36%
20 steps 33%
30 steps 30%
100 steps 24%

These numbers are approximate projections, but they show the main point: a system may look strong on short evaluations while performing poorly on long production workflows.

5. Why are these findings important?

The paper argues that reporting only one overall score—such as “the agent succeeded on 80% of tasks”—is not enough.

That score does not tell us:

  • how many steps the tasks required,
  • how often the agent made mistakes at each step,
  • whether errors could be detected early,
  • or how performance changes as tasks become longer.

The authors recommend testing agents at many different task lengths and measuring their per-step reliability. Organizations should also set a reliability budget. This means estimating how many steps an agent can safely complete before the risk of total failure becomes too high.

Another important recommendation is to check the agent after each step whenever possible. For example, a system could verify that:

  • a tool call used the correct format,
  • a number is still consistent,
  • a file was changed correctly,
  • or the agent’s current state makes sense.

Finding a mistake early is much cheaper and safer than discovering at the end that the entire task failed.

6. Limitations of the study

The study has some important limits.

The tasks were carefully designed and mostly synthetic, meaning they were created for the experiment rather than taken directly from messy real-world work. This made it possible to measure success exactly, but real jobs may include more kinds of difficulty.

The researchers tested nine models, which is a useful sample but not enough to prove that every current or future model behaves in exactly the same way.

The study also used a particular setting for randomness, called the decoding temperature. Different settings might change how often agents make mistakes.

Finally, the ToolQA experiment could not test the very longest tasks because of limits on the available tool-use budget. Therefore, the reported collapse by 16 steps may actually be less severe than what would happen at even longer lengths.

7. Overall conclusion

The paper’s main message is simple:

A small chance of failure at each step can become a very large chance of failure when many steps are connected together.

AI agents are often impressive on short tasks, but their reliability can drop quickly during long workflows. Larger models help, but they do not completely solve the problem. The main cause appears to be the number of dependent actions, not just the length of the conversation.

This could affect how companies build AI systems for customer service, software development, web browsing, office work, and other jobs. In the future, developers may need to focus less on a single impressive benchmark score and more on questions such as:

  • How reliable is the agent at each step?
  • How long can it work before failure becomes likely?
  • Can mistakes be detected and corrected early?
  • Does its performance remain acceptable under real production conditions?

The research suggests that better testing, step-by-step checking, and carefully designed agent workflows could make AI systems safer and more dependable.

Knowledge Gaps

Knowledge gaps, limitations, and open questions

  • External validity to real production workflows remains unresolved. The evaluation uses synthetic, oracle-verifiable tasks, so it is unknown whether the same decay curves and failure mechanisms hold for open-ended software engineering, web navigation, customer support, research, or operational workflows.
  • The ToolQA result is incomplete at longer horizons. API-budget exhaustion prevented evaluation beyond horizon 16, so the study cannot determine whether success continues to decay geometrically, reaches a floor, or exhibits a different regime at substantially longer horizons.
  • The universality of the geometric law is unestablished. Geometric decay was preferred in 28 of 36 model–task cells, while the remaining cells followed threshold or cliff-shaped patterns; the paper does not identify which task, model, or environmental properties determine the functional form.
  • The independence assumption behind per-step reliability is questionable. Errors may be correlated across steps because of shared state, model-specific failure modes, task difficulty, or trajectory history, but the study does not quantify these dependencies or test alternative models such as state-dependent hazards or Markov processes.
  • The causes of the increasing within-trajectory hazard are not isolated. The decline in per-step accuracy from the first to the last trajectory third could reflect context accumulation, error propagation, fatigue-like degradation, changing task difficulty, tool-call drift, or selection effects; the design does not separate these explanations.
  • Recovery after an error is insufficiently measured. The claim that errors are effectively absorbing states is based on observed low recovery, but the study does not systematically test recovery strategies, distinguish recoverable from unrecoverable errors, or estimate recovery probability conditional on error type.
  • The step-count and context-length effects are not fully disentangled. The natural, compressed, and padded regimes differ in more than raw context length, including information availability, formatting, turn structure, and possibly prompt distribution; these confounds limit the causal interpretation that step count is the primary driver.
  • The compressed regime may remove information needed for successful execution. Because the shortened context prunes the agent’s prior reasoning scaffold, its worse performance may reflect loss of essential state rather than a general harmful effect of context truncation. The study does not compare multiple compression methods or provide sufficient-state summaries.
  • The scope of context effects is narrow. Only the particular natural, compressed, and padded regimes and horizons from 2 to 32 are tested; it remains unknown how retrieval augmentation, hierarchical memory, summaries, external state stores, larger context windows, or structured scratchpads affect long-horizon degradation.
  • Decoding sensitivity is unexamined. Results were obtained at a fixed moderate temperature, leaving open how temperature, top-p, sampling versus greedy decoding, repetition controls, and inference-time scaling affect per-step reliability and hazard acceleration.
  • Inference-time compute is not controlled or analyzed. Models may differ in reasoning-token budgets, hidden deliberation, tool-call budgets, latency constraints, and API behavior; the study does not determine whether additional inference-time computation mitigates horizon-related collapse.
  • Model coverage is too limited to support broad scaling claims. Only nine models from six vendor families were tested, and three additional models were excluded for protocol failures. The reported relationship between parameter count and per-step reliability is weakly characterized and may be confounded by training data, architecture, instruction tuning, tool-use training, or release date.
  • The parameter-scaling relationship is statistically underdeveloped. The reported Pearson correlation of r=0.36r=0.36 is based on a small and potentially non-independent sample, and the analysis does not estimate scaling exponents, confidence intervals, nonlinear saturation, or task-specific scaling trends.
  • Proprietary-model comparisons lack sufficient reproducibility detail. Hosted systems may change over time and may differ in hidden system prompts, routing, safety filters, rate limits, and backend versions; the study does not establish whether the results are stable across API snapshots or repeated evaluation periods.
  • Temperature and API nondeterminism may undermine the interpretation of “determinism.” Although temperature and seeds are recorded, proprietary endpoints and tool interactions may remain nondeterministic; the extent of residual stochasticity is not quantified.
  • The task families provide limited coverage of agent capabilities. Ledger, Refchain, and Cipher primarily test controlled state tracking and procedural execution, while ToolQA tests hidden-chain traversal; planning, negotiation, perception, multimodal grounding, uncertainty management, long-term memory, and dynamic environmental change are largely absent.
  • Task difficulty and horizon are not independently varied enough. The study does not determine whether degradation is driven by the number of dependent actions itself or by the cumulative information, branching factor, ambiguity, or reasoning complexity associated with longer tasks.
  • The effect of branching and replanning is unexplored. ToolQA uses a hidden chain, but the paper does not test environments with multiple viable plans, irreversible actions, stochastic outcomes, changing goals, or opportunities to replan after partial failure.
  • Tool-call and formatting failures are not causally separated from reasoning failures. Format/tool-call drift affects 21% of trajectories, but the analysis does not report how much of total degradation would disappear with robust parsers, constrained decoding, retries, schema repair, or tool-call validation.
  • The reported “21% of trajectories” metric is insufficiently granular. Failure rates are not broken down by model, task, horizon, tool type, error category, or trajectory position, preventing researchers from identifying which interface interventions are most valuable.
  • The study does not test reliability interventions. Step-level verification, retries, reflection, planning, checkpointing, external memory, ensemble sampling, verifier models, and human escalation are recommended or implied but not experimentally evaluated.
  • The practical reliability-budgeting framework is not validated in deployment. The projected success values based on r=0.61r=0.61 are not compared with independently observed success on real production workloads, and the paper does not quantify calibration error when the geometric model is applied outside the study tasks.
  • The benchmark-horizon projections may be misleading. Mapping GAIA, WebArena, tau-bench, SWE-bench, and OSWorld to representative step counts assumes comparable definitions of a “step,” but action granularity differs substantially across benchmarks; the paper does not justify or validate these mappings.
  • The projected benchmark gap uses a pooled mean reliability. A single r=0.61r=0.61 may obscure large differences across models, tasks, regimes, and horizons. It is unclear whether pooled projections preserve the distribution of actual task success or substantially misestimate it.
  • Uncertainty around projected long-horizon success is not reported. The paper provides point estimates such as $0.24$ at 100 steps but does not provide confidence or prediction intervals that account for uncertainty in per-step reliability, model heterogeneity, and the nonconstant hazard.
  • The study does not investigate task-instance heterogeneity. Some instances may be intrinsically more difficult than others, but the analysis does not estimate instance-level random effects or determine whether a small subset of difficult trajectories drives the observed decay.
  • The role of prompt design is unresolved. The conclusions may depend on the specific prompts, ReAct formatting, tool descriptions, and state representations used; alternative prompt structures and agent controllers are not compared.
  • The relationship between context truncation and cost is not quantified. Although truncation is described as a production shortcut, the study does not measure the trade-off among context length, token cost, latency, memory, and reliability or identify regimes where truncation could still be beneficial overall.
  • Cross-task transfer and adaptation are untested. It remains unknown whether agents can learn from prior failed trajectories, adapt their policies during deployment, or improve per-step reliability through fine-tuning on long-horizon traces.
  • The effect of external verification is unknown. Because the tasks have exact simulators, future work should test whether exposing intermediate verifiers to the agent changes the hazard curve, rather than using verification only for offline scoring.
  • The paper does not establish whether the observed failures are model-internal or system-level. Failures could arise from the LLM, prompt/controller design, tool interface, serialization, latency, API behavior, or orchestration runtime, but these components are not systematically ablated.
  • Longitudinal model drift is not examined. The study provides a snapshot of model behavior and cannot determine whether repeated deployment, backend updates, prompt changes, or continual adaptation alter per-step reliability and long-horizon degradation over time.

Practical Applications

Immediate Applications

The paper’s findings can be translated into deployment practices without waiting for new model architectures, provided organizations can instrument agent trajectories and define step-level checks.

  • Horizon-aware production evaluation for LLM agents — Software and enterprise AI. Replace single aggregate pass rates with success curves measured at horizons relevant to the intended workflow, such as 8, 16, 32, and 100 dependent steps. Teams can estimate per-step reliability using the released code and fit geometric or threshold models before approving an agent for production.
    • Potential tool: An evaluation dashboard reporting P(success | horizon), per-step reliability, first-error position, recovery rate, and tool-call validity.
    • Dependency: The evaluation tasks must resemble the production workflow; the paper’s synthetic tasks may not fully predict open-ended behavior.
  • Reliability budgets for agent orchestration — Cloud platforms and workflow automation. Treat each dependent action as consuming part of a reliability budget. If an agent’s estimated per-step reliability is rr, an initial planning approximation is P(success)≈rHP(\text{success}) \approx r^H, where HH is the number of dependent steps. Workflows that exceed the acceptable failure threshold can be redesigned, shortened, or routed to human approval.
    • Example: A system requiring at least 80% end-to-end success should not automatically authorize a 20-step workflow based only on a high short-horizon benchmark score.
    • Dependency: The geometric model is an approximation; the observed increase in hazard over time means teams should budget more conservatively than the formula suggests.
  • Step-level verification and early stopping — Software reliability and MLOps. Add validators after high-risk or state-changing actions: schema checks, balance reconciliation, permission checks, tool-result validation, invariant tests, and state consistency checks. If an intermediate result fails validation, the system should stop, retry, roll back, or escalate rather than continue from a corrupted state.
    • Potential products: Agent middleware with checkpointing, rollback, invariant monitoring, and automatic human escalation.
    • Dependency: A cheap and reliable verifier must exist for the relevant step. Verification is easier for structured operations than for subjective outputs.
  • Monitoring for trajectory degradation — Operations and incident response. Monitor intermediate indicators rather than waiting for end-to-end task failure. Useful metrics include invalid tool-call rate, repeated actions, state divergence, increasing latency, unsupported arguments, recovery attempts, and the position of the first anomaly.
    • Operational workflow: Trigger a circuit breaker when the agent exceeds a maximum number of retries or violates a state invariant.
    • Dependency: Monitoring systems need access to complete or sufficiently detailed trajectory logs, raising privacy, security, and storage concerns.
  • Avoid naive context truncation as a reliability intervention — LLM application engineering. The study finds that compressed context performed worse than natural context and that shortening history did not solve step-count degradation. Teams should not assume that reducing context alone will make long-running agents more reliable.
    • Practical alternative: Preserve compact, verified state summaries while retaining critical provenance, tool outputs, and decisions rather than indiscriminately deleting history.
    • Dependency: The result was obtained under the tested tasks and context regimes; specialized summarization or memory systems may behave differently.
  • Workflow decomposition and bounded sub-agents — Business process automation. Split a long workflow into shorter independently verifiable stages, such as planning, retrieval, execution, and reconciliation. Each stage can produce a checked artifact before the next stage begins.
    • Example: In software maintenance, separate issue interpretation, code modification, testing, and deployment approval rather than allowing one agent to perform all actions continuously.
    • Dependency: Decomposition introduces handoff overhead and may create additional interfaces where information can be lost.
  • Safer tool-use interfaces — API platforms and developer tools. Since format and tool-call drift affected 21% of trajectories, production systems can use strict JSON schemas, typed function calls, enumerated arguments, server-side validation, idempotent APIs, and explicit error messages.
    • Potential product: A tool gateway that validates every call, rejects malformed actions, records provenance, and offers only authorized operations.
    • Dependency: Interface validation prevents syntax and contract failures but cannot guarantee that a semantically valid action is appropriate.
  • Human-in-the-loop thresholds for consequential tasks — Healthcare, finance, legal services, and public administration. Require approval at predefined horizons or before irreversible actions. For example, an agent may gather information and prepare a recommendation but require a person to authorize a financial transfer, clinical order, benefits decision, or production deployment.
    • Dependency: The organization must define acceptable risk, escalation latency, audit requirements, and responsibility for the final decision.
  • Benchmark reporting resolved by horizon — Academia and internal model selection. Researchers and procurement teams can report success at multiple horizons, repeated-trial variance, cost, latency, and recovery behavior rather than a single pass-rate number.
    • Potential artifact: A standardized “agent reliability curve” accompanying model cards, benchmark leaderboards, and vendor evaluations.
    • Dependency: Benchmarks need reproducible task definitions, exact success criteria, and enough trials to distinguish a genuine curve from sampling noise.
  • Reproducible reliability testing — Research engineering and model governance. The released prompts, seeds, raw trajectories, task generators, and analysis scripts can be used as a starting point for reproducing the study or adapting its protocol to internal agents.
    • Dependency: Reproduction requires compatible model interfaces, stable decoding settings, and careful separation of benchmark-specific behavior from general capability.
  • Safer personal productivity assistants — Daily life and consumer software. For email, scheduling, shopping, or file-management assistants, users can limit autonomous action counts, require confirmation for irreversible operations, and preserve an audit trail of intermediate actions.
    • Example: An assistant can draft and organize a calendar change but ask for confirmation before sending invitations or canceling appointments.
    • Dependency: Consumer systems need understandable failure messages and privacy-preserving logs; users may not tolerate frequent interruptions.

Long-Term Applications

These applications require further validation on real-world workflows, improved recovery mechanisms, or infrastructure that can maintain reliability across much longer horizons.

  • Reliability-aware agent planners — Autonomous software and enterprise orchestration. Future planners could estimate the reliability cost of candidate plans and select the shortest safe execution path. A planner might prefer a five-step verified workflow over a nominally cheaper 30-step sequence with a higher probability of compounding failure.
    • Potential tool: A risk-sensitive planner that assigns each action an estimated failure probability and optimizes expected task completion, cost, and reversibility.
    • Dependencies: Reliable per-action estimates, calibrated uncertainty, task-specific failure models, and the ability to compare alternative plans.
  • Active error recovery and state reconstruction — Agent architectures and robotics. Because the paper finds that errors often behave like absorbing states, future agents could detect when their internal state has diverged and reconstruct it from external ground truth, logs, or checkpoints.
    • Potential methods: Periodic state reconciliation, rollback to the last verified checkpoint, independent critic agents, tool-based state queries, and recovery policies trained on failure trajectories.
    • Dependency: Recovery must be demonstrably better than continuing or restarting. Multi-agent critics may themselves introduce additional compounding steps and costs.
  • Adaptive horizon control — Robotics, industrial automation, and autonomous operations. Agents could dynamically reduce autonomy when estimated hazard rises—for example, switching from autonomous execution to supervised mode after repeated tool errors or after a trajectory reaches a calibrated risk boundary.
    • Example: A warehouse robot or infrastructure agent can pause and request operator guidance before performing a long sequence of uncertain actions.
    • Dependencies: Reliable online hazard estimation, safe stopping procedures, low-latency human intervention, and domain-specific safety certification.
  • Training objectives focused on long-horizon reliability — AI research. Model training could optimize not only next-step accuracy but also recovery after mistakes, state persistence, valid tool use, and success over extended trajectories. Training data should include realistic error states rather than only successful demonstrations.
    • Potential approaches: Trajectory-level reinforcement learning, process supervision, verifier-guided training, synthetic failure injection, and curriculum learning over increasing horizons.
    • Dependency: Training signals must measure genuine task progress rather than exploit weaknesses in simulators or verifiers. The study does not establish which training method will improve the measured hazard.
  • Persistent, verifiable agent memory — Personal assistants and enterprise knowledge systems. Since compressed context can remove the scaffold of prior reasoning, future systems could maintain an external memory containing verified facts, decisions, provenance, and unresolved assumptions rather than relying solely on raw conversation history.
    • Potential product: A provenance-aware memory store that distinguishes user-confirmed facts, tool-observed state, model inferences, and expired information.
    • Dependencies: Accurate memory updates, conflict resolution, access control, deletion guarantees, and protection against storing hallucinated or sensitive information.
  • Long-horizon benchmarks for consequential workflows — Academia and policy. Benchmark designers could develop suites that vary horizon independently from task difficulty and include recovery, partial success, reversibility, cost, latency, and safety violations. Domains could include clinical administration, cybersecurity, software operations, financial workflows, and public-sector services.
    • Policy value: Regulators and procurement bodies could require horizon-resolved evidence for high-impact autonomous systems instead of accepting short-task benchmark scores.
    • Dependencies: Access to representative tasks, domain experts, robust simulators, privacy-preserving data, and consensus on acceptable failure rates.
  • Reliability-based model and vendor contracts — Enterprise procurement and governance. Service-level agreements could specify maximum supported horizons, tool-call validity rates, recovery behavior, and end-to-end success under defined workloads rather than relying on generic model benchmark scores.
    • Potential contract metrics: Success at specified horizons, p95 time to first error, invalid-action rate, rollback success, and performance under repeated execution.
    • Dependency: Metrics must be auditable across model updates, decoding changes, tool versions, and workload drift.
  • Multi-agent architecture design based on error compounding — Software, robotics, and decision support. The study’s findings suggest that adding more agents or handoffs is not automatically beneficial: every dependent interaction may introduce another failure opportunity. Future systems could allocate responsibilities to independent modules only when the added verification benefit exceeds the additional coordination risk.
    • Potential workflow: Use specialized agents for planning and execution, but require a deterministic verifier or external simulator between stages.
    • Dependency: Multi-agent benefits must be measured end to end; specialization alone does not remove the geometric accumulation of errors.
  • Formal safety cases for autonomous decision-making — Healthcare, energy, transportation, and finance. Long-horizon reliability measurements could become one component of a safety case documenting operating limits, escalation policies, recovery capability, and worst-case behavior.
    • Example: An energy-management agent might be authorized to optimize within a bounded horizon but require human approval for actions affecting grid stability or physical equipment.
    • Dependencies: Domain-specific regulation, formal or statistical guarantees stronger than the paper’s empirical estimates, validated simulators, and clear accountability for failures.
  • Personal assistants that learn individual reliability profiles — Daily life and accessibility technology. Long-term assistants could estimate which classes of tasks a user can safely delegate, automatically limiting autonomy for workflows with high step sensitivity, ambiguous state, or irreversible consequences.
    • Example: The assistant may autonomously reorder routine household items but require confirmation for medical purchases, financial transfers, or communications to new recipients.
    • Dependencies: Continuous calibration, protection against demographic or user-specific bias, transparent explanations, and safeguards against overconfidence from sparse personal data.

Overall, the most immediately actionable implication is to treat horizon and intermediate-step reliability as first-class production metrics. The paper supports using its geometric model as an initial planning tool, but its synthetic tasks, fixed decoding conditions, and limited model sample mean that deployment decisions should also include domain-specific trials, monitoring, verification, and conservative human escalation.

Glossary

  • Absorbing state: A state that, once entered, cannot be exited; here, an error that permanently propagates through the remaining trajectory. “errors behave as absorbing states”
  • AIC (Akaike Information Criterion): A statistical model-selection criterion that balances goodness of fit against model complexity. “Model selection among geometric, threshold, and linear decay shapes uses AIC”
  • Agent orchestration: The coordination and control of agents, tools, prompts, and workflow steps in an AI system. “module readiness or incident response on a production agent-orchestration runtime”
  • Agentic loop: An iterative control process in which an agent reasons, acts through tools, observes results, and continues toward a goal. “a genuinely agentic tool-use loop”
  • Aggregate pass-rate metric: A single overall success percentage that does not distinguish performance by task length or other conditions. “reliability budgeting in place of aggregate pass-rate metrics”
  • API budget: A limit on the number of requests, tokens, or computational resources available through an application programming interface. “The agentic family’s API budget ran out”
  • Benchmark horizon: The number of dependent steps represented in an evaluation benchmark. “Projecting measured reliability onto representative benchmark horizons”
  • Ceiling effect: A measurement limitation in which scores cluster near the maximum possible value, concealing differences in capability. “eight of the nine models solve the task at or near ceiling”
  • Confidence interval: A statistical range estimating the uncertainty around a measured quantity. “Confidence intervals on per-cell success use the Wilson score interval”
  • Compositional complexity: The difficulty created by combining multiple operations, reasoning steps, or components into a single task. “transformer accuracy decays sharply with compositional complexity”
  • Context regime: A defined experimental condition governing how much conversational or generated context an agent receives. “the three context regimes used to separate step count from context length”
  • Context window: The maximum amount of input or conversational history that a LLM can process at once. “Bounding the context window steepens decay”
  • Decoding temperature: A parameter controlling the randomness of a LLM’s token selection during generation. “this study measures reliability under a fixed, moderate decoding temperature”
  • Diminishing returns: A pattern in which additional resources produce progressively smaller improvements in performance. “Scale and frontier engineering buy per-step reliability, but with diminishing returns”
  • Ecological validity: The extent to which experimental findings generalize to realistic environments and real-world tasks. “broader ecological validity against fully open-ended real-world agentic workflows”
  • End-to-end task success: Successful completion of an entire workflow, including all required intermediate steps. “A monitoring pipeline built around end-to-end task success”
  • Evidence trail: Documentation linking reported findings to their underlying data or sources. “The repository includes an evidence trail document”
  • Functional form: The mathematical shape used to describe the relationship between variables. “Selecting among geometric, threshold, and linear functional forms”
  • Geometric decay: A pattern in which a quantity is repeatedly multiplied by a constant factor, producing exponential decline. “Success follows a geometric law”
  • Ground-truth simulator: A programmatic environment that precisely determines whether an agent’s actions satisfy the task requirements. “a tool-using agent loop paired with a ground-truth simulator that verifies success exactly”
  • Hazard rate: The conditional probability or risk of failure at a particular step, given that failure has not occurred earlier. “A pure geometric decay law assumes a constant per-step hazard rate”
  • Horizon: The number of dependent steps an agent must complete correctly for a task to succeed. “A task has a horizon: the number of dependent steps an agent must execute correctly to succeed”
  • Inference-time compute: Computational resources expended while generating an answer or completing a task, rather than during model training. “gains that come from spending more inference-time compute per task”
  • Instruction-tuned model: A LLM further trained to follow natural-language instructions. “Nine instruction-tuned models across six vendor families”
  • Intervention: An experimentally imposed change to a system intended to test a causal explanation. “Both interventions perform worse”
  • Logit slope: The rate of change in the log-odds of an outcome as a predictor changes. “the logit slope per horizon doubling is -0.44 for natural”
  • Long-range retrieval: The ability to locate and use information presented substantially earlier in an input or interaction. “testing long-range retrieval”
  • Lost-in-the-middle: A context-processing failure in which models use information from the beginning and end of a long input more effectively than information in its middle. “contradicting a lost-in-the-middle explanation”
  • Oracle verifiability: The ability to determine task correctness exactly through an objective computational checker rather than subjective judgment. “The task families here are synthetic, chosen deliberately for exact oracle verifiability”
  • Per-step reliability: The probability that an agent performs an individual required step correctly. “Task success follows a geometric law governed by a single per-step reliability parameter”
  • Procedural execution: The ability to carry out a prescribed sequence of operations accurately and in the correct order. “testing procedural execution”
  • ReAct: An agent strategy that interleaves reasoning or planning with actions performed through tools or an environment. “ReAct-style interleaved reasoning-and-acting”
  • Reliability budget: A planned allocation of allowable failure probability across the steps of a multi-step system. “these results argue for horizon-aware evaluation and reliability budgeting”
  • Structured-output probe: An evaluation that tests whether a model produces a required formal output format, often without requiring open-ended interaction. “rather than genuine agentic behavior”
  • Synthetic task family: A deliberately constructed group of related tasks used for controlled experimentation. “Four synthetic task families make up the study”
  • Task horizon: The length of a task measured by the number of dependent operations or steps required. “Benchmarks predominantly sample short-to-medium horizons”
  • Tool-call drift: The gradual failure of an agent to produce valid or appropriate tool invocations during an interaction. “format and tool-call drift”
  • Trajectory: The complete sequence of states, actions, and outputs produced during one agent attempt. “10,664 total analyzed trajectories”
  • Wilson score interval: A confidence-interval method for binomial proportions that performs better than a normal approximation, especially for rates near zero or one. “Wilson score interval rather than a normal approximation”
  • Windowed history: A restricted portion of prior interaction history retained within the model’s available context. “with a shortened context via carried state and windowed history”

Tweets

Sign up for free to view the 3 tweets with 64 likes about this paper.