Papers
Topics
Authors
Recent
Search
2000 character limit reached

AgentTether: Graph-Guided Diagnosis and Runtime Intervention for Reliable LLM Agent Operation

Published 7 Jul 2026 in cs.SE | (2607.06273v1)

Abstract: LLM agents are increasingly used for multi-step, stateful tool-use tasks, yet production reliability remains limited. Unlike static software repair, agent repair must recover dynamic trajectories whose early decisions can propagate into later errors and external state changes. Existing automatic remedies address only part of this problem: blind retry adds no diagnosis, outcome feedback says whether a run failed but not where or why, and self-reflection often lacks grounded evidence to prevent the same failure from recurring. We present AgentTether, a run-time repair framework that automates post-run diagnosis and guided recovery without modifying the underlying agent or environment. AgentTether abstracts each run into Transition Units, links them through a dependency-aware Critical Transition Graph, and localizes failure-critical subtrajectories by combining an offline normal-behavior model with a run-local graph detector. It then converts the localized cause into behavior-scoped guidance backed by cross-iteration Repair Memory, and can optionally apply guarded run-time intervention to keep the correction active during re-execution. The same design can be deployed as an offline diagnostic-and-guidance tool or as an online repair layer. We evaluate AgentTether on 261 tau-bench tasks across three domains with Qwen3.7-max, and test cross-model transfer on Banking with GPT-5.4. On the hardest Banking domain, AgentTether repairs 59.04% (49/83) of initially failed Qwen3.7-max tasks and 65.12% (56/86) of initially failed GPT-5.4 tasks. Overall, AgentTether improves repair effectiveness while reducing agent turns and end-to-end approach tokens, suggesting a practical reliability layer that can wrap existing agent deployments, reduce wasted re-execution, and improve recovery without retraining the agent.

Summary

  • The paper introduces AgentTether, a wrap-around reliability layer that combines dependency-aware Critical Transition Graph diagnosis, repair memory, and guarded runtime intervention to restore failed LLM agent runs.
  • The full system repaired 69.11% of initially failed τ-bench tasks, outperforming blind retry by 26.02 percentage points and improving Banking recovery to 59.04% while using about 13% fewer agent tokens.
  • The results show that accurate diagnosis alone is insufficient because guidance decays during re-execution, while selective intervention helps most in difficult settings but can harm performance when excessive corrections disrupt required actions.

AgentTether addresses a specific gap in LLM agent reliability: the distance between diagnosing why an agent run failed and actually repairing it. The authors, from Nankai University, Tsinghua University, and Microsoft, argue that existing remedies—blind retry, outcome feedback, self-reflection—each cover only part of the repair loop that a human operator performs. The paper's central empirical claim is that even accurate diagnosis does not by itself produce repairs: guidance decays during re-execution, so recovery requires both dependency-aware localization and guarded mid-run intervention (2607.06273).

Motivation: two failure modes of post-hoc repair

The paper grounds its design in a quantitative analysis of 83 initially failed Banking runs under Qwen3.7-max on τ\tau-bench. Two findings drive the architecture:

  • Root causes are far upstream and many-to-many. 94% of failures are behavioral (wrong/missing tool actions), not communication errors. The earliest violated gold action precedes the visible failure by a median of 4 required steps (up to 26) and is strictly upstream in 79% of runs; one root error triggers 3.2 downstream violated checks on average, with 76% of failures violating more than one check. This rules out recency- or frequency-based localization over flat traces.
  • One-shot guidance decays. Across 173 post-feedback reruns, adherence to injected directives starts at 99% but falls to 71% by step 5 and below 50% by step 13 overall—as early as step 8 on Banking. Because deviations surface after the feedback leaves the effective attention window, no improvement to the report itself can prevent them.

These findings motivate the two core mechanisms: graph-based root cause analysis (RCA) and run-time intervention.

Architecture

AgentTether wraps an unmodified agent run with two tracks. Trace acquisition uses monkey-patched LLM-SDK hooks (in the style of distributed tracing) to record full request/response payloads without altering agent behavior.

The post-run track abstracts each trace into Transition Units (TUs), each an Observation–Belief–Action–Feedback cycle, and builds a Critical Transition Graph (CTG) whose nodes are TUs and whose edges include both temporal adjacency and dependency links (shared artifacts, error signatures). Two complementary detectors localize anomalous substructures: an offline heterogeneous graph transformer (HGT) trained self-supervised on 21,143 success-only trajectories from TerminalBench and SWE-smith (deliberately excluding τ\tau-bench domains), scoring subgraph surprise via edge reconstruction and masked-attribute recovery; and a training-free Isolation Forest over 25-dimensional structural, attribute-level, and macro-level TU features. An analyst LLM then converts the pruned evidence packet into a diagnosis (root cause, turning point, recovery hints), which a feedback builder renders as behavior-scoped directives—deliberately general rather than instance-specific to avoid hard-coding answers—plus an injection plan. Cross-iteration Repair Memory records fixed versus unresolved corrections across up to Γ=3\Gamma{=}3 iterations.

The run-time track supervises re-execution at tool-return and text-response hooks via a Check→Decide→Inject pipeline. Check signals cover loop repetition, risk-tiered intent drift (with LLM verification for high-risk candidates), expectation deviation against active guidance, and sparse structural checkpoint reminders. Decide applies three guards—evidence grounding, cooldowns, and minimal-intervention (including suppressing non-loop interventions once expected corrective actions complete)—and Inject delivers corrections either appended to tool results or as synthetic user messages.

Evaluation

Experiments use all three τ\tau-bench domains (261 tasks) with Qwen3.7-max as the repaired agent, DeepSeek-V4-Pro for all auxiliary roles to avoid same-model self-judging, and GPT-5.4 for cross-model transfer on Banking. Key results:

Approach Retail Airline Banking Overall
Blind retry 88.46 57.14 26.51 43.09
Outcome feedback 84.62 64.29 20.48 39.02
Reflexion 88.46 71.43 26.51 44.72
AgentTether (post-run only) 92.31 78.57 46.99 60.16
AgentTether (full) 96.15 78.57 59.04 69.11

Full AgentTether repairs 85/123 initially failed tasks (69.11%), +26.02 points over blind retry overall and +32.53 on Banking. Notably, outcome feedback underperforms blind retry (39.02% vs. 43.09%), and Reflexion adds only 1.63 points overall with zero gain on Banking—the paper reads this as evidence that self-critique rarely identifies upstream causes. Efficiency also improves: 56.36 average agent turns and ~1,197K end-to-end method tokens per task versus 66.52 turns and 1,376K for blind retry (~13% fewer tokens), indicating localized guidance reduces wasted re-execution rather than extending it. Wall-clock time shows a trade-off: full AgentTether is faster than blind retry (15.59 vs. 16.32 min) but slower than post-run-only (13.47 min) due to auxiliary verification latency.

Localization is evaluated conservatively: anomalous substructures cover the earliest violated gold action in 71.8% of behavioral Banking failures, and the peak-scoring TU has median positional error of 5.5 TUs versus 6.0 (recency) and 10.0 (frequency) baselines. Component ablations show both mechanisms matter: removing the offline HGT detector drops overall repair to 46.34% (Banking to 27.71%), and removing Repair Memory drops it to 52.03%.

Cross-model transfer holds: on Banking, post-run-only guidance lifts GPT-5.4 repair by +43.03 points over blind retry (to 60.47%), and full AgentTether reaches 65.12% (56/86). Repair is heavily front-loaded—at iteration 1, AgentTether repairs 43.37% (Qwen3.7-max) and 54.65% (GPT-5.4) of failed Banking tasks, accounting for 73.5% and 83.9% of eventual repairs respectively, supporting the bounded iteration budget.

Guarded intervention contributes significantly only where compliance decay is severe: paired comparison against post-run-only yields a net +10 helped tasks on Banking (p=0.021p{=}0.021, McNemar exact), but is neutral on Airline and marginal on Retail. Helped cases receive sparse interventions (11.3 per task, mostly lightweight tool-return checkpoints); hurt cases receive 29.3 interventions per task on average, suggesting repeated guards around required state-changing actions can induce re-planning instead of commitment.

Limitations and open questions

The paper is explicit about several dependencies. Repair quality rests on auxiliary LLM judgments (analyst, verifier); the authors bound this via CTG-grounded evidence, verifier gating limited to intervention rather than outcomes, and model separation—but do not eliminate it. The offline detector's normal-behavior prior is trained on TerminalBench/SWE-smith telemetry mapped into a behavior-oriented schema, an assumption that execution structure transfers across task semantics; whether this holds in production distributions with different tools and evaluator reliability is untested. Intervention policy remains unresolved: the hurt-case analysis shows cooldowns and caps limit but do not prevent over-control, and adapting intervention strength to domain risk, task phase, and model capability is left open. Evaluation is confined to single-agent τ\tau-bench settings; code generation, multi-agent coordination, and production observability integration remain future work.

Conclusion

AgentTether demonstrates that closing the loop from diagnosis to repair requires more than better failure explanation: attribution must follow information-flow dependencies rather than trace position, and corrections must be carried across iterations and enforced selectively during re-execution. Its strongest results—59.04% and 65.12% repair rates on the most constrained domain across two model families, achieved while reducing agent turns and tokens—support deployment as a wrap-around reliability layer for existing agents. The residual open problem is principled intervention policy: when guards help versus when they derail commitment remains an empirical question this work frames but does not fully answer.

Paper to Video (Beta)

No one has generated a video about this paper yet.

Whiteboard

No one has generated a whiteboard explanation for this paper yet.

Open Problems

We haven't generated a list of open problems mentioned in this paper yet.