Composing Verifiers Across Multi-Step Trajectories

Determine principled methods for composing per-step verifiers across complete execution trajectories of LLM-based AI agents so that local safety checks imply global plan-level policy compliance and bounded cumulative risk, including when tools are nondeterministic or have hidden state.

Background

Agents execute sequences of tool calls under uncertainty; even if each action is locally validated, the aggregate plan may violate policy or accumulate unacceptable risk. This compositional gap is especially acute when tools are nondeterministic or stateful, making stepwise guarantees insufficient.

A solution requires formal notions of trajectory-level safety, policies for cumulative risk, and mechanisms to compose local verifications into global guarantees, potentially with sandboxing, permissions, and trace-level auditing for evidence-based assurance.

References

Another open question is how to compose verifiers across a multi-step trajectory. Even if each step is locally "safe", the global plan may still violate policy or create unacceptable cumulative risk; compositional safety is especially hard when tools are nondeterministic or have hidden state.

— AI Agent Systems: Architectures, Applications, and Evaluation  (2601.01743 - Xu, 5 Jan 2026) in Section 7.1 (Verification and Trustworthy Tool Execution)

Section~\ref{sec:adaptive} takes a first step against a policy-aware adversary; adaptive attackers that learn against the verifier over many rounds, and multi-step agent loops that compose policy-valid legs into a harmful sequence, remain open and are the most important next target.

— PACE: Policy-Attested Contract Execution for Safe AI Agents in Decentralized Finance  (2608.17220 - Karanjai et al., 18 Aug 2026) in Section: Future Work

Evaluating such policies requires policy-in-the-loop execution or simulation; large-scale empirical validation of these settings is left to future work.

— READY or Not: Reliable Enterprise Agent Deployment  (2609.02095 - Chatrath et al., 2 Sep 2026) in Section 6, Scope, Limitations, and Future Work, paragraph “Empirical scope”