Composing Verifiers Across Multi-Step Trajectories
Determine principled methods for composing per-step verifiers across complete execution trajectories of LLM-based AI agents so that local safety checks imply global plan-level policy compliance and bounded cumulative risk, including when tools are nondeterministic or have hidden state.
References
Another open question is how to compose verifiers across a multi-step trajectory. Even if each step is locally "safe", the global plan may still violate policy or create unacceptable cumulative risk; compositional safety is especially hard when tools are nondeterministic or have hidden state.
Section~\ref{sec:adaptive} takes a first step against a policy-aware adversary; adaptive attackers that learn against the verifier over many rounds, and multi-step agent loops that compose policy-valid legs into a harmful sequence, remain open and are the most important next target.
Evaluating such policies requires policy-in-the-loop execution or simulation; large-scale empirical validation of these settings is left to future work.