Agent as Policy (AGP)
- Agent as Policy (AGP) represents an approach where an agent's behavior is explicitly governed by a policy that maps observations to admissible actions, ensuring optimal and controlled task execution in various environments.
- Applications of AGP include robotic manipulation, autonomous grading, and cooperative multi-agent systems, enhancing task efficiency and precision.
- GaaS and Out-of-Band Policy Enforcement (OBPE) implement external policy enforcement, improving security, and auditability by external governance services and infrastructure.
Agent as Policy (AGP) is an architectural and methodological perspective in which an agent’s behavior is represented, generated, or constrained by an explicit policy that maps observations, state, context, or provenance to admissible actions. In its strongest form, the agent itself is a policy: it interprets observations, selects actions, adapts through interaction, and determines execution. In weaker but operationally important forms, an agent proposes actions while an external policy layer authorizes, filters, shapes, or blocks them. The literature therefore distinguishes learned action-selection policies, policy-enforced agents, runtime governance services, information-flow monitors, and execution environments that jointly determine the effective policy of a deployed agent.
1. Conceptual foundations and scope
AGP is not a single formalism shared uniformly across the literature. It denotes a family of architectures centered on the relationship between an agent, its policy, and the environment in which actions occur. The principal distinction is whether policy is internal to the agent, external to it, or distributed across both.
In the strongest interpretation, an agent is a parameterized policy mapping observations to actions:
The policy may be learned through reinforcement learning, behavior cloning, contrastive representation learning, self-supervised training, or runtime interaction. AGPNet exemplifies this interpretation in autonomous dozer grading: its learned high-level policy maps local terrain observations to waypoint actions, while a heuristic low-level controller executes the resulting trajectories (Ross et al., 2021). AgentGFM similarly treats each graph node as an agent whose shared policy determines information reception, signal-channel selection, and halting (Cui et al., 29 Jul 2026). ACPO treats each member of a cooperative multi-agent system as an independently parameterized policy that composes into a joint policy optimized for a shared return (Matsunaga et al., 29 Jun 2026).
A second interpretation treats the agent as a policy-constrained actor. The model proposes an action, while a separate authorization mechanism determines whether that action may execute. In PCAS, the LLM generates an action and a reference monitor evaluates it against a Datalog-derived policy over a causal dependency graph (Palumbo et al., 18 Feb 2026). AgentGuardian learns a context-sensitive authorization envelope for tool calls, including input patterns, attributes, and control-flow transitions (Abaev et al., 15 Jan 2026). Policy Cards define a deployment-bound normative policy that is intended to accompany the agent while being interpreted by runtime enforcement, monitoring, and audit components (Mavračić, 28 Oct 2025).
A third interpretation places policy in infrastructure rather than in the agent. Governance-as-a-Service (GaaS) intercepts proposed actions, evaluates declarative rules, updates a Trust Factor, and returns allow, warn, block, or escalate decisions (Gaurav et al., 26 Aug 2025). The Redpanda Agentic Data Plane moves identity, scope, authorization, approval, and audit metadata into out-of-band infrastructure channels inaccessible to the agent (Akidau et al., 27 May 2026). Out-of-Band Policy Enforcement (OBPE) mediates typed requests and responses at a trusted tool boundary, applying owner policies, agent restrictions, query shaping, field filtering, masking, and semantic gates (Millstone et al., 27 Aug 2026).
The effective policy of a deployed system can consequently be understood as the intersection of several mechanisms:
This formulation is an interpretive synthesis rather than a single equation adopted across the papers. It captures the central AGP issue: a model may generate an action, but the action becomes part of the environment only if the relevant policy and execution layers admit it.
2. Internal policy and learned action selection
Robotic manipulation
“Agent as Policy for Robotic Manipulation” presents the most direct use of the term AGP. A general-purpose multimodal coding agent serves as the robot policy without task-specific or environment-specific training (Jia et al., 11 Sep 2026). Given a task specification, robot interface, conversation history, and persistent workspace, it selects the next tool call:
The agent requests images and robot state, writes or reuses perception and geometry programs, issues motion and gripper commands, observes physical outcomes, diagnoses errors, and revises its programs and subsequent actions. Its model parameters remain fixed during execution; adaptation occurs through new observations, revised programs, updated files, and changed decisions.
The architecture separates a task-preparation agent, a runtime execution agent, and a robot-agent bridge. The bridge exposes images, depth, calibration, joint positions, end-effector pose, gripper state, robot health, target error, and action reports. The agent writes executable local programs for segmentation, geometric estimation, triangulation, grasp computation, robot-call composition, and post-action verification. The low-level runtime retains responsibility for inverse kinematics, trajectory generation, motion limits, motor feedback, and gripper control.
This architecture differs from a conventional visuomotor policy, which directly maps observations to a fixed stream of low-level actions. AGP instead controls observation selection, program construction, action sequencing, recovery, and verification. It also differs from behavior cloning and offline reinforcement learning: demonstrations provide task references rather than state-action trajectories used to fit a task policy, and the agent interacts with the physical robot during deployment rather than optimizing a policy from a fixed offline dataset.
The evaluated tasks include assembly from human videos, block construction from goal images, dice flipping, targeted throwing, and bimanual towel folding. Across assembly, block construction, and dice flipping, AGP succeeds in at least eight of ten trials for each evaluated task configuration (Jia et al., 11 Sep 2026). Experience accumulation allows saved procedures and programs to be reused, shortening execution time across repeated trials.
Autonomous grading
AGPNet represents an intermediate form of internal policy autonomy. The learned policy does not directly output steering, throttle, or track velocities. It maps an ego-centric terrain observation to a composite waypoint action:
A heuristic low-level controller converts each waypoint pair into rotations, forward movements, reversals, and subsequent grading legs (Ross et al., 2021). The architecture therefore combines learned high-level policy autonomy with structured low-level control.
The task is modeled as a Markov Decision Process over terrain error. The policy receives a local field of view of the difference height map and selects waypoint pixels in a down-sampled coordinate frame. Its action distribution is split into destination and starting-point heads. A Gaussian spatial prior favors nearby waypoints, and the policy uses convolutional and fully connected layers with two spatial policy heads.
AGPNet combines behavior cloning from the SnP rule-based expert, PPO-style reinforcement learning, and contrastive representation learning. The reported hybrid method, RL+BC+CL, uses demonstrations to initialize behavior, reinforcement learning to optimize soil-interaction reward, and contrastive learning to improve terrain representations. The experiments evaluate Initial, Continuous, and Edge grading scenarios, each with 50 independent runs. The policy is trained on random simulated scenes and tested on structurally different configurations, including real-world-style terrain evolution (Ross et al., 2021).
The system illustrates a restricted but important AGP form: the learned component replaces high-level domain-specific waypoint selection, while domain knowledge remains in the action hierarchy, Gaussian locality prior, simulator, feasibility assumptions, and low-level controller.
Node-agent information-flow control
AgentGFM treats each graph node as an agent that independently controls information propagation using a shared policy (Cui et al., 29 Jul 2026). The graph is the environment, message propagation is the transition mechanism, and prediction–observation discrepancy is the feedback signal.
At rollout step , node has carrier state and selects:
where controls source reception, selects low- versus high-frequency signal channels, and 0 controls halting. The resulting message flow updates the carrier state and forwarding budget. After the rollout, the node compares a predicted observation with the actual observation and uses the discrepancy to correct its state and influence future decisions.
Unlike independent node-specific models, AgentGFM uses one shared policy across nodes, graph domains, and rollout steps. Node-specific behavior emerges from differences in local state, neighborhood, feedback, budget, and graph structure. The method is trained end-to-end with masked attribute reconstruction, prediction–observation alignment, and information-flow regularization rather than conventional reinforcement learning.
The reported ablations show that source-reception control, channel selection, feedback, and gain-aware halting all contribute to performance. The full model has an average accuracy of 1, compared with 2 without prediction–observation feedback. With a maximum rollout horizon of five, an average of 3 of nodes halt by the second rollout step, while halting distributions differ across graphs (Cui et al., 29 Jul 2026).
Cooperative multi-agent policy composition
ACPO addresses AGP at the level of policy composition. It models a cooperative multi-agent system as an MMDP with a joint policy:
4
Each agent has its own actor, critic, policy-gradient term, and parameters. “Independent” means independently parameterized, not independently motivated: all actors optimize one shared cooperative objective (Matsunaga et al., 29 Jun 2026).
ACPO serializes the simultaneous joint action mathematically into micro-steps. Each agent conditions on a belief over preceding agents’ action distributions. The resulting Agent-Chained Belief MDP provides a decomposition of the global policy gradient into per-agent terms. The decentralized critics are not assumed to factorize the original global value function; instead, they arise from the serialized belief-MDP recursion.
The method is evaluated on Multi-Robot Warehouse, SMACv2, and MA-MuJoCo. In RWARE, ACPO outperforms the reported baselines, with the advantage increasing in 8-agent and 12-agent settings. In SMACv2, it is competitive with or better than the reported on-policy baselines and remains strong as the number of agents grows. ACTD3 extends the approach to continuous control (Matsunaga et al., 29 Jun 2026).
The main AGP implication is that independent policies need not imply independent optimization. Agents can remain locally parameterized while being coordinated through beliefs, chained critics, and a joint-gradient objective.
3. External policy enforcement and governance
Policy cards and deployment identity
Policy Cards define a versioned, machine-readable specification of operational rules, obligations, prohibitions, permissions, exceptions, evidence requirements, monitoring requirements, assurance metrics, change-management procedures, and mappings to governance frameworks (Mavračić, 28 Oct 2025). The Policy Card is intended to sit with the deployed agent and become an integral part of its operational identity.
The architecture can be summarized as:
5
This is an interpretive formulation. The paper does not show that the LLM internally implements the Policy Card. Instead, enforcement may occur through policy gateways, middleware, orchestration layers, monitoring systems, CI/CD pipelines, and audit infrastructure.
Policy Cards use JSON Schema Draft 2020-12. Their controls are represented as ABAC-like tuples:
6
where the effect is allow, deny, or require_escalation. The lifecycle is Declare–Do–Audit: declare and validate a versioned card, enforce it at runtime, and continuously audit actions against the active version.
Policy Cards therefore combine agent-level policy representation with external enforcement. Their principal unresolved issue is formal semantics: the paper does not define a complete condition language, rule-priority algebra, obligation lifecycle, delegation model, policy-composition semantics, or treatment of unknown attributes (Mavračić, 28 Oct 2025).
Runtime governance services
GaaS externalizes policy into an independent enforcement service positioned between agents and the environment (Gaurav et al., 26 Aug 2025). Agents are treated as untrusted, and policies are represented as declarative JSON rules. The service includes a Policy Loader, Violation Checker, Trust Factor mechanism, Enforcement Engine, audit subsystem, and human-oversight interfaces.
The runtime function is:
7
GaaS distinguishes coercive, normative, mimetic, and adaptive governance. Coercive rules can block immediately; normative rules can warn; adaptive enforcement uses violation history and Trust Factor thresholds to escalate or block.
The paper evaluates writing and financial-trading scenarios with LLaMA-3, Qwen-3, and DeepSeek-R1. It reports precision 8, recall 9, and 0 for GaaS in a comparison involving keyword filtering, OpenAI moderation, a constitutional agent, and GaaS (Gaurav et al., 26 Aug 2025). The paper also reports adversarial bypass rates after rule patches of 10% for prompt injection, 15% for ambiguous phrasing, and 20% for mimic compliance.
GaaS’s principal AGP significance is the distinction between policy as an external institutional control and policy as the agent’s internal cognition. It supplies access-control enforcement and auditability but does not teach agents ethics or establish that agents internally reason according to the governing policy.
Out-of-band metadata
The Redpanda Agentic Data Plane moves security-critical metadata outside the agent’s normal input/output path (Akidau et al., 27 May 2026). The metadata includes identity, tenant scope, data classification, tool permissions, budgets, rate limits, approval thresholds, behavioral constraints, and audit context.
The architecture uses an AI Gateway, MCP Gateway, identity providers, data adapters, agent sandboxes, message brokers, approval services, and structured trace capture. Agents cannot read or write the policy channel. Client scope is resolved from infrastructure credentials rather than from agent-generated identifiers, and W3C Trace Context headers are injected at infrastructure hops.
The portfolio-rebalancing demonstration includes Signal, Decision, and Execution Agents plus a human Approval App. Signals and orders use client-specific message channels. Trade recommendations below a threshold are routed toward automatic execution, while higher-value recommendations are routed for human approval. The paper demonstrates per-client isolation, channel isolation, approval gating, credential separation, and infrastructure-controlled audit capture, but does not provide quantitative security measurements, latency benchmarks, or formal proofs (Akidau et al., 27 May 2026).
OBPE applies a similar principle at a trusted tool boundary. It mediates both requests and responses, using typed Cedar authorization, query narrowing, result caps, write shaping, record filtering, field removal, masking, and semantic gates (Millstone et al., 27 Aug 2026). A data-owner policy establishes the maximum grant, while an agent policy may only narrow it:
1
The evaluation reports a reduction in trace failure from 2 with prompt-only controls to 3 with prompt plus OBPE across 3,621 trials. Task fulfillment decreases from 4 to 5, while paired safe-useful completion increases by 21.8 percentage points under one judge (Millstone et al., 27 Aug 2026).
4. Stateful, information-flow, and provenance-aware policy
Execution-level authorization
AgentGuardian learns policies from benign staging traces and enforces them over tool calls (Abaev et al., 15 Jan 2026). Its policy combines control-flow authorization, input patterns, attributes, temporal constraints, token limits, idle intervals, and processing duration.
For each tool, benign execution traces are represented as a directed graph whose nodes are tool occurrences and whose edges are observed valid transitions. A proposed call is authorized only if both the current control-flow state and the tool input satisfy learned rules. AgentGuardian therefore implements an execution-level authorization policy rather than the underlying LLM’s task policy.
The evaluation uses a knowledge assistant and an IT-support application. Across both applications, the reported aggregate false-acceptance rate is 6, false-rejection rate is 7, and benign-execution-failure rate is 8 (Abaev et al., 15 Jan 2026). The principal limitations are the clean-staging assumption, coverage gaps, regex overgeneralization and under-generalization, dependence on the orchestrating LLM, and absence of a formal security guarantee beyond the mediated workflow.
Causal dependency graphs
PCAS replaces linear message histories with dependency graphs that represent causal relationships among messages, tool calls, tool results, and external effects (Palumbo et al., 18 Feb 2026). Policies are written in a Datalog-derived language and evaluated by a reference monitor before actions execute.
A dependency graph is:
9
where vertices are event nodes and edges represent causal dependence. Authorization is evaluated over the backward slice of a proposed action. Recursive Depends relations allow policies to detect transitive influence, such as an external email that depends on both untrusted input and sensitive data.
PCAS uses Differential Datalog and compiles policy evaluation into native Rust code. The reference monitor returns allow or deny, and denied actions do not materialize as executed action nodes. The system therefore provides a conditional safety guarantee: under complete mediation, trusted infrastructure, correct graph construction, reliable authentication, correct policy specification, and absence of side channels, unauthorized mediated actions do not execute.
The reported evaluation includes prompt-injection information-flow defense, customer service workflows, and a pharmacovigilance system. In the customer-service evaluation, overall compliance or pass performance improves from approximately 48% to 93%. In the pharmacovigilance workflow, the instrumented system blocks unauthorized FDA accesses and enables recovery through required registration and retry (Palumbo et al., 18 Feb 2026).
Privacy-preserving execution
GAAP treats the LLM as an untrusted code generator and mediates private-data access and disclosure through an information-flow-controlled execution environment (Stanley et al., 21 Apr 2026). The architecture contains a private-data database, permissions database, persistent disclosure log, service annotations, and an IFC execution core.
The user’s policy is represented as persistent allow/deny decisions over private-data items and external parties:
0
When a code artifact attempts a disclosure for which no permission exists, GAAP issues a directed user prompt. Approved decisions persist across tasks, while one-time approvals do not become standing permissions. The disclosure log tracks direct and transitive flows across execution steps and tasks, including data sent to the model provider.
GAAP’s privacy evaluation reports a zero percent success rate for SSN-leak, phone-leak, and SSN-swap attacks. Utility on the authors’ 20-task suite is 76.0%, compared with 81.0% for the unconstrained agent. The system therefore demonstrates a privacy-focused AGP architecture in which the model may propose code and disclosures but cannot authorize the release of private data (Stanley et al., 21 Apr 2026).
Taint confinement and branching
APPA extends information-flow control with prospective acquisition enforcement and engine-managed context branching (Kravchenko et al., 27 Jul 2026). A trajectory label combines a reader set and trust level:
1
Tool outputs produce label descent through a meet operation:
2
A child trajectory inherits the parent’s current label, performs restrictive reasoning locally, and returns only a labeled value or sanitizer-derived result. The parent label after merge is:
3
Consequently, a child cannot silently widen or degrade the parent’s permissions. A raw restrictive return taints the parent; a derivative whose label is adequate for the parent can preserve the parent label.
APPA also separates label state from a shared append-only event log. Branch abandonment does not roll back external effects, so APPA provides trajectory isolation rather than transactional rollback. Its benchmark contains 14 corporate-assistant scenarios and 17 tools. Across four models, open-arm attack success rates of 31%–50% fall to 0%–7% with APPA, while branching recovers substantial utility for three of the four models (Kravchenko et al., 27 Jul 2026).
Flow-centric policy
AgentFlow represents policies over labeled runtime edges, data provenance, path history, task-scoped capabilities, releases, and delegation (Shivakumar et al., 24 Aug 2026). Its runtime state includes a taint map, execution trace, lineage map, task scope, runtime requirements, and capabilities:
4
Labels contain sensitivity, categories, and trust. Flow rules constrain individual transitions, while path rules constrain sequences such as sensitive-data reads followed by external email. Delegation is an authority boundary: an agent may not delegate data whose sensitivity exceeds the receiving agent’s maximum sensitivity.
The runtime reference monitor returns Allow, Deny, or Pause. A bounded SMT verifier checks structured policy fragments and searches for unsafe symbolic traces. Seven evaluated properties are reported as UNSAT at bound 5, each in under 6 seconds, and all 12 seeded unsafe policy variants are detected.
On 949 AgentDojo injected cases, AgentFlow reduces confirmed compromise from 33.0% to 0.0% while increasing aggregate utility from 46.7% to 63.3%. On 200 AgentDyn Dailylife cases, confirmed compromise decreases from 73.5% to 0.0%, while utility changes from 44.5% to 43.5% (Shivakumar et al., 24 Aug 2026). These results are scoped to modeled, policy-visible behavior and do not establish unbounded security or full semantic noninterference.
5. Stateful governance, concurrency, and protocol integrity
Policy-state serializability
Provenact addresses stale authorization in concurrent agentic systems (Peng et al., 3 Aug 2026). A stateful policy depends on mutable shared facts such as budgets, inventory, risk scores, approval counts, or quotas. If two agents evaluate the same stale state and then commit independently, both actions may be individually authorized at decision time but jointly violate the policy.
The paper defines policy-state serializability (PSS): every concurrent history must have a real-time-respecting serial explanation in which each terminal decision is evaluated against the policy state immediately before its effect, denied operations have no effect, and allowed operations produce the same governed effects and final policy state.
Provenact uses bounded policy programs, certified provider-owned views, logical scopes, effect contracts, transactions, reservations, scoped holds, revalidation, idempotency, and audit records. Its central distinction is between request-time authorization and authorization coupled to the effect that changes policy state.
In the reported 256-client full-conflict workload, naive and Cedar request-local systems produce stale allows, while Global, Manual transaction, Provenact-Tx, and Provenact-Res produce exactly 50 committed transfers, 206 denied transfers, zero stale allows, and PSS satisfaction. In the procurement workflow, Provenact-Tx and Provenact-Res produce 70.6 committed and valid operations, zero stale authorizations, zero stale approvals, and zero reported user, team, or item violations (Peng et al., 3 Aug 2026).
PSS does not require global serialization. It protects policy-induced conflicts, including cases in which application write sets are disjoint but one effect modifies facts read by another operation’s policy.
Agent-mediated payment authorization
AP2 illustrates the difference between cryptographic authorization artifacts and policy-grounded agent authorization (Aviv et al., 24 Aug 2026). AP2 v0.2 uses signed Checkout and Payment Mandates, including open and closed forms, checkout hashes, agent keys, parent-child bindings, receipts, and verification procedures.
In Human-Present mode, the user approves closed mandates. In Human-Not-Present mode, the user approves open mandates and the Shopping Agent later creates and signs closed mandates within those constraints. The open mandate can therefore be interpreted as a delegated policy or capability grant, while the closed mandate is an agent-generated execution decision.
The paper identifies 48 threats across five deployment architectures and five attack families. Eight threats reach the High band in at least one architecture: pre-signing context poisoning, tool-path or verifier-tool manipulation, caller-role confusion, malicious MCP server image, sub-agent prompt drift, protocol downgrade, risk_data poisoning, and cross-tenant credential or state-isolation failure (Aviv et al., 24 Aug 2026).
The central AGP conclusion is that a valid signature authenticates signed bytes but does not guarantee that those bytes reflect the user’s intent. Catalogs, A2A messages, MCP tool calls, RAG outputs, AgentCards, prompts, and delegated agents can shape policy interpretation before signing. Therefore, authorization artifacts require context integrity, provenance, role binding, tool mediation, delegation constraints, and evidence of the decision path.
6. Common distinctions, limitations, and research directions
Agent as policy versus policy-enforced agent
The literature distinguishes at least four notions that should not be conflated:
| Notion | Primary mechanism | Representative systems |
|---|---|---|
| Internal action policy | Learned or runtime mapping from state to actions | AGP robotic manipulation, AGPNet, AgentGFM |
| Joint policy component | Independently parameterized policies composed under a shared objective | ACPO |
| Authorization policy | External allow/deny or constrained-action relation | PCAS, AgentGuardian, AgentFlow |
| Governance infrastructure | External enforcement, audit, approval, and state coordination | GaaS, ADP, OBPE, Provenact |
An external policy layer can alter the system’s observable action policy without transforming the model’s internal policy. A model may continue proposing forbidden actions, while the environment observes only permitted transitions. Conversely, an internal policy may be behaviorally capable but lack hard guarantees against prompt injection, stale state, untrusted tools, or unauthorized data flow.
Policy selection versus policy filtering
Most governance-oriented AGP systems filter proposed actions rather than selecting the best legal action. A monitor commonly returns allow, deny, or pause; it does not necessarily generate a useful alternative. This can produce repeated denials, task abandonment, or reduced utility. AgentFlow, PCAS, AgentGuardian, and OBPE all expose this distinction.
ACPO and the robotic AGP architecture are closer to action-selection policies. They generate actions, programs, or policy distributions, although they remain dependent on execution constraints and, in the robotic case, a structured low-level controller.
Static versus learned policy
Policies may be:
- Learned from interaction or demonstrations, as in AGPNet, AgentGuardian, and ACPO.
- Self-supervised and differentiable, as in AgentGFM.
- Autoformalized, as in the Cedar-based policy compiler (Mondl et al., 25 Jun 2026).
- Declaratively authored, as in Policy Cards, GaaS, AgentFlow, and PCAS.
- User-specified dynamically, as in GAAP.
- Stateful and executable, as in Provenact.
Autoformalization introduces a translation gap. The Cedar-based policy compiler uses an LLM generator-critic loop to translate agent instructions, MCP tool descriptions, and natural-language policy documents into Cedar policies (Mondl et al., 25 Jun 2026). Hard critics establish syntax, schema compliance, type consistency, and static properties, while soft critics assess semantic alignment. The resulting runtime decision is deterministic conditional on the correctness of the generated policy, but Cedar validation does not prove that natural-language intent was translated correctly.
Complete mediation and trusted infrastructure
Nearly all externally enforced AGP systems depend on complete mediation. The guarantees require that relevant tool calls, data accesses, side effects, and state mutations pass through the trusted boundary. Direct filesystem access, unmonitored network paths, compromised hosts, unregistered tools, side channels, or external services outside the mediation boundary can invalidate the guarantees.
The trusted computing base varies by system:
- GAAP trusts its IFC core, databases, interception layer, and annotations.
- PCAS trusts instrumentation, the dependency graph, the policy engine, and the execution substrate.
- AgentFlow trusts the reference monitor, tool metadata, labels, and policy configuration.
- Provenact trusts provider-certified views, effect contracts, coordination, reservations, and recovery.
- ADP and OBPE trust gateways, identity systems, adapters, sandboxes, and audit infrastructure.
Structural guarantees versus semantic guarantees
A recurring limitation is the difference between structural enforcement and semantic safety. Blocking an unauthorized tool call does not guarantee that allowed actions are wise, truthful, or aligned with the user’s intent. Removing a field from one response does not establish noninterference: values may be reconstructed from visible names, counts, timing, ranking, pagination, or repeated queries. A signed mandate does not establish that the signer’s interpretation was faithful to the user’s intent. A policy-state serializability guarantee does not prove that the policy itself is substantively correct.
The strongest defensible claims are therefore conditional:
- mediated actions satisfy the configured policy;
- denied actions do not execute through the protected boundary;
- policy-relevant state is coordinated according to the stated mechanism;
- labeled or restricted data cannot cross the modeled boundary without authorization;
- policy composition is deterministic under specified algebraic assumptions.
These claims do not establish complete safety of the model, environment, organization, or policy specification.
Stateful and temporal policy
Several systems identify state and history as essential. AgentFlow uses traces, lineage, path expressions, capabilities, and task scopes. APPA uses a shared event log and committed-effect predicates. Provenact introduces PSS to bind decisions to mutable policy state. AP2 emphasizes mandate replay, stale state, receipt integrity, and transaction binding.
Stateful policy introduces additional requirements:
- durable policy state;
- explicit policy-visible dependencies;
- commit-time revalidation or reservations;
- idempotency;
- approval binding;
- policy-version handling;
- recovery after partial failure;
- protection against stale authorizations;
- coordination of concurrent operations.
A request-local policy decision is insufficient when the effect changes the state on which policy decisions depend.
Open research directions
The literature identifies several unresolved problems:
- Unified semantics: a general AGP calculus must integrate action selection, authorization, provenance, obligations, delegation, temporal state, and external effects.
- Policy synthesis: autoformalization must bridge natural-language intent and executable policy without relying solely on probabilistic semantic critics.
- Policy composition: independently authored policies require conflict resolution, priority semantics, delegation rules, and version compatibility.
- Semantic provenance: systems must track not only data flow but also how external content influences plans, rationales, and authorization artifacts.
- Complete mediation: tool calls, model calls, filesystem operations, side channels, message brokers, and delegated agents require a unified enforcement boundary.
- Noninterference and inference control: direct field removal is insufficient against relational inference, query-count oracles, timing, and repeated probing.
- Policy-aware planning: external monitors should reject unsafe actions while enabling agents to construct useful alternatives rather than repeatedly failing.
- Adaptive policy governance: learned policies require versioning, provenance, calibrated uncertainty, rollback, and safeguards against policy drift.
- Concurrent external effects: stateful policies need serializability, reservations, durable approvals, idempotency, and recovery protocols across provider boundaries.
- Human governance: user approvals, organizational authority, policy ownership, tenant identity, and sub-agent delegation must remain auditable and semantically bound.
- Evaluation methodology: benchmarks should measure confirmed compromise, context exposure, safe usefulness, policy coverage, false denials, latency, overhead, and behavior under adaptive attacks.
- Formal verification scope: bounded SMT results, typed policy validation, and conformance testing must be connected to unbounded or production-level guarantees.
AGP is consequently best understood not as a single algorithm but as a design space spanning learned decision policies, policy-centered agent architectures, and trusted runtime enforcement. The central architectural lesson is that an agent’s capability to propose an action should be separated from the authority to execute it. Internal policy reasoning can provide adaptability, planning, and task competence; external policy layers can provide authorization, provenance, information-flow control, concurrency correctness, approval, and auditability. A robust AGP system combines these roles while making explicit which guarantees belong to the model, which belong to the policy, and which depend on the trusted execution environment.