Papers
Topics
Authors
Recent
Search
2000 character limit reached

Agent as Policy for Robotic Manipulation

Published 11 Sep 2026 in cs.CL | (2609.12541v2)

Abstract: We demonstrate that a general-purpose agent can directly drive a physical robot throughout task execution without any task-specific or environment-specific training. We introduce Agent as Policy (AGP), which places task planning and execution under the agent's control. Given a task and a robot interface, the agent interprets visual evidence, writes executable programs, issues motion commands, and revises its actions in response to physical outcomes. This brings the agent's reasoning and programming capabilities into continuous interaction with the physical world. We study AGP across multiple real-world manipulation tasks spanning precision manipulation, dynamic motions, and deformable objects. These include assembly from human videos, block construction from goal images, dice flipping, targeted throwing, and bimanual towel folding. Across assembly, block construction, and dice flipping, AGP succeeds in at least eight of ten trials for each evaluated task configuration. We further study efficiency through task experience accumulation and find that reusing saved procedures and programs shortens execution time across repeated trials. These findings support a path for general-purpose agents to act as robot policies, extending their autonomy to physical manipulation through runtime reasoning, programming, and interaction.

Summary

  • The paper introduces 'Agent as Policy for Robotic Manipulation,' utilizing a general-purpose multimodal coding agent for robotic control, improving task flexibility and runtime adaptability with zero-shot capabilities.
  • The method achieves a high success rate across various tasks, such as an 8/10 success rate in assembly and 29/30 in block construction, while maintaining at least 80% success in seven out of eight reported task configurations.
  • Runtime programming allows the agent to learn and adapt from physical feedback, reducing repeated execution costs significantly, with a 30% time reduction observed in the model-dependent experiments with experience reuse.

Problem formulation and central claim

“Agent as Policy for Robotic Manipulation” proposes a control architecture in which a general-purpose multimodal coding agent serves as the robot policy itself, rather than selecting among separately trained skills or producing a fixed program before execution (2609.12541). The central claim is deliberately stronger than conventional language-conditioned planning: the agent controls not only task decomposition but also runtime perception, geometric estimation, motion generation, observation scheduling, error recovery, and program construction.

The paper positions AGP against two common agentic robotics paradigms. In program-synthesis systems, an agent generates code that is subsequently executed, but runtime contingencies must be encoded in advance through explicit logic. In policy-orchestration systems, an agent selects among learned skills while the invoked policies retain control over low-level motion generation. AGP instead keeps the agent in the action loop and permits it to write or revise local programs whenever new physical evidence changes the appropriate response. This distinction is closely related to earlier code-generation approaches such as “Code as Policies” (Liang et al., 2022), but AGP emphasizes continual runtime interaction rather than principally precomputed program synthesis.

The policy is formalized at the tool-decision level. Given the task specification, robot interface, interaction history, and persistent workspace, the fixed-parameter agent samples a tool call. A call may request an observation, inspect a saved image or video, execute locally generated Python, or submit a motion and gripper target. The model weights remain unchanged throughout a trial; adaptation is implemented through new observations, generated programs, modified geometric estimates, and accumulated files.

Figure 1

Figure 1: AGP separates task preparation from runtime execution while allowing the execution agent to construct programs, request observations, issue robot actions, and revise its behavior from physical feedback.

This formulation makes the workspace part of the effective policy state. It contains scripts, measurements, observations, execution reports, and notes, allowing the agent to reuse validated procedures within a session or across repeated executions. The paper’s strongest architectural claim is therefore that a general-purpose agent can operate as a policy without a task-specific action model, provided that the robot interface exposes sufficiently expressive sensing, geometric grounding, and motion primitives.

Robot interface and runtime programming

The robot bridge provides overhead RGB observations, wrist-mounted RGB-D observations, proprioceptive state, calibrated camera information, end-effector poses, gripper state, and target-achievement errors. The agent may command joint targets, Cartesian end-effector poses, gripper opening, and additional observations. Cartesian actions are converted to joint targets using inverse kinematics and then executed under workspace, velocity, acceleration, and step-size limits.

Figure 2

Figure 2: The physical platform uses one or two I2RT YAM arms, a parallel gripper, an overhead RGB camera, and a wrist-mounted RGB-D camera in a shared calibrated workspace.

The interface is intentionally intermediate between high-level symbolic planning and direct torque control. AGP does not synthesize arbitrary feedback controllers at the motor level. Instead, it computes geometric targets and timed motion programs that are passed to an existing robot controller. This division is important for interpreting the results: the agent determines what motion to request and when, while trajectory tracking, inverse kinematics, and low-level constraint enforcement remain delegated to the robot runtime.

The runtime programs are not limited to simple coordinate arithmetic. Across 65 executions, the authors identify 37 retained Python working files, of which 29 implement perception or geometry, seven compose robot calls, and one provides visual annotation. The documented computations include depth measurement, color segmentation, ray-plane intersection, triangulation, surface fitting, grasp-offset estimation, camera-image projection, and concurrent bimanual execution. In one assembly trace, the agent estimates a planar surface from wrist depth, infers the orientation and position of a held pin, generates a corrected insertion target, examines the returned target error, and modifies the subsequent motion.

This mechanism supplies a concrete interpretation of “reasoning with physical feedback.” The agent does not merely reflect verbally after failure; it changes the computational procedure used to estimate state or action targets. However, the paper does not isolate the contribution of runtime programming from that of the underlying model, calibration, controller, or task-specific interface design. The program traces demonstrate capability and mechanism, but not causal necessity.

Task suite and experimental design

The evaluation covers five manipulation categories selected to stress different capabilities:

Task Instruction Primary challenge Successful trials
Four-pair assembly Human video Long-horizon sequencing and precision insertion 8/10
Block construction Goal image Support relations and placement order 29/30
Dice flipping Language Grasp-dependent reorientation 10/10
Targeted throwing Language Timed dynamic release 2/2
Sequential towel folding Human video Bimanual deformable manipulation 5/5
Simultaneous towel folding Human video Coordinated bimanual transport 3/5

The standard configuration uses GPT-6 Astra at high thinking effort with an I2RT YAM arm, overhead RGB, wrist RGB-D, and fixed task budgets. Trials begin without prior task-execution experience or simulation rehearsal. The task preparation agent creates a reusable task definition, while a separate execution agent handles each runtime instance. This separation prevents the main runtime agent from being directly initialized with a hand-designed execution trace, although the task definition itself contains goals, constraints, reference materials, and success criteria.

The task suite is intentionally heterogeneous. Assembly requires preserving previously completed pairs while inferring mating relationships from a demonstration video. Block construction requires recovering object identities, orientations, and support constraints from static goal images. Dice flipping requires selecting grasps and rotations that yield a specified upward face after release. Throwing introduces a timed, joint-driven swing and gripper release. Towel folding tests coordinated manipulation of a deformable object under substantial state uncertainty.

The figures illustrate that the evaluation is not restricted to nominal pick-and-place. In assembly, the robot must execute eight-part, four-pair construction while maintaining earlier progress.

Figure 3

Figure 3: Four-pair assembly requires video interpretation, precise insertion, retry behavior, and preservation of completed subassemblies.

For block construction, the agent must infer an order that maintains structural stability rather than simply matching the final image through unconstrained placements.

Figure 4

Figure 4: Pyramid, two-tower, and six-block constructions require the agent to infer support relationships from static goal images.

The protocol evaluates physical outcomes after gripper release and arm withdrawal. This criterion exposes failures that would be missed by measuring only controller arrival error. For example, the six-block tower can be correctly assembled and subsequently collapse during withdrawal; the trial is therefore unsuccessful.

Main manipulation results

AGP achieves at least 80% success in seven of eight reported task configurations, with 100% observed success in five. The strongest results are 29/30 for block construction and 10/10 for dice flipping. The assembly result of 8/10 is notable because the task includes submillimeter-scale nominal radial clearances for some mating parts and requires preserving multiple completed pairings. The result supports the claim that a general-purpose agent can perform nontrivial long-horizon manipulation without task-specific policy training.

The block-construction results show sensitivity to task complexity. Pyramid construction and two-tower construction both succeed in 10/10 trials, whereas the six-block tower succeeds in 9/10. Mean successful-trial completion times are 21.6, 20.6, and 28.2 minutes, respectively. The six-block tower also consumes more tokens and inference cost than the simpler structures. This pattern is consistent with the additional visual verification and stability management required as the structure grows, although the experiments do not provide a controlled decomposition of which source of complexity dominates.

Dice flipping achieves 10/10 despite requiring object-specific orientation reasoning and repeated grasp selection.

Figure 5

Figure 5: Dice flipping requires choosing grasp and rotation strategies that produce the requested upward face after release.

The result is important because the agent is not merely transporting objects to specified locations; it must reason about how a grasped die’s orientation changes under manipulation. The paper reports examples in which the agent combines a tilt and low release with subsequent verification and regrasping. Nevertheless, the evaluation uses six dice under predefined conditions, so it does not establish robustness to substantially different die geometries, friction properties, or clutter distributions.

Targeted throwing succeeds in both trials. The agent performs motion probes, modifies the swing and gripper-opening schedule, and achieves separation into the target bowl.

Figure 6

Figure 6: Targeted throwing combines a circular joint trajectory with timed release and post-release verification.

The reported mean elapsed time is 22.8 minutes, but the throw itself completes after 13.9 minutes; approximately 40% of the reported duration is attributed to trajectory analysis and verification. This distinction illustrates a central limitation of the approach: successful physical action may be substantially faster than the full agent-mediated reasoning and reporting loop.

The towel experiments are more demanding in a different way. Sequential folding succeeds in 5/5 trials, but takes 50.8 minutes on average and costs USD 24.14 per successful trial. Simultaneous folding is faster, at 21.4 minutes and USD 7.46 on successful trials, but succeeds in only 3/5 trials. The apparent efficiency advantage of simultaneous execution therefore comes with lower reliability and cannot be interpreted as a uniformly superior strategy. Qualitative sequences show recovery from slipped corners, staged placement of flaps, and residual skew or buckling.

Figure 7

Figure 7: Sequential folding achieves the required arrangement after recovery from slips, but retains visible buckling and edge offsets.

Figure 8

Figure 8: Simultaneous folding completes the required inward folds in successful trials but exhibits residual skew and alignment error.

These outcomes support the paper’s claim that deformable-object manipulation remains a significant bottleneck. The agent can recover from local errors and satisfy the relatively permissive success criterion, but the long execution time and 3/5 simultaneous-folding success rate indicate that general-purpose reasoning does not remove the need for robust deformable-object state estimation and contact control.

Efficiency, model dependence, and experience reuse

AGP’s success results are accompanied by substantial inference and execution costs. Successful assembly trials average 37.2 minutes, 13.02 million tokens, and USD 16.62. Dice flipping averages 37.9 minutes, 17.08 million tokens, and USD 21.07. Sequential towel folding is the most expensive main task in the reported set, averaging 50.8 minutes, 18.12 million tokens, and USD 24.14. The time-allocation analysis attributes large fractions of execution to the visual loop and remaining model-response latency, particularly for assembly and towel folding.

The model comparison demonstrates that performance depends strongly on the underlying agent-model combination. On a two-pair assembly task, GPT-6 Astra succeeds in 5/5 trials at low, medium, and high thinking effort, with mean completion times of 9.9, 9.2, and 9.2 minutes. Increasing effort therefore provides little observed benefit for this task. GPT-5.6 Sol also succeeds in 5/5 but takes 14.2 minutes on average. GPT-5.6 Terra succeeds in only 1/5 trials, while GPT-5.6 Luna succeeds in 0/5. Within Claude Code, Claude Opus 5 succeeds in 5/5, whereas Claude Fable 5.1 succeeds in 3/5.

These results contradict any interpretation of AGP as model-agnostic. The interface and runtime architecture are transferable across agents, but task reliability remains highly sensitive to model capability, agent scaffolding, and tool-use behavior. The results also show that inference cost and physical completion time need not move together: Sol has lower reported inference cost than Astra despite longer execution time, while effort increases do not materially improve Astra’s success on this task.

The experience studies provide the paper’s clearest evidence for improving efficiency without changing model parameters. In repeated two-pair assembly, saved measurements, procedures, corrections, and scripts reduce task execution time by 29.3% from the first to the fifth execution. The reasoning-and-programming component falls by 47.0%, while tool execution increases by 9.3%. This pattern indicates that reuse primarily reduces repeated deliberation and program construction rather than physical motion time.

The result should be interpreted as evidence of experience-conditioned efficiency, not as online learning in the model-parameter sense. Each trial begins with a fresh agent context, and improvements are mediated by files that encode prior work. Moreover, the repeated trials are dependent and share a fixed task configuration. The paper itself notes that comparisons against repeated trials with empty experience would be needed to isolate the causal effect of reuse from ordering and familiarity effects.

The strong-to-weak transfer experiment is more consequential. GPT-6 Astra first constructs frozen experience files; GPT-5.6 Terra then performs the same assembly task with or without access to those files. Terra’s success increases from 1/5 without experience to 4/5 with Astra’s experience. Among successful trials, experience reduces mean completion time by 34.3% and token usage by 11.4%, while also lowering inference cost. This result suggests that a stronger agent can amortize exploration for a weaker, less expensive agent.

The evidence remains limited by the small sample size and unequal successful-trial counts: the resource comparison uses one successful no-experience trial and four successful experience trials. Consequently, the magnitude of the reduction should not be generalized beyond the evaluated task without larger matched experiments. The study nevertheless establishes a specific design principle: persistent procedural artifacts can transfer operational competence across model instances and even across model capability levels.

Cross-embodiment evidence

The appendix reports a qualitative cross-embodiment experiment using a seven-joint P7 arm and a dexterous RealHand L6 hand rather than the YAM platform.

Figure 9

Figure 9: The cross-embodiment platform combines a P7 arm, RealHand L6 dexterous hand, overhead RGB-D sensing, and wrist-mounted RGB-D sensing.

The agent reorients the wrist, approaches a bottle, closes the fingers in stages, and commands a 2 cm lift.

Figure 10

Figure 10: AGP coordinates wrist reorientation, staged finger closure, and bottle lifting on a different arm and end-effector embodiment.

This experiment supports the narrower claim that the AGP interaction pattern can be instantiated through a different robot interface. It does not provide a quantitative cross-embodiment success rate, systematic comparison, or evidence that the same task definition and programs transfer without adaptation. The bottle subsequently slips during holding, further limiting the result to qualitative evidence of interface-level portability rather than reliable dexterous manipulation.

Limitations and open questions

The principal limitation is scale. Most main-task configurations contain only five or ten trials, and targeted throwing contains two. The reported percentages therefore have wide uncertainty, particularly for the 3/5 simultaneous-folding result and the 2/2 throwing result. The experiments establish feasibility on selected tabletop tasks, not population-level reliability across environments.

The evaluation also relies on substantial infrastructure. Camera calibration, predefined workspace envelopes, fixed motion limits, an existing inverse-kinematics and trajectory controller, and task-specific interface support are essential components. Targeted throwing uses a specialized timed-motion runtime, so it is not generated solely through ordinary Cartesian and gripper commands. The claim that AGP itself is a complete robot policy should therefore be understood as a claim about policy-level decision generation over a constrained robot interface, not replacement of the full control stack.

The agent’s latency and cost are practical constraints. Many successful trials require tens of minutes and millions of tokens. Visual observation, model response, code generation, and verification dominate several tasks. The paper demonstrates that persistent experience reduces this burden, but it does not establish how experience should be represented, validated, versioned, or prevented from propagating erroneous procedures.

Safety is addressed through interface-level motion and workspace limits, but the experiments do not evaluate human-robot interaction, adversarial instructions, hardware faults, calibration drift, or recovery from unsafe agent-generated programs. The agent can issue motion commands and construct executable code; independent protective mechanisms remain necessary for deployment beyond the controlled tabletop setting.

Several scientific questions remain open within the paper’s scope. It is unclear how AGP compares with trained VLA policies or hybrid systems under equal hardware, time, and compute budgets. It is also unclear whether runtime-written programs remain robust under substantial scene variation, whether accumulated experience transfers across robot embodiments rather than merely across model instances, and how much reliability is contributed by the coding-agent scaffolding relative to the multimodal model itself. Finally, the permissive success criteria for towel folding accept residual skew and buckling, leaving open how AGP would perform under tighter geometric or cosmetic requirements.

Conclusion

AGP demonstrates that a general-purpose multimodal coding agent can control real robot manipulation through a calibrated interface, runtime-generated programs, and iterative physical feedback. Across assembly, construction, dice flipping, throwing, and towel folding, it obtains strong zero-shot results, including 29/30 block-construction trials and 10/10 dice-flipping trials. The same experiments expose substantial execution latency, inference cost, model dependence, and difficulty with coordinated deformable manipulation.

The paper’s most substantive contribution is the treatment of the agent as a runtime policy rather than merely a planner, skill selector, or one-shot code generator. Persistent programs and measurements further provide a mechanism for reducing repeated execution cost and transferring operational experience between agents. The evidence supports AGP as a viable architecture for research-scale physical manipulation, while leaving reliability, efficiency, safety, and cross-embodiment transfer as unresolved empirical questions (2609.12541).

Paper to Video (Beta)

No one has generated a video about this paper yet.

Whiteboard

No one has generated a whiteboard explanation for this paper yet.

Explain it Like I'm 14

1. What is this paper about?

This paper introduces a system called Agent as Policy, or AGP. It explores whether a general-purpose AI agent can control a real robot without being specially trained for every single task.

Usually, a robot is trained to do one particular thing, such as picking up a cup. AGP works differently. The AI can:

  • Look at pictures or videos.
  • Understand what it is supposed to do.
  • Write small computer programs.
  • Tell the robot how to move.
  • Check whether the movement worked.
  • Change its plan if something went wrong.

In simple terms, the AI acts like the robot’s brain, planner, programmer, and problem-solver all at once.

2. What questions did the researchers ask?

The researchers wanted to find out:

  1. Can a general AI control a robot without special training for each new task?
  2. Can the AI understand different kinds of instructions, such as videos, pictures, and written commands?
  3. Can it react to mistakes and unexpected results?
  4. Can it perform different kinds of movements, including careful assembly, throwing, flipping objects, and folding fabric?
  5. Can the AI become faster by remembering what it learned from earlier attempts?
  6. Can one stronger AI pass its useful experience to another, less capable AI?

The main idea is that the AI should not simply follow a fixed list of instructions. Instead, it should behave more like a person solving a problem: observe, act, check, and try again.

3. How did the researchers test AGP?

The robot and its “senses”

The researchers connected the AI to a robot arm with a gripper. The robot used cameras to see its workspace:

  • An overhead camera gave a view from above.
  • A camera on the robot’s wrist gave a closer view.
  • Depth information helped the AI estimate how far away objects were.

The robot also reported information about its own body, such as the positions of its joints. This is similar to how people know where their arms and hands are without looking directly at them.

The robot interface

The AI communicated with the robot through a special connection called a robot interface. This worked like a translator:

  • The AI requested pictures or information.
  • The robot sent back camera images and movement information.
  • The AI sent commands such as “move the gripper here” or “close the fingers.”
  • The robot carried out the movement and reported what happened.

Runtime programming

The AI wrote programs while it was working. For example, it might write a program to:

  1. Find an object in a camera image.
  2. Estimate the object’s position.
  3. Calculate where the gripper should move.
  4. Open or close the gripper.
  5. Check the result.

This is called runtime programming because the programs are created during the task, rather than being prepared completely beforehand.

An everyday comparison would be a person writing down a new plan while assembling furniture after discovering that one piece does not fit.

The tasks

The researchers tested AGP on several real-world tasks:

Task What the robot had to do
Assembly Put parts together by following a human demonstration video
Block construction Build shapes, such as a pyramid or towers, from a goal image
Dice flipping Turn six dice so that the correct number appeared on top
Targeted throwing Throw an object toward a particular target
Towel folding Fold a towel using one or two robot arms

The researchers measured success, time, the amount of text the AI processed, and the estimated cost of using the AI.

4. What did they find?

AGP often succeeded without special training

The robot performed well on many tasks even though it had not been specially trained for each one.

Some important results were:

  • Assembly: 8 successful trials out of 10.
  • Block construction: 29 successful trials out of 30.
  • Dice flipping: 10 successful trials out of 10.
  • Potato throwing: 2 successful trials out of 2.
  • Sequential towel folding: 5 successful trials out of 5.
  • Simultaneous towel folding: 3 successful trials out of 5.

Overall, the results suggest that a general-purpose AI can control a robot in several unfamiliar situations.

The AI could adapt after mistakes

AGP did not always follow one unchanging plan. If an object was in a slightly different place or an insertion failed, the AI could:

  • Ask for another camera image.
  • Measure the object again.
  • Change the gripper’s position.
  • Try a different movement.
  • Use feedback from the robot to improve the next attempt.

This is important because real environments are messy. Objects may move, cameras may be imperfect, and movements may not work exactly as expected.

Experience made repeated tasks faster

The researchers allowed the AI to save useful information in files, including:

  • Measurements.
  • Successful movement plans.
  • Mistakes and corrections.
  • Programs for processing camera images.

When the AI repeated the same task, it could reuse this information rather than starting from nothing.

For one assembly task, the time needed for the task decreased by about 29% between the first and fifth attempts. The AI became more efficient because it remembered useful details while still checking the objects’ current positions.

A stronger AI could help a weaker AI

The researchers also tested whether one more capable AI could prepare instructions and experience for another AI.

Without stored experience, the weaker AI succeeded only 1 time out of 5. With experience from the stronger AI, it succeeded 4 times out of 5. Its average completion time and amount of computer processing also decreased on successful attempts.

This is similar to an experienced student giving a beginner a helpful set of notes and tips before the beginner tries the assignment.

The system was still slow and expensive

Despite its successes, AGP had important weaknesses:

  • Some tasks took a long time.
  • The AI used a large amount of computer processing.
  • Using a powerful AI can be expensive.
  • Folding a towel was difficult, especially when both arms had to move together.
  • The experiments used relatively small numbers of trials, so the results do not prove that AGP will work reliably in every situation.

For example, sequential towel folding took about 51 minutes on average, even though all five trials succeeded.

5. Why are these results important?

Traditional robots are often very good at repeating one carefully programmed action, but they may struggle when something changes. AGP shows a possible way to build robots that are more flexible.

Instead of giving the robot every instruction in advance, people could give it a goal, such as:

“Build this structure from the picture.”

The AI could then decide what to look at, write the needed calculations, move the robot, and correct mistakes.

This could eventually help robots perform more useful jobs in places where conditions change, such as:

  • Homes.
  • Hospitals.
  • Factories.
  • Laboratories.
  • Disaster areas.
  • Warehouses.

Simple conclusion

The paper argues that a general-purpose AI can serve as a robot’s policy, meaning the system that decides what the robot should do next. AGP combines seeing, reasoning, programming, acting, and learning from feedback in one continuing process.

The experiments show promising results: the robot completed many tasks without task-specific training, and it became faster when it reused past experience. However, the system is still too slow, costly, and sometimes unreliable for widespread everyday use.

The biggest long-term possibility is that robots may become less like machines that only repeat fixed instructions and more like adaptable helpers that can understand new goals, try solutions, and learn from their experiences.

Knowledge Gaps

Knowledge Gaps, Limitations, and Open Questions

The paper leaves the following issues unresolved:

  • Limited statistical evidence: Most evaluations use only five or ten trials per configuration, making success-rate estimates highly uncertain and preventing reliable significance testing.
  • Restricted hardware validation: AGP is evaluated primarily on one robot platform, with limited evidence that the approach transfers across robot morphologies, grippers, cameras, controllers, or workspace scales.
  • Dependence on a narrow model ecosystem: The main experiments rely heavily on proprietary, high-capability models, so it remains unclear whether AGP is viable with open-weight, smaller, cheaper, or locally deployable models.
  • Insufficient comparison with strong baselines: The paper does not provide systematic comparisons against trained vision-language-action policies, classical task-and-motion planners, code-generation baselines, or hybrid agent-policy systems under matched hardware, task, and budget conditions.
  • Unclear source of performance gains: The experiments do not isolate how much success comes from the underlying MLLM, runtime programming, visual feedback, robot-interface design, task preparation, or persistent experience.
  • No component-level ablation of the AGP interface: The contributions of wrist images, depth, proprioception, geometric queries, motion feedback, program execution, and action-result monitoring are not separately evaluated.
  • Limited task and object diversity: The benchmark contains a small set of tabletop tasks and does not establish performance on cluttered scenes, articulated objects, fragile objects, transparent or reflective materials, tool use, insertion under occlusion, or contact-rich manipulation.
  • Weak evidence for generalization to unseen environments: The experiments use predefined initial scenes and fixed task configurations; robustness to novel layouts, object instances, lighting, camera poses, workspace arrangements, and distractors remains unexplored.
  • Unclear generalization beyond prepared task definitions: A preparation agent supplies reusable task definitions, constraints, and completion criteria, but the study does not measure how errors or omissions in these definitions affect runtime execution.
  • Limited evaluation of instruction ambiguity: The paper does not test ambiguous, incomplete, contradictory, compositional, multilingual, or naturally varying user instructions.
  • Insufficient analysis of perception failures: The system’s ability to detect and recover from incorrect depth, occlusion, segmentation, pose estimation, camera calibration, or object-identity errors is not quantified.
  • No formal reliability or uncertainty mechanism: AGP appears to act on model-generated interpretations and motion programs without calibrated confidence estimates or explicit uncertainty propagation from perception through control.
  • Unresolved safety guarantees: The interface validates motion requests, but the paper does not establish formal guarantees against collisions, unsafe contact forces, self-damage, object ejection, workspace violations, or hazardous recovery behavior.
  • Limited treatment of execution-time safety: The study does not analyze what happens when the robot encounters an unexpected human, moving obstacle, hardware fault, communication delay, grasp failure, or controller deviation during an action.
  • No systematic failure taxonomy: Reported failures, especially in simultaneous towel folding and low-performing model configurations, are not categorized by cause, severity, recoverability, or whether they originate in reasoning, perception, programming, planning, or control.
  • Incomplete evaluation of closed-loop adaptation: Although the agent receives feedback, the paper does not compare closed-loop AGP with open-loop execution or quantify how often feedback changes the plan and whether those changes improve outcomes.
  • Unclear action granularity: The appropriate division between high-level model reasoning and low-level feedback control is not established; the paper does not determine which motions should be generated by the agent versus delegated to conventional or learned controllers.
  • Weak evidence for deformable-object manipulation: Towel folding is evaluated on a very narrow set of procedures, and the simultaneous condition has only three successes out of five; robustness to different fabrics, sizes, wrinkles, folds, friction, and initial configurations remains unknown.
  • Limited bimanual coordination analysis: The study does not characterize synchronization accuracy, force sharing, collision avoidance, or failure recovery for independently controlled arms.
  • Targeted throwing is under-evaluated: Throwing results are based on only two potato trials, and the reported time includes post-throw analysis; generalization to different object masses, shapes, release targets, trajectories, and environmental conditions is unresolved.
  • Potential evaluation leakage through task construction: The tasks are capability-driven and designed around behaviors expected to challenge AGP, but the paper does not establish whether they are representative of real deployment workloads or how performance changes on independently collected tasks.
  • No long-term reliability assessment: Repeated experience studies cover only five executions and a small number of assembly tasks; degradation, memory corruption, procedural drift, and performance over dozens or hundreds of executions are not examined.
  • Unclear transferability of stored experience: Experience transfer is tested only from Astra to Terra on one assembly task, leaving open whether saved measurements and scripts transfer across models, robots, task instances, environments, or incompatible interface conventions.
  • No controlled causal analysis of experience accumulation: The repeated-trial design conflates memory reuse with familiarity with fixed scenes and repeated object configurations, so it does not establish whether experience improves generalizable skill rather than scene-specific memorization.
  • Risk of stale or incorrect memory: The paper does not evaluate how the system detects and corrects outdated geometric estimates, invalid scripts, contradictory notes, or experience recorded after a failed execution.
  • Efficiency remains poorly characterized: Reported latency and cost are aggregated across reasoning, tool calls, observation capture, action execution, retries, and verification; the study does not identify which components dominate across task types or how much can be reduced without lowering reliability.
  • No real-time feasibility analysis: The paper does not define the latency requirements for dynamic manipulation or determine whether model inference delays are compatible with tasks involving fast contact, moving objects, or narrow timing windows.
  • Cost estimates may not reflect deployment economics: API pricing, caching, proprietary model availability, network latency, and repeated experience-writing costs may differ substantially from on-robot or production deployment conditions.
  • Unclear reproducibility: Although the interface and setup are described, the paper does not establish whether the full prompts, model versions, tool definitions, runtime traces, programs, calibration parameters, videos, and failed trials will be released.
  • No robustness study across model updates: Because the system depends on fixed but proprietary model versions, it remains unknown how performance changes after model updates, API changes, altered reasoning behavior, or different sampling settings.
  • Limited human-in-the-loop analysis: The protocol excludes human intervention except when it terminates a trial; the paper does not investigate supervisory correction, clarification requests, shared autonomy, or how humans could safely recover failed executions.
  • No assessment of user-level task authoring burden: The effort required to create task prompts, reference materials, robot-interface documentation, success criteria, and preparation-agent outputs is not measured.
  • Completion verification is under-specified: The paper relies on predefined physical criteria and recorded evidence, but does not evaluate false positives, false negatives, ambiguous completion states, or disagreements between agent-reported and externally verified success.
  • No formal policy-learning comparison: AGP uses fixed model parameters and runtime memory, but the paper does not determine whether consolidating successful programs into trained policies would provide better reliability, speed, cost, or transfer.
  • Open question about scalability: It remains unclear whether a single runtime agent can reliably manage substantially longer horizons, larger numbers of objects, branching task plans, multiple robots, or tasks requiring persistent state over extended periods.

Practical Applications

Immediate Applications

The paper supports near-term applications primarily as supervised pilots, laboratory systems, and low-risk industrial automation, rather than unrestricted deployment. The strongest evidence concerns tabletop manipulation with calibrated cameras, a robot arm, and human-defined task constraints.

  • Rapid robot-task prototyping for manufacturing and research laboratories — Robotics/Industry
    • Use case: An engineer provides a video, image, or language description of a new assembly task; AGP generates runtime scripts, estimates object geometry, commands the robot, and revises actions after failed grasps or insertions.
    • Potential tool or workflow: A “natural-language robot programming” interface that connects an MLLM coding agent to robot APIs, camera feeds, depth data, inverse kinematics, grippers, and motion controllers.
    • Practical value: Reduces the need to manually write task-specific perception and motion code for low-volume or frequently changing assembly operations.
    • Dependencies: Requires calibrated cameras, a documented and safety-constrained robot interface, reliable inverse kinematics, collision checking, reachable workspace, and human supervision. The reported experiments were small-scale, with success rates ranging from 80% to 100% on selected tasks; they do not establish production-level reliability.
  • Small-batch and high-mix assembly — Manufacturing
    • Use case: Assembly of varied parts, fixtures, kits, or product configurations where collecting a large demonstration dataset for every variant is impractical.
    • Potential product: A robot cell that interprets assembly videos or reference images and automatically creates a task definition, grasp strategy, insertion sequence, and verification routine.
    • Evidence from the paper: AGP assembled multiple part pairs from a human video and adapted alignment and insertion behavior using visual and motion feedback.
    • Dependencies: Parts must be visually distinguishable and sufficiently rigid; tolerances, lighting, occlusion, and force requirements must remain within the system’s capabilities. Safety-rated execution and deterministic fallback controllers would be necessary for factory use.
  • Image-guided packaging, kitting, and product arrangement — Logistics/Retail
    • Use case: Constructing a target arrangement of objects from a goal image, such as arranging components in a kit, stacking products, or preparing educational and promotional displays.
    • Potential workflow: A camera captures the workspace; the agent compares the current scene with the target image, plans object placements, executes them, and verifies the final configuration.
    • Evidence from the paper: Block construction from goal images achieved 29 successes in 30 trials across pyramid, tower, and multi-block configurations.
    • Dependencies: The objects should be graspable and stable, and the target arrangement must be visually unambiguous. More complex objects, clutter, reflective packaging, and fragile products require additional validation.
  • Robot programming education and laboratory instruction — Academia/Education
    • Use case: Students can specify manipulation tasks in natural language or code and observe how perception, geometry, planning, control, and feedback interact.
    • Potential tool: An educational “robot coding copilot” that exposes camera observations, robot state, geometric queries, and safe motion commands while recording the agent’s scripts and corrections.
    • Practical value: Makes robotics experimentation more accessible to students who lack extensive expertise in motion planning or robot middleware.
    • Dependencies: The platform must operate in simulation or under strict speed, force, workspace, and emergency-stop limits. Generated programs require instructor review because language-model reasoning can produce incorrect or unsafe commands.
  • Benchmarking embodied agents and robot interfaces — Academia
    • Use case: Researchers can use AGP-style interfaces to evaluate multimodal agents on long-horizon reasoning, action generation, recovery, and deformable-object manipulation without training a separate policy for every task.
    • Potential research workflow: Record task specifications, observations, tool calls, action requests, failures, execution time, token usage, and physical outcomes in a standardized benchmark.
    • Evidence from the paper: The task suite spans video-guided assembly, image-guided construction, dice flipping, targeted throwing, and bimanual towel folding.
    • Dependencies: Results must be replicated across robots, models, environments, and larger trial counts. Current findings may depend on the particular model, prompt, interface design, and highly controlled tabletop scenes.
  • Reusable robot skill and experience libraries — Robotics/Software
    • Use case: Store successful geometric estimates, grasp poses, scripts, corrections, and recovery procedures for reuse in later executions.
    • Potential product: A version-controlled “robot task memory” repository containing task definitions, validated programs, calibration metadata, and execution logs.
    • Evidence from the paper: Reusing experience reduced two-pair assembly time by 29.3% between the first and fifth execution; transferring experience from a stronger agent increased a weaker agent’s success rate from 1/5 to 4/5.
    • Dependencies: Stored procedures must be checked against the current object pose, robot configuration, calibration, and environment. Memory reuse should not bypass fresh visual verification or safety checks.
  • Supervised manipulation of household or office objects — Daily life/Assistive robotics
    • Use case: A robot could perform constrained tasks such as arranging blocks, flipping objects, folding simple textiles, or preparing items on a table based on spoken instructions.
    • Potential workflow: A user supplies a language instruction or demonstration video; the robot asks for confirmation, executes a bounded action sequence, and requests assistance when uncertain.
    • Dependencies: This is feasible mainly for low-risk, tabletop tasks. Human oversight, slow execution, reliable object detection, physical safeguards, and clear uncertainty reporting are essential. The paper does not demonstrate safe interaction with people, children, pets, hot objects, sharp tools, or fragile items.
  • Policy and workplace training for human oversight of embodied AI — Policy/Industry
    • Use case: Use the AGP workflow to develop operating procedures for robot supervision, logging, intervention, and incident review.
    • Potential policy artifact: A standard requiring robot agents to retain task definitions, generated programs, sensor observations, motion commands, controller feedback, and completion evidence.
    • Practical value: The paper’s persistent workspace and execution records provide a model for auditability and post-task analysis.
    • Dependencies: Logs must be tamper-resistant, privacy-preserving, and sufficient to reconstruct failures. Formal standards for responsibility, approval thresholds, and human intervention are still needed.

Long-Term Applications

The following applications require further work in reliability, latency, safety, hardware integration, generalization, and regulatory validation before broad deployment.

  • General-purpose flexible manufacturing cells — Manufacturing/Robotics
    • Use case: A single robot cell autonomously switches among assembly, inspection, packing, rework, and material-handling tasks from natural-language work orders.
    • Potential product: An agent-controlled manufacturing operating system combining task planning, perception, motion generation, skill memory, quality verification, and scheduling.
    • Why long term: The paper demonstrates selected tabletop tasks, but industrial deployment requires near-perfect reliability, cycle times of seconds rather than tens of minutes, robust force control, multi-object clutter handling, and formal safety certification.
    • Dependencies: High-quality sensors, standardized robot interfaces, deterministic low-level controllers, integration with manufacturing execution systems, and validation under distribution shifts.
  • Autonomous recovery and rework in production — Manufacturing/Quality control
    • Use case: After detecting a misalignment, dropped part, incomplete insertion, or defective placement, the robot independently gathers additional views, changes its grasp or approach direction, and retries.
    • Potential workflow: Camera-based verification triggers AGP recovery; the agent records the failure mode and adds a validated correction to the task memory.
    • Evidence from the paper: AGP explicitly revises estimates and actions after physical outcomes, rather than selecting only among fixed learned skills.
    • Dependencies: Recovery must be bounded and formally verified. An agent must distinguish recoverable errors from conditions that require immediate shutdown, such as collisions, damaged parts, or human intrusion.
  • Robotic manipulation of deformable materials — Apparel, laundry, food, and soft goods
    • Use case: Folding garments, towels, linens, packaging films, cables, or flexible medical materials; potentially sorting and preparing textiles in commercial facilities or homes.
    • Potential product: A deformable-object manipulation system that combines visual state estimation, dual-arm coordination, tactile sensing, and agent-level planning.
    • Evidence and limitation: Sequential towel folding succeeded in 5/5 trials, but required an average of 50.8 minutes; simultaneous folding succeeded in only 3/5 trials.
    • Dependencies: Requires improved cloth-state estimation, dynamic modeling, force/tactile feedback, faster planning, and robust bimanual synchronization. Generalization across fabrics, wrinkles, sizes, and workspace layouts remains unresolved.
  • Assistive and eldercare robots — Healthcare/Independent living
    • Use case: Robots could perform individualized tabletop assistance, such as organizing medication packages, folding clothing, preparing objects, or responding to spoken instructions.
    • Potential product: A supervised home-assistance robot that adapts to demonstrations and retains user-specific procedures.
    • Why long term: Healthcare and home environments introduce people, clutter, unpredictable motion, privacy concerns, and high consequences for mistakes.
    • Dependencies: Extensive safety testing, user consent, privacy-preserving sensing, certified hardware, robust human-aware planning, fallbacks, and clinical or regulatory approval. The system should not independently handle medication dosing, sharp objects, or physical transfers without specialized safeguards.
  • Natural-language programming for general-purpose service robots — Robotics/Software
    • Use case: Users specify tasks such as “move these objects into the matching bins,” “prepare the table as shown in the image,” or “fold the towels like in this video.”
    • Potential product: A robot development platform with task preparation agents, runtime execution agents, code generation, simulation checks, and deployment monitoring.
    • Novel contribution: AGP moves beyond selecting pre-trained skills by allowing the agent to decide what to observe, how to interpret evidence, and how to generate motion requests at runtime.
    • Dependencies: Better grounding between generated code and physical constraints; typed, capability-aware robot APIs; sandboxing; simulation or digital-twin validation; and mechanisms that prevent hallucinated functions, invalid geometry, or unsafe trajectories.
  • Distributed agent specialization and model transfer — Robotics infrastructure
    • Use case: A powerful, expensive model explores a new task and creates a validated procedure; a cheaper model then performs routine repetitions using the resulting experience.
    • Potential workflow: exploration agent → validated task memory → economical execution agent → monitoring and update.
    • Evidence from the paper: Strong-to-weak experience transfer improved success and reduced successful-trial time, token usage, and inference cost.
    • Dependencies: Experience must be portable across models and robot embodiments, and transferred procedures need provenance, confidence scores, and regression testing. Cost savings may disappear if every environment requires expensive re-exploration.
  • Multi-robot and warehouse coordination — Logistics/Industry
    • Use case: Agents assign and adapt manipulation tasks across multiple arms or mobile manipulators, coordinating observation, grasping, timed release, packing, and exception handling.
    • Potential product: A fleet-level embodied-agent scheduler that shares task memories and reallocates work after failures.
    • Why long term: The paper demonstrates concurrent motion for two arms but not large-scale multi-robot coordination, shared-space collision avoidance, or operational scheduling.
    • Dependencies: Real-time communication, centralized safety supervision, shared world models, resource allocation, deterministic collision avoidance, and integration with warehouse control systems.
  • Autonomous scientific and engineering experimentation — Academia/R&D
    • Use case: Agents perform repeatable physical experiments involving object arrangement, assembly, manipulation, and iterative adjustment based on measured outcomes.
    • Potential workflow: A researcher specifies an experimental protocol; the agent prepares the task, executes trials, records observations, modifies parameters, and produces a structured report.
    • Dependencies: Experimental validity requires calibrated instruments, strict protocol compliance, reproducible randomization, provenance of all actions, and human approval for changes. Language-model-generated procedures cannot replace domain-specific experimental safeguards.
  • Regulatory frameworks for agent-controlled physical systems — Policy
    • Use case: Establish requirements for certification, audit logs, human override, model updates, incident reporting, and operational limits for robots controlled by general-purpose agents.
    • Potential policy outputs: Risk-tiered approval schemes, mandatory simulation and physical test suites, minimum intervention latency, traceability of generated programs, and restrictions on autonomous operation in public or clinical settings.
    • Dependencies: Standards must account for changing model behavior, external API updates, prompt or memory contamination, cybersecurity threats, and differences between low-risk tabletop tasks and high-energy industrial systems.
  • Everyday household robots capable of adapting to demonstrations — Daily life
    • Use case: A household robot learns tasks from videos or demonstrations, such as sorting laundry, arranging groceries, folding clothes, or preparing simple items.
    • Why long term: Household environments are much less structured than the paper’s tabletop scenes, and success requires safe interaction with humans, varied objects, narrow spaces, and continuously changing conditions.
    • Dependencies: Low-cost robust hardware, tactile sensing, real-time inference, privacy protections, reliable uncertainty estimation, affordable maintenance, and highly dependable recovery behavior. The paper’s execution times and inference costs currently make widespread consumer deployment impractical.

Glossary

  • Action diversity: Variety of distinct physical actions a system can perform. “Dice flipping and targeted throwing follow language instructions and test action diversity”
  • Agent as Policy (AGP): An approach in which a general-purpose agent directly performs the role of the robot’s control policy. “We introduce Agent as Policy (AGP), which places task planning and execution under the agent's control.”
  • Affordances: The action possibilities that an environment or object provides to an agent. “A second line selects and sequences learned skills according to task goals, affordances, and environment feedback”
  • Bimanual manipulation: Robotic manipulation performed using two arms or hands. “Bimanual towel folding uses two arms with a common coordinate frame”
  • Camera calibration: The process of estimating a camera’s imaging parameters and its geometric relationship to a robot or scene. “The bridge grounds these computations in camera calibration and robot coordinates”
  • Code generation: Automatic creation of executable programs from a high-level specification or instruction. “Code Generation for Free-Form Manipulation Tasks across Real and Simulation”
  • Computation graph: A graph-structured representation of operations and their dependencies. “expressed as code or computation graphs that the robot runs during task execution”
  • Deformable object: An object whose shape can change substantially during manipulation. “Bimanual towel folding follows video instructions and tests object diversity through coordinated manipulation of deformable material.”
  • Diffusion model: A generative model that produces data by progressively transforming noise into a structured sample. “with diffusion models capturing action sequences”
  • Dynamic manipulation: Manipulation involving motion, timing, momentum, or changing physical states. “We study AGP across multiple real-world manipulation tasks spanning precision manipulation, dynamic motions, and deformable objects.”
  • End-effector pose: The position and orientation of a robot’s tool or terminal component. “the corresponding end effector pose computed through forward kinematics”
  • Experience accumulation: Improvement based on storing and reusing information from previous task executions. “We further study efficiency through task experience accumulation”
  • Forward kinematics: Computing a robot end effector’s pose from its joint positions. “the corresponding end effector pose computed through forward kinematics”
  • Foundation model: A broadly pretrained model that can be adapted to many tasks rather than trained for one specific application. “general purpose foundation MLLMs can serve as policies for direct robot arm control”
  • Geometric modeling: Representing the shapes, positions, and spatial relationships of objects and robot components mathematically. “Classical robot manipulation combines geometric modeling, task and motion planning, and feedback control”
  • Gripper: A robotic end effector designed to grasp or hold objects. “a parallel gripper in a tabletop workspace”
  • Inference cost: The computational or financial resources required for a model to generate outputs. “The substantial execution time and inference cost remain barriers to practical deployment.”
  • Inverse kinematics: Computing robot joint configurations required to achieve a desired end-effector pose. “For end effector targets, it uses Mink~\citep{zakka2026mink} for inverse kinematics to compute joint targets.”
  • Joint trajectory: A time-ordered sequence of positions or configurations for a robot’s joints. “It then generates joint trajectories subject to workspace and motion limits”
  • Long-horizon reasoning: Planning and decision-making over many sequential actions or an extended task duration. “We evaluate AGP on real-world robot tasks spanning long horizon reasoning, action diversity, and object diversity”
  • Multimodal LLM (MLLM): A LLM capable of processing and reasoning over multiple modalities, such as text, images, and video. “Multimodal LLMs (MLLMs) demonstrate strong general capabilities”
  • Orchestration: Selecting, sequencing, and coordinating multiple skills, policies, or processes to accomplish a task. “Runtime policy orchestration.”
  • Persistent memory: Information retained across separate interactions or task executions. “Persistent memory carries execution feedback across trials”
  • Policy orchestration: The selection and sequencing of learned policies according to task requirements and feedback. “The second approach uses agents as orchestrators of existing learned policies”
  • Precision manipulation: Robot manipulation requiring highly accurate positioning, alignment, or force application. “We study AGP across multiple real-world manipulation tasks spanning precision manipulation”
  • Proprioceptive state: Information about a robot’s own configuration and motion obtained from internal sensors. “This proprioceptive state describes the robot's current configuration”
  • Recovery: Adaptation or corrective action taken after an error, failure, or unexpected physical outcome. “The agent determines grasp poses, action sequences, observation timing, and recovery during this cycle.”
  • Revolute joint: A robot joint that permits rotation around a fixed axis. “We use an I2RT YAM~\citep{i2rt_yam} arm with six revolute joints”
  • Robot policy: A rule, model, or agent that maps task information and observations to robot actions. “We use a general purpose agent as the robot policy.”
  • Spatiotemporal: Relating to both spatial structure and changes over time. “Spatio-Temporal Reasoning of Relational Keypoint Constraints for Robotic Manipulation”
  • Task and motion planning: Coordinated planning of high-level task steps and the robot motions needed to execute them. “Classical robot manipulation combines geometric modeling, task and motion planning, and feedback control”
  • Tool call: A structured request from an agent to obtain information or execute an operation through an external interface. “A tool call may acquire information or execute a program that submits action targets to the robot controller.”
  • Trajectory: A time-dependent path or sequence of states followed by a moving object or robot. “subsequent trajectory analysis and verification account for approximately 40\% of the reported 22.8 minutes.”
  • Visuomotor policy: A policy that maps visual observations to motor actions. “Learned visuomotor policies predict actions from robot trajectory data”
  • Vision-language-action (VLA) model: A model that connects visual and linguistic inputs to physical actions. “Vision language action models extend this approach to diverse tasks through image and language conditioning”
  • Workspace: The region of physical space that a robot can reach or operate within. “It then generates joint trajectories subject to workspace and motion limits”
  • Zero-shot manipulation: Performing a manipulation task without task-specific training examples or prior task execution. “These results support its effectiveness for zero shot manipulation across the evaluated tasks.”

Open Problems

We haven't generated a list of open problems mentioned in this paper yet.

Tweets

Sign up for free to view the 3 tweets with 73 likes about this paper.