Agent as Policy for Robotic Manipulation
Abstract: We demonstrate that a general-purpose agent can directly drive a physical robot throughout task execution without any task-specific or environment-specific training. We introduce Agent as Policy (AGP), which places task planning and execution under the agent's control. Given a task and a robot interface, the agent interprets visual evidence, writes executable programs, issues motion commands, and revises its actions in response to physical outcomes. This brings the agent's reasoning and programming capabilities into continuous interaction with the physical world. We study AGP across multiple real-world manipulation tasks spanning precision manipulation, dynamic motions, and deformable objects. These include assembly from human videos, block construction from goal images, dice flipping, targeted throwing, and bimanual towel folding. Across assembly, block construction, and dice flipping, AGP succeeds in at least eight of ten trials for each evaluated task configuration. We further study efficiency through task experience accumulation and find that reusing saved procedures and programs shortens execution time across repeated trials. These findings support a path for general-purpose agents to act as robot policies, extending their autonomy to physical manipulation through runtime reasoning, programming, and interaction.
Paper Prompts
Sign up for free to create and run prompts on this paper.
Top Community Prompts
Explain it Like I'm 14
1. What is this paper about?
This paper introduces a system called Agent as Policy, or AGP. It explores whether a general-purpose AI agent can control a real robot without being specially trained for every single task.
Usually, a robot is trained to do one particular thing, such as picking up a cup. AGP works differently. The AI can:
- Look at pictures or videos.
- Understand what it is supposed to do.
- Write small computer programs.
- Tell the robot how to move.
- Check whether the movement worked.
- Change its plan if something went wrong.
In simple terms, the AI acts like the robot’s brain, planner, programmer, and problem-solver all at once.
2. What questions did the researchers ask?
The researchers wanted to find out:
- Can a general AI control a robot without special training for each new task?
- Can the AI understand different kinds of instructions, such as videos, pictures, and written commands?
- Can it react to mistakes and unexpected results?
- Can it perform different kinds of movements, including careful assembly, throwing, flipping objects, and folding fabric?
- Can the AI become faster by remembering what it learned from earlier attempts?
- Can one stronger AI pass its useful experience to another, less capable AI?
The main idea is that the AI should not simply follow a fixed list of instructions. Instead, it should behave more like a person solving a problem: observe, act, check, and try again.
3. How did the researchers test AGP?
The robot and its “senses”
The researchers connected the AI to a robot arm with a gripper. The robot used cameras to see its workspace:
- An overhead camera gave a view from above.
- A camera on the robot’s wrist gave a closer view.
- Depth information helped the AI estimate how far away objects were.
The robot also reported information about its own body, such as the positions of its joints. This is similar to how people know where their arms and hands are without looking directly at them.
The robot interface
The AI communicated with the robot through a special connection called a robot interface. This worked like a translator:
- The AI requested pictures or information.
- The robot sent back camera images and movement information.
- The AI sent commands such as “move the gripper here” or “close the fingers.”
- The robot carried out the movement and reported what happened.
Runtime programming
The AI wrote programs while it was working. For example, it might write a program to:
- Find an object in a camera image.
- Estimate the object’s position.
- Calculate where the gripper should move.
- Open or close the gripper.
- Check the result.
This is called runtime programming because the programs are created during the task, rather than being prepared completely beforehand.
An everyday comparison would be a person writing down a new plan while assembling furniture after discovering that one piece does not fit.
The tasks
The researchers tested AGP on several real-world tasks:
| Task | What the robot had to do |
|---|---|
| Assembly | Put parts together by following a human demonstration video |
| Block construction | Build shapes, such as a pyramid or towers, from a goal image |
| Dice flipping | Turn six dice so that the correct number appeared on top |
| Targeted throwing | Throw an object toward a particular target |
| Towel folding | Fold a towel using one or two robot arms |
The researchers measured success, time, the amount of text the AI processed, and the estimated cost of using the AI.
4. What did they find?
AGP often succeeded without special training
The robot performed well on many tasks even though it had not been specially trained for each one.
Some important results were:
- Assembly: 8 successful trials out of 10.
- Block construction: 29 successful trials out of 30.
- Dice flipping: 10 successful trials out of 10.
- Potato throwing: 2 successful trials out of 2.
- Sequential towel folding: 5 successful trials out of 5.
- Simultaneous towel folding: 3 successful trials out of 5.
Overall, the results suggest that a general-purpose AI can control a robot in several unfamiliar situations.
The AI could adapt after mistakes
AGP did not always follow one unchanging plan. If an object was in a slightly different place or an insertion failed, the AI could:
- Ask for another camera image.
- Measure the object again.
- Change the gripper’s position.
- Try a different movement.
- Use feedback from the robot to improve the next attempt.
This is important because real environments are messy. Objects may move, cameras may be imperfect, and movements may not work exactly as expected.
Experience made repeated tasks faster
The researchers allowed the AI to save useful information in files, including:
- Measurements.
- Successful movement plans.
- Mistakes and corrections.
- Programs for processing camera images.
When the AI repeated the same task, it could reuse this information rather than starting from nothing.
For one assembly task, the time needed for the task decreased by about 29% between the first and fifth attempts. The AI became more efficient because it remembered useful details while still checking the objects’ current positions.
A stronger AI could help a weaker AI
The researchers also tested whether one more capable AI could prepare instructions and experience for another AI.
Without stored experience, the weaker AI succeeded only 1 time out of 5. With experience from the stronger AI, it succeeded 4 times out of 5. Its average completion time and amount of computer processing also decreased on successful attempts.
This is similar to an experienced student giving a beginner a helpful set of notes and tips before the beginner tries the assignment.
The system was still slow and expensive
Despite its successes, AGP had important weaknesses:
- Some tasks took a long time.
- The AI used a large amount of computer processing.
- Using a powerful AI can be expensive.
- Folding a towel was difficult, especially when both arms had to move together.
- The experiments used relatively small numbers of trials, so the results do not prove that AGP will work reliably in every situation.
For example, sequential towel folding took about 51 minutes on average, even though all five trials succeeded.
5. Why are these results important?
Traditional robots are often very good at repeating one carefully programmed action, but they may struggle when something changes. AGP shows a possible way to build robots that are more flexible.
Instead of giving the robot every instruction in advance, people could give it a goal, such as:
“Build this structure from the picture.”
The AI could then decide what to look at, write the needed calculations, move the robot, and correct mistakes.
This could eventually help robots perform more useful jobs in places where conditions change, such as:
- Homes.
- Hospitals.
- Factories.
- Laboratories.
- Disaster areas.
- Warehouses.
Simple conclusion
The paper argues that a general-purpose AI can serve as a robot’s policy, meaning the system that decides what the robot should do next. AGP combines seeing, reasoning, programming, acting, and learning from feedback in one continuing process.
The experiments show promising results: the robot completed many tasks without task-specific training, and it became faster when it reused past experience. However, the system is still too slow, costly, and sometimes unreliable for widespread everyday use.
The biggest long-term possibility is that robots may become less like machines that only repeat fixed instructions and more like adaptable helpers that can understand new goals, try solutions, and learn from their experiences.
Knowledge Gaps
Knowledge Gaps, Limitations, and Open Questions
The paper leaves the following issues unresolved:
- Limited statistical evidence: Most evaluations use only five or ten trials per configuration, making success-rate estimates highly uncertain and preventing reliable significance testing.
- Restricted hardware validation: AGP is evaluated primarily on one robot platform, with limited evidence that the approach transfers across robot morphologies, grippers, cameras, controllers, or workspace scales.
- Dependence on a narrow model ecosystem: The main experiments rely heavily on proprietary, high-capability models, so it remains unclear whether AGP is viable with open-weight, smaller, cheaper, or locally deployable models.
- Insufficient comparison with strong baselines: The paper does not provide systematic comparisons against trained vision-language-action policies, classical task-and-motion planners, code-generation baselines, or hybrid agent-policy systems under matched hardware, task, and budget conditions.
- Unclear source of performance gains: The experiments do not isolate how much success comes from the underlying MLLM, runtime programming, visual feedback, robot-interface design, task preparation, or persistent experience.
- No component-level ablation of the AGP interface: The contributions of wrist images, depth, proprioception, geometric queries, motion feedback, program execution, and action-result monitoring are not separately evaluated.
- Limited task and object diversity: The benchmark contains a small set of tabletop tasks and does not establish performance on cluttered scenes, articulated objects, fragile objects, transparent or reflective materials, tool use, insertion under occlusion, or contact-rich manipulation.
- Weak evidence for generalization to unseen environments: The experiments use predefined initial scenes and fixed task configurations; robustness to novel layouts, object instances, lighting, camera poses, workspace arrangements, and distractors remains unexplored.
- Unclear generalization beyond prepared task definitions: A preparation agent supplies reusable task definitions, constraints, and completion criteria, but the study does not measure how errors or omissions in these definitions affect runtime execution.
- Limited evaluation of instruction ambiguity: The paper does not test ambiguous, incomplete, contradictory, compositional, multilingual, or naturally varying user instructions.
- Insufficient analysis of perception failures: The system’s ability to detect and recover from incorrect depth, occlusion, segmentation, pose estimation, camera calibration, or object-identity errors is not quantified.
- No formal reliability or uncertainty mechanism: AGP appears to act on model-generated interpretations and motion programs without calibrated confidence estimates or explicit uncertainty propagation from perception through control.
- Unresolved safety guarantees: The interface validates motion requests, but the paper does not establish formal guarantees against collisions, unsafe contact forces, self-damage, object ejection, workspace violations, or hazardous recovery behavior.
- Limited treatment of execution-time safety: The study does not analyze what happens when the robot encounters an unexpected human, moving obstacle, hardware fault, communication delay, grasp failure, or controller deviation during an action.
- No systematic failure taxonomy: Reported failures, especially in simultaneous towel folding and low-performing model configurations, are not categorized by cause, severity, recoverability, or whether they originate in reasoning, perception, programming, planning, or control.
- Incomplete evaluation of closed-loop adaptation: Although the agent receives feedback, the paper does not compare closed-loop AGP with open-loop execution or quantify how often feedback changes the plan and whether those changes improve outcomes.
- Unclear action granularity: The appropriate division between high-level model reasoning and low-level feedback control is not established; the paper does not determine which motions should be generated by the agent versus delegated to conventional or learned controllers.
- Weak evidence for deformable-object manipulation: Towel folding is evaluated on a very narrow set of procedures, and the simultaneous condition has only three successes out of five; robustness to different fabrics, sizes, wrinkles, folds, friction, and initial configurations remains unknown.
- Limited bimanual coordination analysis: The study does not characterize synchronization accuracy, force sharing, collision avoidance, or failure recovery for independently controlled arms.
- Targeted throwing is under-evaluated: Throwing results are based on only two potato trials, and the reported time includes post-throw analysis; generalization to different object masses, shapes, release targets, trajectories, and environmental conditions is unresolved.
- Potential evaluation leakage through task construction: The tasks are capability-driven and designed around behaviors expected to challenge AGP, but the paper does not establish whether they are representative of real deployment workloads or how performance changes on independently collected tasks.
- No long-term reliability assessment: Repeated experience studies cover only five executions and a small number of assembly tasks; degradation, memory corruption, procedural drift, and performance over dozens or hundreds of executions are not examined.
- Unclear transferability of stored experience: Experience transfer is tested only from Astra to Terra on one assembly task, leaving open whether saved measurements and scripts transfer across models, robots, task instances, environments, or incompatible interface conventions.
- No controlled causal analysis of experience accumulation: The repeated-trial design conflates memory reuse with familiarity with fixed scenes and repeated object configurations, so it does not establish whether experience improves generalizable skill rather than scene-specific memorization.
- Risk of stale or incorrect memory: The paper does not evaluate how the system detects and corrects outdated geometric estimates, invalid scripts, contradictory notes, or experience recorded after a failed execution.
- Efficiency remains poorly characterized: Reported latency and cost are aggregated across reasoning, tool calls, observation capture, action execution, retries, and verification; the study does not identify which components dominate across task types or how much can be reduced without lowering reliability.
- No real-time feasibility analysis: The paper does not define the latency requirements for dynamic manipulation or determine whether model inference delays are compatible with tasks involving fast contact, moving objects, or narrow timing windows.
- Cost estimates may not reflect deployment economics: API pricing, caching, proprietary model availability, network latency, and repeated experience-writing costs may differ substantially from on-robot or production deployment conditions.
- Unclear reproducibility: Although the interface and setup are described, the paper does not establish whether the full prompts, model versions, tool definitions, runtime traces, programs, calibration parameters, videos, and failed trials will be released.
- No robustness study across model updates: Because the system depends on fixed but proprietary model versions, it remains unknown how performance changes after model updates, API changes, altered reasoning behavior, or different sampling settings.
- Limited human-in-the-loop analysis: The protocol excludes human intervention except when it terminates a trial; the paper does not investigate supervisory correction, clarification requests, shared autonomy, or how humans could safely recover failed executions.
- No assessment of user-level task authoring burden: The effort required to create task prompts, reference materials, robot-interface documentation, success criteria, and preparation-agent outputs is not measured.
- Completion verification is under-specified: The paper relies on predefined physical criteria and recorded evidence, but does not evaluate false positives, false negatives, ambiguous completion states, or disagreements between agent-reported and externally verified success.
- No formal policy-learning comparison: AGP uses fixed model parameters and runtime memory, but the paper does not determine whether consolidating successful programs into trained policies would provide better reliability, speed, cost, or transfer.
- Open question about scalability: It remains unclear whether a single runtime agent can reliably manage substantially longer horizons, larger numbers of objects, branching task plans, multiple robots, or tasks requiring persistent state over extended periods.
Practical Applications
Immediate Applications
The paper supports near-term applications primarily as supervised pilots, laboratory systems, and low-risk industrial automation, rather than unrestricted deployment. The strongest evidence concerns tabletop manipulation with calibrated cameras, a robot arm, and human-defined task constraints.
- Rapid robot-task prototyping for manufacturing and research laboratories — Robotics/Industry
- Use case: An engineer provides a video, image, or language description of a new assembly task; AGP generates runtime scripts, estimates object geometry, commands the robot, and revises actions after failed grasps or insertions.
- Potential tool or workflow: A “natural-language robot programming” interface that connects an MLLM coding agent to robot APIs, camera feeds, depth data, inverse kinematics, grippers, and motion controllers.
- Practical value: Reduces the need to manually write task-specific perception and motion code for low-volume or frequently changing assembly operations.
- Dependencies: Requires calibrated cameras, a documented and safety-constrained robot interface, reliable inverse kinematics, collision checking, reachable workspace, and human supervision. The reported experiments were small-scale, with success rates ranging from 80% to 100% on selected tasks; they do not establish production-level reliability.
- Small-batch and high-mix assembly — Manufacturing
- Use case: Assembly of varied parts, fixtures, kits, or product configurations where collecting a large demonstration dataset for every variant is impractical.
- Potential product: A robot cell that interprets assembly videos or reference images and automatically creates a task definition, grasp strategy, insertion sequence, and verification routine.
- Evidence from the paper: AGP assembled multiple part pairs from a human video and adapted alignment and insertion behavior using visual and motion feedback.
- Dependencies: Parts must be visually distinguishable and sufficiently rigid; tolerances, lighting, occlusion, and force requirements must remain within the system’s capabilities. Safety-rated execution and deterministic fallback controllers would be necessary for factory use.
- Image-guided packaging, kitting, and product arrangement — Logistics/Retail
- Use case: Constructing a target arrangement of objects from a goal image, such as arranging components in a kit, stacking products, or preparing educational and promotional displays.
- Potential workflow: A camera captures the workspace; the agent compares the current scene with the target image, plans object placements, executes them, and verifies the final configuration.
- Evidence from the paper: Block construction from goal images achieved 29 successes in 30 trials across pyramid, tower, and multi-block configurations.
- Dependencies: The objects should be graspable and stable, and the target arrangement must be visually unambiguous. More complex objects, clutter, reflective packaging, and fragile products require additional validation.
- Robot programming education and laboratory instruction — Academia/Education
- Use case: Students can specify manipulation tasks in natural language or code and observe how perception, geometry, planning, control, and feedback interact.
- Potential tool: An educational “robot coding copilot” that exposes camera observations, robot state, geometric queries, and safe motion commands while recording the agent’s scripts and corrections.
- Practical value: Makes robotics experimentation more accessible to students who lack extensive expertise in motion planning or robot middleware.
- Dependencies: The platform must operate in simulation or under strict speed, force, workspace, and emergency-stop limits. Generated programs require instructor review because language-model reasoning can produce incorrect or unsafe commands.
- Benchmarking embodied agents and robot interfaces — Academia
- Use case: Researchers can use AGP-style interfaces to evaluate multimodal agents on long-horizon reasoning, action generation, recovery, and deformable-object manipulation without training a separate policy for every task.
- Potential research workflow: Record task specifications, observations, tool calls, action requests, failures, execution time, token usage, and physical outcomes in a standardized benchmark.
- Evidence from the paper: The task suite spans video-guided assembly, image-guided construction, dice flipping, targeted throwing, and bimanual towel folding.
- Dependencies: Results must be replicated across robots, models, environments, and larger trial counts. Current findings may depend on the particular model, prompt, interface design, and highly controlled tabletop scenes.
- Reusable robot skill and experience libraries — Robotics/Software
- Use case: Store successful geometric estimates, grasp poses, scripts, corrections, and recovery procedures for reuse in later executions.
- Potential product: A version-controlled “robot task memory” repository containing task definitions, validated programs, calibration metadata, and execution logs.
- Evidence from the paper: Reusing experience reduced two-pair assembly time by 29.3% between the first and fifth execution; transferring experience from a stronger agent increased a weaker agent’s success rate from 1/5 to 4/5.
- Dependencies: Stored procedures must be checked against the current object pose, robot configuration, calibration, and environment. Memory reuse should not bypass fresh visual verification or safety checks.
- Supervised manipulation of household or office objects — Daily life/Assistive robotics
- Use case: A robot could perform constrained tasks such as arranging blocks, flipping objects, folding simple textiles, or preparing items on a table based on spoken instructions.
- Potential workflow: A user supplies a language instruction or demonstration video; the robot asks for confirmation, executes a bounded action sequence, and requests assistance when uncertain.
- Dependencies: This is feasible mainly for low-risk, tabletop tasks. Human oversight, slow execution, reliable object detection, physical safeguards, and clear uncertainty reporting are essential. The paper does not demonstrate safe interaction with people, children, pets, hot objects, sharp tools, or fragile items.
- Policy and workplace training for human oversight of embodied AI — Policy/Industry
- Use case: Use the AGP workflow to develop operating procedures for robot supervision, logging, intervention, and incident review.
- Potential policy artifact: A standard requiring robot agents to retain task definitions, generated programs, sensor observations, motion commands, controller feedback, and completion evidence.
- Practical value: The paper’s persistent workspace and execution records provide a model for auditability and post-task analysis.
- Dependencies: Logs must be tamper-resistant, privacy-preserving, and sufficient to reconstruct failures. Formal standards for responsibility, approval thresholds, and human intervention are still needed.
Long-Term Applications
The following applications require further work in reliability, latency, safety, hardware integration, generalization, and regulatory validation before broad deployment.
- General-purpose flexible manufacturing cells — Manufacturing/Robotics
- Use case: A single robot cell autonomously switches among assembly, inspection, packing, rework, and material-handling tasks from natural-language work orders.
- Potential product: An agent-controlled manufacturing operating system combining task planning, perception, motion generation, skill memory, quality verification, and scheduling.
- Why long term: The paper demonstrates selected tabletop tasks, but industrial deployment requires near-perfect reliability, cycle times of seconds rather than tens of minutes, robust force control, multi-object clutter handling, and formal safety certification.
- Dependencies: High-quality sensors, standardized robot interfaces, deterministic low-level controllers, integration with manufacturing execution systems, and validation under distribution shifts.
- Autonomous recovery and rework in production — Manufacturing/Quality control
- Use case: After detecting a misalignment, dropped part, incomplete insertion, or defective placement, the robot independently gathers additional views, changes its grasp or approach direction, and retries.
- Potential workflow: Camera-based verification triggers AGP recovery; the agent records the failure mode and adds a validated correction to the task memory.
- Evidence from the paper: AGP explicitly revises estimates and actions after physical outcomes, rather than selecting only among fixed learned skills.
- Dependencies: Recovery must be bounded and formally verified. An agent must distinguish recoverable errors from conditions that require immediate shutdown, such as collisions, damaged parts, or human intrusion.
- Robotic manipulation of deformable materials — Apparel, laundry, food, and soft goods
- Use case: Folding garments, towels, linens, packaging films, cables, or flexible medical materials; potentially sorting and preparing textiles in commercial facilities or homes.
- Potential product: A deformable-object manipulation system that combines visual state estimation, dual-arm coordination, tactile sensing, and agent-level planning.
- Evidence and limitation: Sequential towel folding succeeded in 5/5 trials, but required an average of 50.8 minutes; simultaneous folding succeeded in only 3/5 trials.
- Dependencies: Requires improved cloth-state estimation, dynamic modeling, force/tactile feedback, faster planning, and robust bimanual synchronization. Generalization across fabrics, wrinkles, sizes, and workspace layouts remains unresolved.
- Assistive and eldercare robots — Healthcare/Independent living
- Use case: Robots could perform individualized tabletop assistance, such as organizing medication packages, folding clothing, preparing objects, or responding to spoken instructions.
- Potential product: A supervised home-assistance robot that adapts to demonstrations and retains user-specific procedures.
- Why long term: Healthcare and home environments introduce people, clutter, unpredictable motion, privacy concerns, and high consequences for mistakes.
- Dependencies: Extensive safety testing, user consent, privacy-preserving sensing, certified hardware, robust human-aware planning, fallbacks, and clinical or regulatory approval. The system should not independently handle medication dosing, sharp objects, or physical transfers without specialized safeguards.
- Natural-language programming for general-purpose service robots — Robotics/Software
- Use case: Users specify tasks such as “move these objects into the matching bins,” “prepare the table as shown in the image,” or “fold the towels like in this video.”
- Potential product: A robot development platform with task preparation agents, runtime execution agents, code generation, simulation checks, and deployment monitoring.
- Novel contribution: AGP moves beyond selecting pre-trained skills by allowing the agent to decide what to observe, how to interpret evidence, and how to generate motion requests at runtime.
- Dependencies: Better grounding between generated code and physical constraints; typed, capability-aware robot APIs; sandboxing; simulation or digital-twin validation; and mechanisms that prevent hallucinated functions, invalid geometry, or unsafe trajectories.
- Distributed agent specialization and model transfer — Robotics infrastructure
- Use case: A powerful, expensive model explores a new task and creates a validated procedure; a cheaper model then performs routine repetitions using the resulting experience.
- Potential workflow:
exploration agent → validated task memory → economical execution agent → monitoring and update. - Evidence from the paper: Strong-to-weak experience transfer improved success and reduced successful-trial time, token usage, and inference cost.
- Dependencies: Experience must be portable across models and robot embodiments, and transferred procedures need provenance, confidence scores, and regression testing. Cost savings may disappear if every environment requires expensive re-exploration.
- Multi-robot and warehouse coordination — Logistics/Industry
- Use case: Agents assign and adapt manipulation tasks across multiple arms or mobile manipulators, coordinating observation, grasping, timed release, packing, and exception handling.
- Potential product: A fleet-level embodied-agent scheduler that shares task memories and reallocates work after failures.
- Why long term: The paper demonstrates concurrent motion for two arms but not large-scale multi-robot coordination, shared-space collision avoidance, or operational scheduling.
- Dependencies: Real-time communication, centralized safety supervision, shared world models, resource allocation, deterministic collision avoidance, and integration with warehouse control systems.
- Autonomous scientific and engineering experimentation — Academia/R&D
- Use case: Agents perform repeatable physical experiments involving object arrangement, assembly, manipulation, and iterative adjustment based on measured outcomes.
- Potential workflow: A researcher specifies an experimental protocol; the agent prepares the task, executes trials, records observations, modifies parameters, and produces a structured report.
- Dependencies: Experimental validity requires calibrated instruments, strict protocol compliance, reproducible randomization, provenance of all actions, and human approval for changes. Language-model-generated procedures cannot replace domain-specific experimental safeguards.
- Regulatory frameworks for agent-controlled physical systems — Policy
- Use case: Establish requirements for certification, audit logs, human override, model updates, incident reporting, and operational limits for robots controlled by general-purpose agents.
- Potential policy outputs: Risk-tiered approval schemes, mandatory simulation and physical test suites, minimum intervention latency, traceability of generated programs, and restrictions on autonomous operation in public or clinical settings.
- Dependencies: Standards must account for changing model behavior, external API updates, prompt or memory contamination, cybersecurity threats, and differences between low-risk tabletop tasks and high-energy industrial systems.
- Everyday household robots capable of adapting to demonstrations — Daily life
- Use case: A household robot learns tasks from videos or demonstrations, such as sorting laundry, arranging groceries, folding clothes, or preparing simple items.
- Why long term: Household environments are much less structured than the paper’s tabletop scenes, and success requires safe interaction with humans, varied objects, narrow spaces, and continuously changing conditions.
- Dependencies: Low-cost robust hardware, tactile sensing, real-time inference, privacy protections, reliable uncertainty estimation, affordable maintenance, and highly dependable recovery behavior. The paper’s execution times and inference costs currently make widespread consumer deployment impractical.
Glossary
- Action diversity: Variety of distinct physical actions a system can perform. “Dice flipping and targeted throwing follow language instructions and test action diversity”
- Agent as Policy (AGP): An approach in which a general-purpose agent directly performs the role of the robot’s control policy. “We introduce Agent as Policy (AGP), which places task planning and execution under the agent's control.”
- Affordances: The action possibilities that an environment or object provides to an agent. “A second line selects and sequences learned skills according to task goals, affordances, and environment feedback”
- Bimanual manipulation: Robotic manipulation performed using two arms or hands. “Bimanual towel folding uses two arms with a common coordinate frame”
- Camera calibration: The process of estimating a camera’s imaging parameters and its geometric relationship to a robot or scene. “The bridge grounds these computations in camera calibration and robot coordinates”
- Code generation: Automatic creation of executable programs from a high-level specification or instruction. “Code Generation for Free-Form Manipulation Tasks across Real and Simulation”
- Computation graph: A graph-structured representation of operations and their dependencies. “expressed as code or computation graphs that the robot runs during task execution”
- Deformable object: An object whose shape can change substantially during manipulation. “Bimanual towel folding follows video instructions and tests object diversity through coordinated manipulation of deformable material.”
- Diffusion model: A generative model that produces data by progressively transforming noise into a structured sample. “with diffusion models capturing action sequences”
- Dynamic manipulation: Manipulation involving motion, timing, momentum, or changing physical states. “We study AGP across multiple real-world manipulation tasks spanning precision manipulation, dynamic motions, and deformable objects.”
- End-effector pose: The position and orientation of a robot’s tool or terminal component. “the corresponding end effector pose computed through forward kinematics”
- Experience accumulation: Improvement based on storing and reusing information from previous task executions. “We further study efficiency through task experience accumulation”
- Forward kinematics: Computing a robot end effector’s pose from its joint positions. “the corresponding end effector pose computed through forward kinematics”
- Foundation model: A broadly pretrained model that can be adapted to many tasks rather than trained for one specific application. “general purpose foundation MLLMs can serve as policies for direct robot arm control”
- Geometric modeling: Representing the shapes, positions, and spatial relationships of objects and robot components mathematically. “Classical robot manipulation combines geometric modeling, task and motion planning, and feedback control”
- Gripper: A robotic end effector designed to grasp or hold objects. “a parallel gripper in a tabletop workspace”
- Inference cost: The computational or financial resources required for a model to generate outputs. “The substantial execution time and inference cost remain barriers to practical deployment.”
- Inverse kinematics: Computing robot joint configurations required to achieve a desired end-effector pose. “For end effector targets, it uses Mink~\citep{zakka2026mink} for inverse kinematics to compute joint targets.”
- Joint trajectory: A time-ordered sequence of positions or configurations for a robot’s joints. “It then generates joint trajectories subject to workspace and motion limits”
- Long-horizon reasoning: Planning and decision-making over many sequential actions or an extended task duration. “We evaluate AGP on real-world robot tasks spanning long horizon reasoning, action diversity, and object diversity”
- Multimodal LLM (MLLM): A LLM capable of processing and reasoning over multiple modalities, such as text, images, and video. “Multimodal LLMs (MLLMs) demonstrate strong general capabilities”
- Orchestration: Selecting, sequencing, and coordinating multiple skills, policies, or processes to accomplish a task. “Runtime policy orchestration.”
- Persistent memory: Information retained across separate interactions or task executions. “Persistent memory carries execution feedback across trials”
- Policy orchestration: The selection and sequencing of learned policies according to task requirements and feedback. “The second approach uses agents as orchestrators of existing learned policies”
- Precision manipulation: Robot manipulation requiring highly accurate positioning, alignment, or force application. “We study AGP across multiple real-world manipulation tasks spanning precision manipulation”
- Proprioceptive state: Information about a robot’s own configuration and motion obtained from internal sensors. “This proprioceptive state describes the robot's current configuration”
- Recovery: Adaptation or corrective action taken after an error, failure, or unexpected physical outcome. “The agent determines grasp poses, action sequences, observation timing, and recovery during this cycle.”
- Revolute joint: A robot joint that permits rotation around a fixed axis. “We use an I2RT YAM~\citep{i2rt_yam} arm with six revolute joints”
- Robot policy: A rule, model, or agent that maps task information and observations to robot actions. “We use a general purpose agent as the robot policy.”
- Spatiotemporal: Relating to both spatial structure and changes over time. “Spatio-Temporal Reasoning of Relational Keypoint Constraints for Robotic Manipulation”
- Task and motion planning: Coordinated planning of high-level task steps and the robot motions needed to execute them. “Classical robot manipulation combines geometric modeling, task and motion planning, and feedback control”
- Tool call: A structured request from an agent to obtain information or execute an operation through an external interface. “A tool call may acquire information or execute a program that submits action targets to the robot controller.”
- Trajectory: A time-dependent path or sequence of states followed by a moving object or robot. “subsequent trajectory analysis and verification account for approximately 40\% of the reported 22.8 minutes.”
- Visuomotor policy: A policy that maps visual observations to motor actions. “Learned visuomotor policies predict actions from robot trajectory data”
- Vision-language-action (VLA) model: A model that connects visual and linguistic inputs to physical actions. “Vision language action models extend this approach to diverse tasks through image and language conditioning”
- Workspace: The region of physical space that a robot can reach or operate within. “It then generates joint trajectories subject to workspace and motion limits”
- Zero-shot manipulation: Performing a manipulation task without task-specific training examples or prior task execution. “These results support its effectiveness for zero shot manipulation across the evaluated tasks.”









