Papers
Topics
Authors
Recent
Search
2000 character limit reached

Foundation Models in Robotics: A Comprehensive Review of Methods, Models, Datasets, Challenges and Future Research Directions

Published 16 Apr 2026 in cs.RO | (2604.15395v1)

Abstract: Over the recent years, the field of robotics has been undergoing a transformative paradigm shift from fixed, single-task, domain-specific solutions towards adaptive, multi-function, general-purpose agents, capable of operating in complex, open-world, and dynamic environments. This tremendous advancement is primarily driven by the emergence of Foundation Models (FMs), i.e., large-scale neural-network architectures trained on massive, heterogeneous datasets that provide unprecedented capabilities in multi-modal understanding and reasoning, long-horizon planning, and cross-embodiment generalization. In this context, the current study provides a holistic, systematic, and in-depth review of the research landscape of FMs in robotics. In particular, the evolution of the field is initially delineated through five distinct research phases, spanning from the early incorporation of NLP and Computer Vision (CV) models to the current frontier of multi-sensory generalization and real-world deployment. Subsequently, a highly-granular taxonomic investigation of the literature is performed, examining the following key aspects: a) the employed FM types, including LLMs, VFMs, VLMs, and VLAs, b) the underlying neural-network architectures, c) the adopted learning paradigms, d) the different learning stages of knowledge incorporation, e) the major robotic tasks, and f) the main real-world application domains. For each aspect, comparative analysis and critical insights are provided. Moreover, a report on the publicly available datasets used for model training and evaluation across the considered robotic tasks is included. Furthermore, a hierarchical discussion on the current open challenges and promising future research directions in the field is incorporated.

Summary

  • The paper presents a six-criteria taxonomy of 435 studies, comparing foundation model types, architectures, learning methods, training stages, robotic tasks, and application domains.
  • Foundation models are advancing from modular perception and planning tools toward multimodal vision-language-action policies, but deployment remains limited by data demands, latency, embodiment gaps, and safety risks.
  • The review identifies tactile and failure-recovery data shortages as critical gaps and recommends physics-informed world models, embodiment-agnostic actions, adaptive safety, and runtime verification.

Scope and methodology of the survey

This paper presents a systematic review of the use of Foundation Models (FMs) in robotics, covering 435 papers selected through a structured screening process across six databases (IEEE Xplore, Google Scholar, Scopus, DBLP, arXiv, and Web of Science), with explicit inclusion/exclusion criteria and iterative refinement. The review is organized along six taxonomic criteria: foundation model type (LLMs, VFMs, VLMs, VLAs), neural network architecture (transformers, state-space models, diffusion models, convolutional/hybrid encoders, graphical models), learning paradigm (pre-training, SSL, fine-tuning, domain adaptation, IL, RL, in-context/prompt learning, world model learning, generative learning), learning stage (pre-training, offline fine-tuning, online adaptation, continuous learning), robotic task (perception, planning, navigation, manipulation, HRI), and application domain (agentic mobility, industrial manipulation, supply operations, service robots, medical robots, agrisystems, crisis agents, maritime robotics, space robotics). For each criterion, the authors provide comparative tables summarizing functions, inputs/outputs, strengths, limitations, and representative models.

The survey distinguishes itself from prior reviews (e.g., (Hu et al., 2023, Stokke et al., 2024, Team et al., 25 Mar 2025)-adjacent VLA surveys) by its breadth across all six criteria simultaneously; prior surveys typically cover one or two criteria or focus on a single task such as manipulation. The comparative positioning against ten earlier surveys is presented explicitly, with each prior work's scope (64–785 papers analyzed) and its stated gaps—most commonly the absence of systematic multi-criteria comparison and of application-domain discussion.

Five-phase evolution of the field

The review delineates five research phases. Phase 1 (2018–2021) used off-the-shelf NLP/CV models (BERT, ViT) in modular pipelines with hand-engineered controllers. Phase 2 (2021–2022) introduced grounded planning via vision-language representations (CLIP, SayCan), though training data remained largely non-embodied proxy sources. Phase 3 (2022–2023) marked the emergence of end-to-end embodied VLA policies trained on large-scale robot demonstration data (RT-1 with over 130,000 real-world trajectories; Gato supporting more than 600 tasks; PaLM-E at 562B parameters). Phase 4 (2023–2024) added memory, autonomous task composition, and web-to-robot transfer (RT-2, OpenVLA, RT-X spanning 527 skills on 22 platforms). Phase 5 (2024–present) targets multi-sensory generalization and real-world deployment, incorporating touch, force, and proprioception alongside vision and language (GR00T N1 at 2B parameters, Gemini Robotics).

A notable observation here is that the trajectory is consistently toward integration: from isolated perception modules to unified policies that map multimodal inputs directly to actions. The implication is that architectural modularity—the dominant paradigm before FMs—is being displaced by monolithic generalist models, with attendant consequences for interpretability that the survey addresses later.

Foundation model types

LLMs function primarily as high-level planners and reasoning engines, translating free-form instructions into goals, code, or plans (SayCan, Code-as-Policies, AutoTAMP). Their documented strengths are robust instruction translation and task decomposition; their limitations include lack of physical grounding, hallucination, latency unsuitable for real-time control, and input sensitivity. The survey notes that LLM-generated plans may be semantically valid but physically infeasible—a recurring theme throughout the taxonomy.

VFMs (SAM, DINOv2, Metric3D v2) provide transferable visual features for recognition, localization, tracking, depth estimation, semantic mapping, and visual-inertial fusion. They exhibit strong zero-shot generalization but suffer from domain specificity, incomplete modeling of physical dynamics, and computational cost.

VLMs bridge language and vision for grounding, open-vocabulary recognition, semantic mapping, and execution-time verification (OmniManip, VoxPoser, Guardian). A key limitation emphasized is that VLMs cannot intrinsically produce precise motor commands and depend on external policy generators.

VLAs (RT-1/RT-2, OpenVLA, GR00T N1, π₀.₅) constitute end-to-end policies mapping multimodal inputs directly to actions. Their advantages are operational simplicity and cross-platform generalization; their limitations are extreme data requirements, inference latency, embodiment-specific performance degradation, and difficulty in failure management. The survey's comparative analysis concludes that VLAs represent the current endpoint of the integration trend, subsuming the roles of the other three types within a single architecture.

Architectures, learning paradigms, and stages

On architectures, transformers dominate for multimodal alignment and reasoning but carry quadratic attention complexity that conflicts with real-time control loops operating above 50 Hz. SSMs (Mamba-based variants such as RoboMamba) offer linear scaling and are positioned as a practical alternative for embedded deployment, though they sacrifice direct cross-token attention and often require hybrid designs. Diffusion models excel at modeling multimodal action distributions but incur iterative sampling latency and lack built-in safety guarantees. Graphical models (scene graphs, GNNs) provide relational structure and explainability but face computational overhead and dynamic-topology challenges.

On learning paradigms, the survey documents concrete scale figures: OpenVLA trains a 7B model on approximately 970K real robot episodes; Octo uses roughly 800K trajectories. Eureka demonstrates LLM-written reward code outperforming expert-designed rewards, and DrEureka extends this to automated sim-to-real pipelines. On learning stages, the analysis identifies a persistent tension: pre-training yields broad generalization but an "embodiment gap"; offline fine-tuning risks catastrophic forgetting and distribution shift; online adaptation and continuous learning promise resilience but remain constrained by stability-plasticity trade-offs and memory overhead.

Datasets and identified gaps

The dataset review covers benchmarks grouped by task, including Open X-Embodiment (1M+ episodes, 60 datasets, 22 platforms), AgiBot World (over 1M trajectories from 100 humanoid robots), DROID (76K trajectories collected by 50 people across 52 buildings), Ego4D (3,670 hours of egocentric video), and RoboVQA (829,502 video-text pairs). Two deficits are highlighted as consequential: tactile and force-sensing data remain critically underrepresented despite being essential for dexterous manipulation, and failure/recovery episodes are almost entirely absent from current benchmarks, limiting the ability to train robust recovery behaviors.

Challenges and open questions

The challenges section identifies bottlenecks across data (scarcity of physical-world trajectories, embodiment heterogeneity without a standardized continuous-action format, domain transfer gap), computation (inference latencies of hundreds of milliseconds to seconds versus control loops above 50 Hz, constrained onboard resources), safety (semantic-physical mismatch producing linguistically sound but dangerous plans, adversarial vulnerability of cloud-hosted models, unreliable uncertainty quantification), embodiment (action-space heterogeneity, sim-to-real gap, limited spatial grounding, haptic sensing as the "final frontier"), reasoning (lack of physical common sense, exponential degradation of long-horizon planning, imprecise behavioral explanations), and evaluation (no unified framework; binary success metrics too coarse to diagnose failures; minimal distribution shifts causing large behavioral changes).

The future directions section proposes embodiment-agnostic action spaces, improved continuous-action tokenization, diffusion/flow-based action modeling, tactile and auditory integration, long-horizon memory frameworks, physics-informed world foundation models, adaptive safety integrated into training loops rather than applied post hoc, and runtime formal verification via control barrier functions and neuro-symbolic specification mining. These are framed as responses to the enumerated challenges rather than independent predictions.

Limitations of the review itself

The survey's claims rest on its selection criteria, which prioritize prominent venues and recent work; the resulting corpus may underrepresent negative results and industrial deployments not published in academic venues. The comparative analyses are qualitative syntheses rather than controlled empirical comparisons—no benchmark results are re-run or normalized across methods—and several quantitative figures cited (e.g., parameter counts, dataset sizes) are taken from the original papers without independent verification. The rapid pace of the field means specific model characterizations may date quickly, even if the structural taxonomy remains applicable.

Conclusion

This survey provides a systematic, six-criteria taxonomy of foundation models in robotics, tracing the field's evolution from modular NLP/CV augmentation to end-to-end multimodal generalist policies. Its principal contributions are the phase-based historical account, per-criterion comparative analyses with explicit strengths and limitations, a consolidated dataset inventory identifying tactile-data and failure-data gaps, and a challenge-to-direction mapping covering data, computation, safety, embodiment, reasoning, and evaluation. The central unresolved tension it documents is between internet-scale pre-training's semantic generality and the physical grounding, latency, and safety constraints of real-world deployment—an imbalance the authors argue will require hardware-agnostic architectures, richer sensory modalities, physics-informed world models, and runtime verification to resolve.

Paper to Video (Beta)

No one has generated a video about this paper yet.

Whiteboard

No one has generated a whiteboard explanation for this paper yet.

Open Problems

We haven't generated a list of open problems mentioned in this paper yet.

Tweets

Sign up for free to view the 2 tweets with 0 likes about this paper.