---
title: 'OmniEVA: Versatile Embodied Planner'
url: https://www.emergentmind.com/topics/omnieva
type: topic
---

# OmniEVA: Versatile Embodied Planner

Searching arXiv for the OmniEVA paper and closely related embodied-planning context.
OmniEVA is an embodied versatile planner for multimodal large language model–based embodied intelligence that is designed to support multimodal understanding, reasoning, interaction, and continuous spatial decision-making through two central mechanisms: Task-Adaptive 3D Grounding and Embodiment-Aware Reasoning [2509.09332]. It is proposed to address two limitations identified in prior embodied systems: the Geometric Adaptability Gap, arising when models are trained solely on 2D inputs or rely on hard-coded 3D geometry injection, and the Embodiment Constraint Gap, arising when plans are semantically valid but physically infeasible for real robots. Within the formulation reported for OmniEVA, these limitations are addressed by selective regulation of 3D fusion based on contextual requirements and by incorporating task goals together with embodiment constraints into the reasoning loop, yielding planning decisions intended to be both goal-directed and executable [2509.09332].

## 1. Problem setting and design rationale

OmniEVA is situated in the context of embodied intelligence, where an agent must reason, plan, and act in the physical world. The motivating claim is that recent multimodal large language models have expanded the scope of embodied systems, but current MLLM-based approaches remain limited by insufficient adaptation to differing spatial demands and by inadequate treatment of robot embodiment constraints [2509.09332].

The Geometric Adaptability Gap is defined in the source material as a failure mode in which models trained only on 2D inputs lack sufficient spatial information, while models with hard-coded 3D geometry injection have restricted 2D generalization. The consequence described is poor adaptability across tasks with diverse spatial requirements. The Embodiment Constraint Gap is defined as a failure mode in which the physical constraints and capacities of real robots are neglected, so task plans may be theoretically valid while remaining practically infeasible [2509.09332].

OmniEVA’s response to these gaps is organized around two innovations. The first is Task-Adaptive 3D Grounding, which uses a gated router for explicit selective regulation of 3D fusion according to contextual requirements. The second is an Embodiment-Aware Reasoning framework that introduces embodiment constraints directly into the reasoning and planning process. Taken together, these components are presented as enabling robust embodied reasoning and planning across both abstract reasoning tasks and tasks requiring real-world execution [2509.09332].

## 2. System architecture and input structure

At the architectural level, OmniEVA builds on standard MLLM components. The reported high-level system includes a Vision Transformer encoder, denoted \(\mathcal{E}_{\text{img}}\), which encodes RGB images into visual tokens, and an autoregressive language model backbone that supports cross-modal communication, optionally through a lightweight mapping network. The input interface includes natural language instruction \(T\), a sequence of images or video frames \((I_1, \dots, I_N)\), and optionally depth maps \((D_1, \dots, D_N)\) for 3D perception, together with camera intrinsics \(K\) and extrinsics \(M_i\) for 3D projection [2509.09332].

The architectural novelty is concentrated in what the paper describes as the Task-Adaptive Gated Router, or TAGR, and a two-stage embodiment-aware training scheme. This arrangement preserves the conventional image-language backbone structure while altering how 3D information is injected and how downstream plans are optimized [2509.09332].

A central implication of this design is that 3D information is not treated as uniformly necessary. The reported system instead allows 3D grounding to be activated or suppressed based on task and scene context. This suggests a hybrid visual processing regime in which the model can operate as a 2D system when geometry is unnecessary and as a 2D+3D system when geometric structure is essential.

## 3. Task-Adaptive 3D Grounding

Task-Adaptive 3D Grounding is motivated by the contrast between two prior strategies: omitting 3D information entirely and statically injecting 3D information into every case. The former is associated with poor geometric generalization, while the latter is described as inefficient or noisy when geometry is irrelevant or unreliable. OmniEVA introduces a dynamic, context-aware alternative [2509.09332].

The mechanism begins with patch-level 3D positional encoding. Each depth image \(D_i\) is projected into 3D world coordinates \(P_i\) using the camera intrinsics and extrinsics. Matching the ViT patch size, average 3D coordinates are computed for each patch to obtain \(P'_i\). A sinusoidal encoding then produces per-patch 3D positional embeddings, denoted \(V^p\) [2509.09332].

Dynamic 3D injection is controlled by the Task-Adaptive Gated Router. The task condition is represented by encoding the instruction into a latent vector \(V^T\) with a sentence transformer. The scene condition is formed by aggregating the ViT patch features \(V^I\), for example through global average pooling, to produce \(V^I_{\text{avg}}\). These are concatenated and processed by an MLP to yield gate logits:
\[
V^g = \mathrm{MLP}_{\psi}(\mathrm{Concat}[V^T, V^I_{\text{avg}}]) \in \mathbb{R}^2
\]
A Gumbel-Softmax then produces a differentiable binary decision \(g \in \{0,1\}\). When \(g = 1\), the model augments the 2D vision tokens with 3D positional embeddings \((V^I + V^p)\); when \(g = 0\), it retains vanilla 2D features \(V^I\). The resulting hybrid visual token is defined as
\[
V^{I}_{\text{hybrid}} = V^I + g \cdot V^{p} = (1-g)V^I + g(V^I + V^p)
\]
and the output is generated as
\[
o = \mathcal{F}_\theta^{\text{llm}}(T, V^{I}_{\text{hybrid}})
\]
[2509.09332]

The paper notes that this formalism resembles a Mixture-of-Experts configuration in which the effective experts are a 2D-only branch and a 2D+3D branch, with routing determined by task and scene context. The gate is pretrained on depth-aware datasets and optimized with cross-entropy loss for the main task together with a KL regularization term to prevent collapse, using a Bernoulli\((0.5)\) prior for unbiased exploration:
\[
\mathcal{L}^{\text{total}} = \mathcal{L}^{\text{CE}} + \alpha \cdot \mathcal{L}^{\text{KL}}
\]
The TAGR parameters are then frozen after pre-training [2509.09332].

The reported gate behavior is tied to linguistic and perceptual context. Quantitative and qualitative analyses indicate that gate activation correlates with spatial or geometric words such as “shape,” “rectangular,” and “go.” The example given is that “how many” produces low gate activation, whereas “shape of the table” produces high activation. This is presented as evidence that 3D information is activated when required for success and suppressed when 2D information is sufficient [2509.09332].

## 4. Embodiment-Aware Reasoning and optimization

The second major component is Embodiment-Aware Reasoning, introduced to address the tendency of prior systems to plan semantically while neglecting affordances, workspace limits, kinematic constraints, and collision constraints. The specific failure mode described is the generation of plans that appear acceptable in simulation or at a semantic level but fail under real robotic execution [2509.09332].

OmniEVA addresses this with a two-stage cascaded training procedure. The first stage is omni-supervised finetuning. In this stage, the model is trained on a mixture of general embodied reasoning tasks spanning spatial, temporal, and

Source: https://www.emergentmind.com/topics/omnieva