Papers
Topics
Authors
Recent
Search
2000 character limit reached

EponaV2: Driving World Model with Comprehensive Future Reasoning

Published 14 May 2026 in cs.CV | (2605.14696v1)

Abstract: Data scaling plays a pivotal role in the pursuit of general intelligence. However, the prevailing perception-planning paradigm in autonomous driving relies heavily on expensive manual annotations to supervise trajectory planning, which severely limits its scalability. Conversely, although existing perception-free driving world models achieve impressive driving performance, their real-world reasoning ability for planning is solely built on next frame image forecasting. Due to the lack of enough supervision, these models often struggle with comprehensive scene understanding, resulting in unsatisfactory trajectory planning. In this paper, we propose EponaV2, a novel paradigm of driving world models, which achieves high-quality planning with comprehensive future reasoning. Inspired by how human drivers anticipate 3D geometry and semantics, we train our model to forecast more comprehensive future representations, which can be additionally decoded to future geometry and semantic maps. Extracting the 3D and semantic modalities enables our model to deeply understand the surrounding environment, and the future prediction task significantly enhances the real-world reasoning capabilities of EponaV2, ultimately leading to improved trajectory planning. Moreover, inspired by the training recipe of LLMs, we introduce a flow matching group relative policy optimization mechanism to further improve planning accuracy. The state-of-the-art (SOTA) performances of EponaV2 among perception-free models on three NAVSIM benchmarks (+1.3PDMS, +5.5EPDMS) demonstrate the effectiveness of our methods.

Summary

  • The paper introduces a 6.7B-parameter perception-free driving world model that forecasts future metric depth and SAM3 semantic features alongside image representations, replacing costly manual annotations with self-supervised foundation-model pseudo-labels.
  • Future structured predictions improve planning more than next-frame image forecasting alone: ablations show a 2.5-point PDMS gain from combined supervision, while future-depth prediction outperforms present-depth prediction by 1.6 points.
  • Flow-matching GRPO fine-tuning adds 1.0 PDMS, helping EponaV2 reach 90.4 PDMS on NAVSIMv1 and achieve a 5.5-point EPDMS gain over prior perception-free methods on NAVSIMv2 navhard.

EponaV2 is a perception-free driving world model that replaces manual perception annotations with self-supervised forecasting of future depth and semantic representations, combined with reinforcement learning fine-tuning of its trajectory planner. The paper's central claim is that next-frame image prediction—the dominant supervision signal in prior perception-free world models—is insufficient for planning, because the planning-relevant information in raw images is entangled and difficult to exploit. By instead supervising future metric depth and future SAM3-derived semantic features, EponaV2 achieves state-of-the-art results among perception-free models on three NAVSIM benchmarks, including a +1.3 PDMS gain on NAVSIMv1 and a +5.5 EPDMS gain on the navhard split of NAVSIMv2.

Motivation and positioning

The paper frames its contribution around data scalability. Perception-planning models—BEV-based systems such as UniAD, VAD, DiffusionDrive, GoalFlow, and VLA-based approaches such as AutoVLA and ReCogDrive—depend on expensive annotations (bounding boxes, segmentations) for auxiliary tasks like occupancy prediction or VQA, which limits the applicability of scaling laws. Perception-free alternatives (DrivingGPT, World4Drive, Epona, DriveVLA-W0, PWM, DriveLaW) remove this dependency but supervise future reasoning only through next-frame image prediction. The authors argue this yields weak scene understanding: inferred future representations inherit the ambiguity of raw pixels rather than encoding actionable geometric and semantic structure. EponaV2's remedy is to forecast decodable future modalities—metric depth and text-aligned semantic features—using pseudo-labels from foundation models (Depth-Anything-V3, SAM3), thereby obtaining perception-like supervision without human annotation.

Architecture

EponaV2 is a 6.7B-parameter model built on the language component of Qwen3-VL 4B as its world model backbone. Input consists of NN front-view frames at 2 Hz (512×1024512\times1024 resolution), encoded by a frozen DINO-Tok tokenizer, together with relative ego movements. A causal attention mask ensures autoregressive inference of future representations {Fi′,ΔAi′}\{F_i', \Delta A_i'\}. Two rectified flow matching heads decode these representations:

  • Trajectory planner: a velocity field vtrajv_{traj} trained to regress the direction from ground-truth trajectories to Gaussian noise, sampled via an ODE integration. Rectified flow converges faster than diffusion alternatives.
  • Future frame predictor: a head vimgv_{img} conditioned on the predicted controlling movement to reconstruct the next encoded frame, providing coarse world-model supervision inherited from Epona.

Comprehensive future reasoning

Two auxiliary heads extend the supervision beyond raw image forecasting:

Future depth prediction. A lightweight decoder predicts next-frame metric depth and a confidence map from the inferred future representation, supervised by Depth-Anything-V3 pseudo-labels scaled to canonical focal length following Metric3D. The loss combines confidence-weighted L1, a log-confidence regularizer, and gradient-matching terms. Notably, the decoding uses the planner-predicted movement ΔA^i+1\Delta\hat{A}_{i+1} during inference, so no extra forward pass through an external depth network is needed at test time—an explicit contrast with approaches that feed pretrained depth features as inputs, which add latency and do not guarantee internalization of the geometric signal.

Future semantic feature prediction. Rather than supervising current-frame segmentation (as in World4Drive), the model predicts next-frame text-fused SAM3 features for prompts "car" and "human", matched with an L2 loss. Text conditioning directs model attention toward driving-critical object categories. During training, both heads use ground-truth movements for faster convergence.

An ablation at downscaled scale (3.5B parameters, Qwen3-VL 2B backbone) supports the design: adding LimgL_{img} alone raises PDMS only marginally (84.4 → 84.8), while depth prediction improves NC, DAC, and TTC (85.2), and combining all losses reaches 86.9 PDMS—a +2.5 gain over trajectory supervision alone. A further ablation shows that predicting future depth outperforms predicting present depth by 1.6 PDMS (85.2 vs. 83.6), directly substantiating the paper's core argument that forecasting, not merely perceiving, drives planning quality.

Flow matching GRPO

Inspired by LLM post-training recipes, the second training stage applies GRPO to the flow matching planner following Flow-GRPO: the deterministic ODE sampler is reformulated as an SDE whose per-step transitions form Gaussian policies, enabling likelihood-based policy gradients over the denoising chain. Group-standardized advantages are computed from rewards defined as the discrepancy between the ground-truth trajectory and the trajectory obtained by executing the generated plan in simulation. PDMS itself cannot serve as reward since it requires manual labels—a constraint the authors state explicitly. An imitation loss regularizes the policy against dataset trajectories, yielding Lgrpo=Lrl+λilLilL_{grpo}=L_{rl}+\lambda_{il}L_{il}. This stage contributes +1.0 PDMS (89.4 → 90.4), concentrated in drivable area compliance (+0.6) and ego progress (+1.2).

Experimental results

Training uses nuPlan and nuScenes clips in stage one and the NAVSIM navtrain split in stage two, across 64 H20 GPUs for roughly 10 days; evaluation runs on a single RTX 4090.

Benchmark Metric Best prior perception-free EponaV2 Gain
NAVSIMv1 navtest PDMS DriveLaW, 89.1 90.4 +1.3
NAVSIMv2 navtest EPDMS DriveLaW, 88.6 88.9 +0.3
NAVSIMv2 navhard EPDMS DriveLaW, 30.6 36.1 +5.5

On NAVSIMv1, EponaV2 attains DAC of 97.9 and EP of 84.8, exceeding all listed perception-free baselines despite using only a single front-view camera. On the two-stage pseudo closed-loop NAVSIMv2 evaluation, gains are largest on the harder navhard split, where stage-two metrics show improvements in DAC (78.0 vs. 67.6 for DriveLaW) and DDC (88.0 vs. 83.5). The authors classify DriveVLA-W0 (Anchor) as perception-based because it uses manual labels for anchor selection, so comparisons are restricted accordingly.

Limitations

The paper concedes that EponaV2 does not surpass the strongest perception-based models, attributing this to the imprecision of pseudo depth and semantic labels produced by the foundation models—an inherent ceiling of the annotation-free approach. Additional unstated dependencies include the restriction to two fixed text prompts ("car", "human") for semantic supervision, reliance on front-view-only input, and the assumption that simulated-trajectory discrepancy is a valid proxy reward in place of PDMS. Whether richer prompt sets, multi-view inputs, or improved foundation-model labelers close the remaining gap to perception-based methods remains open.

Conclusion

EponaV2 demonstrates that perception-free driving world models can substantially improve planning by forecasting structured future modalities—metric depth and foundation-model semantic features—rather than raw frames alone, and that flow matching GRPO provides a further, label-free refinement of trajectory accuracy. The reported gains, particularly the +5.5 EPDMS improvement on navhard, indicate that comprehensive future reasoning is a productive supervision signal for scalable end-to-end driving, while the acknowledged gap to perception-based SOTA leaves open how far pseudo-label quality can be pushed before manual annotation becomes dispensable.

Paper to Video (Beta)

No one has generated a video about this paper yet.

Whiteboard

No one has generated a whiteboard explanation for this paper yet.

Open Problems

We haven't generated a list of open problems mentioned in this paper yet.