- The paper introduces a 6.7B-parameter perception-free driving world model that forecasts future metric depth and SAM3 semantic features alongside image representations, replacing costly manual annotations with self-supervised foundation-model pseudo-labels.
- Future structured predictions improve planning more than next-frame image forecasting alone: ablations show a 2.5-point PDMS gain from combined supervision, while future-depth prediction outperforms present-depth prediction by 1.6 points.
- Flow-matching GRPO fine-tuning adds 1.0 PDMS, helping EponaV2 reach 90.4 PDMS on NAVSIMv1 and achieve a 5.5-point EPDMS gain over prior perception-free methods on NAVSIMv2 navhard.
EponaV2 is a perception-free driving world model that replaces manual perception annotations with self-supervised forecasting of future depth and semantic representations, combined with reinforcement learning fine-tuning of its trajectory planner. The paper's central claim is that next-frame image prediction—the dominant supervision signal in prior perception-free world models—is insufficient for planning, because the planning-relevant information in raw images is entangled and difficult to exploit. By instead supervising future metric depth and future SAM3-derived semantic features, EponaV2 achieves state-of-the-art results among perception-free models on three NAVSIM benchmarks, including a +1.3 PDMS gain on NAVSIMv1 and a +5.5 EPDMS gain on the navhard split of NAVSIMv2.
Motivation and positioning
The paper frames its contribution around data scalability. Perception-planning models—BEV-based systems such as UniAD, VAD, DiffusionDrive, GoalFlow, and VLA-based approaches such as AutoVLA and ReCogDrive—depend on expensive annotations (bounding boxes, segmentations) for auxiliary tasks like occupancy prediction or VQA, which limits the applicability of scaling laws. Perception-free alternatives (DrivingGPT, World4Drive, Epona, DriveVLA-W0, PWM, DriveLaW) remove this dependency but supervise future reasoning only through next-frame image prediction. The authors argue this yields weak scene understanding: inferred future representations inherit the ambiguity of raw pixels rather than encoding actionable geometric and semantic structure. EponaV2's remedy is to forecast decodable future modalities—metric depth and text-aligned semantic features—using pseudo-labels from foundation models (Depth-Anything-V3, SAM3), thereby obtaining perception-like supervision without human annotation.
Architecture
EponaV2 is a 6.7B-parameter model built on the language component of Qwen3-VL 4B as its world model backbone. Input consists of N front-view frames at 2 Hz (512×1024 resolution), encoded by a frozen DINO-Tok tokenizer, together with relative ego movements. A causal attention mask ensures autoregressive inference of future representations {Fi′​,ΔAi′​}. Two rectified flow matching heads decode these representations:
- Trajectory planner: a velocity field vtraj​ trained to regress the direction from ground-truth trajectories to Gaussian noise, sampled via an ODE integration. Rectified flow converges faster than diffusion alternatives.
- Future frame predictor: a head vimg​ conditioned on the predicted controlling movement to reconstruct the next encoded frame, providing coarse world-model supervision inherited from Epona.
Comprehensive future reasoning
Two auxiliary heads extend the supervision beyond raw image forecasting:
Future depth prediction. A lightweight decoder predicts next-frame metric depth and a confidence map from the inferred future representation, supervised by Depth-Anything-V3 pseudo-labels scaled to canonical focal length following Metric3D. The loss combines confidence-weighted L1, a log-confidence regularizer, and gradient-matching terms. Notably, the decoding uses the planner-predicted movement ΔA^i+1​ during inference, so no extra forward pass through an external depth network is needed at test time—an explicit contrast with approaches that feed pretrained depth features as inputs, which add latency and do not guarantee internalization of the geometric signal.
Future semantic feature prediction. Rather than supervising current-frame segmentation (as in World4Drive), the model predicts next-frame text-fused SAM3 features for prompts "car" and "human", matched with an L2 loss. Text conditioning directs model attention toward driving-critical object categories. During training, both heads use ground-truth movements for faster convergence.
An ablation at downscaled scale (3.5B parameters, Qwen3-VL 2B backbone) supports the design: adding Limg​ alone raises PDMS only marginally (84.4 → 84.8), while depth prediction improves NC, DAC, and TTC (85.2), and combining all losses reaches 86.9 PDMS—a +2.5 gain over trajectory supervision alone. A further ablation shows that predicting future depth outperforms predicting present depth by 1.6 PDMS (85.2 vs. 83.6), directly substantiating the paper's core argument that forecasting, not merely perceiving, drives planning quality.
Flow matching GRPO
Inspired by LLM post-training recipes, the second training stage applies GRPO to the flow matching planner following Flow-GRPO: the deterministic ODE sampler is reformulated as an SDE whose per-step transitions form Gaussian policies, enabling likelihood-based policy gradients over the denoising chain. Group-standardized advantages are computed from rewards defined as the discrepancy between the ground-truth trajectory and the trajectory obtained by executing the generated plan in simulation. PDMS itself cannot serve as reward since it requires manual labels—a constraint the authors state explicitly. An imitation loss regularizes the policy against dataset trajectories, yielding Lgrpo​=Lrl​+λil​Lil​. This stage contributes +1.0 PDMS (89.4 → 90.4), concentrated in drivable area compliance (+0.6) and ego progress (+1.2).
Experimental results
Training uses nuPlan and nuScenes clips in stage one and the NAVSIM navtrain split in stage two, across 64 H20 GPUs for roughly 10 days; evaluation runs on a single RTX 4090.
| Benchmark |
Metric |
Best prior perception-free |
EponaV2 |
Gain |
| NAVSIMv1 navtest |
PDMS |
DriveLaW, 89.1 |
90.4 |
+1.3 |
| NAVSIMv2 navtest |
EPDMS |
DriveLaW, 88.6 |
88.9 |
+0.3 |
| NAVSIMv2 navhard |
EPDMS |
DriveLaW, 30.6 |
36.1 |
+5.5 |
On NAVSIMv1, EponaV2 attains DAC of 97.9 and EP of 84.8, exceeding all listed perception-free baselines despite using only a single front-view camera. On the two-stage pseudo closed-loop NAVSIMv2 evaluation, gains are largest on the harder navhard split, where stage-two metrics show improvements in DAC (78.0 vs. 67.6 for DriveLaW) and DDC (88.0 vs. 83.5). The authors classify DriveVLA-W0 (Anchor) as perception-based because it uses manual labels for anchor selection, so comparisons are restricted accordingly.
Limitations
The paper concedes that EponaV2 does not surpass the strongest perception-based models, attributing this to the imprecision of pseudo depth and semantic labels produced by the foundation models—an inherent ceiling of the annotation-free approach. Additional unstated dependencies include the restriction to two fixed text prompts ("car", "human") for semantic supervision, reliance on front-view-only input, and the assumption that simulated-trajectory discrepancy is a valid proxy reward in place of PDMS. Whether richer prompt sets, multi-view inputs, or improved foundation-model labelers close the remaining gap to perception-based methods remains open.
Conclusion
EponaV2 demonstrates that perception-free driving world models can substantially improve planning by forecasting structured future modalities—metric depth and foundation-model semantic features—rather than raw frames alone, and that flow matching GRPO provides a further, label-free refinement of trajectory accuracy. The reported gains, particularly the +5.5 EPDMS improvement on navhard, indicate that comprehensive future reasoning is a productive supervision signal for scalable end-to-end driving, while the acknowledged gap to perception-based SOTA leaves open how far pseudo-label quality can be pushed before manual annotation becomes dispensable.