---
title: 'NAVSIMv2 Benchmark: Extended Autonomous Driving'
url: https://www.emergentmind.com/topics/navsimv2
type: topic
---

# NAVSIMv2 Benchmark: Extended Autonomous Driving

NAVSIMv2 is an extended autonomous-driving planning benchmark that evaluates predicted ego trajectories with the Extended Predictive Driver Model Score (EPDMS). It extends NAVSIMv1’s PDMS-oriented assessment with driving-direction compliance, traffic-light compliance, lane keeping, history comfort, and extended comfort, and uses a two-stage pseudo-simulation protocol for more demanding evaluation of trajectory consequences. The benchmark is derived from real-world driving data associated with OpenScene and nuPlan; its reported configurations include the `navtest` split and the more difficult `navhard` setting. Unlike fully interactive physical simulation, NAVSIMv2 evaluates trajectories using pre-generated or pseudo-simulated future observations rather than requiring complete online rendering and reactive-world simulation. Its central purpose is to measure safety, compliance, progress, and comfort more directly than ADE/FDE-style imitation metrics.

## 1. Development from NAVSIMv1

NAVSIMv2 inherits the evaluation philosophy of NAVSIM, which was introduced as an intermediate regime between inexpensive open-loop trajectory prediction and computationally expensive closed-loop sensor simulation. NAVSIM uses real sensor observations, curated difficult scenarios, short-horizon bird’s-eye-view propagation, and simulation-derived metrics such as collision, progress, time to collision, and comfort [2406.15349].

The original NAVSIM formulation addresses several weaknesses of conventional autonomous-driving benchmarks. Many driving datasets contain large numbers of stationary or nearly constant-speed straight-driving frames, allowing ego-motion-only or constant-velocity policies to obtain strong displacement scores. Moreover, ADE and FDE reward imitation of one recorded human trajectory even when multiple trajectories are valid and another trajectory would be safer, more comfortable, or more progressive. NAVSIM instead evaluates the consequences of a proposed trajectory in a scene abstraction containing the ego vehicle, other agents, drivable area, route geometry, and recorded future motion.

NAVSIMv1 uses the Predictive Driver Model Score (PDMS). Its multiplicative safety terms include No-at-fault Collision (NC) and Drivable Area Compliance (DAC), while Ego Progress (EP), Time to Collision (TTC), and Comfort contribute to a weighted objective. The supplied formulations specify the relative weights \(5:5:2\) for EP, TTC, and Comfort. In contrast, NAVSIMv2 uses EPDMS and adds explicit compliance and temporal-quality dimensions.

The version distinction is important because the literature sometimes uses “closed-loop” imprecisely for NAVSIMv1 or NAVSIMv2. NAVSIMv1 is fundamentally non-reactive: the policy is invoked once, its trajectory is propagated, and other agents do not dynamically respond to it. NAVSIMv2 is described variously as pseudo-closed-loop, two-stage pseudo-simulation, or a closed-loop planning benchmark, depending on the paper. These descriptions generally refer to evaluation under future or perturbed observations, not to a fully interactive physical simulator with unrestricted agent responses.

## 2. Task, data, and pseudo-simulation protocol

In the NAVSIM task, a policy receives an initial driving observation and predicts a future ego trajectory. Depending on the implementation, inputs may include camera images, LiDAR, ego speed and acceleration, navigation commands, historical frames, or language and BEV representations. Outputs are trajectory waypoints or poses rather than direct low-level steering and throttle commands. NAVSIM-based implementations commonly track the predicted trajectory with an LQR controller and propagate the ego vehicle using a kinematic bicycle model.

Several papers report distinct NAVSIMv2 configurations, and no single supplied paper establishes one universal sensor or split specification. Reported settings include:

- `navtest`, with 12,146 scenarios in several benchmark descriptions;
- `navhard`, with stage-specific real and synthetic scenarios in some papers;
- a `navtrain` split, with counts varying by the dataset interpretation or experimental subset;
- camera-only, multi-camera, camera–LiDAR, VLA, and world-model configurations.

The most detailed `navhard` description uses two stages. Stage 1 evaluates a trajectory from the original real observation. Stage 2 evaluates the policy under more difficult follow-up or perturbed observations, including observations produced with 3D Gaussian Splatting or Gaussian-Splatting-based ego-state perturbations. The purpose is to expose failures that are not visible when a model is evaluated only at the logged expert state. These failures include accumulated trajectory error, poor recovery, sensitivity to shifted ego states, and inadequate modeling of future scene consequences.

Other descriptions characterize NAVSIMv2 as using pre-generated follow-up scenes branching from the Stage 1 outcome rather than online rendering. The exact generation procedure, Gaussian parameters, stage aggregation, and scenario inventory are not consistent across the supplied papers and are often not fully specified. Consequently, NAVSIMv2 should be understood as a family of official evaluation protocols centered on two-stage pseudo-simulation, while exact implementation details should be taken from the relevant official evaluator and benchmark release.

The benchmark remains distinct from fully reactive simulation. Future camera or LiDAR observations are generally not rendered online as a consequence of every policy action, and the behavior of other agents is not necessarily recomputed interactively. NAVSIMv2 therefore provides stronger behavioral testing than ordinary open-loop trajectory error while retaining substantially lower computational requirements than complete sensor-rendered closed-loop simulation.

## 3. EPDMS and its component metrics

EPDMS is the primary NAVSIMv2 aggregate metric. Its reported components are:

| Component | Meaning |
|---|---|
| NC | No-at-fault Collision or no-collision score |
| DAC | Drivable Area Compliance |
| DDC | Driving Direction Compliance |
| TLC | Traffic Light Compliance |
| EP | Ego Progress |
| TTC | Time to Collision |
| LK | Lane Keeping |
| HC | History Comfort |
| EC | Extended Comfort |

The benchmark places NC, DAC, DDC, and TLC in multiplicative or penalty terms and combines EP, TTC, LK, HC, and EC through weighted objective terms. One supplied formulation gives the quality-term weights as \(5,5,2,2,2\) for EP, TTC, LK, HC, and EC, respectively, with denominator \(16\). Another source describes the benchmark structure without reproducing the exact weights or aggregation. The official evaluator should therefore be treated as authoritative when comparing scores.

A representative structural form is:

$$
\mathrm{EPDMS}
=
\left(
\prod_{m \in \{\mathrm{NC},\mathrm{DAC},\mathrm{DDC},\mathrm{TLC}\}}
\mathrm{score}_m
\right)
\cdot
\frac{
5\,\mathrm{EP}
+
5\,\mathrm{TTC}
+
2\,\mathrm{LK}
+
2\,\mathrm{HC}
+
2\,\mathrm{EC}
}{16}.
$$

This expression is reported explicitly in some NAVSIMv2-related work, but not consistently reproduced across all papers. Some descriptions additionally apply a human-behavior filter that suppresses penalties when the recorded human trajectory also violates a criterion. DriveStack-VLA reports 87.3 EPDMS without this filter and 91.0 with it, demonstrating that the filter is a metric-protocol effect rather than a model improvement [2606.24051].

The multiplicative structure makes safety and regulatory compliance dominant. A trajectory with high progress but a collision, drivable-area violation, wrong driving direction, or traffic-light violation can receive a sharply reduced aggregate score. EPDMS therefore differs from displacement metrics in three respects: it evaluates consequences rather than imitation distance, it treats several safety and compliance properties as gates, and it explicitly rewards progress and comfort only after basic constraints are satisfied.

NAVSIMv2 results must not be compared directly with NAVSIMv1 PDMS, nuScenes displacement error, nuScenes collision rate, or closed-loop CARLA scores. PDMS omits DDC, TLC, LK, HC, and EC, while nuScenes planning tables use trajectory errors and collision rates rather than either NAVSIM aggregate.

## 4. Policies and architectural approaches

NAVSIMv2 has been used to evaluate conventional end-to-end planners, VLA systems, latent world models, diffusion planners, uncertainty-aware planners, and reinforcement-learning methods.

Several reported results illustrate the range of approaches. DriveWorld-VLA unifies a VLA planner with a latent world model and uses action-conditioned BEV imagination. It reports 86.8 EPDMS on NAVSIMv2, with particularly strong DAC, DDC, and LK values, although the paper does not provide a NAVSIMv2-specific ablation isolating latent imagination [2602.06521].

ELF-VLA introduces Explicit Learning from Failures. It uses structured teacher-generated diagnostic feedback to convert unsuccessful rollouts into corrected high-reward training samples. Its reported NAVSIMv2 result is 87.1 EPDMS, exceeding the listed DriveVLA-W0 result of 86.1 [2603.01063]. The method’s strongest evidence concerns feedback-guided reinforcement learning, but its detailed ablations are primarily on NAVSIMv1.

PaIR-Drive separates imitation learning and reinforcement learning into parallel branches. A tree-structured trajectory sampler generates intention-conditioned alternatives, which are scored with a reward world model. On NAVSIMv2, DiffusionDrive with PaIR-Drive obtains 87.9 EPDMS without best-of-\(N\) selection and 89.6 with best-of-6 selection [2603.13842]. The distinction is essential: 87.9 and 89.6 are different inference protocols, and the latter requires multiple candidate trajectories.

Uncertainty-aware planning is represented by UniUncer, which replaces deterministic vectorized regression heads with probabilistic Laplace regressors for static and dynamic elements. It fuses uncertainty into scene queries and adaptively gates historical ego-status inputs. On the reported Navhard two-stage test, UniUncer improves DiffusionDrive from 25.9 to 28.7 EPDMS, with the main gain concentrated in Stage 2 [2603.07686]. The study models aleatoric uncertainty and does not establish epistemic or out-of-distribution uncertainty estimation.

Other approaches target representation and world-model alignment. EponaV2 predicts future image, depth, and semantic representations and uses flow-matching reinforcement learning; it reports 88.9 EPDMS on NAVSIMv2 `navtest` and 36.1 on `navhard` [2605.14696]. TPS-Drive uses an agent-centric tokenizer whose codebook is supervised by a frozen 3D detection head, followed by scene understanding, future forecasting, and diffusion-based action generation; it reports 86.7 EPDMS [2605.27038]. Metis uses a Mixture-of-Transformers with separate video-generation and action experts; an asymmetric attention mask allows video supervision during training while bypassing video generation at inference. It reports 89.5 EPDMS on `navtest` and 32.2 on `navhard` [2606.15869].

GraphWorld models an ego-centric interaction graph and conditions planning on a latent world state representing nearby-agent interactions. It reports 89.5 EPDMS on `navtest` and 53.6 on `navhard`, although the paper notes that its NAVSIMv2-specific architectural ablations are limited [2606.16274]. SV-WAM preserves six-camera surround-view input while using future-video prediction only as training supervision. It reports 91.0 EPDMS on `navtest`, with action-only inference and a differentiable drivable-area regularizer [2609.03602].

DriveZero obtains behavior supervision from a privileged closed-loop reinforcement-learning teacher, then distills that behavior into a camera-only policy using DriveVFM visual representations and goal-conditioned teacher rollouts. DriveZero reports 51.5 EPDMS on `navhard`, while DriveZero-Scale reaches 57.1, with the largest gain occurring in Stage 2 [2609.06055]. Run-then-Walk instead changes the reinforcement-learning schedule: it first discovers high-progress modes and then repairs safety. On `navtest`, it improves AutoDrive-P³ from 88.7 to 89.6 EPDMS and ReCogDrive from 82.7 to 83.7 [2609.25831].

WALT learns a compact latent trajectory representation aligned with a frozen driving world model. It reports an improvement from 87.3 to 87.9 EPDMS relative to an EponaV2 raw-waypoint baseline, while reducing trajectory-planner FLOPs by 30.5% [2609.30436].

## 5. Benchmark findings and interpretation

The reported NAVSIMv2 results support several broad observations. First, the benchmark is more demanding than conventional open-loop trajectory matching because it exposes degradation under future-state perturbation and evaluates a wider range of behavioral properties. Models that obtain strong Stage 1 results can experience substantial Stage 2 degradation, especially in lane keeping, progress, collision avoidance, or extended comfort.

Second, high aggregate EPDMS does not imply uniform superiority on every component. A method may improve the overall score while reducing EP, HC, or EC, or while trailing another method on NC or DAC. For example, UniUncer’s overall gain is concentrated in difficult Stage 2 scenes, while its Stage 1 score slightly decreases. Run-then-Walk improves AutoDrive-P³’s EP by 6.0 points but slightly reduces several safety and comfort components. WALT’s largest gain is EC, while EP and LK decrease slightly relative to its baseline.

Third, the benchmark favors methods that combine multiple objectives rather than optimizing trajectory imitation alone. Candidate generation and ranking, latent future-state prediction, explicit safety regularization, dynamic-agent uncertainty, structured failure feedback, and reinforcement-learning objectives all target different parts of the EPDMS evaluation. No single mechanism is established as universally sufficient.

Fourth, metric configuration strongly affects comparison. Human penalty filtering can change an EPDMS score substantially, and best-of-\(N\) inference can increase performance by evaluating multiple candidate trajectories. Comparisons are therefore meaningful only when sensor inputs, data splits, human-filter settings, candidate counts, training data, and evaluator versions are aligned. Several papers mark older results with an asterisk because they were produced with an earlier evaluator or before a human-penalty aggregation update.

Finally, NAVSIMv2 is not a substitute for fully reactive simulation or real-world testing. It does not guarantee robustness to independently reacting traffic, long-duration feedback, rendering artifacts outside the benchmark distribution, or real vehicle-dynamics mismatch. The benchmark’s pseudo-simulation improves behavioral relevance and scalability, but its validity remains conditional on the quality of follow-up scene generation, collision responsibility rules, route and map annotations, and metric implementation.

## 6. Limitations, reproducibility, and research directions

The supplied NAVSIMv2 literature leaves several protocol details incompletely specified. These include exact train/validation/test cardinalities, scenario taxonomies, stage aggregation equations, human-filter implementation, Gaussian-Splatting perturbation distributions, trajectory coordinate conventions, planning horizons, and evaluator versions. Some papers report incompatible `navhard` scenario counts or use different interpretations of the NAVSIMv2 split. Exact reproduction therefore requires the official dataset and evaluator rather than relying on headline descriptions in individual papers.

Reported methods also differ substantially in computational requirements. Some use large VLMs or video world models, while others use lightweight end-to-end planners or latent action models. Inference procedures range from single trajectory prediction to best-of-6 selection, test-time LoRA adaptation, candidate reranking, and multiple flow or diffusion steps. A score without its inference budget is therefore incomplete.

Several methodological directions follow from the benchmark’s current design:

- **More explicit interaction modeling**: future agent behavior and ego–agent coupling remain only partially modeled in pseudo-simulation.
- **Improved collision responsibility**: more complete at-fault logic would better assess rearward and interaction-heavy sensing.
- **Traffic-rule completeness**: traffic-light and driving-direction metrics could be extended with richer stop-sign and route-rule semantics.
- **Failure-centric reporting**: aggregate EPDMS should be accompanied by zero-score counts, per-stage failure rates, and scenario-level distributions.
- **Uncertainty and calibration**: aleatoric uncertainty is increasingly used, but epistemic, calibration, and out-of-distribution evaluations remain limited.
- **Protocol standardization**: comparisons should report evaluator version, human-filter status, candidate count, sensor modality, training split, and computational budget.
- **Reactive validation**: NAVSIMv2 should be complemented by fully reactive simulators and real-world evaluation to test replanning, feedback, and long-horizon compounding errors.

NAVSIMv2’s principal significance is therefore methodological. It extends data-driven trajectory evaluation beyond imitation error while avoiding the full cost of sensor-rendered closed-loop simulation. Its EPDMS metric makes safety, regulatory compliance, progress, lane keeping, and comfort jointly measurable, and its two-stage pseudo-simulation exposes weaknesses that can remain hidden in standard open-loop benchmarks. The resulting scores are useful for comparative research, but they should be interpreted as benchmark evidence rather than as direct guarantees of interactive driving safety.

Source: https://www.emergentmind.com/topics/navsimv2