AD-E2E-JEPA: A Joint-Embedding Predictive Architecture For End-to-End Autonomous Driving
Abstract: Autonomous driving requires \textit{world models} that can understand the physical world, reason and plan, and operate safely. In this paper, we first systematically evaluate existing action-conditioned joint-embedding predictive architecture (JEPA) world models, including LeWM, DINO-WM, and JEPA-WM for end-to-end autonomous driving (E2EAD). To isolate world-model quality from policy learning, we employ a goal-conditioned zero-shot planning setting that evaluates these models using ground-truth future observations as goals, without training any driving policy. We find that existing JEPA-based world models are either accurate for driving but computationally expensive, or computationally efficient but insufficient for planning. To address this trade-off, we propose \textbf{AD-E2E-JEPA}, which introduces a SIGReg-regularized learnable projector applied to projected patch embeddings. The projector reduces the number of planning patches by and the embedding dimension by , achieving a inference speedup while retaining planning performance, with a 0.8-second runtime for an 8-frame rollout over 256 candidate trajectories. \textit{Without} training any driving policy, the world model itself reaches the goals located 20 meters away on average within the displacement of respectively 4.0/2.8 meters, using world-model rollouts over trajectory vocabularies of respectively 256/8,192 candidates. On the NAVSIMv2 benchmark, it achieves 67.3/72.9 EPDMS with multiplicative safety metrics and 84.1/86.5 EPDMS without them in goal-conditioned zero-shot planning. Experiments further show that the self-supervised pretrained projector improves downstream imitation learning performance from 80.2 to 85.4 EPDMS. The source code is available at https://github.com/HaoranZhuExplorer/AD-E2E-JEPA
Paper Prompts
Sign up for free to create and run prompts on this paper.
Top Community Prompts
Explain it Like I'm 14
1. What is this paper about?
This paper presents AD-E2E-JEPA, a new artificial intelligence system for autonomous driving.
Most self-driving systems learn by copying human drivers. The car sees what a human did in a situation and tries to do the same thing. This is called imitation learning.
The researchers wanted to build something more thoughtful. Instead of only copying human actions, their system tries to build a kind of mental model of the driving world. This model predicts what might happen if the car chooses different actions, such as turning left, turning right, or continuing straight.
The system can then choose the action that is most likely to lead to a desired future situation.
2. What questions did the researchers ask?
The paper focuses on several main questions:
- Can a JEPA-based world model help a self-driving car plan its movements?
- Can the car plan successfully without first being trained to copy a driving policy?
- Can the system make accurate plans while working quickly enough for driving?
- Can the modelโs learned knowledge also improve ordinary imitation-learning systems?
- How does AD-E2E-JEPA compare with other world models?
A major problem with earlier systems was a trade-off:
- Some models planned accurately but took more than a minute to evaluate one scene.
- Other models were fast but made poor plans.
The researchers wanted a model that was both accurate and fast.
3. How does the method work?
What is a world model?
A world model is an internal picture of how the world works. For example, a person crossing the road will probably keep moving forward, while a parked car will probably stay still.
For a self-driving car, a world model should understand things such as:
- Where the road is
- How the car is moving
- What may happen after turning or braking
- Which future situations are safe or useful
An analogy is a video game. Before moving a character, you might imagine what will happen if you press different buttons. A world model does something similar for a car.
What is JEPA?
JEPA, or Joint-Embedding Predictive Architecture, does not try to recreate every future camera image pixel by pixel. Instead, it converts images into shorter descriptions called embeddings.
An embedding is like a compact summary of an image. Instead of remembering every detail, it might represent information such as:
โThere is a road ahead, a vehicle on the right, and the car is approaching an intersection.โ
The model learns to predict the embedding of a future scene based on:
- Recent camera images
- The carโs recent movements
- A possible future driving path
This allows the model to compare different possible actions without generating complete future images.
Compressing the information
Camera images contain a large amount of information. The model divides each image into small areas, or patches, and creates an embedding for each patch.
Using all these patches would make planning slow. AD-E2E-JEPA adds a learnable projector, which acts like a smart compression tool. It reduces:
- The number of patches by 16 times
- The size of each embedding by 4 times
The researchers also use a technique called SIGReg. Its purpose is to make sure the compressed information does not become too similar everywhere. In simple terms, SIGReg helps the model keep useful differences between scenes instead of turning every scene into nearly the same description.
How the car plans
The system is given a goal, represented by a future camera image. It then considers many possible driving paths.
For each path, the world model predicts what the future scene would look like in embedding form. The system compares each prediction with the embedding of the goal image.
It chooses the path whose predicted future is most similar to the goal.
This is called zero-shot goal-conditioned planning:
- Goal-conditioned means the system plans toward a specified goal.
- Zero-shot means it does this without separately training a driving policy for that task.
An everyday analogy would be choosing a route in a city. You imagine what each route will lead to, compare the results with where you want to go, and choose the best route.
Rollout training
The researchers also tested rollout training. Instead of predicting only one future step, the model repeatedly predicts several steps into the future.
This is similar to thinking:
โIf I turn now, where will I be next? And after that, where will I be?โ
This helps the model make more consistent long-term plans.
4. What were the main findings?
The new system was much faster
With 256 possible driving paths, AD-E2E-JEPA needed about 0.8 seconds per scene.
The other accurate models, DINO-WM and JEPA-WM, needed about 92 to 101 seconds per scene.
Therefore, AD-E2E-JEPA was roughly 100 times faster while keeping similar planning quality.
This is important because a self-driving car must make decisions quickly. Waiting over a minute to choose a path would not be practical on a real road.
It planned more accurately than the simple fast model
The older LeWM model was also fast, taking about 0.7 seconds per scene. However, its plans were much less accurate.
On the full test set:
| Model | Planning time | Final displacement error |
|---|---|---|
| LeWM | 0.7 seconds | 14.6 meters |
| AD-E2E-JEPA with rollout training | 0.8 seconds | 4.0 meters |
The final displacement error measures how far the selected final position was from the correct position. A smaller number is better.
This means AD-E2E-JEPA was almost as fast as LeWM but produced much more accurate paths.
More possible paths improved performance
The researchers tested different numbers of candidate paths. With 8,192 possible trajectories, the best version achieved:
- 72.9 EPDMS with safety measures
- 86.5 EPDMS without those safety measures
- 2.8 meters of average final-position error
- 2.0 degrees of average heading error
Using more candidate paths gave the model more choices, helping it find a path closer to the correct one. However, this also increased planning time, from 0.8 seconds with 256 paths to 18.2 seconds with 8,192 paths.
Training on more driving data helped
The researchers trained some versions on about 10 hours of driving video and others on about 70 hours.
The model trained with more data and rollout training was generally more reliable and accurate. On the full test set, it identified the correct driving path among its best five choices about 83% of the time.
The learned projector helped another driving system
The researchers also tested whether the projectorโs learned knowledge could help a normal imitation-learning system.
They compared:
- A system with a randomly initialized projector
- A system using the projector pretrained by AD-E2E-JEPA
The pretrained projector improved the EPDMS score from 80.2 to 85.4.
This suggests that the projector learned useful information about driving, even when used in a different type of model.
5. Why are these results important?
The results suggest that a self-driving car does not always need to directly copy human drivers. It may instead learn how actions affect the future and use that knowledge to plan.
The main contribution is the balance between:
- Speed
- Accuracy
- Useful understanding of the driving environment
The system also shows that a model can learn helpful knowledge without requiring detailed labels for every possible driving decision.
6. Possible impact and limitations
If developed further, this approach could help create self-driving systems that:
- Make decisions more quickly
- Consider many possible driving choices
- Adapt better to unfamiliar situations
- Use less expensive computation during planning
- Improve other driving systems through transfer learning
However, the results are still based on the NAVSIMv2 benchmark, which is a dataset and computer-based evaluation rather than complete real-world driving. The system mainly uses a front-facing camera and does not prove that it is ready for public roads.
The model also plans by choosing from a fixed collection of candidate trajectories. Real driving may require actions that are not included in that collection. In addition, the reported safety scores do not mean the system is guaranteed to be safe in every unusual situation.
Conclusion
AD-E2E-JEPA is a fast world model for autonomous driving. It learns to predict how different driving actions will change the future scene. Instead of simply copying human drivers, it compares possible paths and chooses one that should lead toward a desired goal.
The researchersโ main innovation is a smart compression system that makes planning about 100 times faster than some earlier accurate methods. The model also produced more accurate driving plans and improved a separate imitation-learning system.
Overall, the paper shows that learning a compact model of how the world changes could be an important step toward faster and more flexible autonomous vehicles.
Knowledge Gaps
Knowledge gaps, limitations, and open questions
- Limited sensor modality: The study evaluates only a front-camera input, leaving it unclear whether AD-E2E-JEPA generalizes to multi-camera, LiDAR, radar, depth, or vehicle-state sensor configurations.
- Restricted environmental diversity: All experiments use NAVSIM/NAVSIMv2 data, so robustness across cities, countries, weather conditions, lighting, road geometries, traffic cultures, and sensor distributions remains untested.
- No real-world or closed-loop validation: The reported results come from offline data and NAVSIM pseudo-simulation; the modelโs behavior under real vehicle dynamics, feedback errors, perception shifts, and interactive traffic is unresolved.
- No safety-oriented planning mechanism: Planning minimizes latent distance to a future image and does not explicitly model collision probability, traffic-rule compliance, emergency maneuvers, uncertainty, or risk-sensitive objectives.
- Safety metrics are not optimized directly: The strongest reported performance relies partly on
EPDMSโ, which excludes multiplicative safety terms, and the paper does not establish whether high latent-goal matching translates into reliable collision avoidance. - Dependence on privileged future observations: Zero-shot planning uses the ground-truth future image as the goal, an assumption unavailable during ordinary autonomous driving. Performance with predicted, noisy, language-based, map-based, or user-specified goals is not evaluated.
- Unclear deployability without future goals: The paper does not specify how goals would be generated online in a complete autonomous-driving system or how the planner would operate when no target future observation is available.
- Action-space coverage is limited: Candidate actions are selected from a fixed vocabulary of 256โ8,192 clustered trajectories. The methodโs ability to discover trajectories outside this vocabulary, including rare evasive or highly maneuverable actions, remains unknown.
- Scalability trade-off is unresolved: Increasing the vocabulary to 8,192 improves EPDMS and displacement accuracy but raises planning time from 0.8 to 18.2 seconds per scene, which may be too slow for real-time control. The practical latencyโperformance frontier is not established.
- No temporal receding-horizon evaluation: The planner selects a trajectory for a future horizon but is not evaluated in a repeated closed-loop receding-horizon setting, where model errors accumulate and previous actions alter subsequent observations.
- Long-horizon prediction remains uncertain: The experiments use a 4-second future horizon. Reliability over longer horizons, especially for intersections, lane changes, merges, and multi-stage maneuvers, is not assessed.
- Autoregressive error accumulation is not characterized: Rollout training improves several metrics, but the paper does not quantify how prediction error grows across each future timestep or identify when autoregressive rollouts become unreliable.
- Limited evaluation of dynamic-agent interaction: The model is trained and evaluated on logged human driving data, but its ability to predict reactions from pedestrians, cyclists, and other vehicles under interventions different from the logged trajectory is not demonstrated.
- Offline data coverage may induce distributional bias: Because the action-conditioned model learns primarily from human driving trajectories, it may not accurately predict consequences of uncommon, unsafe, or counterfactual actions absent from the dataset.
- Causal validity of the learned dynamics is unverified: Matching future latent embeddings does not establish that the model has learned physically or causally correct environment dynamics rather than exploiting visual correlations in the dataset.
- Latent-distance objective may be misaligned with driving quality: The authors note that latent distance appears more directly related to geodesic accuracy and hit rate than to NAVSIM driving metrics. The relationship between embedding similarity, comfort, legality, progress, and safety remains insufficiently understood.
- Representation compression is not fully explained: The paper shows that a projector reducing patch count by and dimension by can preserve performance, but it does not identify which spatial, semantic, or temporal information is discarded.
- SIGRegโs contribution is not isolated comprehensively: The reported gains may result from the projector architecture, dimensionality reduction, patch reduction, stop-gradient training, SIGReg, or their interactions. A complete factorial ablation is needed to attribute improvements reliably.
- Sensitivity to SIGReg hyperparameters is unclear: The SIGReg weight is tuned heuristically according to batch size, but robustness to , batch size, number of random directions, and knot configuration is not systematically reported.
- Potential instability across seeds is unreported: Results lack confidence intervals, multiple random seeds, or statistical significance tests, particularly for the 100-scene subset where the authors acknowledge substantial variance.
- Small-subset comparisons may be unreliable: Comparisons among methods are performed on only 100 test scenes because of computational cost, while the full-set evaluation includes only selected methods. Full-test comparisons against DINO-WM and JEPA-WM remain unavailable.
- Baseline fairness is incomplete: The baselines use different backbones, batch sizes, training resources, and configurations. It is unclear whether all models receive equally optimized hyperparameters, equivalent training data, and comparable compute budgets.
- Backbone dependence is unexamined: The method relies on a frozen DINOv3 ViT-L encoder. Its effectiveness with smaller, independently trained, domain-adapted, or newer visual encoders is unknown.
- Frozen-encoder limitations are unresolved: The study does not test whether jointly adapting the encoder improves robustness to driving-specific visual phenomena, domain shift, or rare traffic events.
- No comparison with strong online planners or policy baselines in the same setting: The paper does not directly compare against modern model-predictive control, diffusion planners, trajectory-ranking systems, or learned policies under identical sensor, data, and compute constraints.
- Transfer-learning evidence is narrow: Projector transfer is evaluated using one simple imitation-learning decoder and a random-projector baseline. It is unclear whether the benefit persists across decoder architectures, training regimes, task objectives, or policy-learning methods.
- The transfer improvement is not disentangled from optimization effects: The pretrained projector may provide better initialization, scale normalization, or regularization rather than predictive driving structure. Additional controls, such as frozen versus trainable components and matched parameterizations, are needed.
- No cross-domain transfer evaluation: The downstream transfer experiment remains within NAVSIM. Transfer to other datasets, cities, sensor configurations, or real-world driving domains is not established.
- Uncertainty estimation is absent: The world model produces point predictions and a single latent-distance score, without calibrated uncertainty over future states, candidate trajectories, or model confidence.
- Planner reliability metrics are incomplete: Top-1 and top-5 hit rates measure ranking of the logged ground-truth trajectory, but do not assess whether alternative trajectories are also safe, feasible, or better than the recorded human action.
- Ground-truth trajectory quality is assumed: Human demonstrations are treated as the target trajectory even though the introduction acknowledges that they may be noisy or suboptimal. The method is not evaluated against expert-optimal, safety-verified, or multi-modal trajectory annotations.
- Heading and pose representation issues remain: The formulation uses planar poses and relative heading changes, but does not discuss angle periodicity, accumulated pose drift, or whether this representation adequately captures vehicle kinematics and nonholonomic constraints.
- Vehicle dynamics and comfort constraints are under-modeled: The candidate trajectory vocabulary and latent planner are not shown to enforce acceleration, jerk, steering-rate, tire-friction, or passenger-comfort limits.
- No analysis of failure cases: The paper provides qualitative examples but does not systematically categorize failures, such as occlusions, ambiguous goals, unusual road layouts, sudden obstacles, mislocalized poses, or visually similar but behaviorally different scenes.
- Runtime measurements are hardware- and implementation-dependent: The reported speedups are measured on an A100 and do not include the full perception, goal-generation, control, and vehicle-interface pipeline. Real-time performance on embedded automotive hardware remains unknown.
- Compute and memory costs are incompletely reported: The paper reports planning time and some training configurations but does not provide a full accounting of model parameters, GPU memory, energy use, preprocessing costs, or deployment constraints.
- Reproducibility is affected by unclear experimental details: The manuscript contains notation and formatting inconsistencies, and some training, evaluation, and benchmark definitions are deferred to appendices or referenced implementations. Exact preprocessing, trajectory clustering, candidate sampling, and checkpoint-selection procedures require fuller specification.
- Generalization to unseen goals is not demonstrated: Although the paper describes zero-shot planning, goals are drawn from ground-truth future observations in the same benchmark distribution. Generalization to novel visual goals, altered target locations, and compositional goals remains open.
- Interactions between goal specification and representation compression are unknown: It is unclear whether the compressed projector preserves information needed for goals involving objects, lanes, traffic signals, or spatial relations that are not represented by the logged future frame.
- The limits of JEPA versus generative world models remain unresolved: The paper discusses generative video world models but does not experimentally determine when latent prediction is preferable to pixel/video generation for physical consistency, planning, interpretability, or uncertainty modeling.
- No interpretability or semantic probing is provided: The paper does not establish which latent dimensions or patches encode road structure, agent motion, traffic signals, obstacles, or action consequences, limiting understanding of why compressed embeddings support planning.
- Policy-free planning performance does not establish complete autonomy: The experiments isolate world-model quality by omitting policy training, but they do not show how AD-E2E-JEPA integrates with a controller or whether its selected trajectories can be executed reliably under perception and actuation delays.**
Practical Applications
Immediate Applications
- Real-time goal-conditioned planning for autonomous vehicles โ automotive and mobility
- Integrate AD-E2E-JEPA as a planning component that evaluates candidate ego-vehicle trajectories against a desired future camera observation or navigation target.
- The reported approximately
0.8 splanning time for an 8-frame rollout over 256 candidates makes the method suitable for offline evaluation, closed-course testing, and potentially low-frequency online replanning. Larger candidate sets improve accuracy but increase latency, reaching approximately18.2 sfor 8,192 candidates on an A100 GPU. - A practical workflow is: encode recent camera frames and poses, roll out the world model for each trajectory in a predefined vocabulary, compare predicted and target latent embeddings, and execute the lowest-cost trajectory.
- Dependencies: reliable camera calibration, ego-pose estimation, a sufficiently representative trajectory vocabulary, GPU or specialized accelerator deployment, and an independent safety layer. The paper evaluates zero-shot planning with ground-truth future images as goals, which are not directly available during ordinary autonomous driving.
- Efficient world-model benchmarking and planner selection โ automotive R&D and academia
- Use the released implementation as a benchmark for comparing latent world models on planning efficiency, final displacement error, heading error, hit rate, and NAVSIM EPDMS.
- The method enables full-test-set evaluation that is impractical for dense-patch baselines: AD-E2E-JEPA evaluates 12,146 scenes at roughly
0.8 sper scene, whereas DINO-WM and JEPA-WM require substantially more computation. - This can support engineering decisions about the trade-off between latent representation size, candidate-trajectory count, safety performance, and inference latency.
- Dependencies: NAVSIM-compatible data and metrics, comparable hardware settings, and careful separation of zero-shot planning performance from policy-learning performance.
- Self-supervised pretraining for end-to-end driving systems โ automotive software
- Use the learned projector as an initialization for an imitation-learning driving model. The reported improvement from
80.2to85.4EPDMS relative to a randomly initialized projector indicates that predictive latent structure can improve downstream trajectory decoding. - A deployable training workflow is to freeze or fine-tune the DINOv3 encoder, initialize the projector from AD-E2E-JEPA, attach a trajectory decoder, and fine-tune on driving demonstrations.
- This may reduce dependence on dense trajectory labels or extensive supervised pretraining, while making existing imitation-learning systems more robust to temporal context.
- Dependencies: transferability beyond NAVSIM, compatibility between sensor configuration and pretrained representations, and sufficient supervised data for the final policy.
- Use the learned projector as an initialization for an imitation-learning driving model. The reported improvement from
- Compute-efficient latent representation compression โ machine-learning infrastructure
- Apply the projector design to reduce dense visual embeddings before downstream planning or control. The method compresses spatial patch count by
16รand embedding dimension by4ร, while preserving much of the planning quality through SIGReg regularization. - The resulting component could be packaged as a reusable latent-compression module for autonomous-driving stacks, robotics systems, or video-based control models.
- Dependencies: SIGReg hyperparameter tuning, adequate batch sizes, preservation of task-relevant spatial information, and validation against representation collapse and rare-event performance degradation.
- Apply the projector design to reduce dense visual embeddings before downstream planning or control. The method compresses spatial patch count by
- Offline trajectory ranking and safety analysis โ fleet operators and testing organizations
- Use the world model to rank candidate trajectories in recorded driving scenarios, identify situations where the ground-truth trajectory is absent from the top-ranked candidates, and generate failure cases for regression testing.
- Top-1 and top-5 hit rates provide a practical diagnostic for whether the learned latent dynamics assign low cost to plausible trajectories.
- This can support dataset curation, planner debugging, scenario prioritization, and comparison of software releases without deploying the model on public roads.
- Dependencies: hit rate is not equivalent to collision avoidance or legal compliance; recorded scenarios must include sufficient environmental diversity, and safety-critical evaluation should include explicit collision, traffic-rule, and uncertainty checks.
- Policy and standards evaluation for efficient autonomous-driving systems โ regulators and public agencies
- Agencies can use latent-world-model benchmarks to evaluate computational efficiency, generalization, reliability, and safety-related planning behavior in a standardized offline setting.
- The paperโs distinction between EPDMS with and without multiplicative safety terms is useful for separating general driving quality from explicit safety performance.
- Dependencies: NAVSIM and pseudo-simulation metrics are not substitutes for real-world validation, certification, or liability assessment. Regulatory use would require scenario coverage, interpretability, uncertainty reporting, and hardware-in-the-loop or closed-track testing.
- Research and teaching tool for model-based decision-making โ academia and education
- AD-E2E-JEPA provides a concrete example for courses and laboratories covering self-supervised learning, JEPA architectures, latent dynamics, model-based planning, trajectory search, and representation regularization.
- Students can reproduce ablations involving projector compression, SIGReg, rollout training, candidate-vocabulary size, and downstream transfer.
- Dependencies: the paperโs code and pretrained models must remain available, and experiments may require high-memory GPUs such as A100-class hardware.
Long-Term Applications
- On-road autonomous driving with closed-loop replanning โ automotive and transportation
- A mature version could serve as a compact model-based planner that repeatedly predicts the consequences of candidate actions and replans as new observations arrive.
- Unlike purely reactive imitation, the architecture could evaluate alternative futures before acting, potentially improving behavior in unfamiliar scenes and reducing dependence on demonstrator quality.
- A production workflow would combine the JEPA world model with perception, localization, route goals, uncertainty estimation, an explicit collision checker, traffic-rule reasoning, and a low-level controller.
- Dependencies: the current evaluation uses a front camera, offline ground-truth future goals, a fixed trajectory vocabulary, and pseudo-simulation. Deployment requires multi-camera or multimodal sensing, real-time worst-case latency guarantees, robustness to weather and sensor corruption, calibrated uncertainty, and extensive closed-loop validation.
- Language- or map-conditioned driving goals โ automotive navigation and humanโvehicle interaction
- The image-goal formulation could be extended to goals specified by HD maps, route waypoints, semantic targets, or natural-language instructions such as โmove to the next safe laneโ or โstop near the marked entrance.โ
- A product could use a goal encoder to convert route or user intent into the same latent space used for trajectory comparison.
- Dependencies: goal representations must be aligned with the predictive latent space; semantic intent, traffic rules, and long-horizon route planning cannot be assumed to emerge from image-latent matching alone.
- Robotics navigation and manipulation โ robotics
- The projector-plus-SIGReg design could be adapted to action-conditioned visual world models for mobile robots, warehouse vehicles, drones, or manipulation systems.
- Candidate action sequences could be ranked by the latent distance between predicted observations and a goal image, enabling reward-free or low-reward planning from offline demonstrations.
- Dependencies: robot dynamics, camera viewpoint, action parameterization, contact physics, and environment changes differ substantially from driving. Real-robot deployment would require uncertainty-aware planning and recovery behavior.
- Cross-domain latent planners for energy and industrial control โ energy and manufacturing
- Similar compressed JEPA world models could forecast latent system states under candidate control sequences for battery management, HVAC optimization, warehouse automation, or process control.
- The approach is particularly relevant where full-fidelity simulation is expensive but historical sensor-action trajectories are available.
- Dependencies: domain-specific safety constraints, reliable action coverage in offline data, observability of system state, and proof that latent distance correlates with operational objectives rather than merely visual similarity.
- Large-scale fleet learning and continual adaptation โ transportation platforms
- Fleet operators could periodically retrain the projector and predictive model on newly collected edge cases, then transfer updated representations to downstream driving policies.
- Adaptive updates could target unusual weather, road layouts, construction zones, or regional driving patterns without retraining an entire end-to-end system from scratch.
- Dependencies: data governance, privacy protection, distribution shift, prevention of catastrophic forgetting, version validation, and mechanisms ensuring that improvements in average metrics do not reduce rare-event safety.
- Safety-aware model-based planning โ autonomous systems and policy
- Future systems could combine latent goal matching with explicit constraints for collision probability, time-to-collision, lane boundaries, pedestrian interactions, comfort, and traffic-law compliance.
- This would address a central limitation of the current method: zero-shot latent-distance planning does not explicitly optimize safety, and the strongest reported EPDMS values without safety terms should not be interpreted as deployment readiness.
- Dependencies: calibrated predictive uncertainty, sufficiently accurate world-model rollouts for rare events, formal or statistical safety guarantees, and validated integration with independent safety monitors.
- Hierarchical long-horizon autonomy โ robotics and autonomous vehicles
- AD-E2E-JEPA could become a low-level or mid-level component in a hierarchical planner: a route planner selects semantic subgoals, while the JEPA model selects short-horizon trajectories that achieve each subgoal.
- This could extend the current 4-second future horizon to complex maneuvers such as intersections, parking, merging, and multi-step navigation.
- Dependencies: temporal abstraction, subgoal generation, error accumulation over repeated rollouts, and mechanisms for recovering when the desired goal is not reachable within the candidate vocabulary.
- Simulation, digital-twin, and synthetic-data generation โ automotive research
- Although the model predicts latent embeddings rather than pixels, it could provide a computationally efficient predictive layer for evaluating hypothetical actions in digital twins or for selecting scenarios that require expensive photorealistic simulation.
- It may also help identify informative trajectories and hard cases for targeted data collection.
- Dependencies: latent predictions are not directly usable as visual simulation outputs; coupling to a generative decoder or external simulator would require further research and could introduce additional model errors.
- Commercial planning and perception software components โ AI tooling
- The projector, SIGReg objective, latent rollout evaluator, trajectory vocabulary search, and reliability metrics could be released as modular libraries or accelerator kernels for model-based control.
- Such tools could support rapid prototyping across autonomous vehicles and robots, particularly where dense embedding evaluation is a computational bottleneck.
- Dependencies: licensing and reproducibility of the DINOv3 backbone, hardware portability, numerical stability, standardized interfaces, and evidence that compression preserves performance under distribution shift rather than only on NAVSIM.
Glossary
- Action-conditioned world model: A model that predicts future states based on a specified sequence of actions. โan action-conditioned world model can be rolled out to predict possible future states under different candidate actionsโ
- AdaLN: Adaptive Layer Normalization, a conditioning mechanism that modulates layer-normalization parameters according to auxiliary inputs. โan AdaLN-style predictor equipped with RoPEโ
- Autoregressive prediction: Sequential prediction in which each predicted output is used as input for later predictions. โbased on autoregressive predictions from frame through frame โ
- CEM (cross-entropy method): An iterative sampling and optimization method for selecting high-performing action sequences. โThe cross-entropy method (CEM) is widely used for JEPA-based world-model planningโ
- CLS token: A special transformer token whose embedding summarizes an entire input, commonly used for classification or global representation. โglobal CLS embeddings across the batch independently at each time stepโ
- CramรฉrโWold theorem: A mathematical result stating that a multivariate probability distribution is determined by all of its one-dimensional projections. โBy the CramรฉrโWold theorem, matching all one-dimensional projected distributions is equivalent to matching the full joint distributionโ
- Dense patch embedding: A separate feature vector representing each local image patch rather than one global image representation. โExisting JEPA-based world model baselines operate on dense patch embeddings produced by the encoderโ
- Diffusion transformer: A transformer architecture used within a diffusion generative model to produce complex data such as images or videos. โcurrent generative world models employ diffusion transformersโ
- Ego pose: The position and orientation of the autonomous vehicle relative to a chosen coordinate frame. โthe corresponding ego pose, consisting of the planar position and heading angle โ
- EppsโPulley test: A statistical test for assessing whether a distribution is consistent with a normal distribution using its empirical characteristic function. โ denotes the univariate EppsโPulley testโ
- EPDMS: A NAVSIM planning metric that combines driving-quality measures with multiplicative safety factors. โEPDMS: NAVSIMv2 uses EPDMS as a pseudo-simulation-based planning metric that combines multiplicative safety terms with a weighted measure of driving quality.โ
- FDE (final displacement error): The distance between the predicted final pose and the ground-truth final pose. โwe report the final displacement error (FDE) and absolute errors in longitudinal positionโ
- Geodesic accuracy: Accuracy measured by the spatial and angular displacement between a predicted trajectory endpoint and the ground-truth endpoint. โwhich measure the geodesic accuracy. of the world model.โ
- Goal-conditioned planning: Planning that selects actions to reach a specified target state or observation. โit enables goal-conditioned zero-shot planning by selecting among candidate driving trajectoriesโ
- Hit rate: The proportion of cases in which the ground-truth trajectory appears among the modelโs best-ranked candidates. โwe use hit rate to measure whether the world model ranks the ground-truth trajectory among the top- lowest-cost candidatesโ
- Imitation learning: Learning a policy by reproducing demonstrated behavior rather than optimizing directly through interaction with the environment. โa reactive policy is trained to imitate human driving trajectoriesโ
- Isotropic Gaussian distribution: A multivariate normal distribution with equal variance in every direction and no correlations between dimensions. โencourage the embeddings to follow an isotropic Gaussian distributionโ
- JEPA (joint-embedding predictive architecture): A self-supervised architecture that predicts a target representation from a contextual representation without reconstructing the input data. โa joint-embedding predictive architecture (JEPA) for end-to-end autonomous drivingโ
- Latent dynamics: The evolution of hidden or learned representations that encode the underlying state of an environment. โlearned latent dynamics using pixel reconstruction objectivesโ
- Latent embedding: A learned vector representation that encodes information about an input in a lower-level feature space. โthe candidate whose predicted future latent embedding is closest to that of the goalโ
- MSE (mean squared error): A loss function that averages the squared differences between predicted and target values. โWe supervise the predictions using only an MSE lossโ
- NAVSIMv2: A benchmark for evaluating autonomous-driving planning using data-driven, pseudo-simulation-based metrics. โWe use the NAVSIM (Cao et al., 2025) dataset and evaluate on the latest NAVSIMv2 benchmark.โ
- One-dimensional projection: The mapping of a multidimensional vector onto a single direction to analyze its distribution. โmatching all one-dimensional projected distributions is equivalent to matching the full joint distributionโ
- Patch embedding: A vector representation generated by encoding a local image patch. โwe apply SIGReg to patch embeddings independently at each patch location and time stepโ
- Policy learning: Training a model that maps observations or states to actions. โTo isolate world-model quality from policy learningโ
- Predictor: The neural-network component of a JEPA that forecasts target embeddings from contextual embeddings and actions. โthe AdaLN predictor predicts the embedding sequence one time step aheadโ
- Projector: A learnable network that transforms and compresses encoder representations into a smaller embedding space. โAD-E2E-JEPA further introduces a learnable projector โ
- Representation collapse: A failure mode in self-supervised learning in which different inputs receive nearly identical, uninformative representations. โwhere distinct DINOv3 embeddings are mapped to nearly constant vectors and thus become uninformativeโ
- RoPE (rotary position embedding): A positional encoding method that incorporates token positions by rotating query and key representations in a transformer. โan AdaLN-style predictor equipped with RoPEโ
- Rollout training: Training a world model on sequences of its own recursively generated predictions to improve long-horizon consistency. โWe optionally add rollout training, as in JEPA-WMโ
- Self-supervised learning: Learning representations from relationships within unlabeled data rather than from manually assigned labels. โThe self-supervised pretrained projector transfers effectively to downstream imitation-learning-based E2E driving.โ
- SIGReg: A statistical regularization method that encourages learned embeddings to match a desired distribution and prevents collapse. โFurthermore, we apply SIGReg to encourage the embeddings to follow an isotropic Gaussian distribution and prevent representation collapse.โ
- Stop gradient: An operation that prevents gradients from propagating through a specified tensor during backpropagation. โWe apply stop gradient to the projected target embedding in the MSE loss.โ
- Teacher forcing: A sequence-model training strategy that conditions each prediction on the ground-truth preceding state rather than solely on the modelโs previous prediction. โThe training objective consists of a teacher-forcing prediction lossโ
- Trajectory vocabulary: A predefined collection of candidate driving trajectories from which a planner selects an action sequence. โwe instead search over the clustered driving trajectory vocabulary as anchorsโ
- ViT (Vision Transformer): A transformer-based vision architecture that processes an image as a sequence of patch tokens. โthe backbone encoder, together with an AdaLN-style predictorโ
- World model: An internal learned representation of an environment that supports prediction, planning, and reasoning about possible future states. โWorld models are internal representations that enable an agent to predict what is likely, plausible, or impossibleโ
- Zero-shot planning: Planning performed without training a task-specific policy for the test scenario or goal. โWithout training any explicit driving policy, the learned world model enables goal-conditioned zero-shot planningโ