Papers
Topics
Authors
Recent
Search
2000 character limit reached

AD-E2E-JEPA: A Joint-Embedding Predictive Architecture For End-to-End Autonomous Driving

Published 28 Sep 2026 in cs.RO, cs.AI, cs.CV, and cs.LG | (2609.34085v1)

Abstract: Autonomous driving requires \textit{world models} that can understand the physical world, reason and plan, and operate safely. In this paper, we first systematically evaluate existing action-conditioned joint-embedding predictive architecture (JEPA) world models, including LeWM, DINO-WM, and JEPA-WM for end-to-end autonomous driving (E2EAD). To isolate world-model quality from policy learning, we employ a goal-conditioned zero-shot planning setting that evaluates these models using ground-truth future observations as goals, without training any driving policy. We find that existing JEPA-based world models are either accurate for driving but computationally expensive, or computationally efficient but insufficient for planning. To address this trade-off, we propose \textbf{AD-E2E-JEPA}, which introduces a SIGReg-regularized learnable projector applied to projected patch embeddings. The projector reduces the number of planning patches by 16ร—16\times and the embedding dimension by 4ร—4\times, achieving a 100ร—100\times inference speedup while retaining planning performance, with a 0.8-second runtime for an 8-frame rollout over 256 candidate trajectories. \textit{Without} training any driving policy, the world model itself reaches the goals located 20 meters away on average within the displacement of respectively 4.0/2.8 meters, using world-model rollouts over trajectory vocabularies of respectively 256/8,192 candidates. On the NAVSIMv2 benchmark, it achieves 67.3/72.9 EPDMS with multiplicative safety metrics and 84.1/86.5 EPDMS<sup>โ€ <sup>{\dagger} without them in goal-conditioned zero-shot planning. Experiments further show that the self-supervised pretrained projector improves downstream imitation learning performance from 80.2 to 85.4 EPDMS. The source code is available at https://github.com/HaoranZhuExplorer/AD-E2E-JEPA

Summary

  • The paper proposes a SIGReg-regularized spatial and channel projector compressing dense visual embeddings to enable zero-shot planning in an autonomous driving world model
  • It introduces a unique approach to predicting future visual representations under ego trajectories, shortening planning time to approximately 0.8 seconds with competitive performance on the NAVSIM benchmark
  • The study demonstrates significant improvements in driving metrics, including EPDMS, geodesic metrics, and hit rates, and shows the benefit of rollout training and data scaling

AD-E2E-JEPA frames end-to-end autonomous driving as a latent-dynamics planning problem rather than exclusively as reactive imitation learning. Its central premise is that an action-conditioned JEPA world model can predict future visual representations under candidate ego trajectories, allowing an autonomous vehicle to select actions by comparing predicted future states with a specified visual goal. The paperโ€™s principal contribution is an efficiency-oriented modification of JEPA-based world modeling: a learnable spatial and channel projector, regularized with SIGReg, compresses dense visual embeddings sufficiently to make zero-shot planning practical while preserving the planning-relevant information of the original representation (2609.34085).

Problem formulation and research objectives

The paper focuses on front-camera-only E2EAD. At each time step, the model receives a temporal context of images and ego poses and predicts future ego poses. Rather than training a policy directly to imitate the demonstrated future trajectory, the proposed system uses the demonstrated ego-motion sequence as action conditioning during world-model training. The distinction is important: human trajectories provide supervision for learning environment dynamics, but the final zero-shot planning experiment does not train a driving policy or trajectory scorer.

The planning task is goal-conditioned. A future observation, specifically an image several frames ahead, is treated as the goal. The model rolls out multiple candidate trajectories and selects the trajectory whose predicted future latent representation is closest to the latent representation of the goal image. This isolates the quality of the learned world model from the quality of a policy trained on expert demonstrations. The evaluation therefore addresses a specific question: whether a JEPA model trained on offline driving data can rank candidate trajectories according to their ability to reach a visual target.

The authors first establish a computational and performance trade-off among existing JEPA-based world models. LeWM uses a global representation and is fast, but its planning accuracy is poor. DINO-WM and JEPA-WM retain dense patch embeddings and provide substantially better planning reliability, but their inference cost is incompatible with efficient autonomous-driving planning. AD-E2E-JEPA is designed to occupy the intermediate regime: it retains dense spatial information in compressed form and uses a trajectory vocabulary instead of iterative CEM optimization.

AD-E2E-JEPA architecture

The baseline world model uses a frozen DINOv3 ViT-L encoder, an AdaLN-style predictor with RoPE, and a linear action encoder. The action at time tt is the relative ego-pose transformation between consecutive frames:

at=[ฮ”xtโ†’t+1,ฮ”ytโ†’t+1,ฮ”ฮธtโ†’t+1]โŠค.\mathbf{a}_t = [\Delta x_{t\rightarrow t+1}, \Delta y_{t\rightarrow t+1}, \Delta\theta_{t\rightarrow t+1}]^\top.

The predictor receives historical visual embeddings and action embeddings and estimates the embeddings of subsequent frames. In the baseline formulation, prediction is performed directly in the DINOv3 embedding space using MSE. Because the target encoder is frozen, the baseline does not require an additional anti-collapse objective.

AD-E2E-JEPA inserts a learnable projector between the frozen DINOv3 encoder and the action-conditioned predictor. The projector consists of two convolutional layers with stride 2ร—22 \times 2. It reduces the spatial number of patch embeddings by 16ร—16\times and reduces the channel dimension from 1024 to 256, yielding a nominal compression of the planning representation by approximately 64ร—64\times in the number of scalar embedding values. The paper attributes the overall planning acceleration primarily to this reduction in the number and dimensionality of tokens.

Compression introduces a collapse risk: distinct DINOv3 embeddings could be mapped to nearly constant projected vectors. The paper addresses this with stop-gradient targets and SIGReg. SIGReg is applied independently at each projected patch location and time step across the batch, using random one-dimensional projections and an Eppsโ€“Pulley normality objective. This differs from formulations that regularize global CLS tokens or aggregate temporal statistics without preserving patchwise structure. The design assumption is that an approximately isotropic Gaussian distribution at each spatial-temporal location maintains sufficient information for prediction while preventing degenerate representations.

The resulting objective combines projected-space prediction loss with SIGReg. An optional rollout objective extends supervision across the full eight-frame future horizon. Teacher-forced predictions are supplemented by autoregressive predictions, with truncated backpropagation through time and stop-gradient through earlier rollout contexts. This explicitly penalizes compounding prediction errors, which is particularly relevant because the planning procedure evaluates multi-step world-model rollouts rather than only one-step predictions.

Zero-shot planning procedure

The planning system uses a fixed vocabulary of up to 8192 clustered driving trajectories. These trajectories are sorted by angular coordinate and can be subsampled to 256, 512, 1024, 2048, or 4096 candidates. For each candidate, the world model predicts the future latent representation under the corresponding action sequence. The selected trajectory minimizes squared latent distance to the latent embedding of the goal image.

This vocabulary-based search replaces CEM, whose iterative evaluation of large candidate populations is considered too expensive for the target setting. The resulting procedure is simple and parallelizable: all candidate rollouts can be evaluated simultaneously on an A100 GPU. It also exposes a direct trade-off between search granularity and inference time. Increasing the vocabulary from 256 to 8192 improves final-pose accuracy and EPDMS, but linearly increases planning latency.

The evaluation reports three classes of metrics. EPDMS measures NAVSIMv2 driving quality and safety, while EPDMSโ€ ^\dagger excludes multiplicative safety terms and therefore reflects the quality component more directly. Geodesic metrics measure final displacement, longitudinal and lateral position errors, and heading error. Top-1 and top-5 hit rates measure whether the ground-truth trajectory is ranked among the lowest-cost candidates. The latter metrics are especially informative for diagnosing the world model itself, because they assess whether latent distance assigns low cost to the demonstrated trajectory even when the selected trajectory is not exactly identical to it.

Comparative planning results

The comparison with prior world models supports the paperโ€™s main efficiency claim. On 100 subsampled NAVSIM test scenes, LeWM requires only 0.7 seconds per scene, but obtains an EPDMS of 48.3, an FDE of 12.4 meters, and Top-1/Top-5 hit rates of 6%/18%. Dense-feature models are much more accurate: DINO-WM reaches 68.3 EPDMS and JEPA-WM reaches 74.2, with Top-1/Top-5 hit rates of 40%/73% and 45%/75%, respectively. However, their planning times are 91.8 and 101.0 seconds per scene.

AD-E2E-JEPA reaches 76.6 EPDMS with 0.8 seconds of planning time in the same 100-scene setting. Its rollout-trained variant reaches 70.4 EPDMS, with improved FDE from 4.2 to 3.5 meters and improved Top-1/Top-5 hit rates from 27%/59% to 34%/67%. The apparent reduction in EPDMS despite better geodesic and reliability metrics is one of the paperโ€™s important non-monotonic results. It indicates that the latent-distance objective is not fully aligned with the multiplicative safety structure of NAVSIMv2. A world model can become better at identifying the ground-truth future and still produce lower safety-weighted performance on a small evaluation subset.

The strongest evidence appears on the full 12,146-scene test set. The trainval-trained rollout variant obtains an EPDMS of 67.3 and EPDMSโ€ ^\dagger of 84.1, with an FDE of 4.0 meters and Top-1/Top-5 hit rates of 53.8%/82.7%. In the same broad setting, LeWM obtains 39.8 EPDMS, 66.7 EPDMSโ€ ^\dagger, 14.6-meter FDE, and 5.7%/11.3% hit rates. Thus, the proposed compressed representation substantially improves both planning accuracy and reliability relative to the efficient baseline.

The candidate-vocabulary ablation shows a consistent accuracy-latency trade-off:

Candidate trajectories Planning time EPDMS EPDMSโ€ ^\dagger FDE Heading error
256 0.8 s 67.3 84.1 4.0 m 3.5ยฐ
512 1.4 s 69.2 85.0 3.6 m 2.9ยฐ
1,024 2.5 s 70.5 85.5 3.2 m 2.5ยฐ
2,048 4.7 s 71.5 86.0 3.0 m 2.2ยฐ
4,096 9.3 s 72.1 86.3 2.9 m 2.1ยฐ
8,192 18.2 s 72.9 86.5 2.8 m 2.0ยฐ

The 8192-trajectory model achieves the paperโ€™s best reported planning result: 72.9 EPDMS, 86.5 EPDMSโ€ ^\dagger, 2.8-meter FDE, and 2.0-degree mean heading error. The implication is that AD-E2E-JEPA is not merely fast at a fixed accuracy level; it provides a tunable compute-performance curve. At 256 candidates, it performs an eight-frame rollout in approximately 0.8 seconds, compared with roughly 92โ€“101 seconds for DINO-WM and JEPA-WM. The paper characterizes this as a at=[ฮ”xtโ†’t+1,ฮ”ytโ†’t+1,ฮ”ฮธtโ†’t+1]โŠค.\mathbf{a}_t = [\Delta x_{t\rightarrow t+1}, \Delta y_{t\rightarrow t+1}, \Delta\theta_{t\rightarrow t+1}]^\top.0 speedup, although the exact ratio relative to the reported dense-model runtimes is closer to at=[ฮ”xtโ†’t+1,ฮ”ytโ†’t+1,ฮ”ฮธtโ†’t+1]โŠค.\mathbf{a}_t = [\Delta x_{t\rightarrow t+1}, \Delta y_{t\rightarrow t+1}, \Delta\theta_{t\rightarrow t+1}]^\top.1โ€“at=[ฮ”xtโ†’t+1,ฮ”ytโ†’t+1,ฮ”ฮธtโ†’t+1]โŠค.\mathbf{a}_t = [\Delta x_{t\rightarrow t+1}, \Delta y_{t\rightarrow t+1}, \Delta\theta_{t\rightarrow t+1}]^\top.2.

The reliability results require more careful interpretation. Although increasing the number of candidates improves EPDMS and final-pose accuracy, the reported Top-1 and Top-5 hit rates decrease from 53.8%/82.7% at 256 candidates to 15.3%/33.8% at 8192 candidates. This is not necessarily a contradiction: the hit-rate definition requires the ground-truth trajectory to rank among the top at=[ฮ”xtโ†’t+1,ฮ”ytโ†’t+1,ฮ”ฮธtโ†’t+1]โŠค.\mathbf{a}_t = [\Delta x_{t\rightarrow t+1}, \Delta y_{t\rightarrow t+1}, \Delta\theta_{t\rightarrow t+1}]^\top.3 candidates in a larger set, making the ranking problem harder as more alternatives are added. The result suggests that aggregate driving metrics and exact candidate-ranking reliability measure different properties of the world model.

Role of rollout training and data scale

Rollout training is a major component of the final system. Without rollout loss, AD-E2E-JEPA trained on trainval obtains 63.2 EPDMS, 80.1 EPDMSat=[ฮ”xtโ†’t+1,ฮ”ytโ†’t+1,ฮ”ฮธtโ†’t+1]โŠค.\mathbf{a}_t = [\Delta x_{t\rightarrow t+1}, \Delta y_{t\rightarrow t+1}, \Delta\theta_{t\rightarrow t+1}]^\top.4, and 6.2-meter FDE on the full test set. Adding rollout training increases performance to 67.3 EPDMS and 84.1 EPDMSat=[ฮ”xtโ†’t+1,ฮ”ytโ†’t+1,ฮ”ฮธtโ†’t+1]โŠค.\mathbf{a}_t = [\Delta x_{t\rightarrow t+1}, \Delta y_{t\rightarrow t+1}, \Delta\theta_{t\rightarrow t+1}]^\top.5, while reducing FDE to 4.0 meters and increasing the Top-1/Top-5 hit rate to 53.8%/82.7%. The improvement supports the claim that one-step latent prediction is insufficient for planning over a four-second horizon.

The authors also scale training data from the 10-hour navtrain split to the 70-hour trainval training portion. The benefit is clearest in combination with rollout training and full-scene evaluation, where the larger configuration attains the strongest overall result. However, the paper does not isolate the effects of data scale, projector design, rollout training, and hyperparameter changes through a complete factorial ablation. The reported trainval configuration also uses different batch sizes and SIGReg weights from the navtrain configuration. Consequently, the performance gain cannot be attributed exclusively to the additional driving data.

Transfer to imitation learning

The pretrained projector is evaluated separately as a representation for conventional imitation learning. The authors discard the world-model predictor, retain the DINOv3 encoder and learned projector, and attach a trajectory decoder based on temporal and spatial positional embeddings, command conditioning, and cross-attention. The entire model is fine-tuned using trajectory MSE.

The pretrained projector improves the reported EPDMS from 80.2 with a randomly initialized projector to 85.4, a 5.2-point gain. The component-wise NAVSIM metrics also improve, including no-collision, drivable-area compliance, driving-direction compliance, time-to-collision, lane keeping, history comfort, and extended comfort. In the benchmark comparison, the pretrained-projector model obtains 97.7 NC, 93.7 DAC, 99.2 DDC, 99.8 TLC, 87.3 EP, 96.8 TTC, 97.0 LK, 98.4 HC, and 88.7 EC, compared with 96.8, 89.8, 98.3, 99.7, 87.1, 95.7, 94.8, 98.3, and 84.1 for the random-projector variant.

This transfer result is conceptually significant because it isolates a component often obscured in end-to-end systems. The projector is not trained with trajectory labels, yet its world-model objective produces a representation that improves downstream supervised driving. The result supports the narrower claim that predictive latent compression can encode task-relevant structure beyond the immediate zero-shot planning task. It does not establish that the projector alone is sufficient for robust driving, since the encoder is DINOv3, the complete model is fine-tuned, and the evaluation remains within the NAVSIM benchmark.

Limitations and open questions

The experimental setting deliberately simplifies autonomous driving. It uses only a front camera, four historical frames, fixed candidate trajectories, and a future image as the goal. The world model therefore does not address multi-camera fusion, LiDAR, explicit map conditioning, dynamic-agent interaction, closed-loop execution, or uncertainty-aware planning. NAVSIM is a pseudo-simulation benchmark rather than a deployment evaluation, so the reported EPDMS values should not be interpreted as evidence of real-world safety.

The goal-conditioned formulation also supplies privileged information: the planner is given the ground-truth future image. In ordinary autonomous driving, a future visual goal is not generally available at planning time. The experiment is valuable for isolating world-model planning capacity, but it does not demonstrate autonomous destination selection or ordinary reactive driving. The safety metrics are likewise not optimized by the latent-distance objective. The paper explicitly reports EPDMSat=[ฮ”xtโ†’t+1,ฮ”ytโ†’t+1,ฮ”ฮธtโ†’t+1]โŠค.\mathbf{a}_t = [\Delta x_{t\rightarrow t+1}, \Delta y_{t\rightarrow t+1}, \Delta\theta_{t\rightarrow t+1}]^\top.6 because zero-shot planning does not directly enforce collision avoidance, drivable-area compliance, traffic-light compliance, or other safety constraints.

Several methodological comparisons remain incomplete. DINO-WM and JEPA-WM are evaluated on only 100 scenes because dense planning is prohibitively expensive, whereas AD-E2E-JEPA and LeWM are evaluated on all 12,146 scenes. This makes the full-test comparison against dense baselines unavailable. In addition, the 100-scene subset produces substantial variance, as reflected by differences between subset and full-test rankings. The paper also does not provide a complete ablation separating spatial compression, channel compression, SIGReg, stop-gradient, predictor architecture, rollout training, and trajectory-vocabulary search. The claim that SIGReg preserves planning information is supported by the aggregate results but not by a systematic collapse or information-retention analysis.

The relationship between latent distance and safety remains open. Increasing candidate density improves geometric accuracy and EPDMS, but decreases hit rates under the paperโ€™s ranking definition. This suggests that the latent metric may support approximate trajectory selection without inducing a stable ordering over increasingly fine candidate sets. Whether alternative distance functions, calibrated uncertainty, safety-aware latent objectives, or hierarchical planning can resolve this discrepancy is not established.

Conclusion

AD-E2E-JEPA presents an action-conditioned JEPA world model for E2EAD in which a SIGReg-regularized projector compresses dense DINOv3 patch embeddings before latent rollout prediction. The design addresses the central computational weakness of dense JEPA planning: DINO-WM and JEPA-WM achieve strong planning quality but require approximately 1โ€“2 minutes per scene, whereas AD-E2E-JEPA performs an eight-frame, 256-candidate rollout in approximately 0.8 seconds while retaining competitive accuracy. On 12,146 NAVSIM test scenes, the rollout-trained model reaches 67.3 EPDMS and 84.1 EPDMSat=[ฮ”xtโ†’t+1,ฮ”ytโ†’t+1,ฮ”ฮธtโ†’t+1]โŠค.\mathbf{a}_t = [\Delta x_{t\rightarrow t+1}, \Delta y_{t\rightarrow t+1}, \Delta\theta_{t\rightarrow t+1}]^\top.7; with 8192 candidates, it reaches 72.9 and 86.5, respectively, with 2.8-meter displacement and 2.0-degree heading error. The pretrained projector also improves imitation-learning performance from 80.2 to 85.4 EPDMS. The results support compressed latent world modeling as a practical planning mechanism, while leaving unresolved the transfer from privileged goal-conditioned evaluation to closed-loop, safety-constrained autonomous driving.

Paper to Video (Beta)

No one has generated a video about this paper yet.

Whiteboard

No one has generated a whiteboard explanation for this paper yet.

Explain it Like I'm 14

1. What is this paper about?

This paper presents AD-E2E-JEPA, a new artificial intelligence system for autonomous driving.

Most self-driving systems learn by copying human drivers. The car sees what a human did in a situation and tries to do the same thing. This is called imitation learning.

The researchers wanted to build something more thoughtful. Instead of only copying human actions, their system tries to build a kind of mental model of the driving world. This model predicts what might happen if the car chooses different actions, such as turning left, turning right, or continuing straight.

The system can then choose the action that is most likely to lead to a desired future situation.

2. What questions did the researchers ask?

The paper focuses on several main questions:

  • Can a JEPA-based world model help a self-driving car plan its movements?
  • Can the car plan successfully without first being trained to copy a driving policy?
  • Can the system make accurate plans while working quickly enough for driving?
  • Can the modelโ€™s learned knowledge also improve ordinary imitation-learning systems?
  • How does AD-E2E-JEPA compare with other world models?

A major problem with earlier systems was a trade-off:

  • Some models planned accurately but took more than a minute to evaluate one scene.
  • Other models were fast but made poor plans.

The researchers wanted a model that was both accurate and fast.

3. How does the method work?

What is a world model?

A world model is an internal picture of how the world works. For example, a person crossing the road will probably keep moving forward, while a parked car will probably stay still.

For a self-driving car, a world model should understand things such as:

  • Where the road is
  • How the car is moving
  • What may happen after turning or braking
  • Which future situations are safe or useful

An analogy is a video game. Before moving a character, you might imagine what will happen if you press different buttons. A world model does something similar for a car.

What is JEPA?

JEPA, or Joint-Embedding Predictive Architecture, does not try to recreate every future camera image pixel by pixel. Instead, it converts images into shorter descriptions called embeddings.

An embedding is like a compact summary of an image. Instead of remembering every detail, it might represent information such as:

โ€œThere is a road ahead, a vehicle on the right, and the car is approaching an intersection.โ€

The model learns to predict the embedding of a future scene based on:

  1. Recent camera images
  2. The carโ€™s recent movements
  3. A possible future driving path

This allows the model to compare different possible actions without generating complete future images.

Compressing the information

Camera images contain a large amount of information. The model divides each image into small areas, or patches, and creates an embedding for each patch.

Using all these patches would make planning slow. AD-E2E-JEPA adds a learnable projector, which acts like a smart compression tool. It reduces:

  • The number of patches by 16 times
  • The size of each embedding by 4 times

The researchers also use a technique called SIGReg. Its purpose is to make sure the compressed information does not become too similar everywhere. In simple terms, SIGReg helps the model keep useful differences between scenes instead of turning every scene into nearly the same description.

How the car plans

The system is given a goal, represented by a future camera image. It then considers many possible driving paths.

For each path, the world model predicts what the future scene would look like in embedding form. The system compares each prediction with the embedding of the goal image.

It chooses the path whose predicted future is most similar to the goal.

This is called zero-shot goal-conditioned planning:

  • Goal-conditioned means the system plans toward a specified goal.
  • Zero-shot means it does this without separately training a driving policy for that task.

An everyday analogy would be choosing a route in a city. You imagine what each route will lead to, compare the results with where you want to go, and choose the best route.

Rollout training

The researchers also tested rollout training. Instead of predicting only one future step, the model repeatedly predicts several steps into the future.

This is similar to thinking:

โ€œIf I turn now, where will I be next? And after that, where will I be?โ€

This helps the model make more consistent long-term plans.

4. What were the main findings?

The new system was much faster

With 256 possible driving paths, AD-E2E-JEPA needed about 0.8 seconds per scene.

The other accurate models, DINO-WM and JEPA-WM, needed about 92 to 101 seconds per scene.

Therefore, AD-E2E-JEPA was roughly 100 times faster while keeping similar planning quality.

This is important because a self-driving car must make decisions quickly. Waiting over a minute to choose a path would not be practical on a real road.

It planned more accurately than the simple fast model

The older LeWM model was also fast, taking about 0.7 seconds per scene. However, its plans were much less accurate.

On the full test set:

Model Planning time Final displacement error
LeWM 0.7 seconds 14.6 meters
AD-E2E-JEPA with rollout training 0.8 seconds 4.0 meters

The final displacement error measures how far the selected final position was from the correct position. A smaller number is better.

This means AD-E2E-JEPA was almost as fast as LeWM but produced much more accurate paths.

More possible paths improved performance

The researchers tested different numbers of candidate paths. With 8,192 possible trajectories, the best version achieved:

  • 72.9 EPDMS with safety measures
  • 86.5 EPDMS without those safety measures
  • 2.8 meters of average final-position error
  • 2.0 degrees of average heading error

Using more candidate paths gave the model more choices, helping it find a path closer to the correct one. However, this also increased planning time, from 0.8 seconds with 256 paths to 18.2 seconds with 8,192 paths.

Training on more driving data helped

The researchers trained some versions on about 10 hours of driving video and others on about 70 hours.

The model trained with more data and rollout training was generally more reliable and accurate. On the full test set, it identified the correct driving path among its best five choices about 83% of the time.

The learned projector helped another driving system

The researchers also tested whether the projectorโ€™s learned knowledge could help a normal imitation-learning system.

They compared:

  • A system with a randomly initialized projector
  • A system using the projector pretrained by AD-E2E-JEPA

The pretrained projector improved the EPDMS score from 80.2 to 85.4.

This suggests that the projector learned useful information about driving, even when used in a different type of model.

5. Why are these results important?

The results suggest that a self-driving car does not always need to directly copy human drivers. It may instead learn how actions affect the future and use that knowledge to plan.

The main contribution is the balance between:

  • Speed
  • Accuracy
  • Useful understanding of the driving environment

The system also shows that a model can learn helpful knowledge without requiring detailed labels for every possible driving decision.

6. Possible impact and limitations

If developed further, this approach could help create self-driving systems that:

  • Make decisions more quickly
  • Consider many possible driving choices
  • Adapt better to unfamiliar situations
  • Use less expensive computation during planning
  • Improve other driving systems through transfer learning

However, the results are still based on the NAVSIMv2 benchmark, which is a dataset and computer-based evaluation rather than complete real-world driving. The system mainly uses a front-facing camera and does not prove that it is ready for public roads.

The model also plans by choosing from a fixed collection of candidate trajectories. Real driving may require actions that are not included in that collection. In addition, the reported safety scores do not mean the system is guaranteed to be safe in every unusual situation.

Conclusion

AD-E2E-JEPA is a fast world model for autonomous driving. It learns to predict how different driving actions will change the future scene. Instead of simply copying human drivers, it compares possible paths and chooses one that should lead toward a desired goal.

The researchersโ€™ main innovation is a smart compression system that makes planning about 100 times faster than some earlier accurate methods. The model also produced more accurate driving plans and improved a separate imitation-learning system.

Overall, the paper shows that learning a compact model of how the world changes could be an important step toward faster and more flexible autonomous vehicles.

Knowledge Gaps

Knowledge gaps, limitations, and open questions

  • Limited sensor modality: The study evaluates only a front-camera input, leaving it unclear whether AD-E2E-JEPA generalizes to multi-camera, LiDAR, radar, depth, or vehicle-state sensor configurations.
  • Restricted environmental diversity: All experiments use NAVSIM/NAVSIMv2 data, so robustness across cities, countries, weather conditions, lighting, road geometries, traffic cultures, and sensor distributions remains untested.
  • No real-world or closed-loop validation: The reported results come from offline data and NAVSIM pseudo-simulation; the modelโ€™s behavior under real vehicle dynamics, feedback errors, perception shifts, and interactive traffic is unresolved.
  • No safety-oriented planning mechanism: Planning minimizes latent distance to a future image and does not explicitly model collision probability, traffic-rule compliance, emergency maneuvers, uncertainty, or risk-sensitive objectives.
  • Safety metrics are not optimized directly: The strongest reported performance relies partly on EPDMSโ€ , which excludes multiplicative safety terms, and the paper does not establish whether high latent-goal matching translates into reliable collision avoidance.
  • Dependence on privileged future observations: Zero-shot planning uses the ground-truth future image as the goal, an assumption unavailable during ordinary autonomous driving. Performance with predicted, noisy, language-based, map-based, or user-specified goals is not evaluated.
  • Unclear deployability without future goals: The paper does not specify how goals would be generated online in a complete autonomous-driving system or how the planner would operate when no target future observation is available.
  • Action-space coverage is limited: Candidate actions are selected from a fixed vocabulary of 256โ€“8,192 clustered trajectories. The methodโ€™s ability to discover trajectories outside this vocabulary, including rare evasive or highly maneuverable actions, remains unknown.
  • Scalability trade-off is unresolved: Increasing the vocabulary to 8,192 improves EPDMS and displacement accuracy but raises planning time from 0.8 to 18.2 seconds per scene, which may be too slow for real-time control. The practical latencyโ€“performance frontier is not established.
  • No temporal receding-horizon evaluation: The planner selects a trajectory for a future horizon but is not evaluated in a repeated closed-loop receding-horizon setting, where model errors accumulate and previous actions alter subsequent observations.
  • Long-horizon prediction remains uncertain: The experiments use a 4-second future horizon. Reliability over longer horizons, especially for intersections, lane changes, merges, and multi-stage maneuvers, is not assessed.
  • Autoregressive error accumulation is not characterized: Rollout training improves several metrics, but the paper does not quantify how prediction error grows across each future timestep or identify when autoregressive rollouts become unreliable.
  • Limited evaluation of dynamic-agent interaction: The model is trained and evaluated on logged human driving data, but its ability to predict reactions from pedestrians, cyclists, and other vehicles under interventions different from the logged trajectory is not demonstrated.
  • Offline data coverage may induce distributional bias: Because the action-conditioned model learns primarily from human driving trajectories, it may not accurately predict consequences of uncommon, unsafe, or counterfactual actions absent from the dataset.
  • Causal validity of the learned dynamics is unverified: Matching future latent embeddings does not establish that the model has learned physically or causally correct environment dynamics rather than exploiting visual correlations in the dataset.
  • Latent-distance objective may be misaligned with driving quality: The authors note that latent distance appears more directly related to geodesic accuracy and hit rate than to NAVSIM driving metrics. The relationship between embedding similarity, comfort, legality, progress, and safety remains insufficiently understood.
  • Representation compression is not fully explained: The paper shows that a projector reducing patch count by 16ร—16\times and dimension by 4ร—4\times can preserve performance, but it does not identify which spatial, semantic, or temporal information is discarded.
  • SIGRegโ€™s contribution is not isolated comprehensively: The reported gains may result from the projector architecture, dimensionality reduction, patch reduction, stop-gradient training, SIGReg, or their interactions. A complete factorial ablation is needed to attribute improvements reliably.
  • Sensitivity to SIGReg hyperparameters is unclear: The SIGReg weight is tuned heuristically according to batch size, but robustness to ฮป\lambda, batch size, number of random directions, and knot configuration is not systematically reported.
  • Potential instability across seeds is unreported: Results lack confidence intervals, multiple random seeds, or statistical significance tests, particularly for the 100-scene subset where the authors acknowledge substantial variance.
  • Small-subset comparisons may be unreliable: Comparisons among methods are performed on only 100 test scenes because of computational cost, while the full-set evaluation includes only selected methods. Full-test comparisons against DINO-WM and JEPA-WM remain unavailable.
  • Baseline fairness is incomplete: The baselines use different backbones, batch sizes, training resources, and configurations. It is unclear whether all models receive equally optimized hyperparameters, equivalent training data, and comparable compute budgets.
  • Backbone dependence is unexamined: The method relies on a frozen DINOv3 ViT-L encoder. Its effectiveness with smaller, independently trained, domain-adapted, or newer visual encoders is unknown.
  • Frozen-encoder limitations are unresolved: The study does not test whether jointly adapting the encoder improves robustness to driving-specific visual phenomena, domain shift, or rare traffic events.
  • No comparison with strong online planners or policy baselines in the same setting: The paper does not directly compare against modern model-predictive control, diffusion planners, trajectory-ranking systems, or learned policies under identical sensor, data, and compute constraints.
  • Transfer-learning evidence is narrow: Projector transfer is evaluated using one simple imitation-learning decoder and a random-projector baseline. It is unclear whether the benefit persists across decoder architectures, training regimes, task objectives, or policy-learning methods.
  • The transfer improvement is not disentangled from optimization effects: The pretrained projector may provide better initialization, scale normalization, or regularization rather than predictive driving structure. Additional controls, such as frozen versus trainable components and matched parameterizations, are needed.
  • No cross-domain transfer evaluation: The downstream transfer experiment remains within NAVSIM. Transfer to other datasets, cities, sensor configurations, or real-world driving domains is not established.
  • Uncertainty estimation is absent: The world model produces point predictions and a single latent-distance score, without calibrated uncertainty over future states, candidate trajectories, or model confidence.
  • Planner reliability metrics are incomplete: Top-1 and top-5 hit rates measure ranking of the logged ground-truth trajectory, but do not assess whether alternative trajectories are also safe, feasible, or better than the recorded human action.
  • Ground-truth trajectory quality is assumed: Human demonstrations are treated as the target trajectory even though the introduction acknowledges that they may be noisy or suboptimal. The method is not evaluated against expert-optimal, safety-verified, or multi-modal trajectory annotations.
  • Heading and pose representation issues remain: The formulation uses planar poses and relative heading changes, but does not discuss angle periodicity, accumulated pose drift, or whether this representation adequately captures vehicle kinematics and nonholonomic constraints.
  • Vehicle dynamics and comfort constraints are under-modeled: The candidate trajectory vocabulary and latent planner are not shown to enforce acceleration, jerk, steering-rate, tire-friction, or passenger-comfort limits.
  • No analysis of failure cases: The paper provides qualitative examples but does not systematically categorize failures, such as occlusions, ambiguous goals, unusual road layouts, sudden obstacles, mislocalized poses, or visually similar but behaviorally different scenes.
  • Runtime measurements are hardware- and implementation-dependent: The reported speedups are measured on an A100 and do not include the full perception, goal-generation, control, and vehicle-interface pipeline. Real-time performance on embedded automotive hardware remains unknown.
  • Compute and memory costs are incompletely reported: The paper reports planning time and some training configurations but does not provide a full accounting of model parameters, GPU memory, energy use, preprocessing costs, or deployment constraints.
  • Reproducibility is affected by unclear experimental details: The manuscript contains notation and formatting inconsistencies, and some training, evaluation, and benchmark definitions are deferred to appendices or referenced implementations. Exact preprocessing, trajectory clustering, candidate sampling, and checkpoint-selection procedures require fuller specification.
  • Generalization to unseen goals is not demonstrated: Although the paper describes zero-shot planning, goals are drawn from ground-truth future observations in the same benchmark distribution. Generalization to novel visual goals, altered target locations, and compositional goals remains open.
  • Interactions between goal specification and representation compression are unknown: It is unclear whether the compressed projector preserves information needed for goals involving objects, lanes, traffic signals, or spatial relations that are not represented by the logged future frame.
  • The limits of JEPA versus generative world models remain unresolved: The paper discusses generative video world models but does not experimentally determine when latent prediction is preferable to pixel/video generation for physical consistency, planning, interpretability, or uncertainty modeling.
  • No interpretability or semantic probing is provided: The paper does not establish which latent dimensions or patches encode road structure, agent motion, traffic signals, obstacles, or action consequences, limiting understanding of why compressed embeddings support planning.
  • Policy-free planning performance does not establish complete autonomy: The experiments isolate world-model quality by omitting policy training, but they do not show how AD-E2E-JEPA integrates with a controller or whether its selected trajectories can be executed reliably under perception and actuation delays.**

Practical Applications

Immediate Applications

  • Real-time goal-conditioned planning for autonomous vehicles โ€” automotive and mobility
    • Integrate AD-E2E-JEPA as a planning component that evaluates candidate ego-vehicle trajectories against a desired future camera observation or navigation target.
    • The reported approximately 0.8 s planning time for an 8-frame rollout over 256 candidates makes the method suitable for offline evaluation, closed-course testing, and potentially low-frequency online replanning. Larger candidate sets improve accuracy but increase latency, reaching approximately 18.2 s for 8,192 candidates on an A100 GPU.
    • A practical workflow is: encode recent camera frames and poses, roll out the world model for each trajectory in a predefined vocabulary, compare predicted and target latent embeddings, and execute the lowest-cost trajectory.
    • Dependencies: reliable camera calibration, ego-pose estimation, a sufficiently representative trajectory vocabulary, GPU or specialized accelerator deployment, and an independent safety layer. The paper evaluates zero-shot planning with ground-truth future images as goals, which are not directly available during ordinary autonomous driving.
  • Efficient world-model benchmarking and planner selection โ€” automotive R&D and academia
    • Use the released implementation as a benchmark for comparing latent world models on planning efficiency, final displacement error, heading error, hit rate, and NAVSIM EPDMS.
    • The method enables full-test-set evaluation that is impractical for dense-patch baselines: AD-E2E-JEPA evaluates 12,146 scenes at roughly 0.8 s per scene, whereas DINO-WM and JEPA-WM require substantially more computation.
    • This can support engineering decisions about the trade-off between latent representation size, candidate-trajectory count, safety performance, and inference latency.
    • Dependencies: NAVSIM-compatible data and metrics, comparable hardware settings, and careful separation of zero-shot planning performance from policy-learning performance.
  • Self-supervised pretraining for end-to-end driving systems โ€” automotive software
    • Use the learned projector as an initialization for an imitation-learning driving model. The reported improvement from 80.2 to 85.4 EPDMS relative to a randomly initialized projector indicates that predictive latent structure can improve downstream trajectory decoding.
    • A deployable training workflow is to freeze or fine-tune the DINOv3 encoder, initialize the projector from AD-E2E-JEPA, attach a trajectory decoder, and fine-tune on driving demonstrations.
    • This may reduce dependence on dense trajectory labels or extensive supervised pretraining, while making existing imitation-learning systems more robust to temporal context.
    • Dependencies: transferability beyond NAVSIM, compatibility between sensor configuration and pretrained representations, and sufficient supervised data for the final policy.
  • Compute-efficient latent representation compression โ€” machine-learning infrastructure
    • Apply the projector design to reduce dense visual embeddings before downstream planning or control. The method compresses spatial patch count by 16ร— and embedding dimension by 4ร—, while preserving much of the planning quality through SIGReg regularization.
    • The resulting component could be packaged as a reusable latent-compression module for autonomous-driving stacks, robotics systems, or video-based control models.
    • Dependencies: SIGReg hyperparameter tuning, adequate batch sizes, preservation of task-relevant spatial information, and validation against representation collapse and rare-event performance degradation.
  • Offline trajectory ranking and safety analysis โ€” fleet operators and testing organizations
    • Use the world model to rank candidate trajectories in recorded driving scenarios, identify situations where the ground-truth trajectory is absent from the top-ranked candidates, and generate failure cases for regression testing.
    • Top-1 and top-5 hit rates provide a practical diagnostic for whether the learned latent dynamics assign low cost to plausible trajectories.
    • This can support dataset curation, planner debugging, scenario prioritization, and comparison of software releases without deploying the model on public roads.
    • Dependencies: hit rate is not equivalent to collision avoidance or legal compliance; recorded scenarios must include sufficient environmental diversity, and safety-critical evaluation should include explicit collision, traffic-rule, and uncertainty checks.
  • Policy and standards evaluation for efficient autonomous-driving systems โ€” regulators and public agencies
    • Agencies can use latent-world-model benchmarks to evaluate computational efficiency, generalization, reliability, and safety-related planning behavior in a standardized offline setting.
    • The paperโ€™s distinction between EPDMS with and without multiplicative safety terms is useful for separating general driving quality from explicit safety performance.
    • Dependencies: NAVSIM and pseudo-simulation metrics are not substitutes for real-world validation, certification, or liability assessment. Regulatory use would require scenario coverage, interpretability, uncertainty reporting, and hardware-in-the-loop or closed-track testing.
  • Research and teaching tool for model-based decision-making โ€” academia and education
    • AD-E2E-JEPA provides a concrete example for courses and laboratories covering self-supervised learning, JEPA architectures, latent dynamics, model-based planning, trajectory search, and representation regularization.
    • Students can reproduce ablations involving projector compression, SIGReg, rollout training, candidate-vocabulary size, and downstream transfer.
    • Dependencies: the paperโ€™s code and pretrained models must remain available, and experiments may require high-memory GPUs such as A100-class hardware.

Long-Term Applications

  • On-road autonomous driving with closed-loop replanning โ€” automotive and transportation
    • A mature version could serve as a compact model-based planner that repeatedly predicts the consequences of candidate actions and replans as new observations arrive.
    • Unlike purely reactive imitation, the architecture could evaluate alternative futures before acting, potentially improving behavior in unfamiliar scenes and reducing dependence on demonstrator quality.
    • A production workflow would combine the JEPA world model with perception, localization, route goals, uncertainty estimation, an explicit collision checker, traffic-rule reasoning, and a low-level controller.
    • Dependencies: the current evaluation uses a front camera, offline ground-truth future goals, a fixed trajectory vocabulary, and pseudo-simulation. Deployment requires multi-camera or multimodal sensing, real-time worst-case latency guarantees, robustness to weather and sensor corruption, calibrated uncertainty, and extensive closed-loop validation.
  • Language- or map-conditioned driving goals โ€” automotive navigation and humanโ€“vehicle interaction
    • The image-goal formulation could be extended to goals specified by HD maps, route waypoints, semantic targets, or natural-language instructions such as โ€œmove to the next safe laneโ€ or โ€œstop near the marked entrance.โ€
    • A product could use a goal encoder to convert route or user intent into the same latent space used for trajectory comparison.
    • Dependencies: goal representations must be aligned with the predictive latent space; semantic intent, traffic rules, and long-horizon route planning cannot be assumed to emerge from image-latent matching alone.
  • Robotics navigation and manipulation โ€” robotics
    • The projector-plus-SIGReg design could be adapted to action-conditioned visual world models for mobile robots, warehouse vehicles, drones, or manipulation systems.
    • Candidate action sequences could be ranked by the latent distance between predicted observations and a goal image, enabling reward-free or low-reward planning from offline demonstrations.
    • Dependencies: robot dynamics, camera viewpoint, action parameterization, contact physics, and environment changes differ substantially from driving. Real-robot deployment would require uncertainty-aware planning and recovery behavior.
  • Cross-domain latent planners for energy and industrial control โ€” energy and manufacturing
    • Similar compressed JEPA world models could forecast latent system states under candidate control sequences for battery management, HVAC optimization, warehouse automation, or process control.
    • The approach is particularly relevant where full-fidelity simulation is expensive but historical sensor-action trajectories are available.
    • Dependencies: domain-specific safety constraints, reliable action coverage in offline data, observability of system state, and proof that latent distance correlates with operational objectives rather than merely visual similarity.
  • Large-scale fleet learning and continual adaptation โ€” transportation platforms
    • Fleet operators could periodically retrain the projector and predictive model on newly collected edge cases, then transfer updated representations to downstream driving policies.
    • Adaptive updates could target unusual weather, road layouts, construction zones, or regional driving patterns without retraining an entire end-to-end system from scratch.
    • Dependencies: data governance, privacy protection, distribution shift, prevention of catastrophic forgetting, version validation, and mechanisms ensuring that improvements in average metrics do not reduce rare-event safety.
  • Safety-aware model-based planning โ€” autonomous systems and policy
    • Future systems could combine latent goal matching with explicit constraints for collision probability, time-to-collision, lane boundaries, pedestrian interactions, comfort, and traffic-law compliance.
    • This would address a central limitation of the current method: zero-shot latent-distance planning does not explicitly optimize safety, and the strongest reported EPDMS values without safety terms should not be interpreted as deployment readiness.
    • Dependencies: calibrated predictive uncertainty, sufficiently accurate world-model rollouts for rare events, formal or statistical safety guarantees, and validated integration with independent safety monitors.
  • Hierarchical long-horizon autonomy โ€” robotics and autonomous vehicles
    • AD-E2E-JEPA could become a low-level or mid-level component in a hierarchical planner: a route planner selects semantic subgoals, while the JEPA model selects short-horizon trajectories that achieve each subgoal.
    • This could extend the current 4-second future horizon to complex maneuvers such as intersections, parking, merging, and multi-step navigation.
    • Dependencies: temporal abstraction, subgoal generation, error accumulation over repeated rollouts, and mechanisms for recovering when the desired goal is not reachable within the candidate vocabulary.
  • Simulation, digital-twin, and synthetic-data generation โ€” automotive research
    • Although the model predicts latent embeddings rather than pixels, it could provide a computationally efficient predictive layer for evaluating hypothetical actions in digital twins or for selecting scenarios that require expensive photorealistic simulation.
    • It may also help identify informative trajectories and hard cases for targeted data collection.
    • Dependencies: latent predictions are not directly usable as visual simulation outputs; coupling to a generative decoder or external simulator would require further research and could introduce additional model errors.
  • Commercial planning and perception software components โ€” AI tooling
    • The projector, SIGReg objective, latent rollout evaluator, trajectory vocabulary search, and reliability metrics could be released as modular libraries or accelerator kernels for model-based control.
    • Such tools could support rapid prototyping across autonomous vehicles and robots, particularly where dense embedding evaluation is a computational bottleneck.
    • Dependencies: licensing and reproducibility of the DINOv3 backbone, hardware portability, numerical stability, standardized interfaces, and evidence that compression preserves performance under distribution shift rather than only on NAVSIM.

Glossary

  • Action-conditioned world model: A model that predicts future states based on a specified sequence of actions. โ€œan action-conditioned world model can be rolled out to predict possible future states under different candidate actionsโ€
  • AdaLN: Adaptive Layer Normalization, a conditioning mechanism that modulates layer-normalization parameters according to auxiliary inputs. โ€œan AdaLN-style predictor equipped with RoPEโ€
  • Autoregressive prediction: Sequential prediction in which each predicted output is used as input for later predictions. โ€œbased on autoregressive predictions from frame t+2t+2 through frame t+Ft+Fโ€
  • CEM (cross-entropy method): An iterative sampling and optimization method for selecting high-performing action sequences. โ€œThe cross-entropy method (CEM) is widely used for JEPA-based world-model planningโ€
  • CLS token: A special transformer token whose embedding summarizes an entire input, commonly used for classification or global representation. โ€œglobal CLS embeddings across the batch independently at each time stepโ€
  • Cramรฉrโ€“Wold theorem: A mathematical result stating that a multivariate probability distribution is determined by all of its one-dimensional projections. โ€œBy the Cramรฉrโ€“Wold theorem, matching all one-dimensional projected distributions is equivalent to matching the full joint distributionโ€
  • Dense patch embedding: A separate feature vector representing each local image patch rather than one global image representation. โ€œExisting JEPA-based world model baselines operate on dense patch embeddings produced by the encoderโ€
  • Diffusion transformer: A transformer architecture used within a diffusion generative model to produce complex data such as images or videos. โ€œcurrent generative world models employ diffusion transformersโ€
  • Ego pose: The position and orientation of the autonomous vehicle relative to a chosen coordinate frame. โ€œthe corresponding ego pose, consisting of the planar position (xj,yj)(x_{j},y_{j}) and heading angle ฮธj\theta_{j}โ€
  • Eppsโ€“Pulley test: A statistical test for assessing whether a distribution is consistent with a normal distribution using its empirical characteristic function. โ€œT(โ‹…)T(\cdot) denotes the univariate Eppsโ€“Pulley testโ€
  • EPDMS: A NAVSIM planning metric that combines driving-quality measures with multiplicative safety factors. โ€œEPDMS: NAVSIMv2 uses EPDMS as a pseudo-simulation-based planning metric that combines multiplicative safety terms with a weighted measure of driving quality.โ€
  • FDE (final displacement error): The distance between the predicted final pose and the ground-truth final pose. โ€œwe report the final displacement error (FDE) and absolute errors in longitudinal positionโ€
  • Geodesic accuracy: Accuracy measured by the spatial and angular displacement between a predicted trajectory endpoint and the ground-truth endpoint. โ€œwhich measure the geodesic accuracy. of the world model.โ€
  • Goal-conditioned planning: Planning that selects actions to reach a specified target state or observation. โ€œit enables goal-conditioned zero-shot planning by selecting among candidate driving trajectoriesโ€
  • Hit rate: The proportion of cases in which the ground-truth trajectory appears among the modelโ€™s best-ranked candidates. โ€œwe use hit rate to measure whether the world model ranks the ground-truth trajectory among the top-kk lowest-cost candidatesโ€
  • Imitation learning: Learning a policy by reproducing demonstrated behavior rather than optimizing directly through interaction with the environment. โ€œa reactive policy is trained to imitate human driving trajectoriesโ€
  • Isotropic Gaussian distribution: A multivariate normal distribution with equal variance in every direction and no correlations between dimensions. โ€œencourage the embeddings to follow an isotropic Gaussian distributionโ€
  • JEPA (joint-embedding predictive architecture): A self-supervised architecture that predicts a target representation from a contextual representation without reconstructing the input data. โ€œa joint-embedding predictive architecture (JEPA) for end-to-end autonomous drivingโ€
  • Latent dynamics: The evolution of hidden or learned representations that encode the underlying state of an environment. โ€œlearned latent dynamics using pixel reconstruction objectivesโ€
  • Latent embedding: A learned vector representation that encodes information about an input in a lower-level feature space. โ€œthe candidate whose predicted future latent embedding is closest to that of the goalโ€
  • MSE (mean squared error): A loss function that averages the squared differences between predicted and target values. โ€œWe supervise the predictions using only an MSE lossโ€
  • NAVSIMv2: A benchmark for evaluating autonomous-driving planning using data-driven, pseudo-simulation-based metrics. โ€œWe use the NAVSIM (Cao et al., 2025) dataset and evaluate on the latest NAVSIMv2 benchmark.โ€
  • One-dimensional projection: The mapping of a multidimensional vector onto a single direction to analyze its distribution. โ€œmatching all one-dimensional projected distributions is equivalent to matching the full joint distributionโ€
  • Patch embedding: A vector representation generated by encoding a local image patch. โ€œwe apply SIGReg to patch embeddings independently at each patch location and time stepโ€
  • Policy learning: Training a model that maps observations or states to actions. โ€œTo isolate world-model quality from policy learningโ€
  • Predictor: The neural-network component of a JEPA that forecasts target embeddings from contextual embeddings and actions. โ€œthe AdaLN predictor predicts the embedding sequence one time step aheadโ€
  • Projector: A learnable network that transforms and compresses encoder representations into a smaller embedding space. โ€œAD-E2E-JEPA further introduces a learnable projector Proj(โ‹…)\mathrm{Proj}(\cdot)โ€
  • Representation collapse: A failure mode in self-supervised learning in which different inputs receive nearly identical, uninformative representations. โ€œwhere distinct DINOv3 embeddings are mapped to nearly constant vectors and thus become uninformativeโ€
  • RoPE (rotary position embedding): A positional encoding method that incorporates token positions by rotating query and key representations in a transformer. โ€œan AdaLN-style predictor equipped with RoPEโ€
  • Rollout training: Training a world model on sequences of its own recursively generated predictions to improve long-horizon consistency. โ€œWe optionally add rollout training, as in JEPA-WMโ€
  • Self-supervised learning: Learning representations from relationships within unlabeled data rather than from manually assigned labels. โ€œThe self-supervised pretrained projector transfers effectively to downstream imitation-learning-based E2E driving.โ€
  • SIGReg: A statistical regularization method that encourages learned embeddings to match a desired distribution and prevents collapse. โ€œFurthermore, we apply SIGReg to encourage the embeddings to follow an isotropic Gaussian distribution and prevent representation collapse.โ€
  • Stop gradient: An operation that prevents gradients from propagating through a specified tensor during backpropagation. โ€œWe apply stop gradient to the projected target embedding in the MSE loss.โ€
  • Teacher forcing: A sequence-model training strategy that conditions each prediction on the ground-truth preceding state rather than solely on the modelโ€™s previous prediction. โ€œThe training objective consists of a teacher-forcing prediction lossโ€
  • Trajectory vocabulary: A predefined collection of candidate driving trajectories from which a planner selects an action sequence. โ€œwe instead search over the clustered driving trajectory vocabulary as anchorsโ€
  • ViT (Vision Transformer): A transformer-based vision architecture that processes an image as a sequence of patch tokens. โ€œthe backbone encoder, together with an AdaLN-style predictorโ€
  • World model: An internal learned representation of an environment that supports prediction, planning, and reasoning about possible future states. โ€œWorld models are internal representations that enable an agent to predict what is likely, plausible, or impossibleโ€
  • Zero-shot planning: Planning performed without training a task-specific policy for the test scenario or goal. โ€œWithout training any explicit driving policy, the learned world model enables goal-conditioned zero-shot planningโ€

Tweets

Sign up for free to view the 2 tweets with 92 likes about this paper.