- The paper introduces DR-NWM, which combines anchor-guided chunked generation with bidirectional epipolar attention to reduce autoregressive perceptual drift and viewpoint-related geometric drift without requiring depth or 3D supervision.
- The method improves long-horizon prediction across RECON, SCAND, HuRoN, and TartanDrive, including a 31.8% FID reduction over NWM at 16 seconds on RECON and substantially lower Sampson errors at extended horizons.
- The paper shows that improved world-model predictions enhance fixed-planner navigation, with CEM+DR-NWM reducing RECON ATE to 2.78 from 3.18 for CEM+NWM, while performance depends on reliable feature matching and adaptive handling of difficult scenes.
Motivation: two coupled failure modes of rollout-based navigation world models
Navigation world models such as NWM forecast future observations conditioned on past frames and candidate actions, enabling planners to compare trajectories by simulating their visual consequences. The paper identifies two failure modes that limit this paradigm. Perceptual drift arises from autoregressive (AR) rollout: each predicted frame is recursively fed back as conditioning, so per-step errors compound over long horizons. Mitigations based on rollout-style training expose the model to its own errors but incur prohibitive computational cost. Geometric drift arises because predictions may fail to track the viewpoint change induced by the commanded action, and because degraded predictions provide unreliable cues for subsequent steps. Prior remedies typically require explicit 3D supervision—depth maps or point clouds—which is costly to acquire at scale and limits generalization.
The paper's central observation is that these two drifts can be addressed jointly by restructuring the rollout itself rather than the supervision signal.
Method: anchor-guided rollout and bidirectional epipolar grounding
DR-NWM decomposes prediction into a two-level hierarchy, mirroring goal-then-steps planning. First, the model predicts a set of sparse temporal anchors x^km at fixed intervals by jumping directly from the observation history using the corresponding action subsequence, bypassing intermediate recursive conditioning entirely. Second, for each chunk between consecutive anchors, intermediate frames are generated jointly conditioned on the past anchor context and the future anchor, together with forward action sequences am→ and inverse action sequences am← defined from the future anchor backward. This partitions the horizon into short chunks, cutting off recursive error accumulation.
The anchors also serve a geometric role. Once both a past view and a future anchor are available, matched feature points induce two epipolar lines in any target frame; their intersection localizes where corresponding content should appear (the paper provides the standard multi-view justification in its appendix). These intersections are discretized onto the transformer token grid (L=196 tokens) as binary masks Mpast,Mfut∈{0,1}L×L applied to cross-attention scores. Source tokens without valid correspondences retain unmasked attention rows, so noisy geometry degrades gracefully to unconstrained attention.
The generative backbone, AC-DiT, extends a DiT initialized from the pretrained CDiT/XL checkpoint of NWM. Scalar conditions (actions, diffusion timestep, relative time offsets, forward and inverse) are combined via zero-initialized gates γcond, γpast, γfut, γτ, preserving pretrained behavior at finetuning start while progressively incorporating future-side information. Training uses the standard DDPM denoising loss; no 3D supervision is required anywhere in the pipeline.
Implementation details worth noting: LoFTR correspondences are filtered by confidence, semantic segmentation (sky/dynamic objects), and border validity; fundamental matrices are estimated with RANSAC; a reliability score based on RANSAC inlier counts disables masking when geometry is unreliable; and exponential-moving-average smoothing stabilizes masks across temporally adjacent chunk frames.
Results: perceptual quality, geometric consistency, and planning
Across four benchmarks (RECON, SCAND, HuRoN, TartanDrive), DR-NWM outperforms NWM and EgoWM on all perceptual metrics at all horizons from 1s to 16s. The gains widen with horizon length, consistent with the drift mechanism: at 16s on RECON, FID improves over NWM by 31.8% (76.9 vs 112.7); on TartanDrive at 8s, FID improves by 29.1%. On HuRoN, LPIPS gains reach 24.9% already at 1s. Qualitatively, NWM deteriorates visibly beyond 4s and EgoWM earlier, whereas the proposed method retains interpretable structure throughout.
Geometric consistency shows similar trends. Epipolar distance and Sampson error are lower than both baselines in nearly all dataset–horizon combinations, with advantages most pronounced at 16s—for example, Sampson error of 17.66 vs 33.09 (NWM) on SCAND and 55.96 vs 94.32 on TartanDrive. MEt3R multi-view consistency is also consistently better across all skip lengths, indicating that reduced geometric drift translates into stronger cross-view coherence.
Because the planner is held fixed, downstream improvements isolate the effect of prediction quality. With CEM and NoMaD planners sampling 32 candidate trajectories over a 4-second horizon, DR-NWM improves ATE, FDE, RPE, and LPIPS on both RECON and SCAND—for instance, CEM+Ours achieves ATE 2.78 vs 3.18 for CEM+NWM on RECON. An additional ablation shows the largest planning gains occur in low-sample regimes (ATE 3.69 vs 3.93 at 10 samples), narrowing as sample count grows; improved prediction is thus most valuable when the planner cannot rely on exhaustive search.
Ablation analysis
Two ablations separate the contributions. Replacing AR rollout with anchor-guided inference alone, keeping the backbone fixed, substantially improves FID at long horizons (76.9 vs 140.3 at 16s) but leaves Sampson error largely unchanged—confirming that the rollout order addresses perceptual drift specifically. Component ablations of AC-DiT show the opposite pattern: FID and LPIPS remain nearly constant across variants, while average Sampson error varies meaningfully, with the full model (future anchor + epipolar mask + chunk attention) achieving the best value (56.37 vs 60.76 for the base). Notably, applying epipolar masking without future conditioning degrades geometric performance relative to no masking at some intervals, because the constraint degenerates from an intersection-based point localization to a weaker line-based one. This supports the paper's claim that bidirectional anchors and epipolar masking are complementary rather than independently beneficial.
Limitations and open questions
Several constraints qualify the results. The method depends on sparse feature matching quality: when filtered matches fall below thresholds or RANSAC fails, masks are disabled and the model reverts to unconstrained attention, so performance in texture-poor or heavily dynamic scenes relies on this fallback. The evaluation uses a lower-resolution HuRoN variant (160×120) because the original high-resolution data is unavailable, which may affect comparability with results reported elsewhere. Anchor intervals are fixed rather than adaptive, leaving open how interval selection should scale with scene complexity or motion speed. Finally, the ablations show that anchor inference alone does not improve geometric consistency—the full architecture is required—so the approach does not offer a lightweight drop-in fix for existing AR world models.
Conclusion
DR-NWM reformulates navigation world-model prediction as anchor-guided chunked generation with bidirectional epipolar attention masking, mitigating perceptual and geometric drift without explicit 3D supervision. Consistent improvements across four benchmarks in long-horizon visual quality, epipolar consistency, multi-view coherence, and fixed-planner trajectory accuracy indicate that rollout structure and geometry-aware conditioning, rather than additional supervision, are levers for more reliable navigation world models.