Papers
Topics
Authors
Recent
Search
2000 character limit reached

GraphWorld: Long-Horizon Planning with World Models for End-to-End Autonomous Driving

Published 15 Jun 2026 in cs.CV | (2606.16274v1)

Abstract: End-to-end autonomous driving has made significant progress by unifying perception, prediction, and planning within a single learning framework, achieving strong performance in short-horizon decision making. However, most existing E2E-AD methods remain confined to short-horizon planning and lack the ability to model long-term temporal dependencies, which severely limits their generalization and security in complex and highly interactive driving scenarios. In this work, we propose GraphWorld, an E2E-AD framework that explicitly enhances long-horizon planning through latent world modeling. We introduce an Ego-Centric Interaction Graph, which adaptively models critical neighboring agents based on spatial proximity, and propagates relational context to planning queries via cross-node cross-attention. We present a World-State-Conditioned Planning that learns ego-centric latent world representations by modeling interactions between an ego vehicle and surrounding agents. This latent world state captures key interaction dynamics and safety-relevant semantics, and serves as a conditioning signal to guide long-horizon, safety-aware trajectory planning. Extensive experiments on Bench2Drive, NAVSIMv1/2, and nuScenes demonstrate that GraphWorld significantly reduces collision rates and improves long-horizon planning performance, validating its effectiveness in complex driving environments.

Summary

  • The paper introduces GraphWorld, an ego-centric latent world model that combines interaction graphs with flow-matching dynamics to improve long-horizon autonomous driving without costly pixel-space rollouts.
  • GraphWorld achieves strong results across nuScenes, Bench2Drive, and NAVSIM, including a 0.70% six-second collision rate, 51.55 Driving Score, and 13.3% faster inference than DiffusionDrive.
  • The framework improves robustness in interaction-heavy, corrupted, and adversarial scenarios, but fixed neighbor selection and single-step world evolution remain important limitations for future research.

GraphWorld (2606.16274) is an end-to-end autonomous driving (E2E-AD) framework that targets long-horizon planning through latent world modeling rather than pixel-space future generation. The paper's central argument is that existing E2E-AD planners are short-sighted: they process temporal information over brief windows and cannot anticipate interaction risks beyond immediate observations, while diffusion-based world-model approaches that do imagine future scenes are too computationally expensive for real-time deployment. GraphWorld instead learns a compact, ego-centric relational world representation and uses it to condition multi-modal trajectory planning, achieving state-of-the-art results on Bench2Drive, NAVSIMv1/v2, and nuScenes.

Rationale: long-horizon planning without rollout

The authors first reframe what "long-horizon planning" means. Since practical planners operate in a receding-horizon regime—predicting a trajectory τt\tau_t and executing only a short prefix before replanning—long-horizon capability should not mean executing a fixed long trajectory, but rather making foresighted decisions under iterative replanning. GraphWorld therefore avoids explicit multi-step rollout entirely. It learns a latent world state Wcur=fθ(O≤t)\mathbf{W}^{\text{cur}} = f_\theta(\mathcal{O}_{\leq t}) encoding interaction dynamics, and models world evolution as a continuous interpolation between current and target latent states, W(t)=(1−t)Wcur+tWtgt\mathbf{W}(t) = (1-t)\mathbf{W}^{\text{cur}} + t\mathbf{W}^{\text{tgt}}, which sidesteps the error accumulation of autoregressive rollout. The authors concede that evaluation still follows standard 6-second horizons, and interpret gains at 5–6 s as evidence of robustness under compounding uncertainty rather than as literal long-horizon execution.

Method: ECIG and WSCP

The framework has two components. The Ego-Centric Interaction Graph (ECIG) selects a fixed-size set of KK neighbors by spatial proximity to the ego vehicle, forms a directed star graph rooted at the ego node, and propagates relational context into multi-modal ego planning queries via cross-attention. Interaction-aware context is combined with a recurrent world encoder over LL-step ego–agent motion histories and pooled map embeddings, producing node-level latent world states through conditioned residual injection. Ablations show the ego-centric star topology outperforms a fully-connected graph (PDMS 90.1 vs. 85.0 on NAVSIMv1), supporting the claim that neighbor–neighbor connections introduce noise.

The World-State-Conditioned Planning (WSCP) module formulates world-state dynamics with conditional Flow-Matching (Lipman et al., 2022): a lightweight velocity field transports the current world state toward a target constructed from mode-aggregated motion and planning latents, trained with an ℓ2\ell_2 velocity-matching objective and integrated at inference with a single Euler step. The refined world state then conditions agent motion queries and is importance-reweighted via a sigmoid MLP that scores each motion mode's reliability, and is injected into ego planning queries as a residual. Notably, only two sampling steps are used; additional steps yield marginal gains while reducing throughput from 51 to 42 FPS, and Euler matches or beats Heun at higher speed.

Training proceeds in two stages: standard multi-task end-to-end training first, then explicit temporal supervision enforcing ℓ2\ell_2 consistency between the world state at time tt and a stop-gradient target from time t+1t{+}1. An ablation shows t+1t{+}1 supervision is optimal; shifting to Wcur=fθ(O≤t)\mathbf{W}^{\text{cur}} = f_\theta(\mathcal{O}_{\leq t})0 or Wcur=fθ(O≤t)\mathbf{W}^{\text{cur}} = f_\theta(\mathcal{O}_{\leq t})1 degrades 6 s collision rate from 1.95% to 2.09% and 2.26%, consistent with the paper's account of rollout uncertainty.

Main results

The strongest headline result is open-loop 6-second planning on nuScenes: GraphWorld achieves an average L2 error of 1.34 m and an average collision rate of 0.70%, a 19.5% relative reduction over World4Drive (0.87%) and over 22% relative improvement in collision rate versus the strongest prior method, with the margin widening at 5–6 s horizons (1.95% vs. 2.14% at 6 s). On 3-second nuScenes planning it remains competitive (0.57 m L2, 0.08% collision) rather than dominant.

Closed-loop results are more consequential for the paper's claims. On Bench2Drive with a SparseDrive-style backbone, GraphWorld raises Driving Score from 44.54 to 51.55 and Success Rate from 16.71% to 25.47%, with gains concentrated in interaction-critical scenarios (overtaking, merging, emergency brake). On NAVSIMv1 navtest it attains a PDMS of 90.1 with NC 99.0 and DAC 97.1, and runs 13.3% faster than DiffusionDrive, supporting the real-time claim. The most demanding evaluation is NAVSIMv2 navhard, where the two-stage protocol (real observations, then 3DGS-perturbed re-evaluation) exposes fragility in prior methods: baselines drop sharply in Stage 2, whereas GraphWorld achieves an EPDMS of 53.6, a clear margin over DiffVLA (45.0) and DriveSuprim (42.1). On NAVSIMv2 navtest it reaches 89.5 EPDMS, narrowly above Latent-WAM (89.3). On nuScenes motion prediction it reduces minADE to 0.55 (11.3% below SparseDrive) with the highest EPA (0.512), though VAD retains a lower miss rate (0.083 vs. 0.112)—a limitation the paper does not discuss.

Robustness evaluations reinforce the interaction-modeling story: on Adv-nuSc the average collision rate is 0.742% versus 1.026% for SparseDrive and 3.95% for UniAD; on nuScenes-C weather corruptions GraphWorld achieves the lowest collision rates (0.17/0.16/0.17 under Snow/Rain/Fog); and on Turning-nuScenes it attains 0.28% average collision rate with the largest gains at 3 s.

Ablation evidence

Module ablations attribute gains to both components: adding ECIG alone to the SparseDrive baseline improves Bench2Drive DS from 44.54 to 49.44, and adding WSCP further raises it to 51.55, with parallel improvements on nuScenes and NAVSIMv1. The Flow-Matching versus Diffusion ablation is consistent across all three benchmarks—for example, Bench2Drive DS improves from 73.22 to 76.71 and NAVSIMv1 PDMS from 88.1 to 90.1—supporting the choice of flow-based latent dynamics over iterative diffusion. The neighbor-radius sweep shows a moderate 10 m threshold is optimal, with both sparse (5 m) and dense (15 m) settings degrading performance.

Limitations and open questions

The paper acknowledges two constraints: neighbor selection is fixed (a static distance threshold with PadOrTrim to Wcur=fθ(O≤t)\mathbf{W}^{\text{cur}} = f_\theta(\mathcal{O}_{\leq t})2 neighbors) rather than adaptive to scene structure, and world-state prediction is single-step rather than multi-step. The Wcur=fθ(O≤t)\mathbf{W}^{\text{cur}} = f_\theta(\mathcal{O}_{\leq t})3 supervision ablation suggests the temporal consistency signal weakens rapidly beyond one step, so extending temporal supervision without incurring rollout-style error accumulation remains an open problem. The evaluation also inherits known caveats of open-loop nuScenes planning, where ego-status shortcuts can inflate scores; the closed-loop Bench2Drive and NAVSIMv2 results mitigate but do not eliminate this concern. Finally, the gap between GraphWorld's moderate open-loop improvements and its substantial closed-loop gains is asserted to reflect interaction modeling but is not isolated experimentally.

Conclusion

GraphWorld demonstrates that replacing pixel-space world-model generation with an ego-centric relational latent world state—built from a sparse interaction graph, refined by Flow-Matching, and supervised for temporal consistency—yields measurable long-horizon and safety gains at real-time cost. The consistent advantage at 5–6 s horizons, in Stage 2 of NAVSIMv2 navhard, and under adversarial and corrupted conditions indicates the latent world state carries genuinely future-relevant information. The framework's reliance on fixed graph construction and one-step world evolution defines the immediate open questions for this line of work.

Paper to Video (Beta)

No one has generated a video about this paper yet.

Whiteboard

No one has generated a whiteboard explanation for this paper yet.

Open Problems

We haven't generated a list of open problems mentioned in this paper yet.