Papers
Topics
Authors
Recent
Search
2000 character limit reached

TPS-Drive: Task-Guided Representation Purification for VLM-based Autonomous Driving

Published 26 May 2026 in cs.RO | (2605.27038v1)

Abstract: Vision-LLMs (VLMs) provide a promising foundation for autonomous driving planning, yet bridging semantic reasoning and precise 3D spatial forecasting remains a critical challenge. Existing representation strategies generally follow two paths: text-aligned methods flatten continuous spatial states into symbols, which compromises geometric structure and induces "spatial hallucinations"; dense visual methods preserve spatial topology but overwhelm standard tokenizers with redundant background textures, leading to "representation interference". To address these limitations, we introduce TPS-Drive, a novel framework centered on Task-Guided Representation Purification that empowers VLMs to Think in Purified Space. At its core, an Agent-Centric Tokenizer utilizes a task-guided vector quantization mechanism supervised by a frozen 3D detection head, which explicitly reallocates limited codebook capacity from pervasive static backgrounds to critical dynamic agents and effectively isolates spatial redundancy. Leveraging this purified spatial vocabulary, TPS-Drive employs a decoupled reasoning pipeline that sequentially performs scene understanding, future forecasting, and action generation. The framework is optimized via a progressive three-stage training paradigm, culminating in reward-driven refinement that surpasses pure imitation learning. Extensive experiments validate our approach: TPS-Drive achieves accurate agent spatial state forecasting and reduces collision rates in open-loop nuScenes evaluations, while establishing new safety records on the rigorous closed-loop NAVSIMv1 and NAVSIMv2 benchmarks.

Authors (3)

Summary

  • The paper introduces Task-Guided Representation Purification, using a CenterPoint-supervised hierarchical VQ tokenizer to prioritize dynamic agents and preserve geometric structure in BEV tokens.
  • TPS-Drive combines compact autoregressive world-token forecasting, parallel residual prediction, scene reasoning, and diffusion-based planning, achieving 34.60% NDS in nuScenes forecasting and 86.7 EPDMS on NAVSIMv2.
  • Ablations show that task-guided tokenization provides the largest gain, while residual codes and reward-driven refinement further improve spatial prediction, collision reduction, and closed-loop driving safety.

TPS-Drive addresses a specific failure mode in VLM-based autonomous driving: the mismatch between language-centric tokenization and the continuous geometric structures required for safe trajectory planning. The paper identifies two competing representation strategies and their characteristic pathologies. Text-aligned methods that serialize 3D states as natural language or flattened box coordinates preserve compatibility with standard VLM tokenizers but destroy geometric proximity, producing "spatial hallucinations." Dense visual methods—future image prediction or 3D occupancy grids—retain spatial topology, but reconstruction-guided Vector Quantization (VQ) allocates limited codebook capacity to static backgrounds and textures rather than dynamic agents, producing "representation interference" that degrades semantic reasoning. The proposed remedy is Task-Guided Representation Purification: a VQ tokenizer supervised by a frozen 3D detection head that reallocates codebook capacity to task-critical dynamic agents, allowing the VLM to "Think in Purified Space" (2605.27038).

Agent-centric task-guided tokenization

The core architectural component is a hierarchical VQ tokenizer over Bird's-Eye-View (BEV) features. An encoder compresses a 200×200200\times200 BEV map to 25×2525\times25 latent positions, each quantized against a primary codebook of 8,192 entries. Unlike conventional VQ-VAE training, the primary codebook objective combines weighted reconstruction, standard codebook/commitment losses, and task-guidance losses—heatmap and 3D bounding box losses—computed by a frozen, pretrained CenterPoint detector. The authors are explicit that reconstruction remains necessary despite partially preserving background detail, because it prevents structural deterioration of the purified features; the tension between purification and reconstruction is managed via loss weighting rather than eliminated.

A single quantization level loses positional precision, so L−1=3L-1=3 residual codebooks are trained in a second phase against frozen primary components, quantizing successive residuals. Notably, only the primary tokens are autoregressively predicted by the VLM; residual tokens are predicted in parallel by lightweight classification heads over the backbone's hidden states. This design keeps the autoregressive sequence compact while retaining fine-grained geometry, an efficiency-oriented compromise rather than a fully autoregressive world model.

Decoupled reasoning and training

Downstream of the tokenizer, TPS-Drive decouples planning into three sequential stages. Scene understanding produces a structured representation covering environmental attributes (area type, road topology, lighting) and physical constraints, supervised by cross-entropy against annotations generated by Qwen3.5-27B with manual verification—a dependency on LLM-generated labels that the paper mitigates but does not eliminate. Future forecasting autoregressively predicts the primary tokens of the future purified BEV map in raster order, conditioned on multi-view images, current BEV features, navigation command, kinematics, and the scene representation. Action generation uses a conditional diffusion planner, conditioned on the hidden states of predicted future tokens, trained with a standard denoising objective independently of the VLM backbone, which the authors report stabilizes optimization.

Training proceeds in three stages: task-guided tokenizer pretraining, joint supervised fine-tuning of all three stages, and reward-driven refinement using grouped relative optimization (GRPO-style, following DeepSeekMath). Multiple stochastic rollouts of world tokens and trajectories are scored by an offline reward based on geometric error and smoothness; group-normalized advantages update the world-model branch via an advantage-weighted objective and the diffusion planner via reward-weighted denoising, avoiding a separate value network.

Empirical results

The evaluation spans open-loop nuScenes planning, agent-centric forecasting, and closed-loop NAVSIMv1/v2.

Benchmark Metric TPS-Drive Best prior
nuScenes (ST-P3, with ego status) Avg. L2 / CR 0.32 m / 0.10% FSDrive*: 0.28 m / 0.10%
nuScenes (UniAD, with ego status) Avg. L2 / CR 0.49 m / 0.14% FSDrive*: 0.45 m / 0.16%
NAVSIMv1 PDMS 89.7 Recogdrive: 89.6
NAVSIMv2 EPDMS 86.7 DriveVLA-W0: 86.1
nuScenes forecasting NDS / mAP 34.60% / 24.03% WoTE: 30.01% / 22.77%

On open-loop nuScenes, the ego-status variant achieves a 0.10% average collision rate under ST-P3 metrics and 0.14% under UniAD metrics, with the paper claiming record-low collision rates among VLM-based methods. Without ego status, the base model records 0.19% CR versus 0.94% for OmniDrive and 0.32% for OccWorld, though its L2 error (0.55 m average) is behind BEV-Planner and FSDrive*, indicating the safety gains come primarily from collision reduction rather than trajectory-matching accuracy. On closed-loop evaluation, the PDMS of 89.7 narrowly exceeds Recogdrive, and the EPDMS of 86.7 on the stricter NAVSIMv2 protocol surpasses both DriveVLA-W0 and HydraMDP++, with 99.8% traffic light compliance and 97.1% lane keeping. The forecasting results—34.60% NDS, a +4.59 point absolute gain over WoTE—substantiate the claim that the purified vocabulary supports accurate spatial state prediction rather than merely improving the planner.

The ablations isolate each contribution. Replacing a reconstruction-guided VQ baseline (17.31% NDS, 81.1 EPDMS) with the task-guided primary tokenizer yields +12.72 points of NDS, the single largest effect, directly supporting the representation-interference hypothesis. Residual layers add +3.75 NDS; reward-driven refinement raises EPDMS from 84.1 to 86.7 and reduces open-loop CR from 0.26% to 0.19%. A notable finding is that reward refinement also improves upstream forecasting (33.78% → 34.60% NDS), suggesting that continuous planning rewards propagate back into discrete representation quality—an interaction the paper observes but does not theoretically explain.

Qualitative comparisons against an implicit image chain-of-thought baseline (FSDrive) show the baseline planning "Forward" into a crossing pedestrian and executing an improper yield in cluttered scenes, while TPS-Drive produces a timely stop and a slow pass, attributed to explicit 3D box forecasting in the purified token space.

Limitations and open questions

The paper concedes two substantive constraints. First, the frozen detection head that supervises purification bounds the representational capacity of the vocabulary: the tokenizer can only purify toward what CenterPoint-style detection defines as task-relevant, and the authors identify end-to-end differentiable spatial purification as future work. Second, the decoupled multi-stage architecture—sequential scene understanding, forecasting, and diffusion planning—restricts real-time inference speed, with no latency figures reported. Two further open questions follow from the results: whether the observed reward-to-representation synergy persists under rewards beyond geometric error and smoothness, and whether the LLM-generated scene annotations remain reliable in distribution-shifted conditions where manual inspection is impractical.

Conclusion

TPS-Drive demonstrates that task-guided vector quantization, supervised by a frozen 3D detection head, can convert dense BEV representations into a compact, agent-centric token vocabulary that mitigates both spatial hallucination and representation interference. Combined with a decoupled reasoning pipeline and reward-driven refinement, the framework achieves strong forecasting accuracy (34.60% NDS) and leading closed-loop safety scores (89.7 PDMS, 86.7 EPDMS), with ablations attributing the largest gain to representation purification itself. The dependence on a frozen detector and the multi-stage inference cost remain the principal open constraints on the approach.

Paper to Video (Beta)

No one has generated a video about this paper yet.

Whiteboard

No one has generated a whiteboard explanation for this paper yet.

Open Problems

We haven't generated a list of open problems mentioned in this paper yet.

Tweets

Sign up for free to view the 1 tweet with 0 likes about this paper.