Papers
Topics
Authors
Recent
Search
2000 character limit reached

ReactSim-Bench: Benchmarking Reactive Behavior World Model Simulation in Autonomous Driving

Published 12 Jun 2026 in cs.RO | (2606.14058v1)

Abstract: Reactive capability is a key property of data-driven behavior world model simulators for autonomous driving simulation systems. With this capability, simulated world agents can respond feasibly to autonomous vehicle (AV) behaviors that differ from the log. However, existing behavior simulation benchmarks do not directly measure reactive capability. They often let the simulator jointly control the AV and surrounding agents and evaluate realism through log similarity or open-loop prediction metrics. In this work, we introduce ReactSim-Bench for evaluating the reactive capability of behavior world model simulation in autonomous driving. We decouple the control of agents and the AV, using AV behaviors that differ from the log and require agents to respond as independent AV inputs. To obtain these AV behaviors, we construct a pipeline that uses an AV planner model to generate candidate behaviors and filters the data using rules and manual verification. Collision metrics, map-based metrics, and kinematic feasibility metrics are used to evaluate the safety and rule compliance of reactive responses. We construct 2,636 test scenarios with three categories and conduct a systematic evaluation of state-of-the-art models across multiple architectures, including Transformer-based, diffusion-based, and next-token-prediction-based models. We further analyze how replan frequency affects performance and provide insights for future studies.

Summary

  • The paper introduces a decoupled benchmark that drives the autonomous vehicle with 2,636 deviated nuPlan trajectories while models simulate surrounding agents in closed loop.
  • The benchmark shows learned models reduce agent–AV collisions five- to seven-fold versus log replay, but realism scores do not reliably predict reactive safety under distribution shift.
  • The results indicate that frequent replanning often improves reactivity, while current imitation-based and next-token models still struggle with collision avoidance when the AV departs from logged behavior.

Motivation and problem statement

Data-driven behavior world models have become a central component of autonomous driving (AD) simulation, replacing log replay and rule-based simulators for closed-loop evaluation and reinforcement learning. A defining property of such simulators is reactive capability: when the vehicle under test (the AV) deviates from the recorded driving log, the simulated surrounding agents should respond feasibly—following, yielding, or avoiding collisions. Existing benchmarks do not measure this property directly. The Waymo Open Sim Agents Challenge (WOSAC) couples the AV and all surrounding agents under the simulator's control and scores realism via log similarity; open-loop protocols inherited from motion prediction cannot assess interaction at all. Under these protocols, a model can score highly by fitting the log while never being tested against an AV whose trajectory diverges from the data.

ReactSim-Bench addresses this gap with a protocol that decouples AV control from agent simulation. The AV is treated as an external input—a fixed, pre-collected trajectory that differs from the log—and the simulator must autoregressively generate surrounding-agent behavior conditioned on the realized AV states. This is the first benchmark, according to the authors, to systematically quantify reactive capability of behavior world models.

Protocol design

The reactive task is formulated as a closed-loop rollout over an HD map M\mathcal{M}, logged history S−Th:0\mathcal{S}_{-T_h:0}, and a fixed deviated AV trajectory S0:Tfdeviated AV\mathcal{S}^{\text{deviated AV}}_{0:T_f}. At each step the simulator qq updates only the world agents:

St+1′agents=q(M,[S−Th:0,S0:t+1′AV,S0:t′agents])\mathcal{S}'^{\text{agents}}_{t+1} = q(\mathcal{M}, [\mathcal{S}_{-T_h:0}, \mathcal{S}'^{\text{AV}}_{0:t+1}, \mathcal{S}'^{\text{agents}}_{0:t}])

This contrasts sharply with WOSAC, where qq controls both the AV and agents. The authors point out two consequences of the coupled setup: realism metrics reward log-fitting rather than reactivity, and because the simulator's own AV prediction is always realized in the rollout, it can implicitly coordinate the AV with other agents to avoid collisions—an advantage unavailable in practical simulation where the tested planner acts independently.

Deviated AV data collection

Constructing valid test inputs requires AV trajectories satisfying three criteria: sufficient deviation from the log, reactive pressure on surrounding agents (log replay would induce collisions or low time-to-collision), and kinematic/map feasibility. Manual trajectory drawing would be prohibitively expensive, so the authors build a four-stage pipeline: (1) scenario pre-selection based on nearby vehicle count, distance, and TTC; (2) candidate generation using Diffusion Planner run 16 times per scenario with elevated temperature for multimodal sampling; (3) rule-based filtering removing candidates too close to the log (ADE), off-drivable-area, kinematically infeasible, lacking TTC-based interaction, or colliding with stationary/slow vehicles; and (4) manual verification with a custom annotation tool.

The resulting benchmark is built on nuPlan and contains 2,636 scenarios split into three deviation categories: longitudinal (937), lateral (900), and directional (799). The reactive pressure is substantial: replaying logged agent behavior against the collected AV trajectories produces at least one AV–agent collision in 83.46% of scenarios and minimum TTC below 0.5 s in 98.14% of scenarios, confirming that log replay is infeasible while conflicts remain avoidable through reasonable agent responses.

Metrics

Evaluation covers seven dimensions adapted from Waymax and nuPlan conventions: Agent-AV collision count (at-fault events only, with fault determined by position, heading, and motion state), Agent-AV risky TTC count (threshold 0.5 s), Agent-Agent collision rate, off-road rate, driving-direction violation rate, and separate acceleration and steering-curvature infeasibility rates following Waymax bounds (∣a∣>6.0|a| > 6.0 m/s2^2, ∣κ∣>0.3|\kappa| > 0.3 m−1^{-1}). Fault attribution and persistence rules are specified precisely in the appendix, which strengthens reproducibility.

Main results

Six state-of-the-art models spanning three paradigms were retrained on a unified nuPlan cache (642,640 training scenarios, 9-second rollouts at 10 Hz, 64 agents): MTR (Transformer regression), CTG and VBD (diffusion), and SMART, CATK, TrajTok (next-token prediction, including WOSAC 2024 and 2025 winners).

Method A-AV Coll. ↓ A-AV Risky ↓ A-A Coll. (%) ↓ Offroad (%) ↓ Acc. (%) ↓ Steer. (%) ↓
Log Replay 0.9829 1.5380 2.25 0.18 0.16 2.51
MTR 0.1457 0.5819 3.29 2.67 0.64 14.29
CTG 0.6195 0.9476 4.88 2.95 10.87 7.08
VBD 0.2276 0.4711 3.19 1.03 0.01 0.18
SMART 0.1419 0.3976 2.23 0.68 9.74 4.83
CATK 0.1426 0.4029 2.22 0.69 10.25 5.02
TrajTok 0.1407 0.4173 2.23 0.61 3.23 3.93

Three findings emerge. First, learned models exhibit genuine but incomplete reactive capability: all reduce Agent-AV collisions roughly five- to seven-fold relative to log replay (e.g., SMART at 0.1419 vs. 0.9829), confirming the scenarios are solvable. Second, realism does not predict reactivity: MTR has worse ADE and kinematic likelihood than CTG under the realism protocol yet far better reactive metrics (0.1457 vs. 0.6195 collision count), and CATK improves realism over SMART while slightly degrading reactive performance. Third, imitation-trained NTP models face an out-of-distribution failure mode: SMART, CATK, and TrajTok achieve very low ADE and near-zero Agent-AV collisions under the realism protocol (0.0095–0.0125), yet their collision counts rise by more than an order of magnitude under the reactive protocol. Qualitative case studies reinforce this: in ramp-merge and directional-deviation scenarios, several simulators keep agents near their logged trajectories and collide with the deviated AV, while others decelerate appropriately.

Replan frequency analysis

Ablating the replanning interval shows that higher replan rates (above 1 Hz) generally improve reactive performance, since the simulator observes the realized AV state more often. This contrasts with prior findings on WOSAC, where lower replan rates yield better realism because accumulated autoregressive error dominates when no external agent exists. ReactSim-Bench therefore imposes no artificial lower bound on replan frequency, which the authors argue better reflects practical simulation. The effect remains model-dependent—frequent replanning can amplify accumulated error for some architectures—so replan strategy is itself a significant confound in reactive evaluation.

Limitations

The paper scopes itself explicitly to reactivity. Controllability and diversity, pursued by other simulation methods, are not benchmarked here. Additionally, the deviated AV trajectories derive from a single planner (Diffusion Planner) plus manual filtering, so the diversity of AV behaviors is bounded by that generator's modes; the benchmark also evaluates only vehicle interactions among the 64 nearest agents, and the fault-attribution heuristics in the collision metric involve threshold choices (stopped-speed 0.05 m/s, rear-collision angle 150°) that could affect comparability with other protocols.

Conclusion

ReactSim-Bench contributes a decoupled AV–agent evaluation protocol, a curated set of 2,636 feasible, interaction-inducing deviated AV behaviors on nuPlan, and a systematic evaluation showing that current SOTA simulators—including WOSAC-winning NTP models—degrade substantially when the AV departs from the log. The central empirical claim, that realism metrics fail to capture reactive capability, is well supported by the protocol-level comparison. The key open question the paper leaves is how to train behavior world models that acquire robust reactive responses rather than log-fitting behavior, given that current diffusion and NTP paradigms both struggle under distribution shift induced by the ego vehicle.

Paper to Video (Beta)

No one has generated a video about this paper yet.

Whiteboard

No one has generated a whiteboard explanation for this paper yet.

Open Problems

We haven't generated a list of open problems mentioned in this paper yet.