- The paper introduces a decoupled benchmark that drives the autonomous vehicle with 2,636 deviated nuPlan trajectories while models simulate surrounding agents in closed loop.
- The benchmark shows learned models reduce agent–AV collisions five- to seven-fold versus log replay, but realism scores do not reliably predict reactive safety under distribution shift.
- The results indicate that frequent replanning often improves reactivity, while current imitation-based and next-token models still struggle with collision avoidance when the AV departs from logged behavior.
Motivation and problem statement
Data-driven behavior world models have become a central component of autonomous driving (AD) simulation, replacing log replay and rule-based simulators for closed-loop evaluation and reinforcement learning. A defining property of such simulators is reactive capability: when the vehicle under test (the AV) deviates from the recorded driving log, the simulated surrounding agents should respond feasibly—following, yielding, or avoiding collisions. Existing benchmarks do not measure this property directly. The Waymo Open Sim Agents Challenge (WOSAC) couples the AV and all surrounding agents under the simulator's control and scores realism via log similarity; open-loop protocols inherited from motion prediction cannot assess interaction at all. Under these protocols, a model can score highly by fitting the log while never being tested against an AV whose trajectory diverges from the data.
ReactSim-Bench addresses this gap with a protocol that decouples AV control from agent simulation. The AV is treated as an external input—a fixed, pre-collected trajectory that differs from the log—and the simulator must autoregressively generate surrounding-agent behavior conditioned on the realized AV states. This is the first benchmark, according to the authors, to systematically quantify reactive capability of behavior world models.
Protocol design
The reactive task is formulated as a closed-loop rollout over an HD map M, logged history S−Th​:0​, and a fixed deviated AV trajectory S0:Tf​deviated AV​. At each step the simulator q updates only the world agents:
St+1′agents​=q(M,[S−Th​:0​,S0:t+1′AV​,S0:t′agents​])
This contrasts sharply with WOSAC, where q controls both the AV and agents. The authors point out two consequences of the coupled setup: realism metrics reward log-fitting rather than reactivity, and because the simulator's own AV prediction is always realized in the rollout, it can implicitly coordinate the AV with other agents to avoid collisions—an advantage unavailable in practical simulation where the tested planner acts independently.
Deviated AV data collection
Constructing valid test inputs requires AV trajectories satisfying three criteria: sufficient deviation from the log, reactive pressure on surrounding agents (log replay would induce collisions or low time-to-collision), and kinematic/map feasibility. Manual trajectory drawing would be prohibitively expensive, so the authors build a four-stage pipeline: (1) scenario pre-selection based on nearby vehicle count, distance, and TTC; (2) candidate generation using Diffusion Planner run 16 times per scenario with elevated temperature for multimodal sampling; (3) rule-based filtering removing candidates too close to the log (ADE), off-drivable-area, kinematically infeasible, lacking TTC-based interaction, or colliding with stationary/slow vehicles; and (4) manual verification with a custom annotation tool.
The resulting benchmark is built on nuPlan and contains 2,636 scenarios split into three deviation categories: longitudinal (937), lateral (900), and directional (799). The reactive pressure is substantial: replaying logged agent behavior against the collected AV trajectories produces at least one AV–agent collision in 83.46% of scenarios and minimum TTC below 0.5 s in 98.14% of scenarios, confirming that log replay is infeasible while conflicts remain avoidable through reasonable agent responses.
Metrics
Evaluation covers seven dimensions adapted from Waymax and nuPlan conventions: Agent-AV collision count (at-fault events only, with fault determined by position, heading, and motion state), Agent-AV risky TTC count (threshold 0.5 s), Agent-Agent collision rate, off-road rate, driving-direction violation rate, and separate acceleration and steering-curvature infeasibility rates following Waymax bounds (∣a∣>6.0 m/s2, ∣κ∣>0.3 m−1). Fault attribution and persistence rules are specified precisely in the appendix, which strengthens reproducibility.
Main results
Six state-of-the-art models spanning three paradigms were retrained on a unified nuPlan cache (642,640 training scenarios, 9-second rollouts at 10 Hz, 64 agents): MTR (Transformer regression), CTG and VBD (diffusion), and SMART, CATK, TrajTok (next-token prediction, including WOSAC 2024 and 2025 winners).
| Method |
A-AV Coll. ↓ |
A-AV Risky ↓ |
A-A Coll. (%) ↓ |
Offroad (%) ↓ |
Acc. (%) ↓ |
Steer. (%) ↓ |
| Log Replay |
0.9829 |
1.5380 |
2.25 |
0.18 |
0.16 |
2.51 |
| MTR |
0.1457 |
0.5819 |
3.29 |
2.67 |
0.64 |
14.29 |
| CTG |
0.6195 |
0.9476 |
4.88 |
2.95 |
10.87 |
7.08 |
| VBD |
0.2276 |
0.4711 |
3.19 |
1.03 |
0.01 |
0.18 |
| SMART |
0.1419 |
0.3976 |
2.23 |
0.68 |
9.74 |
4.83 |
| CATK |
0.1426 |
0.4029 |
2.22 |
0.69 |
10.25 |
5.02 |
| TrajTok |
0.1407 |
0.4173 |
2.23 |
0.61 |
3.23 |
3.93 |
Three findings emerge. First, learned models exhibit genuine but incomplete reactive capability: all reduce Agent-AV collisions roughly five- to seven-fold relative to log replay (e.g., SMART at 0.1419 vs. 0.9829), confirming the scenarios are solvable. Second, realism does not predict reactivity: MTR has worse ADE and kinematic likelihood than CTG under the realism protocol yet far better reactive metrics (0.1457 vs. 0.6195 collision count), and CATK improves realism over SMART while slightly degrading reactive performance. Third, imitation-trained NTP models face an out-of-distribution failure mode: SMART, CATK, and TrajTok achieve very low ADE and near-zero Agent-AV collisions under the realism protocol (0.0095–0.0125), yet their collision counts rise by more than an order of magnitude under the reactive protocol. Qualitative case studies reinforce this: in ramp-merge and directional-deviation scenarios, several simulators keep agents near their logged trajectories and collide with the deviated AV, while others decelerate appropriately.
Replan frequency analysis
Ablating the replanning interval shows that higher replan rates (above 1 Hz) generally improve reactive performance, since the simulator observes the realized AV state more often. This contrasts with prior findings on WOSAC, where lower replan rates yield better realism because accumulated autoregressive error dominates when no external agent exists. ReactSim-Bench therefore imposes no artificial lower bound on replan frequency, which the authors argue better reflects practical simulation. The effect remains model-dependent—frequent replanning can amplify accumulated error for some architectures—so replan strategy is itself a significant confound in reactive evaluation.
Limitations
The paper scopes itself explicitly to reactivity. Controllability and diversity, pursued by other simulation methods, are not benchmarked here. Additionally, the deviated AV trajectories derive from a single planner (Diffusion Planner) plus manual filtering, so the diversity of AV behaviors is bounded by that generator's modes; the benchmark also evaluates only vehicle interactions among the 64 nearest agents, and the fault-attribution heuristics in the collision metric involve threshold choices (stopped-speed 0.05 m/s, rear-collision angle 150°) that could affect comparability with other protocols.
Conclusion
ReactSim-Bench contributes a decoupled AV–agent evaluation protocol, a curated set of 2,636 feasible, interaction-inducing deviated AV behaviors on nuPlan, and a systematic evaluation showing that current SOTA simulators—including WOSAC-winning NTP models—degrade substantially when the AV departs from the log. The central empirical claim, that realism metrics fail to capture reactive capability, is well supported by the protocol-level comparison. The key open question the paper leaves is how to train behavior world models that acquire robust reactive responses rather than log-fitting behavior, given that current diffusion and NTP paradigms both struggle under distribution shift induced by the ego vehicle.