- The paper introduces an LLM-driven framework that integrates OPM-based semantic scene abstraction with language-mediated decision reasoning to improve autonomous driving in mixed traffic.
- It employs a discrete choice model and Monte Carlo trajectory optimization to achieve higher efficiency, comfort, and safety compared to conventional methods.
- The system demonstrates enhanced human-like negotiation at intersections and promises transparent, adaptable, and interactive autonomous driving.
LLM-based Interactive Decision-Making for Autonomous Driving: An Expert Overview
Motivation and Context
The persistent challenge in deploying autonomous driving systems within mixed-traffic environments—where both human-driven vehicles (HDVs) and Connected Automated Vehicles (CAVs) coexist—lies in resolving ambiguities in right-of-way, managing high-conflict scenarios, and fostering proactive, intent-aware interactions. Traditional rule-based and game-theoretic approaches become insufficient due to scale, unpredictable human behaviors, and limited upstream intent modeling. Conventional ML methods, such as imitation and reinforcement learning, face specific deficits in generalizability, interpretability, and intent negotiation. This work addresses these gaps by leveraging LLMs as a unified, semantically robust reasoning core, augmenting perception, decision-making, and external human-machine interfaces (eHMI) within a modular, Object-Process Methodology (OPM)-based framework.

Figure 1: LLM-based interactive autonomous driving system integrating OPM scene abstraction, semantic reasoning, trajectory optimization, and language-based eHMI for closed-loop human-vehicle interaction.
Semantic Scene Abstraction with OPM
Raw sensor data and low-level numerical states, while available in abundance, are ill-suited as direct inputs for LLM-based reasoning architectures due to the mismatch between unstructured feature spaces and the semantic granularity required for intent-aware decision making. The proposed OPM approach overcomes this limitation via explicit semantic decomposition: objects (vehicles/pedestrians), processes (lane changes, stops), and relations (potential conflicts, right-of-way).

Figure 2: Mapping from scene raw information to structured object–process–relation graphs, facilitating compact and causally informative LLM prompting.
Critical to OPM is risk-aware object selection: the scene graph includes only those entities satisfying risk thresholds grounded in time-to-collision (TTC) and spatial proximity, culling irrelevant agents while focusing downstream reasoning on semantically pertinent conflict clusters. This representation admits efficient augmentation with historical and contextual data, enabling both instantaneous and temporally consistent behavioral inference.
LLM-Driven Intent Parsing, Decision Reasoning, and Correction
Unsignalized intersections and other high-interaction scenarios expose the frailty of traditional models, which cannot comprehensively infer latent or evolving intent from heterogeneous agents. To this end, the framework deploys LLMs as semantic parsers and high-level strategists. The LLM integrates explicit communication cues (e.g., turn signals, eHMI utterances) with motion histories processed via an offline-trained attention mechanism, generating agent-wise intent vectors reflecting style (aggressive/conservative), current maneuver, and contextual task priority.
The decision layer employs a discrete choice framework, initializing maneuver preference distributions via logit models conditioned on scenario and agent factors (safety, efficiency, comfort). Consistency and stability are enforced by correction terms encoding behavioral inertia and switch penalties, reminiscent of trending approaches in sequential driving behavior models. LLM-driven natural language reasoning further refines and corrects candidate decisions by incorporating nonparametric, context-conditioned intent parsing, resolving implicit social contracts, and mitigating misinterpretation via confidence thresholds and human-in-the-loop feedback.
Trajectory Generation and Monte Carlo Optimization
Given high-level maneuver selection, trajectory optimization is modeled as a sampling problem over a distribution of terminal states, parameterized by stochastic perturbations sampled from empirically validated state distributions. Each trajectory candidate adheres to fifth-order polynomial (quintic) constraints for kinematic feasibility and is ranked by a cost function balancing safety (conflict and collision avoidance), efficiency (timely intersection clearance), and passenger comfort (minimal jerk). The optimization process is inherently parallelizable but introduces latency challenges as scenario complexity and agent density increase.
Closed-Loop Language-based eHMI via LLM Integration
A key feature distinguishing this system is the tight coupling of decision-making and natural language interaction. Post-decision, the LLM translates selected maneuvers into concise natural language statements aligned with human social norms, which are broadcast to surrounding agents via eHMI. This loop enables not only proactive intent signaling but also dynamic adjustment based on multi-turn human feedback and affective cues, supporting transparent, explainable communication in dense, ambiguous traffic.
Empirical Evaluation: Simulator Results and Human-likeness Assessment
Experiments are conducted on a high-fidelity driving simulator under varied intersection and merging scenarios, reflecting real-world distributions of conflict, efficiency, and comfort demands. Three canonical models are compared: IDM (car-following), a game-theoretic planner, and the proposed LLM-interactive system.

Figure 3: Driving simulator system used for model-in-the-loop evaluation of interactive driving policies.
Key findings are as follows:
- Efficiency: The LLM-based system consistently achieves higher average speeds than both IDM and game-theoretic baselines across all initial velocity regimes.
- Comfort: The LLM model maintains average jerk at acceptable levels—significantly superior to IDM and closely matching the game-theoretic approach.
- Safety/Conflict Resolution: The proposed model notably reduces average conflict duration (by ~35% versus IDM, ~10% versus game-theoretic) under medium-high speed conditions.



Figure 4: Comparative plots of average speed, average jerk, and conflict duration across models under varying initial speeds.
OPM-structured inputs confer measurable advantages in both decision accuracy and inference latency compared to unstructured or naively organized input methods.

Figure 5: OPM input representation yields higher reasoning accuracy and lower latency relative to raw and weakly structured (Simple) alternatives.
Case analysis demonstrates superior human-like negotiation dynamics at intersections, where LLM-driven policies enable proactive and interpretable conflict resolution unattainable with conventional controllers.



Figure 6: Real-time positional evolution of CAV and HDV under IDM, game-theoretic, and LLM control, illustrating deadlock, risky maneuvers, and negotiated cooperation, respectively.
In merging scenarios, the method generalizes to maintain high efficiency, validating the intent-reasoning mechanism beyond simple intersection cases.


Figure 7: Merging scenario evaluation showing consistently superior efficiency of LLM-based policies over IDM and game-theoretic models.
A Turing test employing double-blind drivers in simulated interactive settings indicates that participant accuracy in discriminating human from LLM-driven vehicles does not differ significantly from chance (accuracy ≈ 0.57, Δnaturalness ≈ 0.2 on a 5-point scale).

Figure 8: Experimental protocol for double-blind, two-user driving simulator Turing tests.

Figure 9: Confusion matrix summarizing participants’ ability to distinguish between human and LLM-driven CAV; low discriminability underscores anthropomorphic fidelity.
Implications and Future Directions
The integration of LLMs with OPM scene abstraction and eHMI provides a compelling pathway for bridging the interpretability and adaptability gap in current CAV systems. Practically, this architecture enables transparent, proactive negotiation in complex, stochastic traffic environments, enhancing road safety, efficiency, and public trust. Theoretically, this work elucidates how semantically structured scene models can synergistically interact with powerful pretrained reasoning engines to surpass the bounds of traditional control-theoretic or purely end-to-end learning paradigms.
However, the approach presents latency constraints—especially under high interaction density—due to the computational demands of LLM-based trajectory optimization and rich multi-agent reasoning. Hallucination control and overfitting to domain-specific communication styles remain open concerns, requiring systematic validation as new LLM architectures and multimodal prompts become embedded within hardware-constrained vehicles.
Future progress will demand real-world pilot deployments, chain-of-thought reasoning exploration, and deeper multimodal integration (perception, language, control). Adaptation to personalized driving styles, task-specific fine-tuning, and feedback-driven language grounding between CAVs and human sociocultural patterns are clear research trajectories.
Conclusion
This study delivers a comprehensive LLM-based interactive decision-making framework for autonomous driving in mixed traffic, characterized by OPM-structured scene reasoning, semantic agent intent parsing, sampling-based trajectory optimization, and language-mediated behavioral transparency. Strong empirical results underscore improvements in efficiency, comfort, and human-likeness, with marked reductions in conflict durations compared to classical approaches. These findings evidence a promising research direction in integrating semantic abstraction and large-scale language understanding for socially aligned, interpretable autonomous vehicles, while highlighting necessary advances for robust, scalable deployment in unconstrained real-world environments.