---
title: 'SimCoachCorpus: Multimodal Driving Coaching Data'
url: https://www.emergentmind.com/topics/simcoachcorpus
type: topic
---

# SimCoachCorpus: Multimodal Driving Coaching Data

Searching arXiv for the specified paper and closely related context.
arxiv_search.query({"search_query":"id:2509.14548 OR ti:\"SimCoachCorpus\"","start":0,"max_results":5})
arxiv_search.query({"search_query":"id:2509.14548 OR ti:\"SimCoachCorpus\"","start":0,"max_results":5})
arxiv_search.query: {"search_query":"id:2509.14548 OR ti:\"SimCoachCorpus\"","start":0,"max_results":5}
SimCoachCorpus is a naturalistic, longitudinal dataset for embodied teaching that fuses expert language with simulated race car driving trajectories in a one-on-one coaching setting. It was introduced to address the relative scarcity of datasets in which language and physical action are deeply intertwined over time, particularly for the study of motor skill acquisition through verbal instruction. The dataset centers on high-performance driving education on a static CARLA-based simulator at Thunderhill Raceway and synchronizes vehicle state, driver inputs, map context, cone landmarks, coach and student audio, time-aligned transcripts, coaching annotations, compliance labels, and survey-based measures of cognitive load and affect. Its intended uses include motor learning analysis, linguistic analysis, and computational modeling of teaching, with demonstrated applications in in-context generation, imitation learning, and topic modeling [2509.14548].

## 1. Dataset definition and scope

SimCoachCorpus is described as a first-of-its-kind dataset that combines expert coaching language with embodied action in a longitudinal teaching setting. Its central design objective is to provide a reproducible, richly annotated resource for studying how people acquire motor skills through verbal instruction over time and for building computational models that reason jointly over linguistic and physical modalities. The domain is race car simulator driving, where the coach produces concurrent, action-oriented utterances during driving and terminal feedback between laps [2509.14548].

The corpus captures both coached and unguided learning. Twenty-nine humans were asked to drive in a simulator for approximately ninety minutes. Fifteen participants were given personalized one-on-one instruction from a professional performance driving coach. Fifteen participants were recruited for self-practice, but one participant experienced motion sickness early and was excluded from analyses, with baseline-only data recorded; this leaves 14 participants in analyzed self-practice data. The dataset includes over 40 hours of vehicle driving data, over 20,000 concurrent feedback utterances, and over 400 terminal feedback utterances [2509.14548].

The dataset is organized around several interacting research dimensions. These include motor skill acquisition through trajectories and performance metrics, interactive structure in concurrent versus terminal feedback, linguistic analysis through taxonomy labels and topic modeling, and multimodal prediction problems involving teacher action and student motion. A plausible implication is that SimCoachCorpus is designed not only as a benchmark corpus but also as a measurement instrument for longitudinal human learning under verbal guidance.

## 2. Composition, modalities, and file structure

SimCoachCorpus combines synchronized multimodal streams spanning vehicle dynamics, environment representation, audio-transcript pairs, annotations, and surveys. Vehicle data include position in the 2D track plane, velocity, yaw, and driver inputs comprising brake, throttle, and steering. Environmental data include track boundaries as polylines, the raceline as a polyline, cone landmarks as a point set, and start/end gates. The track is Thunderhill Raceway driven counterclockwise [2509.14548].

The audio and transcript layer is central to the corpus. Coach and student audio were recorded and are released as mp3 files with per-speaker SRT subtitles containing line-level timestamps. The time-synchronization and transcript cleaning process is documented through a Subtitle Composer workflow involving splitting and joining lines, standardizing common spellings, removing proctor speech, and eliminating cross-talk. Utterances are aligned to driving segments and can be paired with trajectory indices; this alignment was used during compliance annotation via synchronized session renderings [2509.14548].

The paper does not report sensor sampling rates. It specifies that trajectories and map features are represented in a 2D track coordinate frame, and that many analyses use the longitudinal coordinate along the raceline for spatial density analysis. In modeling experiments, trajectories are normalized relative to the final step of the past trajectory, and the same transform is applied to local map features [2509.14548].

A concise summary of the dataset scale and composition is given below.

| Component | Content | Reported scale |
|---|---|---|
| Participants | Coached and self-practice drivers | 29 total |
| Concurrent feedback | Coach utterances during driving | 20,938 sentences |
| Terminal feedback | End-of-lap feedback sentences | 411 sentences |
| Driving data | Vehicle trajectory recordings | Over 40 hours |
| Retention phase | Final laps without coaching | 2 laps per group |

The file organization comprises per-session mp3 files, per-speaker SRT transcript files such as `coach.srt` and `driver.srt`, trajectory time series, raceline and boundary data, cone data, and annotation files. Exact schemas are not specified in the paper and are instead delegated to the dataset release [2509.14548].

## 3. Experimental design and simulator environment

The experimental setting is a static racing simulator built on CARLA. The reported hardware consists of a Puget Systems workstation with an Intel Core i9 processor, 128 GB RAM, and dual NVIDIA RTX A6000 Ada GPUs. Sessions were conducted on Thunderhill Raceway, with a provided map including start and end gates. The coach had 10 years of professional performance driving experience and more than 6 years of coaching experience, including 2 years in simulators [2509.14548].

The protocol was structured into baseline, training, and retention phases. Participants first completed 2 baseline laps. They then drove through 4–5 driving blocks of approximately 15 minutes each. In the coached condition, the coach provided concurrent feedback while participants were driving and optional terminal feedback after laps, with between-lap conversations of no more than 90 seconds. Both groups concluded with 2 retention laps without coaching [2509.14548].

The introductory experience differed by condition. In the coached condition, the coach drove a demonstration lap while explaining basics such as braking zones and apexes; seat adjustments were made and motion sickness signs were reviewed. In the self-practice condition, participants were shown a video demonstration matched to wheel motion and an input bargraph. Coaching guidelines restricted intervention to language only: the coach did not touch the wheel or point, gave specific action-oriented instructions during driving, reserved explanations for between laps, sat to the student’s right, observed the screen and driver inputs, adapted instruction to prior behavior, and administered surveys approximately every 15 minutes [2509.14548].

This design separates several pedagogically important regimes: uncoached baseline performance, coached training, and post-training retention without active verbal support. This suggests that the dataset is structured to permit analyses of both immediate instructional effects and persistence of learned behavior.

## 4. Annotation taxonomy and survey instrumentation

The annotation scheme for concurrent feedback uses a manually defined, expert-informed taxonomy developed with the coach. It has three high-level types: instruction, feedback, and commentary. Instruction is further subdivided into six domain-relevant categories: throttle, lateral position, steering, looking ahead, turn, and brake. The taxonomy also includes subcategory examples, such as throttle on/off/stay, left/right lateral positioning, steering guidance tied to landmarks, looking-ahead prompts, turn timing, and brake timing or amount [2509.14548].

The concurrent feedback layer contains 20,938 manually labeled sentences. Its top-level distribution is instruction 72.19%, feedback 13.54%, commentary 9.14%, and other 5.13%. Within instruction, the category distribution is throttle 35.75%, lateral position 22.24%, steering 14.44%, looking ahead 11.90%, turn 9.82%, and brake 5.86%. Terminal feedback comprises 411 sentences with category counts per lap; its distribution is steering 30.21%, throttle 17.21%, looking ahead 15.23%, turn 14.98%, lateral position 11.23%, brake 8.14%, non-instruction 2.88%, and other 0.07%. Terminal feedback was given 75.92% of the time at lap ends [2509.14548].

Compliance annotation was conducted on a subset of 15,116 utterances by five annotators using synchronized session renderings. Labels included actionable status, category, a compliance score from 1 to 7, reaction time in seconds, and a free-text behavior description. After removing non-instruction utterances and impossible-to-tell cases such as looking commands, 8,052 utterances remained for analysis. Inter-annotator agreement is not reported; quality control relied on guidance, examples, and an FAQs protocol for edge cases [2509.14548].

Survey instrumentation links subjective state to session and lap structure. Administered approximately every 15 minutes and at session boundaries, the surveys include NASA TLX on a 1–21 scale for mental demand, physical demand, temporal demand, success, and frustration; PANAS short form with positive items such as Determined and Attentive and negative items such as Afraid and Nervous; a fun rating on a 1–10 scale; the Fast Motion Sickness Scale; and a post-drive Intrinsic Motivation Inventory with 7-point Likert ratings across interest/enjoyment, perceived competence, pressure/tension, and perceived choice. Coach surveys provide initial and mid/post assessments of student skill, focus areas, simulator comfort, receptivity, and motivation [2509.14548].

## 5. Performance measures, learning dynamics, and statistical findings

The paper operationalizes several trajectory-level metrics. Lap time is measured in seconds between crossing the start and end gates. Out-of-bounds percentage is defined as
$$
OOB\% = 100 \cdot (1/T) \sum_{t=1}^{T} 1[p(t) \notin TrackBounds].
$$
Raceline adherence is defined as the average lateral distance to the raceline polyline:
$$
RA = (1/T) \sum_{t=1}^{T} ||p(t) - p_r(s_t)||.
$$
Lateral g-force during cornering is also analyzed qualitatively, although the paper does not provide its exact formula [2509.14548].

The comparative trajectory statistics distinguish baseline, driving laps, and retention. For baseline, coached versus self-practice values are 145.00 (10.30) versus 125.00 (7.50) s for lap time, 11.50% (1.17) versus 14.60% (2.32) for out-of-bounds, and 2.50 (0.08) versus 2.97 (0.30) m for raceline adherence. During driving laps, the values are 101.00 (0.78) versus 99.30 (0.63) s, 10.80% (0.38) versus 16.10% (0.46), and 2.46 (0.04) versus 3.01 (0.06) m. During retention, they are 100.00 (2.23) versus 96.40 (2.38) s, 10.30% (1.29) versus 17.10% (1.83), and 2.32 (0.12) versus 2.92 (0.22) m, respectively. Lower is better for all metrics [2509.14548].

The paper reports that percent improvement in lap time relative to initial baseline exceeded 35% in the coached condition and remained below 20% in self-practice. For raceline adherence across all driving trials, Welch’s t-test gives $t(811.15) = -9.08$, $p < .001$, with 95% CI $[-0.75, -0.48]$. For out-of-bounds, Welch’s t-test gives $t(901.67) = -9.61$, $p < .001$, with 95% CI $[-6.36, -4.20]$. Lateral g-force improved by more than 40% in the coached group and approximately 30% in self-practice, but this difference was not statistically significant [2509.14548].

Student state is linked to performance through a linear mixed-effects model predicting mean lap time per session. Significant predictors are session number, with $b = -2.60$, $SE = 0.53$, $t(94.44) = -4.92$, $p < .001$, and fun rating, with $b = -2.27$, $SE = 0.94$, $t(100.2) = -2.41$, $p = .018$. A marginal effect indicates that self-practice was faster, with $b = -6.66$, $SE = 3.87$, $t(29.05) = -1.72$, $p = .096$. The coached group reported higher fun, with Wilcoxon $W = 2671$, $p < .001$, and higher PANAS Positive, with $W = 2674.5$, $p < .001$ [2509.14548].

A plausible implication is that the dataset separates speed from control quality. The self-practice group is reported as faster in absolute lap-time terms during driving and retention, yet the coached group shows better out-of-bounds and raceline adherence statistics and larger improvement relative to baseline.

## 6. Spatial-temporal structure, modeling benchmarks, and limitations

SimCoachCorpus supports analysis of where and when instruction occurs. Instruction density is estimated along the track’s longitudinal coordinate using one-dimensional kernel density estimation with bandwidth $h = 50$ m:
$$
\hat{f}_c(s) = (1/N_c h) \sum_{i=1}^{N_c} K((s - s_i)/h).
$$
Brake and turn instructions exhibit high location specificity, clustering near turns, and the paper reports similar mode specificity for lateral position, looking ahead, throttle, and steering categories. Temporal evolution in instruction counts is modeled with a linear-Gaussian state space model with latent size 1:
$$
z_t = A z_{t-1} + w_t,\; w_t \sim N(0, Q),
$$
$$
y_t = C z_t + v_t,\; v_t \sim N(0, R).
$$
Fitted by EM in Dynamax over 50 effective epochs, this model yields marginal likelihood $-178.06$ and posterior trajectories interpreted as curriculum-like shifts in category usage over laps [2509.14548].

The dataset also supports unsupervised linguistic structure discovery. Topic modeling is performed with BERTopic using transformer embeddings, PCA to 3 components, K-Means with 20 clusters, and c-TF-IDF topic representation. The embeddings `all-MiniLM-L6-v2` and `all-mpnet-base-v2` are both reported to produce robust topics. Against taxonomy labels on a held-out set, the unsupervised topics achieve RAND = 89.1% and V-measure approximately 50.2%. Representative topics include `[gas, back, on, off]` mapped to “back on the gas” at 9.8% of utterances, `[there, you, go, more]` mapped to “There you go” at 9.6%, `[stay, right, to, side]` mapped to “stay to the right” at 7.6%, `[lift, brake, little, bit]` mapped to “lift, lift, lift” at 7.2%, and `[ahead, hands, it, look]` mapped to “keep looking ahead” at 6.7% [2509.14548].

For compliance, a linear mixed model predicts 1–7 compliance from trial number and category with random intercepts per participant. Trial number has $b = 0.010$, $SE = 0.004$, $t(8039) = 2.089$, $p = .037$, indicating a slight increase over time. Relative to brake as the reference category, lateral position has $b = -0.338$, $SE = 0.092$, $t(8037) = -3.69$, $p < .001$, and steering has $b = -0.264$, $SE = 0.098$, $t(8037) = -2.70$, $p = .007$. The interaction Trial $\times$ Lateral position has $b = -0.023$, $SE = 0.005$, $t(8036) = -4.34$, $p < .001$, indicating slower improvement for lateral positioning. A one-way ANOVA gives $F(4, 8047) = 250.80$, $p < .001$, $\eta^2 = .11$, 95% CI $[.10, 1.00]$, with lateral position showing the lowest compliance and braking the highest [2509.14548].

The paper demonstrates three benchmark-style machine learning uses. First, in-context learning uses GPT-4o with few-shot prompts and JSON outputs to generate terminal feedback from concurrent feedback, optionally augmented with segment times and smoothness metrics; structured inputs produce more specific grounding. Second, a joint imitation-learning and trajectory-prediction setup predicts future student trajectory and multilabel coaching-category presence over the next 5 s. The architecture combines an MLP-to-Transformer trajectory encoder, a PointNet-to-Transformer map encoder, an LSTM trajectory decoder with position embeddings and a Transformer decoder, and an MLP instruction classifier with sigmoid outputs; hidden size is 32 throughout. Inputs include past 5 s trajectories, local map patches, and vehicle inputs, and outputs are future 5 s 2D coordinates plus future coaching categories. Training uses AdamW with learning rate $1 \times 10^{-3}$ and linear decay over 300 epochs, with 5 random seeds. The losses are
$$
L_{traj} = (1/T) \sum_{t=1}^{T} ||x_t - \hat{x}_t||^2
$$
and
$$
L_{BCE} = -(1/K) \sum_{k=1}^{K} [ w_k y_k \log(\sigma(\hat{y}_k)) + (1 - y_k) \log(1 - \sigma(\hat{y}_k)) ].
$$
Reported results averaged over 5 seeds are F1 0.607 (0.040) and ADE 4.89 m (2.637) for single-task baselines, and F1 0.621 (0.043) and ADE 4.83 m (0.756) for the multi-task baseline. Third, the topic-dynamics setup uses lap-wise histograms of instruction categories in the state-space model already described [2509.14548].

The corpus is explicitly framed as complementary to prior education, sports analytics, and driving datasets because it combines longitudinal one-on-one coaching, dense state/action streams, student-state surveys, and compliance annotations. At the same time, the paper emphasizes several limitations: a simulator-to-real-world domain gap; the risk of over-reliance on AI advice in real vehicles; single-track, single-coach, and 29-participant coverage; demographic and geographic concentration; the absence of participant video; potential transcript timing or transcription errors; and the absence of reported inter-annotator agreement. The study was IRB-approved, participants consented, compensation was $150, and the authors state that the dataset should not be used to deploy real-vehicle coaching systems without appropriate safety controls. Public release is scheduled upon publication of the peer-reviewed version, with early access available through the stated registration form; the license is not specified in the paper [2509.14548].

Source: https://www.emergentmind.com/topics/simcoachcorpus