Papers
Topics
Authors
Recent
Search
2000 character limit reached

Trust in AVs Scale Measurement

Updated 14 July 2026
  • Trust in AVs Scale encompasses a range of measurement instruments quantifying users' willingness to rely on autonomous vehicles under uncertainty and vulnerability.
  • It integrates methods such as Likert scales, descriptor inventories, situational assessments, and physiological proxies to capture both cognitive and affective trust dimensions.
  • Practical insights stress the importance of matching instrument design to specific trust layers, user roles, and contextual deployment challenges.

Trust in AVs scales are the family of measurement instruments used to quantify how strongly a person is willing to rely on an autonomous vehicle under uncertainty, vulnerability, and incomplete control. Across the literature, trust is not operationalized by a single canonical instrument. Instead, it is measured through several partially overlapping traditions: attribute-based post-drive questionnaires, trust–distrust descriptor inventories, situational event-level scales, layered trait/state frameworks, role-specific adaptations for pedestrians and bystanders, and real-time physiological or behavioral proxies. This plurality reflects a substantive point: trust in AVs is measured differently depending on whether the target construct is general propensity, initial expectation, immediate situational confidence, behavioral reliance, or public trust in deployment and governance (Zhang et al., 2021, Avetisian et al., 2022, Du et al., 2019).

1. Conceptual foundations

Two definitions recur in the AV trust literature. One defines trust as the willingness to be vulnerable to the actions of another party based on the expectation that the other will perform an important action, irrespective of the ability to monitor or control that party; this formulation is explicitly imported into the AV domain in work separating cognitive and affective trust. A second, widely used automation framing defines trust as the attitude that an agent will help achieve a person’s goals in a situation characterized by uncertainty and vulnerability. In AV studies, that goal is typically instantiated as safe driving or safe mobility (Zhang et al., 2021, Avetisian et al., 2022).

These definitions generate different measurement emphases. Some instruments operationalize trust as perceived competence, predictability, dependability, and safety of the AV. Others treat trust as a relation between user and automation that includes emotional comfort, willingness to delegate, or willingness to perform non-driving-related tasks. Still others extend the construct beyond the vehicle’s driving performance toward legibility, explicit communication, accountability, transparency, accessibility, and governance. The qualitative resistance literature is especially important on this point: it links low trust not only to unpredictable or “unhuman” driving behavior, but also to the absence of eye contact or gestures, unclear liability, inaccessible design, surveillance concerns, and a lack of public engagement (Nordhoff, 2023).

A central distinction is between technical trustworthiness and broader socio-technical trust. The former concerns whether the AV behaves safely, consistently, and predictably in interaction. The latter concerns whether deployment is fair, transparent, accountable, and inclusive. This suggests that a general “Trust in AVs” construct is often broader than a pure performance rating, especially outside the driver’s seat.

2. Questionnaire families and scoring schemes

The most established questionnaire family in the provided literature is the Muir-derived AV trust rating. In a simulator study of explanation timing, the AV was rated on seven positively worded traits—competence, responsibility, reliability over time, faith, predictability, dependability, and overall degree of trust—using a 7-point Likert-type scale from 1 = “none at all” to 7 = “extremely high.” The authors treated the seven items as a unidimensional composite trust score and also added a separate ordinal trust ranking across four explanation conditions (Du et al., 2019).

A second family adapts Jian, Bisantz, and Drury’s trust-in-automation descriptor approach to self-driving cars. In the interaction-gap study, the questionnaire contained 12 aspects: deceptive, underhanded, suspicious, wary, harmful, confident, security, integrity, dependable, reliable, trust, and familiar, each rated on a 7-point Likert scale. The instrument was used both as an overall mean trust score and as an item-level profile of trust and distrust facets, which proved important because different facets changed differently over time without interaction (Hunter et al., 2022).

A third family is a broader attitude composite rather than a narrowly interactional trust scale. A large survey of young adults operationalized trust with 10 items on 5-point agreement scales: dependability, adaptability, goal alignment, danger to self, danger to others, safety, general trust, good idea, positivity, and recommendation. The two danger items were reverse-scored, and the resulting composite was later dichotomized at the scale midpoint into high versus low trust for machine-learning analysis (Kaufman et al., 2024).

A fourth family uses ad hoc scenario-specific trust indices. In an online study on trustworthy interaction and level of autonomy, participants selected among six trust statements after each scene: four positive statements about reliability, safe actions, safe decision-making for wellbeing, and following the car’s instructions, and two negative statements about not feeling safe and not trusting the car’s behavior. Positive choices were coded as +1+1, negative choices as 1-1, and summed into a scene-level trust score ranging from 2-2 to +4+4 (Tanevska et al., 3 Feb 2025).

Family Representative format Typical scope
Muir-derived AV trust rating 7 items, 7-point, trait attributions Post-drive composite trust
Jian-derived trust–distrust inventory 12 aspects, 7-point descriptors Overall trust plus facet-level analysis
Broad AV attitude composite 10 items, 5-point agreement General trust/acceptance outcome
Scenario-specific trust index Signed item selections per scene Short-horizon situational trust

The coexistence of these families shows that “Trust in AVs Scale” is best understood as a measurement domain rather than a single fixed questionnaire.

3. Multidimensional, layered, and profile-based models

Several papers reject a single undifferentiated trust score. One line of work imports the cognitive/affective distinction from interpersonal trust theory. In that formulation, cognitive trust is “trust from the head,” grounded in rational, evidence-based assessment of competence, reliability, and predictability, whereas affective trust is “trust from the heart,” grounded in emotional exchange, care, comfort, and perceived benevolence. The proposed AV measure adapts McAllister’s two-factor trust scale to the AV context on a 7-point Likert scale from “strongly disagree” to “strongly agree,” with planned PCA and internal-consistency checks. Because the paper is a design study, it does not report final items, loadings, or reliabilities, but its main contribution is conceptual: it argues that explanations may affect cognitive and affective trust differently (Zhang et al., 2021).

Another line of work follows Hoff and Bashir’s layered framework and separates dispositional trust, initial learned trust, and situational trust. In the anticipated-emotions study, dispositional trust was measured with Merritt’s six-item scale, initial learned trust with Manchon et al.’s ten-item TiAD questionnaire, and situational trust with the six-item Situational Trust Scale for Automated Driving (STS-AD). Only situational trust responded significantly to failure versus success scenarios, with F(1,104)=19.715,p=.000F(1,104)=19.715, p=.000, while dispositional and initial learned trust did not (Avetisian et al., 2022). A later mediation study retained the same layered logic: dispositional and learned trust entered as covariates, while situational trust measured by STS-AD served as the main dependent variable under performance and risk manipulations (Avetisyan et al., 6 Apr 2025).

The profile-based literature pushes this further by combining trust layers with personality and emotion. In conditionally automated driving, one study measured dispositional trust with six Merritt-derived items, initial learned trust with a 10-item AV-specific scale about safety, delegation, complex situations, and non-driving tasks, and dynamic trust with a single repeated 0–10 question, “How much do you trust the AV?” K-means clustering then identified three trust profiles—believers, oscillators, and disbelievers—and a multinomial logistic model predicted these profiles with F1=0.90F1=0.90 and accuracy =0.89=0.89 using 25 top SHAP-ranked features (Avetisyan et al., 2023).

A further extension treats trust as a latent cognitive state inside a Dynamic Bayesian Network. In a study of a self-driving scooter interacting with delivery robots, the user state at event kk was xkE={wk,tk,ik}x_k^E=\{w_k,t_k,i_k\}, where wkw_k is well-being, 1-10 trust in the AV, and 1-11 user intention. Trust in the scooter was measured by a single event-level item, “Based on the current interaction, I trust my self-driving scooter,” on a 7-point Likert scale, then mapped to a discretized latent trust node. This formalization makes trust explicitly dynamic and action-dependent rather than merely retrospective (Zahedi et al., 21 May 2025).

4. Dynamic, behavioral, and multimodal measurement

Situational trust scales become especially important in conditional automation and takeover settings. In a video-based takeover study, self-reported situational trust was measured with STS-AD, using six items: “I trust the automation in this situation,” “I would have performed better than the AV in this situation,” “In this situation, the AV performs well enough for me to engage in other activities,” “The situation was risky,” “The AV made an unsafe judgment in this situation,” and “The AV reacted appropriately to the environment.” On the overall SST measure, system accuracy had a strong main effect, 1-12, with 95% 1-13 80% 1-14 70% (Ayoub et al., 2021).

That same study distinguished self-reported situational trust from behavioral situational trust. Behavioral trust was operationalized through an initial takeover decision and a post-observation revision, summarized by the agreement fraction

1-15

and the switch fraction

1-16

Here SST tracked current system performance, whereas BST was more sensitive to prior trust preconditions such as experimentally induced overtrust and undertrust. This is a substantive measurement result: reported trust and behavioral reliance need not respond to the same causal variables (Ayoub et al., 2021).

Trust can also decay or rebound between encounters. In the interaction-gap study, participants completed the Jian-derived 12-item trust instrument after repeated simulator drives and again after a one-week gap. The change in overall trust during the first interaction and the change across the interaction gap showed a moderate inverse correlation, 1-17, 1-18, 1-19, 2-20. Suspicion showed an especially strong facet-level effect, 2-21, 2-22. The authors interpret this as partial “forgetting” of gained trust or distrust, with scores tending to revert toward earlier trust levels (Hunter et al., 2022).

A more radical move is to use real-time proxy signals rather than questionnaires alone. In conditionally automated driving, repeated single-item trust ratings on a 0–10 scale were collected every 25 seconds and binarized at 5 for machine learning. Using galvanic skin response, heart-rate indices, and eye-tracking features, XGBoost achieved the best real-time trust prediction with 2-23 and accuracy 2-24 (Ayoub et al., 2022). TRUCE-AV generalizes this logic to fully autonomous simulated driving: it combines 1–10 trust votes after events or neutral segments, continuous 0–100 comfort ratings, heart rate, gaze, facial emotions, and environmental variables. Real-time trust correlated strongly with TiA trust (2-25 for drive 1 and 2-26 for drive 2), while discomfort correlated negatively with trust (2-27 and 2-28, both 2-29) (Bhalla et al., 25 Aug 2025).

5. Role- and context-specific adaptations

Most early AV trust scales were built for drivers or passengers, but later work adapts trust measurement to pedestrians, bystanders, and mixed-role interaction. In a naturalistic robotaxi study at an uncontrolled urban intersection, pedestrians completed a pre- and post-experiment Trust in AVs Scale adapted from prior trust-in-automation work. The instrument was described as a 6-point multi-item questionnaire whose higher mean score indicated higher trust in AVs. Trust increased after real-world interaction, and trust change was associated with the PRQF Interaction component and with the PBQ Error and Positive subscales (Chang et al., 30 Sep 2025).

This pedestrian adaptation matters because the underlying trust target shifts. For pedestrians, trust concerns whether AVs are safe, reliable, and will behave appropriately in traffic at crossings; for drivers, trust often concerns whether the AV can safely manage driving so that the human can delegate monitoring or secondary tasks. A plausible implication is that role-sensitive item wording is not a cosmetic issue but a construct-validity requirement.

Community-facing work broadens the context further. The resistance analysis of San Francisco public comments did not administer a formal trust scale, but it documented stable trust-relevant themes: erratic driving, illegal behavior, conflict situations, lack of explicit communication, blocked emergency vehicles, inaccessible design, unclear liability, lack of transparency, privacy concerns, and perceived unfairness of deployment. For vulnerable road users, the absence of eye contact or gestures was repeatedly linked to low perceived safety and low trust. This suggests that scales for public trust in AV deployment may require dimensions—governance, accessibility, data practices, procedural justice—that are absent from cockpit-centered driver scales (Nordhoff, 2023).

Scenario-based driver studies point in the same direction. Trust in online AV vignettes rose with higher allowed levels of autonomy and was negatively associated with AV-NARS negative attitudes, with +4+40 between AV-NARS and trust, and +4+41 between trust and level of autonomy. Qualitative responses centered on safety and performance, transparency, and competency and control, with repeated emphasis on the possibility of takeover and on understanding what the AV was doing (Tanevska et al., 3 Feb 2025).

6. Psychometric properties, interpretation, and limitations

Psychometric evidence is uneven across the literature. The strongest internal-consistency evidence in the provided set comes from the Muir-derived AV trust scale: Cronbach’s +4+42. The same study also reports convergent/discriminant evidence through AVE and correlations: trust rating correlated strongly with preference (+4+43, +4+44), negatively with mental workload (+4+45, +4+46) and anxiety (+4+47, +4+48), and negatively with ordinal trust ranking (+4+49, F(1,104)=19.715,p=.000F(1,104)=19.715, p=.0000) (Du et al., 2019). The high trust–preference correlation is itself important, because the authors note that preference did not show high discriminant validity from trust.

By contrast, several influential instruments are used without fresh reliability estimation in the study that employs them. The STS-AD-based work reports strong sensitivity to manipulated performance but does not report Cronbach’s F(1,104)=19.715,p=.000F(1,104)=19.715, p=.0001 or factor fit in that dataset (Ayoub et al., 2021). The 10-item large-survey composite similarly does not report internal consistency or factor structure, even though its items span dependability, safety, alignment, positivity, and recommendation (Kaufman et al., 2024). The pedestrian public-road study likewise uses a unidimensional mean score without reporting reliability or factor analysis in that context (Chang et al., 30 Sep 2025). The cognitive/affective trust paper explicitly remains at the design stage and reports no empirical psychometrics at all (Zhang et al., 2021).

Another limitation is construct heterogeneity. Some measures target trust narrowly as perceived competence and reliability; others blend trust with safety, affect, recommendation, willingness to delegate, or even acceptance. Some are explicitly state-like, tied to “this situation” or “the current interaction”; others are general attitudes toward AVs as a technology class. Still others embed trust inside larger state architectures involving well-being, intention, or comfort. This diversity is theoretically productive, but it prevents direct score equivalence across studies.

Context dependence is equally consequential. Many studies are simulator-based and use small or specialized samples, including university-centered cohorts, MTurk workers, or region-specific pedestrian groups. Several studies involve always-safe or deliberately manipulated AV behavior, which constrains calibration range. Real-world work remains rarer, though the robotaxi pedestrian study and the resistance study show that on-road interaction and public deployment raise trust concerns that cockpit-only scales do not fully capture (Chang et al., 30 Sep 2025, Nordhoff, 2023).

Taken together, the literature indicates that a “Trust in AVs Scale” is best treated as a family of measurement strategies organized around at least four axes: unidimensional versus multidimensional scoring, trait versus situational timing, self-report versus behavioral or physiological observables, and driver-centered versus role-specific public interaction. The practical consequence is methodological rather than merely terminological: instrument choice should be matched to the trust layer, user role, and deployment context being studied.

Definition Search Book Streamline Icon: https://streamlinehq.com
References (14)

Topic to Video (Beta)

No one has generated a video about this topic yet.

Whiteboard

No one has generated a whiteboard explanation for this topic yet.

Follow Topic

Get notified by email when new papers are published related to Trust in AVs Scale.