Papers
Topics
Authors
Recent
Search
2000 character limit reached

How Much Trust is Enough? Towards Calibrating Trust in Technology

Published 7 Apr 2026 in cs.HC | (2604.05658v1)

Abstract: The role of trust within Human-Computer Interaction is being redefined. With the increasing omnipresence, autonomy, and opacity of technology, users often struggle to understand the capabilities and limitations of systems. In this article, we present the results of an empirical study designed to provide a practical, evidence-based interpretation of trust propensity assessment using the Human-Computer Trust Scale (HCTS). We outline the process used to develop a guideline for interpreting the instrument's results and explain the rationale for our decisions, advocating for calibrating trust in technology within HCI. Our findings demonstrate that the HCTS is a promising tool for conducting an initial evaluation of propensity to trust, but that such an assessment requires reflection and interpretation that should be considered within the context of the interaction.

Summary

  • The paper develops empirically grounded HCTS interpretation bands—undertrust at ≤2.30, adequate trust at 2.31–3.60, and overtrust at ≥3.61—using SUS-inspired analyses across 938 respondents and two biometric technologies.
  • The proposed framework treats moderate-to-high trust, rather than complete trust, as the calibration target and recommends interpreting aggregate scores alongside competence, benevolence, perceived risk, and structural assurance.
  • The findings provide a practical baseline but require validation across more technologies and contexts, measurement-invariance testing, and studies linking trust propensity with actual reliance behavior.

Motivation and framing

The paper addresses a persistent interpretive gap in trust measurement within Human-Computer Interaction (HCI). Although validated instruments exist for assessing propensity to trust technology, their numeric outputs have historically been interpreted only through relative, context-specific comparisons—between constructs, between groups, or between studies. The authors argue that this limitation has become more consequential as autonomous, opaque systems proliferate and the field's concern shifts from fostering trust toward calibrating it: aligning user trust with a system's actual capabilities so that neither overtrust (misuse) nor undertrust (disuse) degrades interaction quality (2604.05658).

The work centers on the Human-Computer Trust Scale (HCTS), a nine-item psychometric instrument grounded in Mayer's organizational trust model and its socio-technical extension, the Human-Computer Trust Model. The revised HCTS measures four constructs on 5-point Likert items: Competence (COM), Benevolence (BEN), Perceived Risk (PR), and Structural Assurance (SA). Prior work established partial measurement invariance across countries for the scale; the present paper extends it by deriving empirically grounded interpretation ranges for the aggregate trust score.

Methodology

The procedure is explicitly modeled on the development of acceptability ranges for the System Usability Scale (SUS), where adjective ratings were mapped onto score bands to aid interpretation. Two online survey studies were conducted between 2023 and 2025: Study 1 examined facial recognition systems for law enforcement (N = 711, aggregated from seven surveys, stimulus videos of China's Skynet and London's Metropolitan Police), and Study 2 examined biometric payment systems (N = 227, aggregated from five surveys). Both used video stimuli depicting hypothetical system implementations and recruited participants via convenience sampling across multiple world regions.

After completing the HCTS, participants rated their trust using a seven-point adjective scale ("I trust it completely" through "I completely do not trust it"), constructed following Dodd & Gerbrick's criteria for word scales. The analysis proceeded in three steps:

  1. Within-study label estimation: one-way ANOVAs tested whether HCTS scores differed across adjective categories, with 95% confidence intervals computed around each group mean.
  2. Within-study boundary estimation: cut-offs were placed at midpoints between adjacent categories' CI bounds.
  3. Cross-study generalization: boundaries were synthesized into a single interpretation scheme.

A key conceptual commitment shapes the entire procedure: guided by Muir's calibration framework and Lee & See's resolution-based model, the authors treat moderate-to-high trust—not maximal trust—as the adequate target. Consequently, both tails of the adjective scale are treated as problematic: the lowest ratings indicate undertrust, while "I trust it completely" is interpreted as a potential marker of overtrust rather than a desirable outcome.

Results

Study 1 showed significant score differences across all seven original adjective categories immediately; Study 2 did not, requiring two iterations of merging adjacent categories that failed pairwise significance tests. The final recoded scale comprises four categories, with significant ANOVA results in both studies (F(3, 695) = 46.745, p < .001 for Study 1; F(3, 213) = 13.273, p < .001 for Study 2). Because Levene's test indicated heterogeneity of variances in both cases, Games-Howell post-hoc comparisons were used. Cronbach's alpha was 0.814 (Study 1) and 0.860 (Study 2).

Notably, the derived cut-offs were highly consistent across the two studies despite differences in sample size, region, and system, supporting the robustness of the boundaries. Using the pooled CI bounds—and deliberately adopting the wider interval bounds rather than point cut-offs to accommodate trust's contextual variability—the authors propose the following interpretation ranges:

Interpretation HCTS score
Undertrust ≤ 2.30
Adequate 2.31 – 3.60
Overtrust ≥ 3.61

Two observations deserve emphasis. First, the "adequate" range centers near the scale midpoint (3), which the authors flag as an open question: it may reflect genuine neutrality of appropriate trust, or an artifact of the category-merging procedure combined with their own framing that both extremes warrant attention. Second, the ranges are presented visually as a continuum with transition zones rather than hard thresholds, reinforcing that they serve as interpretive guides, not normative verdicts.

Discussion

The authors position the resulting guideline as a baseline for initial interpretation that must be complemented by construct-level breakdowns (the HCTS permits separate COM, BEN, PR, and SA means) and comparative analyses across respondent groups. They invoke Lee & See's notion of resolution to justify treating extreme self-ratings as anchors of poor calibration rather than direct behavioral markers—an important caveat, since propensity measures capture disposition, not reliance behavior, and the translation from disposition to behavior remains contingent on cognitive, social, and cultural factors.

The choice to use CI bounds rather than point estimates broadened the "adequate" band intentionally, prioritizing practical flexibility over precision. The authors also note a methodological counterpoint regarding Study 2's smaller sample: while reduced power limited detectable distinctions, smaller samples can surface subgroup variation that aggregation obscures in larger datasets, making Study 2 a useful exploratory complement rather than merely a weaker replication.

Limitations and open questions

The paper is candid about several constraints. The threshold derivation rests on only two studies covering two technically similar biometric technologies, which scarcely represent the heterogeneity of technological systems and interaction contexts; measurement invariance was not assessed across contexts due to sample-size asymmetries. The procedure itself involves substantial subjectivity—in selecting adjectives, in merging categories, and in deciding what constitutes "adequate" trust—a subjectivity the authors acknowledge mirrors the SUS developers' own concession that usability has no absolute scale. Whether the observed adequate range truly reflects appropriate trust levels, or requires refinement, is left explicitly unresolved, as is the question of whether midpoint scores indicate calibrated neutrality or strategic disengagement. Finally, because the assessment targets pre-interaction propensity, the guideline provides no direct evidence about how these baselines predict or should be updated against actual reliance behavior during interaction.

Conclusion

This paper converts the HCTS from a purely relative comparison tool into an instrument with actionable, empirically derived interpretation bands—undertrust (≤ 2.30), adequate (2.31–3.60), and overtrust (≥ 3.61)—grounded in a SUS-inspired mapping procedure applied to 938 respondents across two studies. Its central contribution is less the specific numbers than the reframing they encode: adequate trust propensity is moderate, not maximal, and any threshold-based reading must be situated within the interaction context and supplemented by construct-level analysis. The proposed ranges are offered as a starting point for reflection rather than a finished standard, with cross-domain validation and behavioral linkage identified as the outstanding work needed to solidify them.

Paper to Video (Beta)

No one has generated a video about this paper yet.

Whiteboard

No one has generated a whiteboard explanation for this paper yet.

Open Problems

We haven't generated a list of open problems mentioned in this paper yet.