Papers
Topics
Authors
Recent
Search
2000 character limit reached

Epistemic Trustworthiness in Generative AI: A Normative Framework for Warranted Reliance in High-Stakes Workflows

Published 6 Aug 2026 in cs.AI and cs.HC | (2608.05602v2)

Abstract: Generative AI systems are increasingly deployed in high-stakes professional contexts, where their outputs shape what users believe, how they reason, and what they treat as settled. This raises a central question for responsible AI: under what conditions is reliance on generative AI outputs epistemically warranted rather than behaviourally induced? Existing frameworks largely ask whether AI outputs are accurate, fair, explainable, safe, or trusted by users. These questions remain necessary, and each can contribute to warranted reliance. However, they do not directly specify warranted reliance as a distinct evaluative target: the conditions under which users are justified in treating AI outputs as inputs into their own reasoning. We argue that this requires an account of epistemic trustworthiness: what makes a system epistemically worthy of reliance. Drawing on philosophical accounts of trustworthiness as competence and audience-orientation, we develop a constitutive normative framework comprising three jointly necessary and non-fungible conditions. First, epistemic humility requires systems to represent and communicate the limits of their competence. Second, epistemic access requires systems to enable users to inspect, question, and contest outputs in context. Third, resistance to epistemic injustice requires systems to recognise users as legitimate epistemic agents and avoid marginalising their knowledge and experience. Through real-world case analyses in legal reasoning, medical reasoning, and hiring, we show how failures of epistemic humility, epistemic access, and resistance to epistemic injustice can produce consequential harms that standard measures of accuracy, fairness, and usability do not address on their own. We conclude by outlining design and evaluation implications for GenAI systems organised around epistemically warranted reliance rather than output correctness alone.

Summary

  • The paper proposes a conjunctive framework in which epistemic humility, epistemic access, and resistance to epistemic injustice are jointly necessary for warranted reliance on generative-AI outputs.
  • The paper shows how legal hallucinations, unsupported citations, identity-sensitive resume rankings, and role-dependent medical refusals can undermine trustworthiness despite apparent accuracy or safety.
  • The paper recommends relational, layer-specific audits and calibrated friction, including actionable uncertainty signals, claim-level provenance, contestation pathways, and participatory evaluation.
  • ],
  • ,

Epistemic Trustworthiness in Generative AI

Central Problem and Contribution

“Epistemic Trustworthiness in Generative AI: A Normative Framework for Warranted Reliance in High-Stakes Workflows” (2608.05602) addresses a specific deficiency in current responsible-AI research: existing evaluations of accuracy, calibration, fairness, explainability, safety, and user trust do not directly determine when reliance on a particular generative-AI output is epistemically warranted. The paper’s central question is therefore relational and situational: what must a GenAI system provide, at the moment of interaction, for a situated user to be justified in relying on, verifying, contesting, or withholding reliance from a particular output?

The paper distinguishes epistemically warranted reliance from reliance that is merely behaviourally induced. Fluent language, authoritative style, conversational presentation, citations, and confidence displays can produce deference even when they do not track evidential quality. This distinction is supported by empirical work showing that conversational presentation can impair users’ ability to detect inaccuracies [10.1038/s41598-024-67829-6] and that authoritative communicative style can increase trust even when limitation disclaimers fail to recalibrate it [10.1145/3613904.3642122]. The paper consequently treats reliance not as an outcome to be maximised, but as a normative relation requiring adequate grounds.

Its principal contribution is a constitutive framework with three jointly necessary and non-fungible conditions:

  1. Epistemic humility: the system represents and communicates the limits of its competence.
  2. Epistemic access: users can inspect, interpret, question, and contest the basis of outputs in context.
  3. Resistance to epistemic injustice: the system recognises users and affected communities as legitimate epistemic agents rather than discounting their testimony, concepts, language, or situated knowledge.

These conditions presuppose first-order accuracy but cannot be reduced to it. The framework is explicitly normative rather than a completed measurement instrument. It identifies the properties that evaluation must address while leaving thresholds and implementation mechanisms to context-sensitive empirical and institutional analysis.

Philosophical Derivation

The framework translates philosophical accounts of trustworthy testimony into the sociotechnical setting of GenAI. Testimonial trustworthiness is understood as involving both competence and audience-orientation. Competence concerns the capacity to produce well-supported claims. Audience-orientation concerns responsiveness to the epistemic situation and needs of those who depend on the source.

The paper argues that first-order competence—measured through accuracy, calibration, or benchmark performance—is insufficient. A system may be accurate on average while failing to identify when a particular answer is outside its reliable operating range. The relevant requirement is second-order competence: the ability to recognise and communicate limitations. In GenAI, this becomes epistemic humility.

Audience-orientation is divided into two dimensions. Its practical-communicative dimension yields epistemic access: the user must be able to evaluate the evidential basis and operational boundaries of an output. Its recognitional dimension yields resistance to epistemic injustice: users must be treated as legitimate knowers whose experience and contextual knowledge can affect the interaction. Figure 1

Figure 1: Derivation of epistemic humility, epistemic access, and resistance to epistemic injustice from competence and audience-orientation.

The derivation is important because it prevents the framework from becoming a list of familiar responsible-AI desiderata. Explainability, provenance, abstention, fairness, and calibration are treated as possible mechanisms or indicators. They are not themselves the constitutive conditions. For example, citations may support access, but citations that do not substantiate the associated claims do not establish epistemic access. Similarly, a fairness metric may identify disparate treatment while failing to establish that a system recognises the epistemic standing of affected communities.

The framework also adopts an instrumental stance. It does not attribute belief, knowledge, self-awareness, or moral agency to LLMs. Instead, it evaluates whether systems function as if they exhibit relevant epistemic properties: whether they communicate uncertainty, support justified belief formation, and respond to users as legitimate participants in epistemic exchange.

The Three Conditions

Epistemic Humility

Epistemic humility is defined as the system’s capacity to make material limitations visible when reliance is being formed. It is not equivalent to generic disclaimers, low confidence, indiscriminate refusal, or occasional abstention. A humble system must help users determine whether an output is grounded, uncertain, underspecified, outside the system’s reliable scope, or in need of external verification.

The paper identifies two components. Actionable limitation-signalling requires the system to explain what the user should do in response to a limitation. Merely stating that “errors may occur” does not tell a professional whether to verify a citation, seek an alternative source, defer action, or escalate the matter. Interactional humility requires limitation-signalling to persist across turns. A system that appropriately qualifies an initial answer but confidently reconfirms it when challenged remains epistemically untrustworthy.

This distinction is reinforced by MetaMedQA, where models achieved high performance on standard medical question-answering tasks but consistently failed to detect absent correct answers and their own knowledge limitations [10.1038/s41467-024-55628-6]. The result is theoretically significant: metacognitive reliability does not automatically emerge from improved first-order accuracy. A model can know many correct answers without reliably knowing when it does not know.

The Mata v. Avianca litigation illustrates the practical consequence. ChatGPT generated fictitious legal authorities and subsequently reaffirmed their existence when asked for confirmation. The lawyers retained responsibility for verification, but the system also failed to communicate that the authorities were unverified and failed to revise its response under direct challenge. The relevant defect was therefore not only hallucination, but a failure of interactional humility.

Epistemic Access

Epistemic access concerns whether users can practically inspect and contest the basis of an output. It includes three elements:

  • Verifiable claim–evidence linkage: users can determine whether a cited source actually supports the claim.
  • Inspectable retrieval: users can examine why specific evidence was selected and identify omissions, assumptions, or conflicting sources.
  • Contestability: users can challenge, correct, override, or escalate an output through an effective interactional or institutional pathway.

The paper rejects the assumption that explanations, documentation, or citations automatically establish access. Access is a practical property of the user–system relation, not a formal property of the interface. It depends on whether representative users can perform meaningful verification within realistic constraints of expertise, time, workload, and authority.

The analysis of legal RAG systems makes this point particularly clearly. In an evaluation of Lexis+ AI, Westlaw AI Assisted Research, and Ask Practical Law AI, hallucinated responses occurred in 17% to 33% of evaluated queries. The failures included fabricated sources and misgrounded responses in which real sources did not support the propositions for which they were cited [magesh2025hallucination]. A real citation can therefore create a stronger appearance of reliability than a fabricated one while requiring substantial legal analysis to detect its misuse.

The implication is that source-grounding systems must be evaluated at the level of claim-to-evidence correspondence, retrieval transparency, and correction workflows. Retrieval augmentation can improve factual grounding, but it does not by itself create epistemic access [10.1145/3637528.3671470; (2005.1145)?]. The system must expose enough of the evidential structure for users to determine whether reliance is justified.

Resistance to Epistemic Injustice

Resistance to epistemic injustice requires systems to avoid systematically downgrading or excluding users’ credibility, concepts, language, and contextual experience. The paper distinguishes this condition from aggregate fairness. A system can satisfy some statistical fairness criteria while still failing to recognise particular people or communities as legitimate sources of knowledge.

The condition has two components. Equal credibility weighting concerns whether identity-signalling features affect the treatment of substantively equivalent evidence. Recognitional adequacy concerns whether users outside institutionally privileged roles are treated as legitimate knowers whose needs and contextual knowledge can shape system behaviour.

The resume-screening case provides a quantitative illustration. Wilson and Caliskan varied names while holding resume content constant across 554 resumes, 571 job descriptions, and nine occupations. Across 27 model–occupation comparisons, models favoured White-associated names in 85.1% of racial comparisons and male-associated names in 51.9% of gender comparisons. White-male-associated names were favoured over Black-male-associated names in all 27 comparisons [10.5555/3716662.3716799]. These results indicate more than unequal rankings. In retrieval-based screening, identity-sensitive ranking controls whether evidence of qualification becomes visible to the recruiter at all.

The paper carefully limits the inference. The benchmark establishes identity-sensitive retrieval but does not by itself identify the causal mechanism or demonstrate how recruiters use the ranking in deployment. Nevertheless, it supports an R1 diagnosis at the retrieval stage. Explanations and contestation mechanisms could make the bias visible, but they would not repair the underlying credibility weighting. Remediation must address the data, embeddings, ranking procedure, and institutional workflow.

The IatroBench analysis extends this concern to medical interaction. Across 3,600 responses generated from 60 clinical scenarios and six frontier models, physician-framed prompts produced more complete guidance than matched patient-framed prompts. The benchmark also reported a substantial evaluator failure: a standard LLM judge assigned an omission-harm score of zero to 73% of responses that physicians rated as harmful, with negligible agreement, κ=0.045\kappa = 0.045 (Gringras, 9 Apr 2026). The result challenges the assumption that refusal is intrinsically safe. A patient may be denied clinically relevant guidance because the system implicitly treats professional status as a proxy for epistemic legitimacy.

Non-Fungibility and Conjunctive Evaluation

A major theoretical claim is that humility, access, and resistance to epistemic injustice are non-fungible. They cannot be averaged into a compensatory trust score. High performance on one dimension cannot offset failure on another.

The paper gives three countermodels:

  • A system may communicate uncertainty and provide inspectable outputs while still marginalising a community’s knowledge.
  • A system may recognise users as legitimate knowers and signal limitations while leaving them unable to inspect or contest outputs.
  • A system may provide extensive access and recognise users’ epistemic standing while expressing confidence that exceeds its evidential support.

The third case is especially important because it challenges the common assumption that transparency and participation can compensate for poor calibration. If a system communicates unsupported claims with unwarranted confidence, users lack the limitation signals necessary to calibrate deference, even if the underlying evidence is technically inspectable.

This structure implies that evaluation should produce a diagnostic profile rather than a single global trustworthiness score. Audits should identify which condition fails, at which sociotechnical layer, for which users, and with what consequences. A system can be highly accurate and well documented yet remain unsuitable for a particular high-stakes workflow because it fails recognitional adequacy or interactional humility.

Sociotechnical Intervention Points

The conditions have distinct primary intervention points, although none can be satisfied at only one layer.

Epistemic humility begins principally at the model layer, where the system must estimate its operating limits, evidence coverage, and uncertainty. It becomes epistemically relevant only when surfaced through the interface in a form users can interpret and act upon. Confidence estimates generated by the same linguistic mechanism responsible for hallucination are especially problematic; verbal confidence often remains high even when factual reliability is low [10.1145/???]. More robust approaches may combine retrieval quality, semantic uncertainty, source agreement, perturbation stability, and external verification status.

Epistemic access begins at the interface layer through claim-level provenance, inspectable retrieval, evidence comparison, correction controls, and escalation mechanisms. It depends on model and data infrastructures capable of preserving traceable relationships between claims and sources. Full chain-of-thought disclosure is neither required nor necessarily desirable. The relevant target is usable evidential access, not unrestricted exposure of internal generation traces.

Resistance to epistemic injustice is distributed across the stack. Dataset composition determines which knowledge forms are represented; retrieval and ranking mechanisms determine whose testimony becomes salient; model objectives influence credibility weighting; and interaction design determines whether user expertise can modify the system’s interpretation. Participatory audits, multilingual evaluation, subgroup-specific stress tests, and deployment studies with affected communities are therefore necessary complements to aggregate fairness metrics.

Calibrated Friction

The paper proposes calibrated friction as a design implication. Interfaces should not treat all outputs as epistemically equivalent. When evidence is incomplete, conflicting, weakly grounded, or consequential, the system should increase the effort required before users accept or act on the output.

Appropriate friction may include claim-level uncertainty markers, source-inspection prompts, disagreement indicators, confirmation steps, and explicit routes for correction or escalation. Its justification is normative rather than merely behavioural: friction is warranted when it protects users’ capacity for epistemic judgement in proportion to the risk of reliance.

The proposal is more demanding than simply adding warnings. Generic disclaimers often fail to change trust, while authoritative style can continue to induce deference [10.1145/3613904.3642122]. Friction must redirect attention toward valid reliability cues rather than encouraging users to inspect misleading confidence scores or polished explanations. It must also avoid excessive interruption, which can make warranted reliance impractical or produce warning fatigue.

Evidence from AVA-AI illustrates the contextual nature of this trade-off. In an early deployment with approximately 50 curated reports, abstention rates reached 40–70%, and users interpreted refusals as evidence of limited capability. After the corpus expanded to more than 4,000 World Bank reports, abstention fell below 10%, and reasoned abstention became legible as a reliability signal. The result suggests that abstention is not intrinsically useful or harmful: its epistemic function depends on evidence coverage, explanation, provenance, and users’ understanding of system scope. The same refusal behaviour can signal responsible limitation or inadequate capability depending on deployment conditions.

Theoretical and Practical Implications

Theoretically, the paper shifts the unit of analysis from system-level trustworthiness and user-level trust to the situated user–output relation. This move clarifies why output correctness is insufficient. A correct answer can be accepted for bad reasons, while an incorrect answer can be accompanied by a clear warning that prevents reliance. Epistemic evaluation must therefore consider the grounds made available to users, not only whether the final proposition happens to be true.

The framework also contributes a normative account of human oversight. Oversight is meaningful only when users possess sufficient epistemic access and practical authority to evaluate and alter system-mediated outcomes. This is consistent with work arguing that effective oversight requires both causal power and access to relevant information [10.1145/363(0106.36590)51]. “Human in the loop” is consequently not an adequate safeguard when the human cannot determine what the system did, why it did it, or how to contest it.

Practically, high-stakes deployments should implement separate audits for:

  • multi-turn limitation signalling and abstention;
  • calibration against evidence quality and domain boundaries;
  • claim-level source support and retrieval completeness;
  • user verification effort under realistic workflow constraints;
  • contestation, correction, and escalation effectiveness;
  • identity-sensitive retrieval and credibility weighting;
  • differential treatment of professional and non-professional users;
  • representational adequacy for affected communities.

Future AI systems are likely to require integrated epistemic control planes rather than isolated confidence displays. Such systems could track provenance, evidence conflicts, retrieval coverage, uncertainty under perturbation, user corrections, and subgroup-specific behavioural variation. The paper’s framework suggests that these signals should be exposed through context-sensitive interaction policies, while governance processes should determine thresholds appropriate to the stakes of a deployment.

Limitations and Future Research

The paper acknowledges that its framework is not yet operationalised as a validated scoring system. The most difficult condition is resistance to epistemic injustice, particularly recognitional adequacy. Existing fairness benchmarks do not adequately measure whether systems treat local, experiential, non-dominant, or community-based knowledge as epistemically relevant.

Several empirical programmes follow naturally. First, researchers should develop task- and domain-specific measures of actionable humility, including whether users make better decisions after receiving limitation signals. Second, epistemic access should be evaluated through realistic professional studies that measure verification time, error detection, successful correction, and escalation outcomes rather than merely the presence of citations. Third, resistance to epistemic injustice requires longitudinal and participatory evaluation involving affected communities. Fourth, future benchmarks should evaluate refusal quality, including omission harm, rather than treating refusal as a uniformly positive safety outcome.

The framework also raises questions about accountability. If a system’s output is accurate but its confidence is miscalibrated, or if its sources are technically available but practically unverifiable, responsibility cannot be assigned solely to the end user. Warranted reliance is produced by the interaction of model, interface, data, institution, and professional norms. Future regulation and procurement standards may therefore need to specify relational evidence of warranted reliance rather than relying exclusively on model cards, aggregate benchmarks, or generic human-oversight requirements.

Conclusion

“Epistemic Trustworthiness in Generative AI: A Normative Framework for Warranted Reliance in High-Stakes Workflows” (2608.05602) argues that responsible GenAI deployment requires more than accurate outputs or calibrated user behaviour. It requires systems to establish the conditions under which users can appropriately defer, verify, contest, or reject particular outputs in context.

The proposed framework identifies three jointly necessary conditions: epistemic humility, epistemic access, and resistance to epistemic injustice. Their non-fungibility prevents strong performance in one dimension from concealing serious failures in another. The case analyses show how hallucinated legal authorities, misgrounded citations, identity-sensitive resume retrieval, and role-contingent medical refusals can undermine warranted reliance even when systems appear useful, source-grounded, safe, or accurate under conventional evaluation.

The paper’s principal implication is methodological: high-stakes GenAI should be evaluated through conjunctive, relational, and layer-explicit audits. Its principal design implication is that systems should introduce calibrated friction grounded in reliable epistemic signals. Future progress will depend on developing technical and institutional mechanisms that preserve human epistemic agency—not by eliminating reliance on generative systems, but by ensuring that reliance is supported by adequate and contestable grounds.

Whiteboard

Open Problems

We found no open problems mentioned in this paper.

Collections

Sign up for free to add this paper to one or more collections.

Tweets

Sign up for free to view the 1 tweet with 1 like about this paper.