Papers
Topics
Authors
Recent
Search
2000 character limit reached

Between Knowledge and Care: A Mixed-Methods Evaluation of Generative AI for T2DM Self-Management from Patient and Physician Perspectives

Published 4 Jul 2026 in cs.HC | (2607.03720v1)

Abstract: Generative AI is increasingly used for everyday health guidance, yet its clinical appropriateness in chronic disease contexts remains poorly understood. This paper presents a two-part mixed-methods study on \revise{Type 2 Diabetes Mellitus (T2DM)}, examining how patients and physicians assess AI-generated health information. \revise{Study~1} analyzes 784 \revise{participant reported} patient queries to characterize seven informational need categories and \revise{develops a structured five dimensional physician rating rubric informed by patient query categories and clinician priorities} (\textit{Accuracy, Safety, Clarity, Integrity, Action Orientation}). \revise{Study~2} engages seven physicians scoring responses from four AI models and discussing evaluative reasoning through in-depth interviews. Models perform well on factual explanation and lifestyle guidance but consistently underperform on medication reasoning and emotional support. Two \revise{analytic concepts} emerge \revise{from the data}. The \textit{pre-visit primer} \revise{frames AI as preparation for clinical encounters rather than as a replacement for physicians}. The \textit{fluency illusion} \revise{describes how polished language may convey epistemic authority that the clinical content does not support}. Patients and physicians converged on three shared limitations (role boundaries, emotional inadequacy, personalization gaps) while diverging in evaluative emphasis, \revise{which informed} four design directions, task-aware orchestration, risk-aware fallback, dynamic personalization, and emotionally attuned interaction.

Summary

  • The paper presents a mixed-methods evaluation assessing generative AI's role in T2DM self-management using patient-derived queries and clinician rubrics.
  • It quantifies model performance, with ChatGPT excelling in accuracy and safety while revealing gaps in personalization and emotional support.
  • Findings underscore the need for task-aware orchestration and risk-aware fallback strategies to improve AI deployment in chronic care management.

Mixed-Methods Evaluation of Generative AI for T2DM Self-Management: An Expert Analysis

Introduction

This paper presents a rigorous mixed-methods evaluation of generative AI systems used for self-management in Type 2 Diabetes Mellitus (T2DM), foregrounding both patient-reported informational needs and structured physician assessment. With the growing utilization of LLMs for chronic care information, understanding their clinical appropriateness, safety, and user alignment is critical. The work is notable for deriving its evaluation framework directly from participant-elicited queries rather than synthetic or guideline-generated prompts, enabling a more precise characterization of real-world chronic care information gaps. The study triangulates quantitative rubric-based assessment with qualitative perspectives from both stakeholders, yielding actionable implications for AI design, boundary-setting, and risk mitigation in the chronic disease domain.

Patient Informational Needs and Attitudes

The initial phase synthesizes 784 patient-elicited T2DM queries from 21 participants, revealing seven stable informational need categories: Factual Knowledge, Diet Management, Sports Advice, Medication Guide, Medication Interpretation, Complications Related, and Life and Psychology. Attitude data indicate that patients rate AI highly for convenience (M=9.05M=9.05), accessibility, and efficient pre-visit information preparation (the "pre-visit primer" role), but substantially lower for psychological support (M=5.10M=5.10) and personalization (M=5.62M=5.62). Qualitative data underscore the distinction between AI as a preparatory educational tool and as a clinical authority, with explicit patient deference to physicians in cases of conflicting advice. Personalization and emotional inadequacy were repeatedly flagged as structural limitations.

Figure 1

Figure 1: Heatmap summarizing System Usability Scale-style ratings, highlighting high marks for convenience but lower scores for personalization and psychological support.

Structured Physician Evaluation: Rubric and Protocol

The second study engages seven endocrinologists and related specialists in a controlled evaluation of responses from four mainstream LLMs (ChatGPT, DeepSeek, Kimi, ERNIE Bot) to a curated dataset of 66 representative patient questions, spanning the seven identified categories. The evaluation rubric operationalizes five dimensions—Accuracy, Safety, Clarity, Integrity, Action Orientation—weighted by clinical priority (Accuracy and Safety: 30 points each; others: 20 points each, maximum 120 points).

Physician recruitment emphasized clinical expertise and independence from model development. Evaluation was followed by semi-structured interviews probing comparative model performance, safety calibration, and practical deployment boundaries.

Model-Wise and Dimension-Wise Performance

Quantitative assessment shows strong differentiation across models. ChatGPT achieved the highest aggregate mean (M=103.07M=103.07, SD=7.61SD=7.61), outperforming DeepSeek (M=97.51M=97.51), ERNIE Bot (M=88.36M=88.36), and Kimi (M=86.46M=86.46). The performance gap is statistically robust (F(3,18)=2315.46F(3,18)=2315.46, p<0.001p<0.001). Notably, Action Orientation is the weakest dimension even for top-performing models, while directive clarity is often decoupled from actual safety, particularly for ERNIE Bot.

Figure 2

Figure 2: Boxplot comparison of model performance across all evaluation dimensions and question types; ChatGPT and DeepSeek feature higher means and greater variance.

ChatGPT maintains superiority across all five dimensions, but particularly in Accuracy and Safety. ERNIE Bot, while globally lowest in composite quality, paradoxically achieves the highest Action Orientation (M=5.10M=5.100), albeit at the expense of omissions in critical safety warnings.

Figure 3

Figure 3

Figure 3

Figure 3

Figure 3: Boxplots of dimension-wise scores, visualizing model-specific distributions for Accuracy, Safety, Clarity, Integrity, and Action Orientation.

Figure 4

Figure 4: Radar chart depicting average scores by dimension for each model; area size indicates overall dimensional performance.

Task and Category Sensitivity

Performance is also strongly category-dependent. Factual Knowledge, Diet Management, and Sports Advice receive the highest quality ratings, consistent with their standardized nature and lower clinical risk. In contrast, Medication Guide and Medication Interpretation yield the lowest scores, registering a ~10-point drop in aggregate, corresponding with increased failure rates in context-sensitive reasoning and safety-critical guidance.

Figure 5

Figure 5: Physician ratings across seven T2DM question categories, demonstrating domain-driven variability in response suitability.

Dimension-level analysis reveals that models excel on standard education but systematically fail on tasks involving clinical inference, individualized recommendations, or emotional support. The heatmap of scores across all model-dimension-category triplets illustrates persistent asymmetries, with no model solving actionability and contextual safety in medication domains.

Figure 6

Figure 6

Figure 6

Figure 6

Figure 6

Figure 6

Figure 6

Figure 6: Distribution of dimension-level ratings for each question category, indicating variable suitability by informational need.

Figure 7

Figure 7: Radar chart comparison of model performance across all question categories.

Figure 8

Figure 8: Comprehensive heatmap of normalized mean scores by combination of model, evaluation dimension, and question category.

Qualitative Insights and Analytic Concepts

Physician interviews corroborate quantitative trends, emphasizing LLM reliability in factual education and lifestyle domains, and pronounced deficiencies in contextual, medication-related, and affect-laden advice. ChatGPT is perceived as safest due to robust disclaimers and guideline-aligned explanations; Kimi is considered unsuitable for any unsupervised deployment due to omissions in critical safety actions (e.g., hypoglycemia management). ERNIE Bot's procedural strength is undermined by lack of factual nuance, sometimes producing dangerous advice through excessive directness.

A central analytic contribution is the formulation of the fluency illusion: the risk that well-structured, confidence-signaling language is misinterpreted by users as authoritative, masking substantive clinical insufficiency. Physicians report that surface-level quality frequently substitutes for genuinely grounded medical guidance—a pattern especially problematic for lay users with limited domain knowledge.

A second concept, the pre-visit primer, describes the preparatory, health-literacy-raising role that patients and clinicians independently endorse for AI, in contrast with substitution for clinical judgment.

Synthesis: Convergences, Divergences, and Boundary Setting

Triangulation reveals strong convergence between patients and clinicians on three points:

  • Role boundaries: Both groups reject AI as a clinical decision-maker, situating its value in health education and pre-consultation scaffolding.
  • Personalization deficit: Both identify failure to adapt guidance to comorbidities and life context as a significant limitation.
  • Emotional inadequacy: Emotional support is lowest-rated by both, with "robotic", "formulaic" affective responses insufficient for psychosocial needs.

Divergences are evident in trust calibration—patients overweight convenience and linguistic quality, while physicians privilege safety, explicit disclaimers, and context-specific warnings—and in risk sensitivity, with clinicians noting population-level safety risks invisible to patients.

Implications for AI Health Tool Design

The findings support four design imperatives:

  1. Task-aware orchestration: Query context-aware routing to select between models and response paradigms, particularly isolating clinical/safety-critical queries for escalation or added safeguards.
  2. Risk-aware fallback: Automated detection of medication management, symptom triage, or high-stakes queries must trigger disclaimers, refusal to provide individualized advice, or redirection to clinicians.
  3. Dynamic personalization: Integration with user physiological data (e.g., CGM, wearables), comorbidity flags, and longitudinal health records to enable context-sensitive responses, bounded by explicit disclosure of personalization limits.
  4. Emotionally attuned interaction: Augmentation of affective response templates to include validation, normalization, practical coping guidance, and clear pathways to professional psychosocial support—eschewing hollow empathy in favor of pragmatic reassurance.

Theoretical and Practical Impact

Practically, the evidence substantiates the deployment of LLMs in factual and lifestyle education roles but cautions against any automated use in high-stakes symptom or medication interpretation without fine-grained boundary guards. Theoretically, the concept of the fluency illusion advances understanding of automation bias in health-AI interfaces, articulating a mechanism by which linguistic quality occludes epistemic risk.

Future directions include longitudinal assessment of LLM performance as models are iteratively tuned, evaluation of multi-turn interactions and dialogical correction protocols, and the development of adaptive interfaces capable of robustly distinguishing between educational and clinical-advisory functions in real-world settings.

Conclusion

This work demonstrates domain- and task-sensitive boundaries for the deployment of generative AI in T2DM self-management. LLMs are validated as scalable health educators but repeatedly fail in context-sensitive and safety-critical domains critical for chronic care. The empirically grounded five-dimensional evaluation rubric provides a replicable model for structured assessment, while identified analytic concepts (pre-visit primer, fluency illusion) will inform future HCI, informatics, and clinical deployment discussions. Addressing the systematically convergent personalization and affective gaps between patients and clinicians remains a central research imperative.

(2607.03720)

Paper to Video (Beta)

No one has generated a video about this paper yet.

Whiteboard

No one has generated a whiteboard explanation for this paper yet.

Open Problems

We haven't generated a list of open problems mentioned in this paper yet.

Tweets

Sign up for free to view the 1 tweet with 0 likes about this paper.