- The paper demonstrates that patient communication styles critically alter triage outcomes in LLM-based systems, with over-triage differences up to 13.5 percentage points.
- Methodologically, a modular patient simulator with 20 behavioral parameters was developed, achieving a macro concordance of 0.89 and robust clinical fidelity.
- Implications highlight the need for realistic simulation benchmarks to address fairness and safety in consumer-facing health chatbots.
Authoritative Summary of "The complexities of patient-centred conversational artificial intelligence" (2607.08625)
Introduction
The proliferation of consumer-facing health chatbots powered by LLMs introduces a new interface between patients and health care delivery. Such systems are increasingly adopted for symptom assessment, triage, and guidance. However, their evaluation and development have systematically relied on idealized, cooperative, and articulate simulated patients, neglecting real-world heterogeneity and communication diversity inherent in actual user interactions. This paper investigates the impact of patient communication styles on LLM-based clinical triage, introduces a modular LLM-driven patient simulator for benchmark creation, and demonstrates the substantial influence of patient communication variation on AI performance and fairness metrics.
Characterization of Real Patient-Chatbot Interactions
Analysis of 2,053 patient-chatbot conversations within the Verily Me mobile app reveals extensive heterogeneity:
- Session attrition is high: only 14% reached triage completion; 77% abandoned after one turn.
- Message brevity prevails: 67% of messages contain fewer than six words.
- Communication idiosyncrasies are prevalent: non-standard capitalization (23%), missing punctuation (52%), grammar errors (64%), and typos (43%) appeared in the majority of sessions.
- Emotional signaling was frequent; anxiousness (21%) and frustration (19%) dominated.
- Health literacy was mostly basic or medium; critical literacy was rare (<1%).
These findings contradict the assumption that patients engage chatbots with the same communication competence and effort as simulated users.
Modular Patient Simulation Framework
To systematically study the effect of patient behavior on conversational AI, the authors designed a multi-channel patient simulator. This architecture decomposes patient responses into four distinct aspects: clinical content, emotional state, conversational strategy, and communication style, each independently parameterizable.
- The simulator is configurable over 20 behavioral parameters—including health literacy, recall, goal orientation, emotional tone, verbosity, English proficiency, and communication irregularities—allowing the generation of diverse patient personae.
- Evaluation of parameter adherence using Claude Opus 4.6 yielded a macro concordance statistic of 0.89, with several parameters (anxiousness, sensitive topics, communication style) reaching c-statistic = 1.0.
- Clinical fidelity was robust: 94.3% of supported claims were rated fully accurate (mean accuracy 3.94/4), with no contradictions.
- Realism in simulation approached parity: human graders classified simulated vs. real conversations at only 55% accuracy—nearly chance.
Impact of Communication Style on LLM Urgency Assessment
Using the patient simulator, the authors generated 1,164 clinical cases for urgency assessment, each presented under five distinct patient personae: Anxious Patient, Dismissive Patient, Informed Advocate, Limited Communicator, and Default. Four LLM-based clinician models were evaluated.
Key findings:
- Communication style significantly altered triage outcomes. For Gemini 3.5 Flash, over-triage for Anxious Patient was 36.8% vs. 23.3% for Dismissive Patient—a 13.5 percentage point difference, holding clinical facts constant.
- Under-triage rates exhibited the opposite trend: Dismissive patients experienced higher under-triage (6.6%) than Anxious patients (2.5%).
- This pattern was consistent across all evaluated LLMs, though effect size varied.
A detailed calibration analysis revealed that discrimination (as measured by c-statistic) was stable across personae, but calibration-in-the-large was heavily influenced by communication style. Models systematically overestimated urgent triage probability for anxious presentations and underestimated self-care recommendations for dismissive ones.
Demographic stratification (sex, age) did not materially alter persona-driven triage error patterns; confidence interval overlap indicated no subgroup interaction.
Implications for AI Evaluation and Healthcare Equity
This study explicitly demonstrates that conversational AI evaluation based on cooperative, literate, and emotionally neutral simulated patients produces inflated performance metrics and fails to reflect real-world system effectiveness. The magnitude of triage disparity attributable to communication style—up to 13.5 percentage points—far exceeds the effect sizes documented in prior studies examining sociodemographic bias in LLMs.
The results have substantial implications:
- System benchmarks should systematically vary patient communication parameters, not solely clinical facts, to evaluate fairness and safety.
- Calibration disparities driven by patient behavior necessitate potential recalibration of LLM decision thresholds based on communication features. However, reliable and ethical measurement of such features at inference remains a challenge.
- Design of consumer-facing conversational AI should focus on minimizing dependence on articulate, sustained, and motivated engagement, as these are not universally achievable and may compound health inequities.
- More realistic and robust patient simulation is essential, but remains limited by the correspondence between simulated and true patient behavior; continuous empirical and interdisciplinary refinement is required.
Limitations and Future Directions
The study is limited to urgency assessment as a clinical task and utilizes LLMs accessed through commercial APIs. The scope of personae sampled, and the use of synthetic and de-identified data, restricts the generalizability to other settings and tasks. Prospective evaluation with real users is optimal but economically and ethically challenging. The authors advocate for ongoing, systematic attention to patient-centered and communication-diverse evaluation in all consumer-facing conversational medical AI deployments.
Future research ought to extend simulation methods to incorporate multimodal inputs, explore open-ended outcome spaces, measure system adaptation to dynamically shifting communication behavior, and continuously audit AI calibration relative to marginalized and digitally underserved groups.
Conclusion
Patient communication style exerts a strong and measurable influence on the triage decisions of LLM-based conversational medical AI, even when clinical content is held fixed. Conventional benchmarks relying on cooperative simulated patients fail to capture this axis of variance and risk amplifying health disparities. Patient-centered simulation frameworks that parameterize communication diversity are necessary for credible evaluation, equitable deployment, and safety of consumer-facing health chatbots. The findings call for recalibrated system design and evaluation paradigms that account for the real-world complexity of patient-AI interaction.