- The paper analyzes over 500,000 de-identified Copilot health conversations using a clinician-designed 12-category taxonomy and privacy-preserving, LLM-based classification.
- Personal symptom and condition queries account for nearly one in five conversations, while one in seven such queries concern a child, parent, partner, or another person.
- Health questions increase during evenings and nights, mobile users focus more on personal concerns, and healthcare navigation requests expose persistent access and administrative friction.
Overview and motivation
This paper presents a large-scale observational analysis of health-related usage of Microsoft Copilot, a consumer conversational AI assistant. Drawing on more than 500,000 de-identified health conversations sampled from January 2026, the authors construct a hierarchical intent taxonomy of 12 primary categories with fine-grained topic clusters within each, and use it to characterize what people ask about health, on whose behalf they ask, and how usage varies by device and time of day. The work extends the methodology of the Copilot Usage Report 2025 [costagomes2025] and addresses a gap the authors identify explicitly: while prior studies have established that LLMs perform well on medical licensing exams and clinical reasoning benchmarks [nori2023capabilitiesgpt4medicalchallenge, nori2025sequentialdiagnosislanguagemodels], and that users sometimes prefer AI-generated health responses to physician responses [lizee2024conversational, ruben2025artificial], a systematic characterization of what people actually ask at scale has been missing.
The paper's framing is careful. The authors note that strong benchmark performance does not guarantee real-world reliability — chatbots have been shown to fail at triage severity assessment [ramaswamy2026chatgpt], and LLM-assisted users do not always outperform controls in condition identification [bean2026reliability]. This motivates their central argument: understanding the distribution of user intents is a prerequisite for knowing when an AI assistant should provide information versus direct users to professional care.
Data pipeline and privacy model
The dataset consists of a random sample of Copilot conversations from January 2026 classified as "Health and Fitness," excluding enterprise, educational, and commercial accounts. Approximately 22% of conversations originate in the United States and roughly 45% are in English; the remainder are global and multilingual.
The privacy architecture is notable and worth describing precisely. Raw transcripts pass through automated PII scrubbing (names, contact details, government identifiers, financial data), after which an LLM generates a short English-language summary capturing topic and intent without reproducing the user's original words. All downstream analysis — including clustering — operates exclusively on these summaries. No human researcher accessed raw conversation content; the pipeline follows an "eyes-off" model in which only machine classifiers interact with scrubbed data. The authors state that no re-identification or individual-level inference was attempted, all processing occurred within Microsoft-controlled systems under internal privacy review, and results are reported only at aggregate level. They also note that these evaluation activities were not classified as human subjects research because they are not designed to generate generalizable knowledge or evaluate clinical hypotheses — a classification choice some readers may wish to scrutinize, given that the paper does generate broadly applicable empirical findings.
Taxonomy construction and validation
The 12-category intent taxonomy was developed prior to the study by in-house clinician scientists, informed by earlier observational analyses. It spans general information and education; symptom questions and health concerns; condition information and care questions; fitness, lifestyle, and coaching; emotional wellbeing; healthcare navigation and access to care; coverage and benefits; research and academic support; medical paperwork; digital tools and fitness apps; plus two telemetry categories ("Other Health/Fitness" and "Not Health") excluded from substantive analysis.
An LLM-based classifier assigns each conversation a health-specific intent and extracts structured attributes, including who the query is about and any explicitly mentioned symptoms or conditions. Validation consisted of independent labeling of a human-reviewable sample by clinical scientists, with reported agreement between classifier and annotators — though the paper does not report quantitative agreement statistics such as Cohen's kappa, which is a gap in methodological transparency. Thematic structure within each intent was derived via LLM-driven clustering following TnT-LLM [wan2024], producing named clusters ranked by prevalence.
Five principal findings
Personal health intents are prevalent, and likely underestimated. Nearly one in five conversations involve personal symptom assessment or condition discussion. The largest category, "Health Information & Education," accounts for over 40% of conversations, but its top topic clusters are dominated by queries about specific treatments, medications, and conditions rather than abstract knowledge — the leading cluster ("how specific treatments, medications, or medical procedures work") alone comprises 18.8% of that category. The authors argue this makes the 40% figure a lower bound on personal health intent, since generally framed queries (e.g., "what are the side effects of metformin") may reflect the user's own medication concerns but default to the educational label when context is insufficient. This is a defensible claim, but it rests on an assumption the authors acknowledge: the boundary between general curiosity and personal concern is often unobservable from conversation content alone.
Conversational AI functions as a caregiving tool. Across symptom questions and condition information intents, one in seven conversations concern someone other than the user — a child, an aging parent, or a partner. This reframes the user population: the person typing is not always the person the query is about. The authors draw out both design implications (caregivers may need different contextual cues and follow-up recommendations) and safety implications (third-party queries may affect the accuracy and completeness of information needed).
Personal queries peak when healthcare access is lowest. Symptom and emotional wellbeing queries rise markedly in evening and nighttime hours. The authors connect the emotional wellbeing pattern to the independently documented diurnal rhythm of negative affect, which peaks in the evening [golder2011], and treat this convergence as suggestive evidence for the construct validity of their classification approach. They are appropriately cautious about causal interpretation, noting the pattern is likely multiply determined — evening reflection time, reduced professional availability, and affective rhythms may all contribute.
Device is a strong signal for intent. Mobile concentrates on personal health concerns (symptom questions are the second most common mobile intent), while desktop is dominated by research, academic support, and medical paperwork, which peak during working hours. The authors argue device choice reflects fundamentally different modes of engagement rather than mere convenience, with practical implications for platform-specific response design.
Healthcare navigation reveals system friction. A substantial share of queries address navigating healthcare systems rather than health itself: finding nearby providers, clinics, or specialists is the single largest cluster (43.1%) within the Healthcare Navigation intent, followed by paperwork assistance (11.5%), provider comparison (7.7%), insurance comprehension (6.7%), and appointment booking (6.4%). The authors interpret this volume as a signal of friction in existing care delivery — users asking AI for help with tasks that should be straightforward.
Topical concentrations
The topic-cluster analysis shows clear concentration around core needs within narrower intents. Within "Symptom Questions & Health Concerns," understanding new or unexpected symptoms dominates at 32.7%, with recurring symptoms (14.6%) and plain-language explanations of lab or imaging results (12.8%) following. Within "Emotional Wellbeing," understanding personal emotional or behavioral health challenges leads at 30.5%. Within "Fitness, Lifestyle & Coaching," tailored meal plans and calorie targets lead at 26.6% — a figure plausibly inflated by the January sampling window, given New Year's resolution effects, which the authors acknowledge as a limitation.
Limitations
The paper is explicit about its constraints, and several bear directly on the strength of its claims. First, the analysis covers a single platform (consumer Microsoft Copilot) and a single month, so seasonal effects — particularly January fitness resolutions — may distort intent distributions, and generalization to other platforms, clinical settings, or populations is uncertain. Second, the study observes queries but not outcomes: there is no evidence on whether users sought clinical care afterward, how they interpreted responses, or whether the information improved their decisions. Third, and most consequentially for interpretation, the taxonomy captures expressed intent rather than underlying clinical need; the dominance of the education category partly reflects the inherent difficulty of separating general from personal information seeking, meaning the reported one-in-five personal health share should be read as a floor rather than a point estimate. Fourth, classifier validation is described qualitatively without reported agreement metrics. Open questions the paper leaves include whether the education category can be further subdivided, how intent distributions shift geographically across healthcare systems with differing primary care access, and whether intent patterns can be linked to response quality and downstream outcomes.
Conclusion
This paper provides a systematic baseline characterization of health-related conversational AI usage at scale, built on a privacy-preserving pipeline and a clinician-informed taxonomy. Its most consequential findings — the prevalence of personal symptom and condition queries, the one-in-seven caregiving share, the after-hours concentration of personal and emotional health intents, the sharp device-based divergence, and the volume of healthcare navigation requests — collectively indicate that consumer AI assistants already function as personal health tools, caregiver aids, and healthcare system guides, often outside the hours and contexts in which professional care is available. The authors position these categories as defining where the consequences of conversational AI responses are highest, and where investment in response quality and safety measures should accordingly concentrate.