KnowGuard: Evidence-Driven Clinical Reasoning
- KnowGuard is a knowledge-driven abstention framework that enables safe, multi-round clinical reasoning by grounding decisions in external medical evidence.
- It systematically combines graph expansion and direct retrieval from a WHO-based medical knowledge graph to evaluate evidence sufficiency in incremental consultations.
- Evaluations show KnowGuard improves diagnostic accuracy by up to 3.93% while reducing unnecessary interaction rounds by approximately 7 turns.
KnowGuard is a knowledge-driven abstention framework for multi-round clinical reasoning that treats abstention as a safety mechanism rather than as a failure mode. It is introduced in "KnowGuard: Knowledge-Driven Abstention for Multi-Round Clinical Reasoning" (Dang et al., 29 Sep 2025) and is designed for interactive diagnostic settings in which patient information is disclosed incrementally across rounds. Its central proposal is an investigate-before-abstain paradigm: instead of relying only on model self-assessment or confidence scores, the system grounds the stop-or-continue decision in systematic exploration of external medical knowledge, using a medical knowledge graph, a shared contextualized evidence pool, and multi-factor evidence ranking.
1. Problem setting and abstention formulation
KnowGuard studies a clinical consultation process in which a Patient Agent holds the full patient information
but only the initial presentation is visible at the beginning. A Doctor Agent receives additional information round by round. After the patient answers a targeted question with response , the accumulated information is updated as
At each round, the doctor makes a binary abstention decision
where means abstain and continue gathering information, and means enough evidence exists to diagnose (Dang et al., 29 Sep 2025).
This formulation is motivated by the observation that recent investigations have reported the application of LLMs in medical scenarios, but existing LLMs struggle with abstentions and frequently provide overconfident responses despite incomplete information. The paper attributes this to conventional abstention methods that rely only on model self-assessments and therefore lack systematic strategies to identify knowledge boundaries with external medical evidences. KnowGuard accordingly reframes abstention as a problem of evidence sufficiency under interaction, rather than as confidence calibration alone.
The benchmark setting is explicitly open-ended and multi-round. Original records are parsed into structured atomic facts such as age, gender, chief complaint, and additional evidence. Initially, only age, gender, and chief complaint are visible to the Doctor Agent; the remaining facts are hidden and can be revealed only through targeted questioning. This makes the main optimization target an accuracy-efficiency trade-off: answering too early risks unsafe misdiagnosis, whereas asking too many unnecessary questions increases interaction burden and delays care (Dang et al., 29 Sep 2025).
2. Core architecture and knowledge substrate
KnowGuard operates through two stages that share a persistent evidence memory: an evidence discovery stage and an evidence evaluation stage. Both stages are built around a shared contextualized evidence pool
a bounded top- priority queue of candidate knowledge triplets, where 0 is a medical triplet and 1 is its priority score (Dang et al., 29 Sep 2025).
The external knowledge source is a medical knowledge graph
2
constructed from over 300 WHO guidelines. The graph contains about 22k medical entities and over 100k clinical relationships. Each triplet is augmented with source text descriptions and document page images, so the evidence source is explicitly multi-modal. The graph is further organized with demographic and disease-specific features derived from guideline titles and abstracts, which later support patient population reasoning (Dang et al., 29 Sep 2025).
At round 3, the doctor’s final decision is conditioned jointly on accumulated patient information, the evidence pool, and the retrieved multi-modal evidence:
4
The paper does not define a fixed candidate diagnosis set or an explicit closed-form stopping threshold. Instead, abstention emerges from the model’s judgment over the evolving evidence state. This suggests that KnowGuard is closer to evidence-grounded decision support than to a thresholded confidence model (Dang et al., 29 Sep 2025).
3. Evidence discovery stage
The evidence discovery stage is designed to explore the medical knowledge space systematically rather than to perform one-shot retrieval. It uses two retrieval mechanisms.
The first mechanism is graph expansion-based retrieval, which expands from entities already present in the current evidence pool:
5
where 6 denotes the entities appearing in the current evidence triplets. This supports path tracing, because once a symptom, condition, medication, or demographic factor becomes salient, the system can continue expanding through nearby medical relations.
The second mechanism is direct retrieval, which uses the current patient response 7 to generate a new query and search the graph:
8
This allows the system to jump to new regions of the graph when the latest patient disclosure changes the diagnostic direction.
Candidate evidence is the union of the two retrieval sets:
9
Because the evidence pool is shared across rounds, discovery is history-sensitive rather than reset-based: prior evidence remains available, graph expansion preserves continuity, and direct retrieval reacts to newly disclosed facts (Dang et al., 29 Sep 2025).
The paper’s canonical case study illustrates this process in a patient presenting with abdominal pain. KnowGuard investigates contextual evidence, explores the graph structure, discovers a connection between NRTI-class drugs used for HIV and acute pancreatitis, and then reaches the correct diagnosis with treatment recommendations. This example is important because it shows that the system is intended not merely to retrieve supporting facts, but to expose hidden causal structure that defines whether current evidence is sufficient (Dang et al., 29 Sep 2025).
4. Evidence evaluation and abstention decision
After retrieval, KnowGuard ranks each candidate triplet using five factors: embedding similarity, LLM relevance, graph coherence, round decay, and patient population reasoning. The first two form a dual validation of relevance.
The hard relevance score is embedding similarity:
0
The soft relevance score is an LLM-based clinical relevance estimate:
1
To preserve consistent reasoning paths across rounds, KnowGuard also uses graph coherence:
2
where the count term is the cumulative frequency of the entity in the evidence pools accumulated over the conversation.
A separate patient population reasoning module infers the likely demographic or clinical population category from the conversation:
3
and then upweights triplets that belong to the corresponding population-specific subgraph:
4
Temporal adaptation across rounds is handled by a decayed carryover rule:
5
The final priority score is
6
and the next evidence pool is updated by
7
In the sensitivity study, the best-performing values are reported as 8, 9, 0, 1, and a patient population reasoning weight of 2 (Dang et al., 29 Sep 2025).
A notable design feature is that the paper does not define explicit standalone formulas for reliability, sufficiency, or novelty. Sufficiency is instead treated operationally through the interaction of retrieval, reranking, and the doctor model’s final abstention decision. This suggests that KnowGuard’s notion of a knowledge boundary is procedural rather than threshold-based (Dang et al., 29 Sep 2025).
5. Benchmarks, metrics, and empirical results
KnowGuard is evaluated on three medical datasets converted into open-ended, interactive, multi-round reasoning benchmarks: MEDQA → ioMEDQA, CRAFT-MD → ioCRAFT-MD, and AFRIMEDQA → ioAFRIMEDQA. The overall benchmark contains 3,061 cases across the three datasets. Because the task is open-ended, free-text predictions are evaluated by an LLM judge. For originally multiple-choice items, the judge matches the prediction against the candidate options; for originally open-ended items, it decides whether the predicted answer semantically matches the ground-truth answer (Dang et al., 29 Sep 2025).
The primary reported metrics are Accuracy (ACC) and average conversation rounds (avg. Turn). These jointly quantify diagnostic correctness and interaction efficiency. The baseline methods are Basic, Binary Decision, Numerical Score, Scale Rating, and Long Context, with enhanced variants using rationale generation and self-consistency.
The main results show that KnowGuard attains the highest accuracy across all benchmarks in both the basic and enhanced settings. The paper’s overall summary states that it improves diagnostic accuracy by 3.93% while reducing unnecessary interaction by 7.27 turns on average (Dang et al., 29 Sep 2025).
| Benchmark | Basic setting: ACC / avg. Turn | Enhanced setting: ACC / avg. Turn |
|---|---|---|
| ioAFRIMEDQA | 68.70 / 5.26 | 73.20 / 5.30 |
| ioMEDQA | 70.98 / 5.41 | 74.12 / 5.40 |
| ioCRAFT-MD | 66.47 / 4.89 | 71.96 / 6.51 |
In the basic setting, KnowGuard achieves 68.70 on ioAFRIMEDQA, 70.98 on ioMEDQA, and 66.47 on ioCRAFT-MD, all with substantially fewer turns than the abstention-heavy Binary Decision baseline and with better accuracy than Long Context. In the enhanced setting, the corresponding values rise to 73.20, 74.12, and 71.96 (Dang et al., 29 Sep 2025).
The ablation study is equally informative. Removing patient population reasoning mainly hurts efficiency, while removing evidence evaluation causes a marked accuracy drop and fewer turns, indicating more premature answering. On ioAFRIMEDQA, for example, the full model yields 73.20 accuracy and 5.30 turns; removing evidence evaluation and patient population reasoning reduces performance to 66.22 accuracy and 2.69 turns. This supports the paper’s claim that the evidence evaluation stage is central to the accuracy-efficiency trade-off (Dang et al., 29 Sep 2025).
The qualitative analysis shows three additional patterns. First, KnowGuard’s accuracy continues to improve as conversations get longer, unlike confidence-only baselines. Second, its confidence rises faster than Scale Rating over normalized conversation progress. Third, the gains are stronger on rare disease diagnosis, where external medical knowledge is especially valuable (Dang et al., 29 Sep 2025).
6. Significance, limitations, and relation to broader guardrail research
KnowGuard’s main conceptual contribution is to shift medical abstention from self-confidence estimation to evidence-grounded knowledge boundary detection. In that respect, it belongs to a broader family of guard and guardrail systems that externalize structure or reasoning instead of relying only on end-to-end prediction. Examples include "3-Guard: Robust Reasoning Enabled LLM Guardrail via Knowledge-Enhanced Logical Reasoning" (Kang et al., 2024), which combines category-specific detectors with probabilistic reasoning over logical rules; "Qwen3Guard Technical Report" (Zhao et al., 16 Oct 2025), which separates full-context moderation from token-level streaming moderation; and "GuardReasoner-Omni: A Reasoning-based Multi-modal Guardrail for Text, Image, and Video" (Zhu et al., 3 Feb 2026), which uses explicit reasoning traces for multi-modal moderation. KnowGuard differs from those systems in that it targets abstention in multi-round clinical reasoning and grounds the stop-or-continue decision in medical knowledge graph exploration (Dang et al., 29 Sep 2025).
Its limitations are also explicit. The paper states that KnowGuard is a research prototype only and is not intended for clinical deployment. The system may inherit biases from GPT-4, benchmark datasets, and source guidelines, and the authors identify comprehensive fairness evaluation as future work. The method also depends on the coverage and quality of the WHO-derived knowledge graph and on the quality of graph retrieval and evidence reranking. Moreover, the paper does not specify an explicit abstention threshold, a formal stopping utility, a maximum-round stopping rule, or detailed retrieval-model and embedding-model implementations (Dang et al., 29 Sep 2025).
A broader implication is that KnowGuard is best understood as a blueprint for safe clinical reasoning under incomplete evidence rather than as a confidence-calibration wrapper. Its persistent evidence pool, graph expansion plus direct retrieval, multi-factor evidence reranking, and demographic guidance together define a procedural account of when a model should continue investigating and when it should answer. This suggests a general design principle for safety-critical LLM systems: abstention quality can improve when uncertainty is evaluated against structured external evidence instead of being inferred solely from the model’s internal belief state (Dang et al., 29 Sep 2025).