Papers
Topics
Authors
Recent
Search
2000 character limit reached

ZhiFangDanTai: TCM Formula Generation Framework

Updated 10 July 2026
  • ZhiFangDanTai is an AI framework that generates complete Traditional Chinese Medicine formulas with detailed multi-section explanations, including herbal roles and contraindications.
  • It leverages a novel GraphRAG retrieval method combined with SFT and DPO fine-tuning on an 80k-formula dataset to enhance prediction accuracy and reduce hallucination.
  • The system ensures structural completeness by decomposing outputs into clinically relevant sections such as formula recommendation, herbal composition, tongue/pulse diagnosis, and preparation methods.

Searching arXiv for the specified paper to ground the article. Searching arXiv for "(Zhang et al., 6 Sep 2025)". ZhiFangDanTai is an AI framework for Traditional Chinese Medicine (TCM) formula generation and explanation. It takes a symptom description xx as input and produces a complete TCM prescription plus detailed, structured explanations yy, including formula recommendation, herbal composition, decomposition into 君、臣、佐、使 roles, efficacy, indications, tongue and pulse diagnosis, contraindications, and preparation or processing methods. The framework combines Graph-based Retrieval-Augmented Generation (GraphRAG) with LLM fine-tuning, and is presented as a response to limitations in traditional formula mining, deep learning prescription generation, generic LLM prompting, and earlier instruction-tuned TCM systems that do not provide full, fine-grained formula sheets (Zhang et al., 6 Sep 2025).

1. Definition and task scope

ZhiFangDanTai is explicitly designed for TCM formula generation and explanation rather than only formula name prediction or herb-list completion. For a given symptom description, the system is intended to recommend one or more formulas and list their herbal composition; decompose herbs into 君 (sovereign), 臣 (minister), 佐 (assistant), 使 (courier) roles with rationale; explain efficacy and therapeutic principle; provide indications or applicable population; analyze tongue and pulse diagnosis corresponding to the pattern; state contraindications and cautions, including compatibility rules such as “18 incompatibles” and “19 antagonisms”; and describe preparation and processing methods (炮制法) (Zhang et al., 6 Sep 2025).

The paper situates this task against several prior methodological families. Traditional TCM formula mining based on association rules, clustering, complex networks, or GNNs is described as focusing on pairwise herb co-occurrence or network structure, without generating complete formulas from symptoms or providing natural-language explanation. Deep learning approaches based on CNN/GCN combined with BERT, Seq2Seq, GAN, or RL can produce prescriptions from multi-modal clinical indicators, but their explanations are absent or shallow and do not cover fine-grained decomposition into 君臣佐使, tongue and pulse, and related dimensions. Generic LLM prompting without retrieval is characterized as having limited TCM formula knowledge in training corpora and a high hallucination rate, including invented formulas, incorrect herb roles, and fabricated indications. Instruction-tuned TCM LLMs such as TCMLLM and TCM-FTP are described as improving formula recommendation and simple rationale, but still lacking depth in efficacy explanations, contraindications, tongue and pulse, and detailed preparation or processing (Zhang et al., 6 Sep 2025).

A central claim of the framework is that a clinically useful TCM output must be structurally complete. This suggests that the contribution is not merely improved prescription prediction accuracy, but a shift toward full-format explanation aligned with the conceptual schema used by TCM practitioners. The paper repeatedly frames this as the difference between generating a plausible prescription and generating a coherent, multi-section formula sheet (Zhang et al., 6 Sep 2025).

2. System architecture and end-to-end pipeline

The overall architecture combines a formula-centric TCM knowledge graph, GraphRAG retrieval, instruction dataset construction, and two-stage fine-tuning. The high-level pipeline begins with knowledge-base construction from 80k formula records collected from authoritative TCM sources. LLM-based extraction and augmentation are used to add fine-grained fields, and a knowledge graph G\mathcal{G} is constructed over formulas, herbs, diseases, symptoms, and related entities. Community detection is then performed to obtain domain-specific subgraphs covering diseases, formulas, herbs, tongue and pulse, contraindications, and other aspects (Zhang et al., 6 Sep 2025).

Given an input symptom description xx, the retrieval stage expands the query to {x,x′}\{x, x'\} and retrieves local answers AiA_i from each of seven communities. These local answers are aggregated by an LLM into a global textual summary cc, which represents the relevant subgraph Gx∗\mathcal{G}_x^*. The dataset-construction stage then forms instruction triples whose instruction is a fixed prompt such as “Recommend a TCM formula and provide detailed explanations based on the symptoms,” whose input consists of symptom text xx plus retrieved context cc, and whose output is a curated detailed explanation yy0. This process yields 60,000 instruction examples for supervised fine-tuning and additional preference pairs for DPO (Zhang et al., 6 Sep 2025).

The fine-tuning stage uses LLaMA3-Chinese-Chat / LLaMA3.2-7B as the main base model, with Qwen2.5-7B also tested, and applies LoRA for efficient parameter tuning. Stage 1 is SFT with the loss

yy1

Stage 2 is DPO using preference pairs yy2 conditioned on yy3 (Zhang et al., 6 Sep 2025).

At inference time, new symptoms yy4 undergo query expansion, GraphRAG retrieves local answers yy5 and global context

yy6

and the fine-tuned LLM yy7 generates the final formula plus full explanation yy8 on the basis of yy9 (Zhang et al., 6 Sep 2025).

3. GraphRAG and formula-centric knowledge organization

A major design distinction in ZhiFangDanTai is the replacement of standard chunk-based vector retrieval with graph-structured retrieval over a TCM ontology. Standard RAG is described as storing documents as text chunks in a vector index and retrieving top-G\mathcal{G}0 chunks via embedding similarity, without using explicit entity-relation structure. By contrast, the GraphRAG variant in ZhiFangDanTai builds a TCM knowledge graph G\mathcal{G}1 with typed entities and relations, performs community detection with the Leiden algorithm, and retrieves local answers separately from semantically coherent subgraphs corresponding to TCM dimensions (Zhang et al., 6 Sep 2025).

The knowledge graph is constructed from source documents whose fields include Disease, Recommended Formulas, Herbal Ingredients, Applicable Symptoms and Population, Pulse and Tongue, Contraindications, and Preparation Methods. Documents are split into 512-token chunks. An LLM such as DeepSeek-7B is used for zero-shot information extraction to identify entities including herbs, formulas, diseases, symptoms, body signs, preparation steps, roles (君臣佐使), and contraindication types. It also extracts relations within and across categories, such as herb–herb synergy, formula–herb inclusion, symptom–formula indicated-for, disease–formula recommended-for, syndrome–tongue/pulse pattern, and formula–contraindication. Neo4j is used to construct the graph G\mathcal{G}2 with nodes as entities and edges as typed relations (Zhang et al., 6 Sep 2025).

The paper notes two edge types, “Inter-category” and “Intra-category,” while also stating that the naming is slightly flipped in the text but conceptually corresponds to within-community versus between-community structure. Community detection is performed hierarchically using the Leiden algorithm until no further division is possible, after which seven high-level communities G\mathcal{G}3 are produced. Each community has an entity set, a community description, and a vector representation G\mathcal{G}4 via an encoder such as mxbai-embed-large (Zhang et al., 6 Sep 2025).

Retrieval is formulated at both local and global levels. After query expansion from G\mathcal{G}5 to G\mathcal{G}6, a score for candidate local answer G\mathcal{G}7 in community G\mathcal{G}8 is computed as

G\mathcal{G}9

where xx0 denotes the encoder and xx1 denotes concatenation. Top-xx2 local answers are selected per community with beam search for diversity. A MapReduce-style aggregation then produces the global answer

xx3

Top-xx4 global answers are also collected, with xx5 tuned experimentally between 1 and 3 (Zhang et al., 6 Sep 2025).

The paper argues that this community-wise retrieval preserves TCM structure by almost always covering all seven aspects: diseases, recommended formulas, herbs with roles, symptoms and population, tongue and pulse, contraindications, and preparation. A plausible implication is that the ontology is functioning not only as an indexing device but as a structural prior that constrains the generation space toward clinically organized outputs (Zhang et al., 6 Sep 2025).

4. Dataset construction and optimization objectives

The raw formula corpus contains 80,000 formulas from the Chinese Formula Database, Ancient Formula DB, and official catalogues. Augmentation is performed using DeepSeek with 50 manually annotated examples as in-context seeds to extract missing fine-grained fields, after which outputs are cleaned through redundancy processing and human validation. The instruction dataset contains 60,000 examples for SFT and an additional 3,000 preference pairs for DPO (Zhang et al., 6 Sep 2025).

Each instruction instance has a fixed instruction template, an input comprising symptom description xx6 and GraphRAG global answer xx7, and an output xx8 that is a well-structured multi-section response. The target output includes recommended formula(s) and exact herbal composition with dosages; role assignments and synergy explanations for 君、臣、佐、使; indications or applicable symptoms and population; pulse and tongue patterns; contraindications, including pregnancy, Yin deficiency fire, spleen–stomach deficiency cold, and forbidden combinations; and preparation methods such as honey-frying, charring, decoction steps, and pill formation. The appendix is described as containing detailed case studies illustrating typical outputs (Zhang et al., 6 Sep 2025).

The SFT objective is given as

xx9

In the paper’s interpretation, this trains the model to use retrieved context {x,x′}\{x, x'\}0 effectively and to produce consistent, structured explanations across all seven aspects. DPO then aligns model preferences with human or LLM-derived pairwise judgments over candidate responses generated by the SFT reference policy {x,x′}\{x, x'\}1 (Zhang et al., 6 Sep 2025).

The DPO loss is written as

{x,x′}\{x, x'\}2

with {x,x′}\{x, x'\}3 the sigmoid and {x,x′}\{x, x'\}4. The paper states that this stage focuses on coherence, TCM correctness, and reduced hallucinations, and does so without introducing an explicit reward model (Zhang et al., 6 Sep 2025).

5. Theoretical analysis of generalization and hallucination

The theoretical analysis formalizes expected negative log-likelihood as a generalization error: {x,x′}\{x, x'\}5 for SFT alone, and

{x,x′}\{x, x'\}6

for GraphRAG+SFT (Zhang et al., 6 Sep 2025).

Proposition 1 introduces conditional mutual information

{x,x′}\{x, x'\}7

Under the assumptions that {x,x′}\{x, x'\}8 and the loss is {x,x′}\{x, x'\}9-smooth, the paper gives the bound

AiA_i0

Its stated intuition is that informative GraphRAG context AiA_i1 provides additional predictive information for AiA_i2 beyond AiA_i3 alone, and thus decreases error more effectively under gradient descent (Zhang et al., 6 Sep 2025).

Proposition 2 extends the analysis to DPO. Defining preference strength as

AiA_i4

the paper states

AiA_i5

Combining this with Proposition 1 yields

AiA_i6

The interpretation provided is that expert-preferred answers provide stronger alignment signals and further reduce error (Zhang et al., 6 Sep 2025).

Hallucination is formalized through AiA_i7, the probability that generated AiA_i8 is not supported by factual TCM knowledge given context AiA_i9. Proposition 3 assumes that the relevant fact set cc0 is covered by GraphRAG context with probability at least cc1,

cc2

and lets cc3 be the probability that the model ignores context and relies on internal prior. It then gives

cc4

The paper’s intuition is that hallucination becomes unlikely when retrieval is accurate and the model learns to trust context (Zhang et al., 6 Sep 2025).

Proposition 4 further states that under DPO, the probability of generating a hallucinated answer cc5 is bounded by

cc6

so that hallucination is exponentially suppressed as preference strength cc7 grows. In the TCM setting, the relevant preferences are said to align with expert judgments on correct formula composition and herb roles, valid indications and contraindications, valid tongue and pulse patterns, and the absence of fabricated herbs or formulas (Zhang et al., 6 Sep 2025).

6. Experimental evaluation, deployment, and research context

The experimental setup uses three datasets. The collected formula dataset contains 80,000 records, from which 60,000 examples are used for SFT, 5,000 for testing, and 3,000 for DPO pairs. A clinical dataset from Haodf.com contains 22,800 training and 7,600 test samples with fields including age, disease name, title, description, hopeHelp, visit direction labels (0–19), and secondary labels (0–60); TCM experts provide ground-truth formulas for evaluation in a zero-shot setting. A conflict knowledge dataset contains synthetic instructions with conflicting medical theories, sources, or legal constraints and is used to train the model to detect conflicts and output warnings rather than choosing a side blindly (Zhang et al., 6 Sep 2025).

Evaluation covers formula prediction quality, structural explanation accuracy, and hallucination analysis. Reported text metrics are BLEU and ROUGE-1/2/L averaged as ROUGE-S. The six TCM-oriented metrics are CCR, CSCR, CCHR, FS, SCR, and LR. Their definitions are as follows (Zhang et al., 6 Sep 2025):

Metric Definition
CCR Compatibility Compliance Rate
CSCR Correct Sovereign-minister-assistant-messenger Compatibility Rate
CCHR Counter Coarse Hallucination Rate
FS FactScore
SCR Structural/Content Clarity Rate
LR Logical Rate

CCR is defined by

cc8

CSCR is

cc9

with weights Gx∗\mathcal{G}_x^*0. CCHR is

Gx∗\mathcal{G}_x^*1

FS is

Gx∗\mathcal{G}_x^*2

SCR is

Gx∗\mathcal{G}_x^*3

with

Gx∗\mathcal{G}_x^*4

where Gx∗\mathcal{G}_x^*5 is the number of six key components present: formula, ingredients, indications, tongue/pulse, contraindications, and preparation. LR is the proportion of logically coherent contextual sentences among all contextual sentences (Zhang et al., 6 Sep 2025).

The baselines include plug-and-play LLMs such as LLaMA3.2-7B (Chinese Chat), Kimi-7B, DeepSeek-7B, and Qwen2.5-7B; the FT-only TCM LLM TCMLLM; standard RAG using FAISS-based LlamaIndex retrieval with LLaMA3.2-7B as generator; GraphRAG with the same base LLaMA; RAG+SFT and RAG+SFT+DPO; GraphRAG+SFT; and the full model ZhiFangDanTai, defined as GraphRAG + SFT + DPO. Top-Gx∗\mathcal{G}_x^*6 retrieval is typically 3 in the main experiments, with 1 and 2 also tested (Zhang et al., 6 Sep 2025).

On the 5,000-sample test set with top-Gx∗\mathcal{G}_x^*7, ZhiFangDanTai is reported to achieve the best scores across BLEU, ROUGE-S, CCR, CSCR, CCHR, FS, SCR, and LR, while GraphRAG+SFT is second best. The paper reports that GraphRAG consistently outperforms plain RAG, that fine-tuning improves over retrieval-only variants, that DPO further improves coarse and fine-grained hallucination measures, and that vanilla LLMs substantially underperform on all TCM-specific metrics. TCMLLM is reported to improve over plain LLaMA but to remain limited by its instruction data, especially on tongue and pulse, contraindications, and preparation, which lowers SCR and CSCR relative to ZhiFangDanTai. Training FLOPs are higher for the full system because of SFT+DPO, whereas inference FLOPs are only slightly higher than pure LLM calling (Zhang et al., 6 Sep 2025).

The top-Gx∗\mathcal{G}_x^*8 ablation indicates that ZhiFangDanTai is already strong at Gx∗\mathcal{G}_x^*9, that BLEU and ROUGE improve as xx0 increases, and that TCM metrics improve slightly from xx1 to xx2 or xx3. The authors describe xx4–3 as a good trade-off and use xx5 in some later experiments for efficiency. A separate ablation on an added GPT pre-training stage reports that GraphRAG+GPT, GraphRAG+GPT+SFT, and GraphRAG+GPT+SFT+DPO perform worse than ZhiFangDanTai, and the paper attributes this to mismatch between pre-training data distribution and task distribution, as well as high Rademacher complexity xx6 for a 7B model under relatively small domain data. The conclusion given is that GPT is unnecessary and even harmful in this setting (Zhang et al., 6 Sep 2025).

Backbone ablation with Qwen2.5-7B shows that the method still outperforms all baselines under Qwen, while LLaMA-based variants perform slightly better overall except for minor BLEU differences. On 7,600 real clinical cases with top-xx7, and without retraining on that data, ZhiFangDanTai attains the highest TCM metrics and the lowest hallucination rates, which the paper interprets as strong zero-shot generalization to real-world clinical records (Zhang et al., 6 Sep 2025).

The model is open-sourced at the Hugging Face repository tczzx6/ZhiFangDanTai1.0, with a corresponding dataset release. The implementation uses LlamaIndex for GraphRAG, retrieval, and MapReduce; LLaMA-Factory for SFT, DPO, and LoRA; and vLLM for efficient inference. A WebUI accepts symptoms and returns formulas plus explanations, and includes safety disclaimers that present the system as a decision-support tool rather than a substitute for a practitioner. The paper states that it is best suited as an educational tool for teaching formula design and rationale, and as clinical decision support for practitioner review rather than direct patient self-medication. It also notes safety mechanisms for contradictory knowledge and warnings involving conflicting sources or illegal substances such as rhinoceros horn (Zhang et al., 6 Sep 2025).

The stated limitations include incomplete knowledge graph coverage for rare syndromes, regional practices, or modern clinical innovations; data bias toward classical and official Chinese sources; possible weakness on rare or mixed syndromes; ongoing clinical risks such as dose errors and overlooked comorbidities or drug interactions; optimization primarily for Chinese and English graph translations; and the absence of multimodal integration for tongue images or pulse waveforms. Future directions mentioned include knowledge graph expansion, multilingual support, multimodal integration, stronger theoretical frameworks for generalization and hallucination, and safer deployment through dynamic guideline updates, EMR integration, and stricter safety constraints (Zhang et al., 6 Sep 2025).

Within TCM informatics and TCM NLP, ZhiFangDanTai is positioned as building on traditional data mining, neural prescription generation from EMR or tongue–pulse images, and recent LLM-based TCM systems such as TCMLLM and TCM-FTP. Within graph-based RAG and hallucination reduction, it is described as building on GraphRAG by Edge et al. and connecting to broader RAG literature including Lewis et al., Xiong et al., and corrective RAG. Its stated uniqueness lies in integrating GraphRAG with TCM formulas, providing systematic theoretical analysis of generalization and hallucination in this domain, and targeting full-format formula explanation including composition, roles, indications, tongue and pulse, contraindications, and processing (Zhang et al., 6 Sep 2025).

Definition Search Book Streamline Icon: https://streamlinehq.com
References (1)

Topic to Video (Beta)

No one has generated a video about this topic yet.

Whiteboard

No one has generated a whiteboard explanation for this topic yet.

Follow Topic

Get notified by email when new papers are published related to ZhiFangDanTai.