---
title: 'ZhiFangDanTai: TCM Formula Generation Framework'
url: https://www.emergentmind.com/topics/zhifangdantai
type: topic
---

# ZhiFangDanTai: TCM Formula Generation Framework

Searching arXiv for the specified paper to ground the article.
Searching arXiv for "2509.05867".
ZhiFangDanTai is an AI framework for Traditional Chinese Medicine (TCM) formula generation and explanation. It takes a symptom description $x$ as input and produces a complete TCM prescription plus detailed, structured explanations $y$, including formula recommendation, herbal composition, decomposition into 君、臣、佐、使 roles, efficacy, indications, tongue and pulse diagnosis, contraindications, and preparation or processing methods. The framework combines Graph-based Retrieval-Augmented Generation (GraphRAG) with LLM fine-tuning, and is presented as a response to limitations in traditional formula mining, deep learning prescription generation, generic LLM prompting, and earlier instruction-tuned TCM systems that do not provide full, fine-grained formula sheets [2509.05867].

## 1. Definition and task scope

ZhiFangDanTai is explicitly designed for TCM formula generation and explanation rather than only formula name prediction or herb-list completion. For a given symptom description, the system is intended to recommend one or more formulas and list their herbal composition; decompose herbs into 君 (sovereign), 臣 (minister), 佐 (assistant), 使 (courier) roles with rationale; explain efficacy and therapeutic principle; provide indications or applicable population; analyze tongue and pulse diagnosis corresponding to the pattern; state contraindications and cautions, including compatibility rules such as “18 incompatibles” and “19 antagonisms”; and describe preparation and processing methods (炮制法) [2509.05867].

The paper situates this task against several prior methodological families. Traditional TCM formula mining based on association rules, clustering, complex networks, or GNNs is described as focusing on pairwise herb co-occurrence or network structure, without generating complete formulas from symptoms or providing natural-language explanation. Deep learning approaches based on CNN/GCN combined with BERT, Seq2Seq, GAN, or RL can produce prescriptions from multi-modal clinical indicators, but their explanations are absent or shallow and do not cover fine-grained decomposition into 君臣佐使, tongue and pulse, and related dimensions. Generic LLM prompting without retrieval is characterized as having limited TCM formula knowledge in training corpora and a high hallucination rate, including invented formulas, incorrect herb roles, and fabricated indications. Instruction-tuned TCM LLMs such as TCMLLM and TCM-FTP are described as improving formula recommendation and simple rationale, but still lacking depth in efficacy explanations, contraindications, tongue and pulse, and detailed preparation or processing [2509.05867].

A central claim of the framework is that a clinically useful TCM output must be structurally complete. This suggests that the contribution is not merely improved prescription prediction accuracy, but a shift toward full-format explanation aligned with the conceptual schema used by TCM practitioners. The paper repeatedly frames this as the difference between generating a plausible prescription and generating a coherent, multi-section formula sheet [2509.05867].

## 2. System architecture and end-to-end pipeline

The overall architecture combines a formula-centric TCM knowledge graph, GraphRAG retrieval, instruction dataset construction, and two-stage fine-tuning. The high-level pipeline begins with knowledge-base construction from 80k formula records collected from authoritative TCM sources. LLM-based extraction and augmentation are used to add fine-grained fields, and a knowledge graph $\mathcal{G}$ is constructed over formulas, herbs, diseases, symptoms, and related entities. Community detection is then performed to obtain domain-specific subgraphs covering diseases, formulas, herbs, tongue and pulse, contraindications, and other aspects [2509.05867].

Given an input symptom description $x$, the retrieval stage expands the query to $\{x, x'\}$ and retrieves local answers $A_i$ from each of seven communities. These local answers are aggregated by an LLM into a global textual summary $c$, which represents the relevant subgraph $\mathcal{G}_x^*$. The dataset-construction stage then forms instruction triples whose instruction is a fixed prompt such as “Recommend a TCM formula and provide detailed explanations based on the symptoms,” whose input consists of symptom text $x$ plus retrieved context $c$, and whose output is a curated detailed explanation $y$. This process yields 60,000 instruction examples for supervised fine-tuning and additional preference pairs for DPO [2509.05867].

The fine-tuning stage uses LLaMA3-Chinese-Chat / LLaMA3.2-7B as the main base model, with Qwen2.5-7B also tested, and applies LoRA for efficient parameter tuning. Stage 1 is SFT with the loss
\[
\mathcal{L}_{\text{GraphRAG+SFT}
= -\mathbb{E}_{(x, c, y) \sim D} \log P_\theta(y \mid x, c).
\]
Stage 2 is DPO using preference pairs $(y_w, y_l)$ conditioned on $x, c$ [2509.05867].

At inference time, new symptoms $x$ undergo query expansion, GraphRAG retrieves local answers $\{A_1,\dots,A_7\}$ and global context
\[
c = \pi(x, x', A_1,\dots,A_7),
\]
and the fine-tuned LLM $\pi_\theta$ generates the final formula plus full explanation $y$ on the basis of $(x, x', c)$ [2509.05867].

## 3. GraphRAG and formula-centric knowledge organization

A major design distinction in ZhiFangDanTai is the replacement of standard chunk-based vector retrieval with graph-structured retrieval over a TCM ontology. Standard RAG is described as storing documents as text chunks in a vector index and retrieving top-$k$ chunks via embedding similarity, without using explicit entity-relation structure. By contrast, the GraphRAG variant in ZhiFangDanTai builds a TCM knowledge graph $\mathcal{G}$ with typed entities and relations, performs community detection with the Leiden algorithm, and retrieves local answers separately from semantically coherent subgraphs corresponding to TCM dimensions [2509.05867].

The knowledge graph is constructed from source documents whose fields include Disease, Recommended Formulas, Herbal Ingredients, Applicable Symptoms and Population, Pulse and Tongue, Contraindications, and Preparation Methods. Documents are split into 512-token chunks. An LLM such as DeepSeek-7B is used for zero-shot information extraction to identify entities including herbs, formulas, diseases, symptoms, body signs, preparation steps, roles (君臣佐使), and contraindication types. It also extracts relations within and across categories, such as herb–herb synergy, formula–herb inclusion, symptom–formula indicated-for, disease–formula recommended-for, syndrome–tongue/pulse pattern, and formula–contraindication. Neo4j is used to construct the graph $\mathcal{G}$ with nodes as entities and edges as typed relations [2509.05867].

The paper notes two edge types, “Inter-category” and “Intra-category,” while also stating that the naming is slightly flipped in the text but conceptually corresponds to within-community versus between-community structure. Community detection is performed hierarchically using the Leiden algorithm until no further division is possible, after which seven high-level communities $C=\{C_1,\dots,C_7\}$ are produced. Each community has an entity set, a community description, and a vector representation $\vec{C}_i$ via an encoder such as mxbai-embed-large [2509.05867].

Retrieval is formulated at both local and global levels. After query expansion from $x$ to $\{x,x'\}$, a score for candidate local answer $A_i$ in community $C_i$ is computed as
\[
p(A_i \mid \{x, x'\}) \propto \exp\big(E(A_i) \cdot E(x \| x')\big),
\]
where $E(\cdot)$ denotes the encoder and $\|$ denotes concatenation. Top-$k$ local answers are selected per community with beam search for diversity. A MapReduce-style aggregation then produces the global answer
\[
c = \pi(x, x', A_1,\dots,A_7) = \pi(x, x', \mathcal{G}_x^*).
\]
Top-$k$ global answers are also collected, with $k$ tuned experimentally between 1 and 3 [2509.05867].

The paper argues that this community-wise retrieval preserves TCM structure by almost always covering all seven aspects: diseases, recommended formulas, herbs with roles, symptoms and population, tongue and pulse, contraindications, and preparation. A plausible implication is that the ontology is functioning not only as an indexing device but as a structural prior that constrains the generation space toward clinically organized outputs [2509.05867].

## 4. Dataset construction and optimization objectives

The raw formula corpus contains 80,000 formulas from the Chinese Formula Database, Ancient Formula DB, and official catalogues. Augmentation is performed using DeepSeek with 50 manually annotated examples as in-context seeds to extract missing fine-grained fields, after which outputs are cleaned through redundancy processing and human validation. The instruction dataset contains 60,000 examples for SFT and an additional 3,000 preference pairs for DPO [2509.05867].

Each instruction instance has a fixed instruction template, an input comprising symptom description $x$ and GraphRAG global answer $c$, and an output $y$ that is a well-structured multi-section response. The target output includes recommended formula(s) and exact herbal composition with dosages; role assignments and synergy explanations for 君、臣、佐、使; indications or applicable symptoms and population; pulse and tongue patterns; contraindications, including pregnancy, Yin deficiency fire, spleen–stomach deficiency cold, and forbidden combinations; and preparation methods such as honey-frying, charring, decoction steps, and pill formation. The appendix is described as containing detailed case studies illustrating typical outputs [2509.05867].

The SFT objective is given as
\[
\mathcal{L}_{\text{GraphRAG+SFT}
= -\mathbb{E}_{(x,c,y)\sim D}\log P_{\theta}(y\mid x,c). \tag{1}
\]
In the paper’s interpretation, this trains the model to use retrieved context $c$ effectively and to produce consistent, structured explanations across all seven aspects. DPO then aligns model preferences with human or LLM-derived pairwise judgments over candidate responses generated by the SFT reference policy $\pi_{\text{ref}}$ [2509.05867].

The DPO loss is written as
\[
\mathcal{L}_{\text{DPO}(\pi_{\theta}, \pi_{\text{ref}) = - \mathbb{E}_{p \sim D} \left[ \log \sigma \left( \beta \log \frac{\pi_{\theta}(y_w \mid x)}{\pi_{\text{ref}(y_w \mid x)} - \beta \log \frac{\pi_{\theta}(y_l \mid x)}{\pi_{\text{ref}(y_l \mid x)} \right) \right], \tag{2}
\]
with $\sigma$ the sigmoid and $p=(x,c,y_w,y_l)$. The paper states that this stage focuses on coherence, TCM correctness, and reduced hallucinations, and does so without introducing an explicit reward model [2509.05867].

## 5. Theoretical analysis of generalization and hallucination

The theoretical analysis formalizes expected negative log-likelihood as a generalization error:
\[
\mathcal{E}(\theta_{\text{SFT})
= \mathbb{E}\left[-\log P_\theta(y \mid x)\right],
\]
for SFT alone, and
\[
\mathcal{E}(\theta_{\text{GraphRAG+SFT})
= \mathbb{E}\left[-\log P_\theta(y \mid x, c)\right],
\]
for GraphRAG+SFT [2509.05867].

Proposition 1 introduces conditional mutual information
\[
I(y; c \mid x) = \mathbb{E}_{(x,y,c)} \left[ \log \frac{P(y \mid x, c)}{P(y \mid x)} \right].
\]
Under the assumptions that $I(y; c \mid x) \ge \gamma > 0$ and the loss is $\beta$-smooth, the paper gives the bound
\[
\mathcal{E}(\theta_{\text{GraphRAG+SFT}) \le \mathcal{E}(\theta_{\text{SFT}) - \frac{\gamma}{\beta}. \tag{3}
\]
Its stated intuition is that informative GraphRAG context $c$ provides additional predictive information for $y$ beyond $x$ alone, and thus decreases error more effectively under gradient descent [2509.05867].

Proposition 2 extends the analysis to DPO. Defining preference strength as
\[
\Delta = \log \frac{P_{\text{ref}(y_w\mid x,c)}{P_{\text{ref}(y_l\mid x,c)},
\]
the paper states
\[
\mathcal{E}(\theta_{\text{DPO}) \le \mathcal{E}(\theta_{\text{SFT}) - \frac{\mathbb{E}[\Delta]}{\beta}. \tag{4}
\]
Combining this with Proposition 1 yields
\[
\mathcal{E}(\theta_{\text{Final}) \le \mathcal{E}(\theta_{\text{SFT}) - \frac{\gamma}{\beta} - \frac{\mathbb{E}[\Delta]}{\beta}. \tag{5}
\]
The interpretation provided is that expert-preferred answers provide stronger alignment signals and further reduce error [2509.05867].

Hallucination is formalized through $P_{\text{hall}(y \mid x, c)$, the probability that generated $y$ is not supported by factual TCM knowledge given context $c$. Proposition 3 assumes that the relevant fact set $F_x$ is covered by GraphRAG context with probability at least $1-\varepsilon$,
\[
P(F_x \subseteq c(x, G)) \ge 1 - \varepsilon,
\]
and lets $\delta$ be the probability that the model ignores context and relies on internal prior. It then gives
\[
P_{\text{hall}(y \mid x, c(x, G)) \le \varepsilon + \delta. \tag{6}
\]
The paper’s intuition is that hallucination becomes unlikely when retrieval is accurate and the model learns to trust context [2509.05867].

Proposition 4 further states that under DPO, the probability of generating a hallucinated answer $y_l$ is bounded by
\[
P_\theta(y_l \mid x, c) \le P_{\text{ref}(y_l \mid x, c) \cdot e^{-\beta^{-1}\Delta}, \tag{7}
\]
so that hallucination is exponentially suppressed as preference strength $\Delta$ grows. In the TCM setting, the relevant preferences are said to align with expert judgments on correct formula composition and herb roles, valid indications and contraindications, valid tongue and pulse patterns, and the absence of fabricated herbs or formulas [2509.05867].

## 6. Experimental evaluation, deployment, and research context

The experimental setup uses three datasets. The collected formula dataset contains 80,000 records, from which 60,000 examples are used for SFT, 5,000 for testing, and 3,000 for DPO pairs. A clinical dataset from Haodf.com contains 22,800 training and 7,600 test samples with fields including age, disease name, title, description, hopeHelp, visit direction labels (0–19), and secondary labels (0–60); TCM experts provide ground-truth formulas for evaluation in a zero-shot setting. A conflict knowledge dataset contains synthetic instructions with conflicting medical theories, sources, or legal constraints and is used to train the model to detect conflicts and output warnings rather than choosing a side blindly [2509.05867].

Evaluation covers formula prediction quality, structural explanation accuracy, and hallucination analysis. Reported text metrics are BLEU and ROUGE-1/2/L averaged as ROUGE-S. The six TCM-oriented metrics are CCR, CSCR, CCHR, FS, SCR, and LR. Their definitions are as follows [2509.05867]:

| Metric | Definition |
|---|---|
| CCR | Compatibility Compliance Rate |
| CSCR | Correct Sovereign-minister-assistant-messenger Compatibility Rate |
| CCHR | Counter Coarse Hallucination Rate |
| FS | FactScore |
| SCR | Structural/Content Clarity Rate |
| LR | Logical Rate |

CCR is defined by
\[
\text{CCR}
= \left( 1 - \frac{\text{\# non-compliant herb pairs}}{\text{\# total herb pairs}} \right)\times 100\%.
\]
CSCR is
\[
\text{CSCR}
= w_s r_s + w_{mi} r_{mi} + w_a r_a + w_{me} r_{me},
\]
with weights $w_s=w_{mi}=w_a=w_{me}=0.25$. CCHR is
\[
\text{CCHR}
= \left(1 - \frac{\text{\# hallucinated responses}}{\text{\# total responses}}\right)\times 100\%.
\]
FS is
\[
\text{FS}
= \frac{\text{\# supported atomic facts}}{\text{\# total asserted facts}}\times 100\%.
\]
SCR is
\[
\text{SCR} = 0.5 \times \text{CR} + 0.5 \times \text{CPR},
\]
with
\[
\text{CR} = \frac{a}{6}\times 100\%,
\]
where $a$ is the number of six key components present: formula, ingredients, indications, tongue/pulse, contraindications, and preparation. LR is the proportion of logically coherent contextual sentences among all contextual sentences [2509.05867].

The baselines include plug-and-play LLMs such as LLaMA3.2-7B (Chinese Chat), Kimi-7B, DeepSeek-7B, and Qwen2.5-7B; the FT-only TCM LLM TCMLLM; standard RAG using FAISS-based LlamaIndex retrieval with LLaMA3.2-7B as generator; GraphRAG with the same base LLaMA; RAG+SFT and RAG+SFT+DPO; GraphRAG+SFT; and the full model ZhiFangDanTai, defined as GraphRAG + SFT + DPO. Top-$k$ retrieval is typically 3 in the main experiments, with 1 and 2 also tested [2509.05867].

On the 5,000-sample test set with top-$k=3$, ZhiFangDanTai is reported to achieve the best scores across BLEU, ROUGE-S, CCR, CSCR, CCHR, FS, SCR, and LR, while GraphRAG+SFT is second best. The paper reports that GraphRAG consistently outperforms plain RAG, that fine-tuning improves over retrieval-only variants, that DPO further improves coarse and fine-grained hallucination measures, and that vanilla LLMs substantially underperform on all TCM-specific metrics. TCMLLM is reported to improve over plain LLaMA but to remain limited by its instruction data, especially on tongue and pulse, contraindications, and preparation, which lowers SCR and CSCR relative to ZhiFangDanTai. Training FLOPs are higher for the full system because of SFT+DPO, whereas inference FLOPs are only slightly higher than pure LLM calling [2509.05867].

The top-$k$ ablation indicates that ZhiFangDanTai is already strong at $k=1$, that BLEU and ROUGE improve as $k$ increases, and that TCM metrics improve slightly from $k=1$ to $k=2$ or $3$. The authors describe $k=2$–3 as a good trade-off and use $k=2$ in some later experiments for efficiency. A separate ablation on an added GPT pre-training stage reports that GraphRAG+GPT, GraphRAG+GPT+SFT, and GraphRAG+GPT+SFT+DPO perform worse than ZhiFangDanTai, and the paper attributes this to mismatch between pre-training data distribution and task distribution, as well as high Rademacher complexity $\mathfrak{R}_N(\mathcal{F})$ for a 7B model under relatively small domain data. The conclusion given is that GPT is unnecessary and even harmful in this setting [2509.05867].

Backbone ablation with Qwen2.5-7B shows that the method still outperforms all baselines under Qwen, while LLaMA-based variants perform slightly better overall except for minor BLEU differences. On 7,600 real clinical cases with top-$k=2$, and without retraining on that data, ZhiFangDanTai attains the highest TCM metrics and the lowest hallucination rates, which the paper interprets as strong zero-shot generalization to real-world clinical records [2509.05867].

The model is open-sourced at the Hugging Face repository `tczzx6/ZhiFangDanTai1.0`, with a corresponding dataset release. The implementation uses LlamaIndex for GraphRAG, retrieval, and MapReduce; LLaMA-Factory for SFT, DPO, and LoRA; and vLLM for efficient inference. A WebUI accepts symptoms and returns formulas plus explanations, and includes safety disclaimers that present the system as a decision-support tool rather than a substitute for a practitioner. The paper states that it is best suited as an educational tool for teaching formula design and rationale, and as clinical decision support for practitioner review rather than direct patient self-medication. It also notes safety mechanisms for contradictory knowledge and warnings involving conflicting sources or illegal substances such as rhinoceros horn [2509.05867].

The stated limitations include incomplete knowledge graph coverage for rare syndromes, regional practices, or modern clinical innovations; data bias toward classical and official Chinese sources; possible weakness on rare or mixed syndromes; ongoing clinical risks such as dose errors and overlooked comorbidities or drug interactions; optimization primarily for Chinese and English graph translations; and the absence of multimodal integration for tongue images or pulse waveforms. Future directions mentioned include knowledge graph expansion, multilingual support, multimodal integration, stronger theoretical frameworks for generalization and hallucination, and safer deployment through dynamic guideline updates, EMR integration, and stricter safety constraints [2509.05867].

Within TCM informatics and TCM NLP, ZhiFangDanTai is positioned as building on traditional data mining, neural prescription generation from EMR or tongue–pulse images, and recent LLM-based TCM systems such as TCMLLM and TCM-FTP. Within graph-based RAG and hallucination reduction, it is described as building on GraphRAG by Edge et al. and connecting to broader RAG literature including Lewis et al., Xiong et al., and corrective RAG. Its stated uniqueness lies in integrating GraphRAG with TCM formulas, providing systematic theoretical analysis of generalization and hallucination in this domain, and targeting full-format formula explanation including composition, roles, indications, tongue and pulse, contraindications, and processing [2509.05867].

Source: https://www.emergentmind.com/topics/zhifangdantai