Papers
Topics
Authors
Recent
Search
2000 character limit reached

Fraud-Aware Selective CoT Distillation

Updated 6 February 2026
  • The paper introduces a novel fraud-aware selective CoT distillation mechanism that autonomously generates and selects multi-hop reasoning chains to reveal fraud cues in text-attributed graphs.
  • The methodology combines semantic and structural cues through an integrated LLM-GNN pipeline, achieving up to +8.8% improvement in AUPRC over leading baseline models.
  • The approach employs asymmetric co-training and caching strategies to reduce LLM calls and scale batch sizes efficiently, enabling significant computational speedups.

Fraud-aware selective Chain-of-Thought (CoT) distillation is a targeted mechanism developed to enhance graph-based fraud detection in text-attributed graphs (TAGs) by autonomously generating and selecting diverse, semantically-relevant reasoning paths. Integrated within the FraudCoT framework, it addresses the limitations of predefined prompting and decoupled LLM-GNN training pipelines by aligning semantic and structural cues from textual and relational data. This process involves autonomous LLM generation of multi-hop reasoning chains, their selective distillation based on fraud-relevance and diversity, and their integration into node representations for downstream graph neural network (GNN) classification—enabling state-of-the-art detection performance and scalable end-to-end optimization (Li et al., 30 Jan 2026).

1. Problem Formulation and Notation

Fraud-aware selective CoT distillation operates on a text-attributed graph (TAG) G=(V,E,X)G = (V, E, X), where VV is the set of nodes (e.g., reviews, transactions, users), E⊆V×VE \subseteq V \times V represents binary relations, and X={xv ∣ v∈V}X = \{x_v\,|\,v \in V\} comprises raw textual attributes. Each node vv is assigned a binary fraud label yv∈{0,1}y_v \in \{0, 1\}, and A∈{0,1}∣V∣×∣V∣A \in \{0, 1\}^{|V|\times|V|} denotes the adjacency matrix, with N(v)N(v) defining the set of neighbors of vv. The objective is to learn a scoring function fθ(v)f_\theta(v) approximating VV0 for accurate node-level fraud detection (Li et al., 30 Jan 2026).

2. Autonomous Chain-of-Thought Generation

The distillation process begins with teacher LLM-based generation of graph-aware reasoning chains for selected nodes. For each distillation node VV1, a composite prompt VV2 is constructed by concatenating its raw text VV3 and textual descriptions of its 1-hop or multi-hop neighbors VV4. The teacher LLM VV5 receives VV6 and samples VV7 chain-of-thoughts:

VV8

Each VV9 is a free-form path that references both local semantic features and relational (neighbor) signals, aiming to uncover multi-hop fraud cues embedded in the graph structure (Li et al., 30 Jan 2026).

3. Selective Distillation: Relevance Scoring and Diversity Enforcement

A fraud-aware CoT scoring function E⊆V×VE \subseteq V \times V0 is proposed to select high-value chains:

E⊆V×VE \subseteq V \times V1

where E⊆V×VE \subseteq V \times V2 is cosine similarity and E⊆V×VE \subseteq V \times V3 is a lightweight encoder (e.g., BERT-tiny), balancing semantic relevance and structural alignment. Scores are normalized to E⊆V×VE \subseteq V \times V4 across samples.

To ensure that distilled CoTs are both fraud-relevant and diverse, the procedure selects:

  • E⊆V×VE \subseteq V \times V5: the top-E⊆V×VE \subseteq V \times V6 CoTs with scores E⊆V×VE \subseteq V \times V7,
  • E⊆V×VE \subseteq V \times V8: the bottom-E⊆V×VE \subseteq V \times V9 CoTs with scores X={xv ∣ v∈V}X = \{x_v\,|\,v \in V\}0, retaining only those with pairwise Levenshtein edit distance X={xv ∣ v∈V}X = \{x_v\,|\,v \in V\}1 to enforce diversity.

The student model X={xv ∣ v∈V}X = \{x_v\,|\,v \in V\}2 is distilled by minimizing:

X={xv ∣ v∈V}X = \{x_v\,|\,v \in V\}3

where X={xv ∣ v∈V}X = \{x_v\,|\,v \in V\}4 and X={xv ∣ v∈V}X = \{x_v\,|\,v \in V\}5 are sets of positive and negative chains, and X={xv ∣ v∈V}X = \{x_v\,|\,v \in V\}6 is an unlikelihood loss coefficient. LoRA (Low-Rank Adaptation) adapters with rank X={xv ∣ v∈V}X = \{x_v\,|\,v \in V\}7–X={xv ∣ v∈V}X = \{x_v\,|\,v \in V\}8 are employed for efficient parameterization during student finetuning. Empirically, X={xv ∣ v∈V}X = \{x_v\,|\,v \in V\}9 nodes, vv0 samples, vv1–3, and vv2 are optimal (Li et al., 30 Jan 2026).

4. Integration of Distilled CoTs into GNN Representation

After student finetuning, a distilled CoT is generated for each node vv3 as vv4 and concatenated with the raw text to yield vv5. This extended textual attribute is processed by an LLM-style encoder vv6 to obtain initial node embeddings:

vv7

These enriched embeddings encode both direct and multi-hop semantic-structural fraud cues, providing a more informative input for the downstream GNN (Li et al., 30 Jan 2026).

5. Asymmetric LLM–GNN Co-Training and Efficiency

A key component is the asymmetric co-training pipeline. In each batch vv8, only target nodes vv9 are freshly encoded:

yv∈{0,1}y_v \in \{0, 1\}0

while neighbor nodes yv∈{0,1}y_v \in \{0, 1\}1 utilize cached embeddings yv∈{0,1}y_v \in \{0, 1\}2 computed once at initialization. Message passing employs yv∈{0,1}y_v \in \{0, 1\}3 heterogeneous GraphSAGE layers:

yv∈{0,1}y_v \in \{0, 1\}4

The total loss is a weighted sum of binary cross-entropy node classification and distillation loss:

yv∈{0,1}y_v \in \{0, 1\}5

Typically, yv∈{0,1}y_v \in \{0, 1\}6. Only target nodes propagate gradients through yv∈{0,1}y_v \in \{0, 1\}7, greatly reducing backpropagation cost.

In terms of efficiency, the asymmetric strategy decreases per-epoch LLM calls from yv∈{0,1}y_v \in \{0, 1\}8 (naive) to yv∈{0,1}y_v \in \{0, 1\}9, with a one-time A∈{0,1}∣V∣×∣V∣A \in \{0, 1\}^{|V|\times|V|}0 cache build. Using the DigitalMusic dataset, this yields epoch time reduction from A∈{0,1}∣V∣×∣V∣A \in \{0, 1\}^{|V|\times|V|}1 s (naive) to A∈{0,1}∣V∣×∣V∣A \in \{0, 1\}^{|V|\times|V|}2 s (asymmetric), corresponding to a A∈{0,1}∣V∣×∣V∣A \in \{0, 1\}^{|V|\times|V|}3 speedup, and permits batch size scaling from A∈{0,1}∣V∣×∣V∣A \in \{0, 1\}^{|V|\times|V|}4 to A∈{0,1}∣V∣×∣V∣A \in \{0, 1\}^{|V|\times|V|}5 (Li et al., 30 Jan 2026).

6. Empirical Evaluation and Baselines

FraudCoT, incorporating fraud-aware selective CoT distillation, is experimentally validated on InstantVideo (37,126 nodes, 9.9M edges), DigitalMusic (64,706 nodes, 7.7M edges), and PromotionAbuse (371K nodes, 1.3M edges). Comparisons are drawn against:

  • GNNs: GraphSAGE, HGT, ConsisGAD, PMP, GAAP
  • Pure LLMs: Qwen-8B, InstructGLM
  • Graph-enhanced LLMs: LLaGA, GraphGPT, HiGPT
  • LLM-enhanced GNNs: TAPE, FLAG

Evaluation metrics include Macro-F1, AUROC, and AUPRC (using scikit-learn implementations). FraudCoT achieves up to A∈{0,1}∣V∣×∣V∣A \in \{0, 1\}^{|V|\times|V|}6 improvement in AUPRC over the best baseline (PromotionAbuse), with consistent performance gains across all public and industrial dataset splits. Ablation studies confirm that both the selective distillation and asymmetric co-training components are essential for optimal results (Li et al., 30 Jan 2026).

Dataset Nodes Edges Best AUPRC Gain
InstantVideo 37,126 9.9M up to +8.8%
DigitalMusic 64,706 7.7M up to +8.8%
PromotionAbuse 371,000 1.3M up to +8.8%

7. Implementation Strategies and Practical Considerations

Implementation guidelines include:

  • Use LoRA adapters with rank A∈{0,1}∣V∣×∣V∣A \in \{0, 1\}^{|V|\times|V|}7–A∈{0,1}∣V∣×∣V∣A \in \{0, 1\}^{|V|\times|V|}8.
  • Set unlikelihood loss A∈{0,1}∣V∣×∣V∣A \in \{0, 1\}^{|V|\times|V|}9.
  • Sample N(v)N(v)0 neighbors, with hop number N(v)N(v)1.
  • Cache node embeddings on CPU or host memory to minimize GPU LLM calls.
  • Apply early stopping based on validation AUROC; batch sizes N(v)N(v)2–N(v)N(v)3 provide optimal throughput.
  • Precompute all text encodings for CoT scoring to accelerate distillation selection.

A plausible implication is that this design allows practitioners to scale LLM-enhanced GNNs to industrial TAGs with both superior predictive performance and orders-of-magnitude throughput improvements, without requiring expensive end-to-end LLM inference on all neighbors during training (Li et al., 30 Jan 2026).

Definition Search Book Streamline Icon: https://streamlinehq.com
References (1)

Topic to Video (Beta)

No one has generated a video about this topic yet.

Whiteboard

No one has generated a whiteboard explanation for this topic yet.

Follow Topic

Get notified by email when new papers are published related to Fraud-Aware Selective CoT Distillation.