---
title: 'VFLAIR-LLM: Privacy-Preserving Split LLM'
url: https://www.emergentmind.com/topics/vflair-llm
type: topic
---

# VFLAIR-LLM: Privacy-Preserving Split LLM

VFLAIR-LLM is an extensible and lightweight split learning framework for large language models (LLMs) that enables privacy-preserving LLM inference and fine-tuning in resource-constrained environments. Developed to address data privacy concerns and the computational demands of private LLM deployments, VFLAIR-LLM provides two core model partition strategies, supports a broad array of benchmarks, and includes comprehensive modules for simulating and countering information leakage attacks. Its codebase and pre-configured benchmarks are available at https://github.com/FLAIR-THU/VFLAIR-LLM, facilitating real-world adoption and reproducible evaluation [2508.03097].

## 1. Model Partitioning Schemes in VFLAIR-LLM

VFLAIR-LLM formalizes two principal approaches to partitioning pretrained LLMs with $n$ transformer layers for split learning: Head–Tail (HT) and Head–Body–Tail (HBT). These allow flexible trade-offs between privacy, communication, and computational overhead.

### Head–Tail (HT) Split
The model is divided into:
- **Data Party (Client):** $M_{\text{head}}$ (embedding + first $n_\text{head}$ layers)
- **Model Party (Server):** $M_{\text{tail}}$ (remaining $n_\text{tail} = n - n_\text{head}$ layers + output head)

**Data flow:**
- Forward: $H_1 = M_{\text{head}}(X)$, $\hat{Y} = M_{\text{tail}}(H_1)$
- Backward: $G_1 = \partial L/\partial H_1$; $G_1$ used to update $M_{\text{head}}$ via backpropagation

**Resource metrics per sample:**
- Client compute: $T_{\text{comp,client(HT)}} = \sum_{\ell=1}^{n_\text{head}} t_{\text{layer}}(\ell)$
- Communication: $C_{\text{comm}} = \Vert H_1\Vert\,\cdot\,\text{sizeof(float)}$

### Head–Body–Tail (HBT) Split
Here, the model is divided into:
- **Data Party (Client):** $M_{\text{head}}$ (first $n_\text{head}$ layers) and $M_{\text{tail}}$ (last $n_\text{tail}$ layers)
- **Model Party (Server):** $M_{\text{body}}$ (middle $n_\text{body}$ layers; $n = n_\text{head} + n_\text{body} + n_\text{tail}$)

**Data flow:**
- Forward: $H_1 = M_{\text{head}}(X)$; $H_2 = M_{\text{body}}(H_1)$; $\hat{Y} = M_{\text{tail}}(H_2)$
- Backward: $G_2 = \partial L/\partial H_2$ updates $M_{\text{tail}}$ and $M_{\text{body}}$, backward $G_1$ updates $M_{\text{head}}$

**Resource metrics:**
- Client compute: $T_{\text{comp,client(HBT)}} = \sum_{\ell=1}^{n_\text{head}} t_{\text{layer}}(\ell) + \sum_{\ell=n_\text{head} + n_\text{body} + 1}^{n} t_{\text{layer}}(\ell)$
- Communication: $H_1$ (client→server) and $G_1$ (server→client) per sample

### Compute–Communication–Privacy Trade-offs
Client compute $\sim f(n_\text{head})$ increases linearly with added head layers. Communication cost $C_\text{comm} \sim O(d_\text{hidden}\times\text{batch size})$. Privacy leakage—measured by attack performance (AP)—decreases empirically as $n_\text{head}$ increases, but this incurs higher local compute and communication [2508.03097, Figure 8].

## 2. Task and Dataset Scope

VFLAIR-LLM benchmarks three major NLP task categories across 18 datasets, ensuring coverage of both practical and adversarially-relevant scenarios:

| Task Type                    | Datasets (examples)                     | Primary Metrics         |
|------------------------------|-----------------------------------------|------------------------|
| Classification/Regression    | SST-2, CoLA, MRPC, MNLI, QNLI, RTE, Yelp, STS-B | Accuracy, Pearson’s $\rho$ |
| Span QA                      | SQuAD v1.1                              | EM, F1                 |
| Generation/CausalLM/QA       | Lambada, Alpaca, Dolly, CodeAlpaca, MATH, GSM8K, TextVQA | Rouge, CodeBLEU        |

These datasets represent a diverse spectrum, including discrete-label and sequence-level outputs, as well as specialty domains (code, math, vision-language) [2508.03097, Table 3]. The inclusion facilitates comprehensive assessment of privacy–utility trade-offs in practical and high-risk adaptation settings.

## 3. Attack and Defense Framework

VFLAIR-LLM implements standard modules for five adversarial attacks and nine defense mechanisms, encompassing both perturbation- and learning-based approaches.

### Attack Methods

- **Model Inversion Attacks (MIA):** Aim to reconstruct $X$ from $H_1$.
  - VMI (Vanilla Model Inversion): $\min_{X'} \Vert M_\text{head}(X') - H_1\Vert_2^2$
  - RMI: continuous relaxation and regularization, $\min_{z}\Vert M_\text{head}(E(z))- H_1\Vert^2 +\lambda\,\text{reg}(z)$
  - BiSR: two-phase, semi-white-box, achieves highest empirical AP
  - Attack Performance (AP): recall/exact match rate of reconstructed tokens

- **Label Inference Attacks (LIA):** Aim to infer $Y$ from $G_2$
  - BLI (Batch-level Label Inversion): gradient-statistics to label inversion
  - NS (Norm Scoring): infer $Y_i$ by thresholding $\Vert\nabla_{x_i}\mathcal{L}\Vert$

### Defense Strategies

- **Perturbation-Based (6):** Differential Privacy (DP), Sparsification (SP), token-level MLDP (SanText, CusText, RanText), Split-N-Denoise (SnD)
  - Hyperparameters: $\epsilon$ for DP/MLDP ($0.01$–$500$); $r$ for SP ($95$–$98$\%)
- **Learning-Based (3):**
  - MID (Mutual Information Defense): bottleneck $f_\theta(H_1)$, minimize $L_\text{task} + \lambda I(f_\theta(H_1); X)$, $\lambda \in [1\text{e-5},0.5]$
  - AT (Adversarial Training): privatizer $P$ adversarial to inverter $A$, tradeoff parameters $\alpha,\beta$
  - TO (TextObfuscator): embedding perturbations clustered; $n_\text{cluster} \in \{100,150,200,250\}$

## 4. Empirical Benchmarks and Insights

Benchmarking combines main task performance (MP: accuracy, Rouge, CodeBLEU, EM, etc.), inversion/attack performance (AP), and composite defense capability scores (DCS).

### Fine-Tuning Regimes

- **Full-Vanilla**, **Full-LoRA**, **Local-Vanilla**, **Local-LoRA** are evaluated
- Full-LoRA halves training epochs on SST-2 (BERT) for matched accuracy and gives $>2\times$ compute reduction on GPT-2/Mistral with only $3{-}5$\% MP drop
- Local tuning benefits from HBT as it allows more trainable parameters client-side

### Defense Efficacy

- Without defense, BiSR AP $\approx$ 0.30 (SST-2), 0.15 (Alpaca), 0.12 (GSM8K)
- Perturbation methods reduce both utility (MP) and information leakage (AP) with stronger noise
- Learning-based defenses (MID, AT) consistently achieve best privacy–utility trade-off: MID yields lowest AP for $MP>0.85$; AT is close
- On SQuAD v1.1, MID achieves composite C-DCS $\approx 0.85$ (SST-2), $0.83$ (CoLA); AT achieves $0.81/0.79$; DP or SP require $\epsilon<1$ or $r>97\%$ to approach, but with $>10\%$ MP loss

### Partitioning and Privacy

- Increasing $n_\text{head}$ ($2 \leq n_\text{head}\leq 5$): reduces AP by $20{-}40\%$ at $<2\%$ MP loss, but increases local computation proportionally
- Learning-based MID/AT atop LoRA fine-tuning provides optimal privacy–utility; recommended: MID $\lambda\approx0.1{-}0.5$, AT $\alpha\approx1.0$, $\beta\approx0.1$ 

## 5. Best Practices and Deployment Recommendations

- **Partitioning:** HT is straightforward; HBT preferred if the server must not access ground-truth labels or outputs
- **Layer assignment:** $n_\text{head}\sim3{-}5$ (about $10{-}20\%$ of model) gives a robust privacy–utility compromise
- **Defenses:** MID or AT on LoRA-finetuned splits optimally balances MP and AP; DP/SP are viable for lightweight protection with careful tuning, but significant MP decrease occurs for strong privacy settings
- **Resource management:** Reduce $n_\text{head}$ or use Local tuning if client GPU memory is critical; batch size $16{-}32$, LoRA LR $1\text{e-4}$, early stopping is recommended
- **Operational mode:** Standalone mode is preferred for research (e.g., Alpaca on Llama3-8B: $23$ tokens/s standalone vs. $7$ tokens/s distributed)

## 6. Scope, Limitations, and Reproducibility

VFLAIR-LLM is designed for flexible, privacy-conscious LLM deployment under real-world resource and threat constraints. Benchmarks and code are publicly available, supporting customization for model splits, attacks, defenses, and hyperparameter sweeps. The platform’s task, dataset, and threat coverage enable rigorous experimental comparisons across the privacy–efficiency spectrum [2508.03097]. The framework does not extend to scenarios outside the split learning paradigm, nor to privacy threats not captured by inversion or label inference attacks as modeled in the included modules. A plausible implication is that privacy guarantees are tightly coupled to the implemented attack/defense taxonomy. 

For implementation and experiment recipes, practitioners are referred to the open-source codebase.

Source: https://www.emergentmind.com/topics/vflair-llm