- The paper introduces TIGER, a continuous embedding-subspace optimization attack that reconstructs private text from transformer gradients across batched decoder and encoder models.
- TIGER achieves 98.8 ROUGE-1 without defenses and retains 71.1 under Gaussian noise, while BF16 quantization causes only modest decoder degradation from 99.1% to 94.5%.
- The method shows that utility-preserving noise and reduced-precision training may not prevent leakage, although recovery weakens when token counts approach hidden dimensions or encoder sequence order must be reconstructed.
Overview
TIGER (Transformer Data Inversion from Gradients via Embedding-subspace Reconstruction) is a gradient inversion attack that reconstructs private client text from transformer model updates shared in federated learning (FL) (2606.18312). The attack targets the honest-but-curious server setting and is designed for the regimes where prior text inversion methods are most brittle: batched updates, encoder-only architectures with non-causal attention, quantized gradients, and gradients perturbed with differential-privacy (DP)-style noise. Its central move is to convert the low-rank subspace signal previously used for discrete token membership tests—most prominently in DAGER—into a differentiable objective over continuous token embeddings, avoiding brittle threshold-based filtering entirely.
For a linear layer with input Zl∈RT×d (T non-padding tokens, d hidden dimension), the query-gradient factors as GQl=(Zl)⊤ΔQl, so rank(GQl)≤T. When T<d and the backpropagated query gradient has full rank, the row span of the block inputs equals the column span of the observed gradient. DAGER exploits this by testing whether candidate token embeddings lie in this gradient-induced subspace Sl, using a distance-to-subspace test with layer-dependent thresholds. This discrete membership test is exact in high precision but collapses under numerical noise: the paper shows DAGER's ROUGE-1 drops to 0.0 at even σ=10−5 additive Gaussian gradient noise, and under BF16 gradient quantization.
Method
TIGER replaces the membership test with continuous optimization of a dummy embedding matrix Z^0. The core forward span-distance loss sums, over the first Lmax transformer layers, the squared distance between normalized recovered hidden states and their projections onto T0:
T1
Normalization is essential: the subspace constraint identifies directions only, and without it the optimizer can trivially drive hidden-state norms to zero. After optimization, embeddings are mapped to tokens by cosine similarity against the vocabulary.
Decoder attack: causal masking permits token-by-token recovery, optimizing a single position at a time given the recovered prefix. Because first tokens causally determine the rest of each sequence and admit multiple valid global minima across batch elements, TIGER adds a deduplication loss that projects the current candidate's hidden state onto the subspace directions already occupied by previously recovered first tokens (constructed via an SVD-based basis rotation, described in the appendix). This component is worth roughly 23 ROUGE points in the ablation.
Encoder attack: with non-causal attention, TIGER jointly optimizes the entire batch embedding matrix and adds a backward span-distance loss requiring the observed subspace basis T2 to lie in the span of the recovered hidden states, computed via a jitter-regularized projection to avoid differentiating through an SVD. Removing this loss drops encoder ROUGE-1 from 72.2% to 41.4%, with failures driven by cross-example collapse: without it, none of ten attacks recovered four distinct sequences, and in three of ten all four reconstructions collapsed to a single batch member.
Initialization: the non-convex objective depends heavily on initialization. TIGER fits per-position full-covariance Gaussians over raw embeddings from a public corpus (WikiText-103 by default) and runs T3 restarts of 3000 Adam steps each. An IMDb-fitted prior performs nearly identically (55.4 vs. 54.9 ROUGE-1 for the decoder), indicating corpus choice is not critical, whereas vocabulary-only and random initializations degrade performance substantially.
Experimental results
The evaluation covers Gemma-3-4B-IT (decoder, next-token prediction) and EmbeddingGemma-300M (encoder, sequence classification), on WikiText-103 batches, measured by token-level ROUGE-1/ROUGE-L with optimal one-to-one batch matching.
Robustness to noise is the headline decoder result. At batch size 1, TIGER achieves 98.8 ROUGE-1 undefended and retains 71.1 ROUGE-1 at T4, while DAGER scores 0.0 at every nonzero noise level across all batch sizes. Noise levels were validated for utility preservation: fine-tuning accuracy on FictionalQA is largely unaffected until T5, so reconstruction remains feasible in the utility-preserving regime. Under BF16 quantization, TIGER degrades only modestly (99.1% → 94.5% decoder; 72.2% → 65.5% encoder) while DAGER again drops to zero.
Encoder results are larger in relative terms but more noise-sensitive. Undefended, TIGER reaches 90.8 ROUGE-1 at T6 and 60.3 at T7, versus LAMP's 9.8–23.3 across the same range. However, at T8 encoder performance falls sharply (e.g., 90.8 → 48.1 at T9), and the paper concedes this indicates joint batch optimization provides a weaker signal than sequential decoder recovery when subspace estimates are noisy.
Scaling behavior: decoder recovery is consistently higher for fewer, longer sequences at a fixed token budget d0, reflecting the benefit of recovered prefixes. Encoder ROUGE-1 stays comparatively strong at larger d1 (e.g., 60.3 at d2, d3) while ROUGE-L falls faster, indicating many tokens are recovered but ordering degrades in non-causal settings.
Ablations: performance saturates around d4–d5 layers while runtime grows roughly linearly, justifying d6. Replacing sequential decoder recovery with joint optimization plus the backward loss costs nearly 25 ROUGE points, confirming that exploiting causal structure is the dominant factor in decoder stability.
Limitations and open questions
The paper identifies three principal constraints. First, the attack inherits the d7 requirement: as total token count approaches the hidden dimension, the gradient subspace becomes full-rank and the objective loses discriminative power, bounding applicability to large-batch regimes on small models. Second, the defense evaluation is limited to additive Gaussian noise; the interaction with gradient clipping, masking, secure aggregation, and compression is unexamined, and these mechanisms may degrade the subspace signal in qualitatively different ways. Third, TIGER recovers token content but not reliably ordered sequences in the encoder case, and the authors leave open hybrid pipelines combining continuous subspace recovery with discrete refinement (e.g., language-model-prior reordering in the style of LAMP). An additional methodological caveat is that DAGER's second-layer threshold was relaxed to d8 for these experiments, so baseline comparisons embed a tuning choice; and the label-free LAMP adaptation for decoders failed entirely, so no optimization-based decoder baseline with a language prior is reported.
Conclusion
TIGER demonstrates that the low-rank structure of transformer linear-layer gradients supports a robust continuous inversion objective, extending gradient inversion to encoder-only models and to defended decoder settings where discrete algebraic attacks fail completely. The practical implication is direct: modest gradient noise or reduced-precision gradient computation, at levels that do not harm training utility, should not be assumed to eliminate reconstruction risk in federated LLM training. The residual open problems—full-rank subspaces at large d9, non-noise defenses, and ordered encoder recovery—define the boundary of the current attack's applicability.