Token-based Model Inversion (TMI)
- Token-based Model Inversion (TMI) is a paradigm that organizes inversion operations at token granularity, exploiting localized signals for reconstruction, editing, and selective unlearning.
- Core techniques include token-local optimization, saliency-based token selection, and continuous token-space search followed by discrete projection to enhance inversion fidelity.
- TMI finds dual-use applications ranging from data-free learning in federated settings to high-fidelity privacy attacks, with evaluations using metrics like BLEU, F1, and exact match.
Token-based Model Inversion (TMI) designates inversion procedures in which the recovered object, the optimization variable, or the supervisory signal is organized at token granularity rather than treated as a monolithic image or sequence. In recent work, the relevant tokens may be ViT patch tokens and class-token attention, answer tokens in autoregressive vision-LLMs, discrete text tokens or continuous token embeddings in LLMs, or newly learned pseudo-tokens in diffusion text encoders. Some papers use the term explicitly, notably in vision-LLM inversion, while others describe the same pattern under names such as sparse model inversion, token relabel, selective unlearning, textual inversion, or internal-state inversion (Nguyen et al., 6 Aug 2025, Shen et al., 11 May 2026, Zhou et al., 22 Jan 2026, Kalikman et al., 16 Jun 2026, Fei et al., 2023, Dong et al., 22 Jul 2025).
1. Conceptual scope and token granularity
Across the literature, TMI is unified less by a single threat model than by a shared granularity assumption: the decisive inversion operations are token-wise. In ViTs, the token is a patch embedding whose contribution is indexed by class-token attention. In autoregressive VLMs and LLMs, the token is an output or input subword whose probability, hidden state, or gradient footprint can be manipulated or inverted. In diffusion personalization and analogy generation, the token is often a learned pseudo-token or a small set of trainable token embeddings inserted into a frozen text encoder.
| Setting | Token unit | Inversion target |
|---|---|---|
| Vision transformers in one-shot FL | Patch tokens and class-token attention | Synthetic images and token relabeling for global ViT training |
| VLMs | Autoregressive answer tokens | Image latent producing a target textual answer |
| LLMs and transformer leakage | Text tokens, next-token distributions, hidden states, or embeddings | Prompt recovery or token reconstruction |
| Diffusion personalization and analogy | Learned pseudo-tokens or DIFF tokens | Subject, style, or transformation encoded in token embeddings |
This scope includes both adversarial and constructive uses. Adversarial instances reconstruct private training inputs, hidden prompts, client text, or visual examples from logits, hidden states, or gradients. Constructive instances synthesize pseudo-data for one-shot federated learning, remove memorized PII without original training data, or encode a subject, style, or relational edit into reusable text tokens. This suggests that “token-based” in this literature refers less to discreteness than to the level at which inversion decisions are made.
A recurrent contrast is with dense or sequence-level inversion. Dense inversion updates every pixel or patch uniformly; sequence-level inversion aggregates all token losses into one objective. TMI instead privileges localized signals: attention-ranked foreground patches, masked PII spans, per-token autoregressive losses, continuous token embeddings later projected back to the vocabulary, or token-specific subspace constraints.
2. Core algorithmic patterns
A first pattern is token-local optimization under a global model. In VLM inversion, token-based MI treats answer generation as a sequence of token predictions and updates the generator latent after every token using
followed by gradient descent on (Nguyen et al., 6 Aug 2025). In hidden-state inversion for decoder-only LMs, a continuous proxy embedding is optimized position-wise so that the predicted hidden state matches the leaked target , and only after the inner loop is a discrete token committed by nearest-neighbor verification (Słowikowski et al., 1 Jul 2026). In TIGER, the optimized variable is a continuous token-embedding matrix , and the loss is the distance of recovered hidden trajectories to gradient-induced embedding subspaces rather than direct full-gradient matching (Kalikman et al., 16 Jun 2026).
A second pattern is token selection by saliency or structure. FedMITR computes class-token attention in a ViT and masks low-attention tokens during inversion, thereby separating semantic foreground from low-information background. The same attention scores are then used for token relabeling: high-information tokens retain the pseudo-label, whereas low-information tokens are relabeled by an ensemble teacher via KL distillation (Shen et al., 11 May 2026). DFSU likewise introduces token-level privacy masks over synthetic pseudo-PII; masked tokens receive privacy-maximizing pressure, while unmasked tokens are optimized for utility preservation (Zhou et al., 22 Jan 2026).
A third pattern is continuous token-space search followed by discrete projection. Gradient-free textual inversion parameterizes a learned token embedding as
with optimized by CMA-ES in a low-dimensional subspace defined by PCA or prior normalization, so only forward passes are required (Fei et al., 2023). TIGER and hidden-state inversion follow the same broad logic: optimize in embedding space, then map back to discrete tokens by cosine similarity or nearest-neighbor testing (Kalikman et al., 16 Jun 2026, Słowikowski et al., 1 Jul 2026). This suggests a common design choice across TMI systems: continuous search is often used to avoid brittle discrete optimization, but the recovered artifact remains token-level at evaluation time.
3. Vision, diffusion, and multimodal generation
In generative imaging, TMI first appeared in forms that optimize text-conditioning tokens rather than pixels. “Gradient-Free Textual Inversion” formulates the task as learning an embedding $e^\*$ for a pseudo-token $x^\*$ so that prompts containing $x^\*$ reconstruct and generalize a subject or style from a few images. The contribution is a gradient-free framework based on CMA-ES, CLIP-based initialization, and low-dimensional subspace optimization; the resulting performance is reported as comparable to gradient-based textual inversion while retaining forward-only deployment advantages (Fei et al., 2023).
BRAT extends this line by introducing bonus tokens and enforcing orthogonality, thereby increasing the token-space capacity available to represent a concept. The method is architecture agnostic in the narrow sense used by the paper: it does not rely on UNet-specific layers and is demonstrated with both a UNet backbone and a diffusion transformer. Its reported finding is that the bonus token improves adherence to the source images and the vision transformer improves adherence to the prompt (Baker, 2024).
Difference Inversion reorients token-space inversion from subject identity to relational editing. The transformation 0 is encoded as learnable DIFF tokens 1, guided by an interpolated CLIP delta and trained with a Token Consistency Loss that enforces reversibility between 2 and 3. Zero initialization of token embeddings is central: the DIFF tokens begin as a no-op and are forced to encode only the difference rather than residual content from 4 (Kim et al., 9 Jun 2025).
In VLMs, TMI is explicit rather than implicit. “Model Inversion Attacks on Vision-LLMs” defines TMI as stepping through answer tokens 5 and updating the image latent after every token. The paper also introduces TMI-C, which fully converges each token before advancing, and compares both to sequence-based methods SMI and SMI-AW. The empirical result is not that token-wise inversion is universally superior: TMI works, but TMI-C performs worse than TMI, and both are outperformed by sequence-based methods, especially SMI-AW with logit-maximization (Nguyen et al., 6 Aug 2025). This corrects a possible misconception: token granularity offers a strong local signal, but it does not automatically dominate sequence-level optimization.
4. LLMs, hidden states, and prompt recovery
In language modeling, TMI often targets hidden prompts rather than training data. “LLM Inversion” studies the problem of recovering a prompt 6 from the next-token probability vector
7
using a T5-based inverter that unrolls the full probability vector into a sequence of pseudo-tokens for the encoder. On Llama-2 7B, the reported results are BLEU 8, token-level F1 9, and exact recovery of 0 of prompts, showing that next-token distributions contain substantial information about preceding text (Morris et al., 2023).
PILS sharpens this result by using sequences of next-token logprob vectors rather than a single step. Its key theoretical claim is that, after additive log-ratio transformation, next-token distributions lie in a 1-dimensional subspace determined by the hidden size, which enables lossless compression of the full distribution sequence into a lower-dimensional representation for inversion. Empirically, PILS yields 2–3 times higher exact recovery than prior methods; one highlighted case increases exact recovery from 4 to 5, and the method generalizes upward in test-time generation length, with an inverter trained on 6 steps improving when evaluated on 7 steps (Nazir et al., 20 Jun 2025).
A separate branch targets leaked internal states rather than logits. “Recovering Input Text from Hidden States” studies inversion from last-layer hidden states of GPT-2 using continuous embedding-space optimization and delayed token commitment. On 8-token C4 prompts, exact match rises from 9 to 0 and mean similarity from 1 to 2 as the candidate window is widened, indicating that many failures are near-miss neighborhood errors rather than deep ambiguities. The paper also identifies a categorical asymmetry: space-prefixed, high-frequency function words dominate the failures, while content-bearing tokens are recovered almost perfectly (Słowikowski et al., 1 Jul 2026).
“Depth Gives a False Sense of Privacy” extends hidden-state inversion to deep layers and long prompts in LLMs up to 3B parameters. It introduces two white-box optimization attacks, a transfer-based black-box attack, and a generation-based attack that treats inversion as translation from internal states to text. The headline result is that a 4-token medical consulting prompt can be nearly perfectly inverted with 5 F1 token matching from the middle layer of a Llama-3 model, directly challenging the assumption that deeper states are too abstract to invert (Dong et al., 22 Jul 2025).
5. Federated learning, privacy, and security
TMI also appears as a mechanism for data-free learning. FedMITR addresses one-shot federated learning with ViTs by synthesizing data from uploaded client models and then training a server model using sparse model inversion and token relabel. Sparse inversion masks low-attention tokens so that only semantic foreground patches continue to change, while token relabel uses pseudo-labels for high-information tokens and ensemble soft labels for low-information tokens. The paper gives an algorithmic-stability analysis showing that sparse inversion eliminates gradient instability from background noise and token relabel reduces gradient variance, yielding a tighter generalization bound than dense inversion (Shen et al., 11 May 2026).
Other work uses TMI defensively against memorization. DFSU synthesizes pseudo-PII by LLM inversion, annotates token-level privacy masks, and applies Privacy-Selective Contrastive Unlearning in a LoRA subspace. Its loss separates privacy and utility streams:
6
so minimizing 7 preserves performance on non-PII tokens while maximizing loss on masked PII tokens. On AI4Privacy-based experiments with Pythia-410M, DFSU reduces E-Hit from 8 to 9 on WikiText+PII with PPL changing from 0 to 1, and reduces E-Hit from 2 to 3 on MNLI+PII with accuracy changing from 4 to 5 (Zhou et al., 22 Jan 2026).
In federated attack settings, TIGER shows that token-level leakage persists even under perturbations often assumed to be protective. It optimizes continuous token embeddings to minimize their distance to gradient-induced subspaces, then projects the recovered embeddings back to vocabulary tokens. On decoder models, TIGER remains effective under BF16 gradients and under DP-style Gaussian noise; for example, when DAGER’s ROUGE-1 drops to 6 at 7, TIGER remains high and still reports non-zero reconstruction at 8 (Kalikman et al., 16 Jun 2026). A similar conclusion appears in internal-state inversion: quantization, dropout, noisy input embeddings, and Laplace-style DP noise degrade utility before they prevent inversion, so simple perturbative defenses do not produce a satisfactory privacy-utility tradeoff (Dong et al., 22 Jul 2025).
These results support a broader interpretation of TMI as dual-use. The same token-level machinery can improve data-free learning or selective unlearning, yet it can also reconstruct prompts, hidden states, and client data with high fidelity when exposed signals remain sufficiently informative.
6. Evaluation, comparative findings, and open problems
TMI is evaluated with markedly different metrics across domains. Language-model prompt recovery uses BLEU, ROUGE, exact match, token-level F1, and embedding cosine similarity (Morris et al., 2023, Nazir et al., 20 Jun 2025, Słowikowski et al., 1 Jul 2026). Privacy-preserving unlearning uses Exact Reconstruction Rate, Fractional Reconstruction Similarity, Sample-Level Exposure Rate, and Entity-Level Hit Rate (Zhou et al., 22 Jan 2026). VLM inversion combines conventional DNN-based attack accuracy, MLLM-based attack accuracy, human evaluation, and feature distances, with human evaluation reaching an attack accuracy of 9 on reconstructed images (Nguyen et al., 6 Aug 2025). Diffusion analogy generation uses CLIP and DINO-v2 directional scores, human preference, and VLM judgments; Difference Inversion reports a DINO-v2 directional score of 0 and a 1 human preference rate over baselines (Kim et al., 9 Jun 2025). One-shot FL reports downstream classification accuracy, where FedMITR improves Mini-ImageNet Dir2 from DeepInversion’s 3 to 4 and OfficeHome Dir5 from 6 to 7 (Shen et al., 11 May 2026).
Several comparative findings recur. First, richer output traces usually help: PILS benefits from multiple generation steps rather than a single logprob vector, and hidden-state inversion benefits from larger candidate windows (Nazir et al., 20 Jun 2025, Słowikowski et al., 1 Jul 2026). Second, token-level treatment is not inherently best at the task level: in VLM inversion, TMI and especially TMI-C are weaker than sequence-based SMI and SMI-AW, indicating that local token supervision can be noisier than sequence-level aggregation (Nguyen et al., 6 Aug 2025). Third, access assumptions remain decisive. White-box gradient or hidden-state access generally yields stronger token recovery than black-box transfer, and black-box generative inversion depends more on the match between attacker data and victim inputs (Dong et al., 22 Jul 2025).
The limitations reported across papers are also structurally aligned. Many methods require white-box logits, hidden states, gradients, or equal-weight client ensembles; others rely on a public generator, a public corpus, or a distributionally similar attacker dataset. Several are architecture-specific in practice even when conceptually broader: FedMITR is tailored to ViTs, DFSU assumes white-box logit access, and gradient-free textual inversion benefits from CLIP/text-encoder alignment in Stable Diffusion (Shen et al., 11 May 2026, Zhou et al., 22 Jan 2026, Fei et al., 2023). Future directions named in the papers include better token importance metrics beyond attention, more robust PII annotators, extensions to CNN-transformer hybrids and multimodal foundation models, stronger formal privacy guarantees, and defenses that do more than add light noise or quantization (Shen et al., 11 May 2026, Zhou et al., 22 Jan 2026, Dong et al., 22 Jul 2025).
Taken together, the recent literature establishes TMI as a general inversion paradigm rather than a single algorithm. Its defining move is to recast inversion around token-local structure—patches, subwords, pseudo-tokens, masked spans, or continuous token embeddings—and then to exploit that structure for reconstruction, editing, distillation, or forgetting. The resulting systems differ sharply in objective, threat model, and modality, but they converge on a shared conclusion: token-level signals in modern architectures are often sufficiently informative to support both high-utility constructive procedures and high-fidelity privacy attacks.