Membership Leakage via Tokenizers
- The paper demonstrates that released tokenizer artifacts, such as vocabulary and merge rules, can reveal dataset membership even without access to model weights.
- Token-level leakage techniques identify sensitive tokens via prediction probabilities, exposing sparse yet significant privacy signals over specific token positions.
- Boundary instability caused by incomplete tokens highlights behavioral vulnerabilities that affect both data compression efficacy and multilingual tokenization.
Membership leakage through tokenizers denotes privacy leakage that is mediated either by tokenizer artifacts themselves or by the tokenized prediction units on which LLMs are trained and queried. In the strongest direct form, released tokenizer artifacts such as vocabulary, merge rules, and merge order can support dataset-level membership inference without access to model weights or logits. In broader token-level formulations, membership signals become more visible after text has been segmented into tokens, because leakage is often concentrated in a small subset of token positions rather than in whole-sequence averages. A related, more indirect line of evidence shows that byte-level tokenization artifacts can create unstable boundary-sensitive representations, suggesting that tokenization is not a neutral preprocessing step but a privacy-relevant component of the modeling pipeline (Tong et al., 7 Oct 2025, Jawad et al., 27 Jan 2026, Tao et al., 7 Oct 2025, Jang et al., 2024).
1. Conceptual scope and levels of analysis
The literature distinguishes three analytically different, but connected, notions of leakage through tokenizers. The first is tokenizer-level leakage, where the tokenizer itself is the attack surface. The second is token-level leakage on tokenized sequences, where membership is inferred from next-token probabilities or tokenwise scores after text has already been segmented. The third is boundary-level instability, where unusual token boundaries expose brittle model behavior that may create or confound membership signals.
| Level | Mechanism | Representative evidence |
|---|---|---|
| Tokenizer artifact leakage | Vocabulary, merge rules, and merge order encode corpus-specific provenance | "Membership Inference Attacks on Tokenizers of LLMs" (Tong et al., 7 Oct 2025) |
| Token-level membership leakage | Leakage is concentrated in token probabilities or tokenwise scores rather than sequence averages | "What Hard Tokens Reveal" (Jawad et al., 27 Jan 2026); "(Token-Level) InfoRMIA" (Tao et al., 7 Oct 2025) |
| Boundary-level instability | Incomplete tokens and improbable bigrams induce unstable behavior on legal but unusual tokenizations | "Improbable Bigrams Expose Vulnerabilities of Incomplete Tokens in Byte-Level Tokenizers" (Jang et al., 2024) |
This tripartite view is important because “membership leakage through tokenizers” is not a single experimental paradigm. In one setting, the attacker asks whether a dataset was used to train the tokenizer. In another, the attacker asks whether a sequence was used to fine-tune a LLM, but discovers that the answer is most visible at specific token positions. In a third, tokenization artifacts do not themselves establish membership inference, yet they change model behavior strongly enough that they become plausible probes or confounders for membership analysis.
2. Tokenizers as a direct attack surface
The clearest direct evidence comes from dataset-level membership inference on released BPE tokenizers. In this setting, the attacker is given a target dataset and asks whether was part of the corpus used to train the target tokenizer. The formal threat model is
where $1$ means . The motivating claim is that tokenizers are much cheaper to retrain from scratch than full LLMs, their training data is typically representative of LLM pretraining data, and providers often release tokenizer artifacts openly for token counting and billing transparency (Tong et al., 7 Oct 2025).
For BPE tokenizers, the leakage mechanism is corpus compression. Training starts from primitive symbols and repeatedly merges the most frequent adjacent pair in the corpus. The final vocabulary, merge ordering, and segmentation behavior therefore reflect corpus-specific compression preferences. If a dataset introduces substrings that are frequent enough within that dataset and sufficiently rare elsewhere, BPE may allocate dedicated tokens to them. Those tokens become provenance clues. The paper identifies several usable signals: merge order or rank correlation, vocabulary overlap of distinctive tokens, relative token frequency within , self-information of a token, and compression rate.
Five attacks are evaluated. MIA via Merge Similarity compares Spearman rank correlation between target merge order and IN versus OUT shadow tokenizers, but is essentially ineffective, with AUC around 0.49–0.50 across vocabulary sizes. MIA via Vocabulary Overlap is the strongest method and relies on distinctive tokens that appear in tokenizers trained with but not without it. MIA via Frequency Estimation uses the paper’s Relative Token Frequency with Self-information: with
Two lighter baselines, Probability Estimation and Compression Rate, are weaker (Tong et al., 7 Oct 2025).
The empirical setup uses real web data from C4: 1,681,296 web pages from 4,133 websites, with each website treated as one dataset . Half of the websites are randomly selected as member datasets for tokenizer training. The blind baseline recommended to test member/non-member separability without seeing the tokenizer reaches AUC 0.513, and the t-SNE visualization shows no clear separation, so the attack signal is not attributable to obvious distribution shift. Across vocabulary sizes from 80,000 to 200,000, Vocabulary Overlap improves from AUC 0.693, BA 0.666, TPR@1%FPR 26.77% to AUC 0.771, BA 0.711, TPR@1%FPR 34.61%. Frequency Estimation improves from AUC 0.610, BA 0.614, TPR 21.30% to AUC 0.740, BA 0.705, TPR 27.88% (Tong et al., 7 Oct 2025).
Leakage scales with both vocabulary size and target dataset size. For datasets with 800–1200 samples, Vocabulary Overlap reaches AUC 0.882 and TPR@1%FPR 62.50% at 200k, while Frequency Estimation reaches AUC 0.843 and TPR@1%FPR 53.75% at 200k. The paper frames this as a privacy downside of tokenizer scaling: larger vocabularies improve compression efficiency but give BPE more capacity to memorize distinctive strings. The reported power-law fit for token frequency by merge rank,
0
has estimated 1 values 1.717, 1.604, 1.537, 1.493, and 1.460 for vocabulary sizes 80k, 110k, 140k, 170k, and 200k respectively, with standard errors from 0.003 down to 0.001 (Tong et al., 7 Oct 2025).
A central implication follows directly: releasing only the tokenizer does not eliminate privacy risk. Vocabulary, merge rules, and merge order are not merely metadata; they can encode dataset-level training membership.
3. Token-level membership inference on tokenized sequences
A second line of work shifts attention from the tokenizer artifact to the tokenized prediction process. The core criticism is that sequence-level MIAs aggregate over all token positions and thereby mix generalization with memorization. In an autoregressive model,
2
so the operational unit of training is the next-token prediction event, not the whole string. This motivates attacks and analyses at token granularity rather than sequence granularity (Jawad et al., 27 Jan 2026, Tao et al., 7 Oct 2025).
HT-MIA formalizes this by selecting hard tokens, defined as the positions with the lowest target-model confidence under teacher forcing. For each position 3,
4
and the hard-token index set consists of the 5 smallest 6, where
7
The attack computes 8 and aggregates
9
The result is a count of positive token-level improvements on selected hard tokens, rather than a whole-sequence loss or ratio. The paper argues that generalization improves many token predictions for both members and non-members, whereas memorization yields unusually strong gains on difficult tokens for seen examples (Jawad et al., 27 Jan 2026).
The experiments use GPT-2, Qwen-3-0.6B, and LLaMA-3.2-1B fine-tuned on Asclepius, Clinicalnotes, Wikipedia, and IMDB, with all datasets split 50:50 into members and non-members. On Asclepius, HT-MIA achieves AUCs of 0.7348, 0.7312, and 0.5679 for LLaMA-3.2-1B, Qwen-3-0.6B, and GPT-2, outperforming all seven baselines; at TPR@FPR=0.01 it reaches 0.0817, 0.0477, and 0.0181. On Clinicalnotes, the AUCs are 0.8659, 0.8843, and 0.6978, with TPR@FPR=0.01 of 0.4253, 0.3839, and 0.0524. On Wikipedia, HT-MIA reaches AUC 86.2% on LLaMA and 91.0% on Qwen, beating Ratio by 4.7% and 7.3%. On IMDB, however, Ratio reaches 95.9% and 97.4%, while HT-MIA obtains 92.3% and 94.7%, indicating that token-level hard-token analysis is not universally superior in every distribution (Jawad et al., 27 Jan 2026).
InfoRMIA makes a related, but distinct, argument. It reformulates RMIA in information-theoretic terms: 0 and then applies the same form at token level, using vocabulary alternatives as the null population. For a sequence 1, the model yields token-level scores 2, which can then be averaged or aggregated by min-3. The paper emphasizes that token-level InfoRMIA localizes membership signals to exact token positions and can identify which tokens are memorized within generated outputs. It also argues that sequence-level attacks are lossy compressions because private information is often sparse within a sentence (Tao et al., 7 Oct 2025).
The empirical pattern is not that token-level AUC is always larger than sequence-level AUC, but that token-level analysis exposes leakage structure that sequence averages can hide. On AG News with GPT-2 fine-tuning, sequence InfoRMIA reaches AUC 0.843 and TPR@1%FPR 23.0% after 1 epoch, while token InfoRMIA reaches AUC 0.836 and TPR@1%FPR 20.2%; after 4 epochs, sequence InfoRMIA reaches AUC 0.945 and TPR 16.2%, while token InfoRMIA reaches AUC 0.942 and TPR@1%FPR 20.6%. On ai4privacy, token-level scores are also used for localized privacy analysis. The paper reports low correlation between average sequence-level membership scores and average private-token scores, and notes that some top-ranked sequence-level “memorized” examples contain no private tokens at all, whereas private-token signals can be strong but sparse (Tao et al., 7 Oct 2025).
For the tokenizer question, both papers support a narrow but important conclusion: membership leakage is exposed through token-level probabilities on tokenized sequences, and the observed leakage depends on the token units defined by the tokenizer. Neither paper directly proves that one tokenizer family is inherently leakier than another.
4. Boundary artifacts, incomplete tokens, and indirect evidence
A third line of work is not a membership-inference study, yet it is highly relevant because it shows that token boundary artifacts can produce large behavioral discontinuities. The relevant objects are incomplete tokens, defined as “undecodable tokens with stray bytes resulting from byte-level BPE tokenization.” Byte-level BPE can merge byte subsequences without respecting Unicode character boundaries, so multilingual UTF-8 characters may be split across tokens. The result is a token whose meaning is not internally self-contained, because adjacent tokens must supply the missing bytes needed to complete a valid character (Jang et al., 2024).
The paper decomposes incomplete tokens into prefixes, which end with stray bytes and require bytes from the following token, and suffixes, which begin with stray bytes and rely on bytes from the preceding token. It then constructs improbable bigrams: legal pairs of incomplete tokens whose stray bytes complete each other into a valid character boundary, but whose combination is highly unlikely in training. Multilinguality is used as a heuristic for improbability, especially when the composed bytes yield a phrase mixing scripts from different languages. A crucial control is the decode-encode check, which retains only pairs that decode to a valid string and then re-encode back to the intended prefix/suffix tokenization. The attack therefore depends on preserving the exact fragile token boundary (Jang et al., 2024).
The experiments cover five instruction-tuned models with byte-level BPE tokenizers: Meta-Llama-3.1-8B-Instruct, EXAONE-3.0-7.8B-Instruct, Qwen2.5-32B-Instruct, Mistral-Nemo-Instruct-2407, and C4AI-Command-R-v01. Reported tokenizer statistics are 128k vocab, 1224 incomplete tokens, 71k incomplete bigrams for Llama-3.1; 102k, 1222, 36k for Exaone-3.0; 151k, 1320, 39k for Qwen2.5; 131k, 1307, 135k for Mistral-Nemo; and 255k, 2956, 1479k for Command-R-v01. The evaluation uses three prompt templates to elicit verbatim repetition of a target phrase, and a phrase is counted as hallucinatory only if the model fails to correctly repeat it in all three prompts. The construction also controls for simple undertraining by selecting only incomplete tokens in the upper half of an embedding-based degree-of-training ranking (Jang et al., 2024).
The reported effects are large. Hallucinatory rates for improbable bigrams versus baseline bigrams are 43/100 versus 0/100 for Llama 3.1, 79/100 versus 26/100 for Exaone, 38/100 versus 2/100 for Qwen2.5, 52/71 versus 0/71 for Mistral-Nemo, and 56/100 versus 9/100 for Command-R. When the same decoded phrase is re-expressed through alternative tokenization by pre-segmenting at character boundaries, hallucination frequencies change from 0.43 to 0.03 for Llama 3.1, from 0.79 to 0.52 for Exaone, from 0.38 to 0.14 for Qwen2.5, from 0.73 to 0.01 for Mistral-Nemo, and from 0.56 to 0.57 for Command-R. The Llama 3.1 reduction is 93%, and the Mistral-Nemo reduction is 98% (Jang et al., 2024).
This does not prove membership leakage. The paper does not show higher likelihood for seen examples than unseen examples, confidence gaps, per-example loss differences, canary extraction rates, or attack AUC against a member/non-member benchmark. Its contribution is different: it shows that surface-identical strings can produce very different behavior solely because of tokenization. This suggests that tokenizer-specific decomposition can act both as a probe for memorization-sensitive behavior and as a confounder in membership evaluation.
5. Attack regimes, observables, and common misconceptions
The papers operate under different threat models, and conflating them leads to several common misconceptions.
| Setting | Required access | Observable signal |
|---|---|---|
| Tokenizer-only dataset MIA | Public vocabulary, merge order, auxiliary datasets, often shadow tokenizers | Distinctive-token overlap, merge order, self-information, compression (Tong et al., 7 Oct 2025) |
| Token-level LLM MIA | Teacher-forced token probabilities or logits; reference model or reference distribution | Hard-token improvements or token-level InfoRMIA scores (Jawad et al., 27 Jan 2026, Tao et al., 7 Oct 2025) |
| Boundary artifact probing | Black-box generation over instruction-tuned models | Copy failure, hallucination, instability across tokenizations (Jang et al., 2024) |
One misconception is that all token-level analyses are studies of tokenizer design. This is false in the present literature. HT-MIA does not compare BPE versus SentencePiece versus WordPiece; InfoRMIA does not vary segmentation algorithms; both papers instead show that once text has been tokenized, membership signals are often localized at token positions. The evidence therefore concerns leakage through tokenized structure, not direct tokenizer-family comparisons (Jawad et al., 27 Jan 2026, Tao et al., 7 Oct 2025).
A second misconception is that any tokenizer-induced behavioral instability is already membership inference. The incomplete-token work demonstrates hallucination vulnerabilities, not canonical MI. Its relevance is inferential: if token-boundary artifacts create unusually brittle representations, then training exposure to exact boundary patterns could matter more than surface text alone. A plausible implication is that rare or out-of-distribution token boundaries may produce distinguishable responses between in-training and out-of-training forms of otherwise similar text, but that implication remains speculative until tested on explicit member/non-member benchmarks (Jang et al., 2024).
A third misconception is that sequence-level metrics are a sufficient gold standard for privacy in LLMs. Both HT-MIA and token-level InfoRMIA argue that whole-sequence scores can dilute privacy signals because many easy token predictions improve for both members and non-members. The token-level perspective is therefore not merely a diagnostic refinement; it changes what is visible to the auditor. At the same time, the evidence does not support a universal ranking in which token-level methods always dominate. IMDB in HT-MIA is a clear counterexample, where Ratio already captures much of the signal (Jawad et al., 27 Jan 2026, Tao et al., 7 Oct 2025).
A fourth misconception is that tokenizers are harmless metadata. The tokenizer-only attacks directly contradict that assumption. The released vocabulary and merge artifacts can leak whether a website-scale dataset was included in training, with stronger leakage as vocabulary size and target dataset size grow (Tong et al., 7 Oct 2025).
6. Mitigations, tradeoffs, and open problems
Existing mitigations reflect the level at which leakage appears. For incomplete-token boundary artifacts, the main mitigation is alternative tokenization via pre-segmentation, specifically avoiding tokenization that crosses character boundaries. More generally, the paper motivates tokenization constrained to Unicode character boundaries, subword tokenizers that are not byte-greedy across character boundaries, preprocessing that normalizes or pre-segments multilingual text, and vocabulary construction that penalizes merges creating undecodable tokens. The Command-R result, where alternative tokenization changes 0.56 to 0.57 and therefore shows no improvement, indicates that mitigation is model- and tokenizer-dependent (Jang et al., 2024).
For tokenizer-only membership leakage, the paper proposes the min count mechanism: remove token 4 from 5 if its aggregated frequency over member datasets is below 6, with 7. This weakens the rare, distinctive-token signals used by Vocabulary Overlap and Frequency Estimation. At 200k vocabulary size, Vocabulary Overlap changes from AUC 0.771, TPR 34.61% without defense to AUC 0.717, TPR 26.48% at 8. Frequency Estimation changes from AUC 0.740, TPR 27.88% to AUC 0.668, TPR 23.72%. The privacy gain comes with reduced compression utility: at 200k, WikiText changes from 5.025 to 5.000 bytes per token, GitHub from 4.009 to 3.965, MGSM from 3.740 to 3.716, and GPQA from 4.117 to 4.093 (Tong et al., 7 Oct 2025).
For token-level LLM leakage, DP-SGD is studied as a defense during fine-tuning. The reported result is a substantial collapse in HT-MIA performance: on Asclepius, AUC drops from 0.5679 to 0.5008 on GPT-2 and from 0.7312 to 0.4957 on Qwen-3-0.6B. The privacy gain has a utility cost on downstream medical QA benchmarks; on PubMedQA, Qwen drops from 0.5580 to 0.4620 and GPT-2 from 0.4800 to 0.3240. The tokenizer paper discusses differential privacy as a plausible future defense for tokenizer training, but states that there is no established DP tokenizer-training mechanism yet (Jawad et al., 27 Jan 2026, Tong et al., 7 Oct 2025).
Token-level localization also motivates more targeted interventions. InfoRMIA explicitly frames token-level membership scores as a substrate for exact unlearning, targeted redaction of sensitive tokens, privacy auditing at known PII locations, and token-guided data reconstruction and extraction. This suggests that future privacy work may need to separate at least three questions: whether a dataset was used to train the tokenizer, whether a sequence was used to train or fine-tune the model, and which exact token positions within that sequence carry the leakage (Tao et al., 7 Oct 2025).
The main open problem is therefore not simply “whether tokenizers leak,” which is already established at dataset level, but how tokenizer design, token granularity, vocabulary growth, and boundary artifacts jointly shape privacy. The current evidence shows that tokenization can be a direct attack vector, a measurement lens that reveals token-localized membership signals, and a source of behavioral instability that can enable or contaminate membership inference.