- The paper demonstrates that demographic identity in Mistral-7B has three dissociable properties—readability, faithful inter-group arrangement, and causal use—using RSA, survey ground truth, and causal interventions.
- Attention heads encode demographic geometry with held-out fidelity as high as ρ = 0.63, while political and socioeconomic attributes are more stable than race-containing categories across prompts and model components.
- The paper finds that faithful representations do not predict behavioral influence: causal effects are small, one low-fidelity pathway survives correction, and probes improve answers without replacing real survey data.
This paper asks where demographic group identity lives inside a LLM, whether its internal geometry mirrors real inter-group opinion structure, and whether the model actually uses what it encodes. Using representational similarity analysis (RSA) against Pew American Trends Panel ground truth over 169 intersectional demographic cells, and causal interventions across all layers of Mistral-7B, it demonstrates that three properties of demographic identity—readability, faithful arrangement, and causal use—are dissociable. The central finding is that representational fidelity does not predict causal use: the clearest causal pathway sits in one of the least faithful attribute types, while the most faithful type shows no correction-surviving single-layer effect.
Motivation and framing
LLMs are increasingly deployed as simulated survey respondents, yet behavioral evidence shows their outputs are excessively uniform: simulated users are overly agreeable, opinion predictions barely differentiate demographic groups, and general capability does not predict simulation fidelity (2608.18768). Prior interpretability work localizes demographic information down to individual attention heads but validates only against the model's own behavior; survey-simulation work possesses real population ground truth but treats the model as a black box. The paper supplies the missing bridge by scoring internal inter-group geometry against real survey opinion structure—the "second-order" complement to per-group probe accuracy—and then asking causally whether that geometry drives answers.
Methodology
Ground truth consists of OpinionQA-style per-group response distributions from 15 Pew ATP waves, organized into 169 intersectional cells across six attribute types (e.g., AGE×POLPARTY, RACE×RELIG). For each pair of cells, the real distance is the mean Wasserstein distance between answer distributions over shared questions. On the model side, representations are extracted at 1,089 locations in Mistral-7B—residual streams after every layer, all 1,024 attention-head outputs before out-projection, and FFN block outputs—under four prompt templates plus a template-averaged read-out. Fidelity at each location is the Spearman correlation between the upper triangles of the two representational dissimilarity matrices within an attribute type.
The statistical apparatus is unusually careful. Because any "best location" claim over 1,089 candidates inflates under selection, the authors apply max-statistic permutation tests, held-out-template selection, and split-half validation; headline numbers are held-out values rather than selected maxima. Causal experiments use cluster-robust inference at the level of 12 donor–recipient pairs per type (exact sign-flip permutation), since items within a pair share persona idiosyncrasies. A lexical-baseline control partials out sentence-encoder similarity of the identity words to rule out the possibility that the map merely echoes word similarity.
The standard read-out understates the model
The last-token residual-stream read-out used implicitly across the LLM-survey literature yields weak, type-uneven fidelity (pooled ≈0.17) under severe anisotropy: cell embeddings share a dominant common direction (norm ratio 0.95). Mean-centering doubles the pooled correlation but decreases every within-type correlation—a between-type scale artifact. Top-5 neighbor precision is roughly twice chance but only 0.35–0.48 absolute, which explains why a consistency-regularization pilot built on this read-out failed. Crucially, weakness of one read-out does not establish weakness of the model, motivating a component-wise search.
The fidelity map
Attention heads dominate the residual stream in five of six types and never lose. Selected head maxima reach +0.74 for AGE×POLPARTY, with selection-corrected held-out fidelity up to ρ = 0.63 (split-half median)—roughly 70% of the measurement-reliability ceiling of 0.84–0.96 estimated from both survey-side and model-side reliability. Random 128-dimensional projections of the residual stream track the residual stream rather than the heads, so the head advantage is not a dimensionality artifact. The lexical confound is real (lexical baseline alone reaches ρ up to 0.64 for EDUCATION×INCOME) but does not explain the map: L11's partial fidelity remains 0.35–0.59 with p < 0.0005 in five of six types.
A single attention head, L11, fixed in advance without per-type re-selection, achieves significant fidelity in all six types (ρ = 0.33–0.67, surviving a conservative ×1024 Bonferroni). The authors candidly note a provenance caveat: L11 was noticed because it recurred in per-type top-10 lists, so the fixed-location test rules out per-type selection but not the initial noticing. Fidelity is sharply heterogeneous across attribute types: political and socio-economic structure is deeply encoded and template-stable, while every race-containing type is weaker and prompt-fragile everywhere tested (e.g., RACE×RELIG ranges from ρ = 0.16 to 0.45 depending on template). This hierarchy parallels output-level findings that ideological variables dominate persona effects.
Causal use does not follow fidelity
Under natural prompting, replacing the entire demographic identity moves predictions by less than 2% of total prediction error (~0.004 Wasserstein against ~0.37 error), and the small residual variation is directionally uninformative—sign tests are approximately coin flips. This near-invariance replicates on Mistral-7B-Instruct-v0.2 under its chat template (3.9% of error), though instruction tuning nearly doubles absolute error while making the small movement reliably directionally correct.
With a strong patching instrument (identity localized to a QA answer token, all 32 heads of layer 11 patched jointly, questions selected for maximal disagreement), a real but small causal effect emerges: for AGE×POLPARTY, 59% of trials move correctly versus 32% under random-head control, recovering 41% of the achievable ceiling (pair-level p = 0.019, marginal and not Holm-surviving). Most strikingly, patching L11 alone in RACE×RELIG shifts predictions in the wrong direction with the most consistent effect in the analysis (pair-level p = 1/4096); donor controls show this backfire is carried by coherent identity content, not generic corruption. Amplifying patches 5× makes effects negative in all types.
An exhaustive sweep patching all 32 heads at every layer for all six types yields the paper's central dissociation. Under cluster-robust inference, no type survives the pair-level max-statistic correction across 32 layers. The clean result lives at the a-priori locus L11: RACE×POLIDEOLOGY—one of the least faithful types (ρ = 0.30)—reaches t = 3.22, exact p = 0.0020, surviving Bonferroni, recovering 42% of its ceiling. Meanwhile the top-fidelity types show no correction-surviving single-layer locus under any fidelity estimator. Across six types, rank correlations between fidelity and causal strength are nowhere significantly positive (+0.26 at L11, −0.77 at best layers). The authors also test whether the dissociation is an artifact of measuring fidelity on identity-only prompts while intervening in QA context: the map does degrade in QA context (e.g., EDUCATION×INCOME falls from 0.67 to 0.10), but the fidelity–causality correlation is unchanged on the QA-context ruler. Three separate times, the layer with the largest raw shift fails correction while a smaller, steadier layer passes—an object lesson in why mean effect sizes without cluster-corrected variance tests mislead.
What direct read-out recovers
A ridge probe (128 PCA components) on layer 11, evaluated leave-one-cell-out, lands 22–30% closer to survey truth than the model's own answers in all six types, better on 65–68% of items; temperature calibration of the letter-probability read-out barely closes the gap. The single 128-dimensional head matches or beats the full-layer read-out in four of six types, and unreduced ridge on all 4,096 dimensions is worse than both. However, a group-blind baseline predicting other cells' mean answers beats every read-out by a wide margin, and a half-split group-ordering test returns a flat null: probes recover per-question group ordering correlations of −0.05 to +0.07 against a reliability ceiling of +0.81 to +0.86. The precise supported claim is therefore narrow: the model holds a group map faithful on average, but that map cannot say who answers what on any particular question. Probing improves on the model's mouth; it does not substitute for survey data.
Cross-model replication
On three checkpoints of Qwen3-8B differing only in human-simulation training (base, midtrained on 10B tokens, post-RL), the map phenomenon replicates: heads dominate in all six types at all checkpoints, the type hierarchy replicates including the race weakness, and a family-specific general-geometry head (L33 H9) scores 0.35–0.58 in five types. Notably, ten billion tokens of behavioral simulation training barely move either the map or the mouth (fidelity changes −0.07 to +0.09; output accuracy changes ≤0.03). The causal side has not been replicated cross-family, so the map–use dissociation remains a one-model result.
Limitations
The paper states its limits plainly. The causal results exist only for Mistral-7B. Output claims scope to the letter-probability interface, not free-text generation. Ground truth is single-institution and US-centric. The dissociation claim rests on six attribute types—enough for two unambiguous counterexamples, not enough to estimate the fidelity–causality relationship precisely—with 12 effective clusters per type bounding causal power. Every patch is single-layer, so redundant encoding across layers remains partially unruled out (only marginally tested via a two-layer patch for one type). L11's discovery provenance leaves open the need for preregistered replication. The lexical control uses a single small encoder. Layer 1's mechanistic role remains an explicitly open question, as the layer-0 read-out is degenerate in the prompt design.
Conclusion
By scoring internal demographic geometry against real survey structure and then intervening causally, this paper decomposes "can LLMs simulate populations?" into three properties that come apart: readable (established previously), faithfully arranged (demonstrated here, concentrated in individual attention heads and replicating across model families), and actually used (real but small, and located independently of fidelity). The practical consequences are concrete: check second-order fidelity before any application relating groups; treat demographic attributes as non-interchangeable, since race-based structure is weak and prompt-fragile; do not steer through a single apparently-faithful unit without corruption controls; and prefer reading internal states over asking the model, while recognizing that neither replaces survey data. The open questions the paper leaves—whether the map–use dissociation holds in a second model family, what layer 1 encodes mechanistically, and whether reading the faithful location transfers where surface fine-tuning does not—define the immediate agenda for this line of work.