- The paper introduces semantic transfer entropy and semantic partial information decomposition, using LLM probabilities and attention masking to quantify directed influence and multi-source information structure.
- The paper validates the framework across four experiments and three language models, finding strong directional asymmetry in persuasion, lower therapist-to-client transfer in higher-quality motivational interviewing, and positive premise synergy in arguments.
- The paper shows that methodological choices are critical: masking and physical omission can produce radically different estimates, while model calibration, hidden confounders, and Granger-style assumptions limit causal interpretation.
Overview
This paper introduces an information-theoretic framework for quantifying how semantic content flows between interlocutors, implemented entirely through the conditional probabilities of LLMs. The authors—Goodall, Luppi, and Mediano—propose two measures. Semantic transfer entropy (STE) quantifies directed predictive influence from a source speaker to a target speaker by comparing an LLM's log-likelihood of the target's next utterance with and without access to the source's prior content, operationalised via attention masking in a pretrained autoregressive transformer (2606.30096). Semantic partial information decomposition (SPID) extends this to multi-source settings, decomposing the information multiple sources jointly provide about a target into redundant, unique, and synergistic atoms using the minimum mutual information (MMI) redundancy functional, with marginalisation achieved by physically omitting sources from the input sequence.
The framework is validated across four experiments spanning synthetic dialogue with manipulated cognitive styles, naturalistic persuasion, motivational interviewing (MI) psychotherapy sessions, and argumentative essays. All key effects replicate in direction and significance across three LLMs spanning two architectures and two scales (LLaMA 3.2-3B, Phi-3-mini-4k, Mistral-7B-v0.3). The implementation is released as PSIDyn, an open-source Python package.
Methodological foundations
The core estimator exploits the fact that autoregressive transformers provide direct access to conditional text probabilities. For STE, the quantity of interest is
STEY→X=E[logp(X∣X−,Y−,Z−)−logp(X∣X−,Z−)],
where X is the target, Y the source, and Z− all remaining conversational context. The counterfactual conditioning is realised by constructing a modified attention mask that severs edges from the target's tokens to the source's lagged tokens while leaving the rest of the sequence untouched. This preserves positional encodings and formatting—an important design choice, since physical omission of a participant in multi-turn dialogue would cascade representational changes through downstream messages that themselves attended to the omitted speaker.
For SPID, the decomposition follows Williams and Beer's lattice, with redundancy defined as min(I(Y;X1),I(Y;X2)) under MMI. A source-independence constraint blocks inter-source attention so that neither source's representations contaminate the other, ensuring symmetry under permutation. For more than two sources, where full PID becomes unwieldy (4, 18, and 166 atoms for n=2,3,4), the paper employs the redundancy-synergy index (RSI), the whole-minus-sum scalar summary. An alternative redundancy functional based on common change in surprisal (CCS) is also implemented; MMI and CCS redundancy estimates correlate at r=.90, though CCS yields lower synergy estimates (M=0.15 vs. $0.32$).
A unit test on small transformers trained on canonical logic gates confirms correct recovery of ground truths: XOR produces maximal synergy, COPY maximal redundancy, and UNIQ concentrates information in one source.
Experiment 1: Cognitive rigidity as a controlled sensitivity test
Using 3,000 synthetic dyadic conversations generated by GPT-5-nano, with cognitive style (rigid vs. flexible) experimentally manipulated and orthogonalised against speaker position, STE proved sensitive to source cognitive style. Flexible sources generated higher STE than rigid sources (M=0.52 vs. X0; main effect X1, X2), while target type showed no independent effect (X3). Critically, within mixed dyads, flexible-to-rigid transfer exceeded rigid-to-flexible transfer within the same conversations (X4, X5), robust to controlling for turn order.
The authors are explicit that this is a discriminative-capacity demonstration, not evidence about cognitive flexibility as a psychological construct—the conversations were LLM-generated rather than produced by humans with verified cognitive profiles. Notably, the target-type effect observed with LLaMA did not replicate under Phi-3 or Mistral, which the authors attribute to model-specific calibration; the source effect is the replicable finding.
Experiment 2: Directional asymmetry in persuasion
On PersuasionForGood (300 dyadic donation-solicitation dialogues, 10,785 exchanges), STE revealed a stark asymmetry: persuader-to-persuadee transfer averaged 1.70 bits/token versus 0.05 bits/token in the reverse direction, with 98% of conversations showing this pattern. This asymmetry was not reducible to predictability—the surprisal baseline showed no relationship with donation outcomes—and it replicated across models with large effect sizes (X6–X7).
Strategy analysis yielded significant negative associations between several persuasion strategies and persuader-to-persuadee STE: personal stories (X8), credibility appeals (X9), foot-in-the-door (Y0), and donation information (Y1). The interpretation offered is that such strategies introduce self-contained content about the persuader that contributes little incremental signal for predicting the persuadee's next turn—not that they are persuasive failures. Indeed, STE itself did not predict donation success in logistic regression (Y2, Y3); only donation-information count did (Y4, Y5). The authors flag this strategy-level interpretation as exploratory.
The most clinically consequential result concerns motivational interviewing. Drawing on the established causal chain from therapist verbal style to client change talk to outcomes, the authors hypothesised that high-quality MI would show lower therapist-to-client STE, since effective therapists follow rather than direct the client. The data confirmed this: therapist-to-client STE was substantially higher in low-quality sessions (Y6) than high-quality ones (Y7; Welch's Y8, Y9, Z−0), and therapist-to-client STE correlated negatively with client change talk proportion (Z−1, Z−2).
The implication is notable: STE indexed therapeutic quality without access to any coded behavioural categories, purely from the directional structure of semantic influence. The direction of the effect replicated across all three models, though the primary contrast was marginal under Phi-3 (Z−3) with a smaller session sample of low-quality sessions (Z−4) limiting power.
Experiment 4: Synergistic premise contributions in argumentation
Applying SPID to 325 two-premise claim triplets from the Argument Annotated Essays corpus, the authors found positive average synergy (Z−5 bits/token; Z−6, Z−7, Z−8): premises jointly predicted claims better than the sum of their individual contributions. Redundancy was moderate (Z−9) and unique information modest and symmetric (~0.20 per premise). A permutation null confirmed construct validity: real pairings showed far higher redundancy than shuffled pairs (min(I(Y;X1),I(Y;X2))0 vs. min(I(Y;X1),I(Y;X2))1; min(I(Y;X1),I(Y;X2))2), and SPID atoms correlated sensibly with embedding-based semantic similarity (unique information with premise-claim similarity at min(I(Y;X1),I(Y;X2))3).
Extending to three and four premises via RSI revealed a monotonic shift toward redundancy dominance (min(I(Y;X1),I(Y;X2))4, min(I(Y;X1),I(Y;X2))5, min(I(Y;X1),I(Y;X2))6 for min(I(Y;X1),I(Y;X2))7; Kruskal–Wallis min(I(Y;X1),I(Y;X2))8, min(I(Y;X1),I(Y;X2))9), alongside declining per-premise information yield (n=2,3,40, n=2,3,41, n=2,3,42). Aggregating additional premises about a fixed claim thus produces increasingly overlapping rather than complementary evidence—a quantitative characterisation of diminishing rhetorical returns.
One subtlety deserves emphasis: shuffled premise pairs actually showed higher raw synergy than real pairs (n=2,3,43 vs. n=2,3,44). The authors explain this as an artefact of the MMI partition—when both marginals collapse toward zero in the null, residual generic joint structure is absorbed into synergy by construction—rather than weaker semantic synergy in real arguments.
Masking versus omission: a consequential methodological finding
The supplementary analysis comparing attention masking with physical omission on Experiment 4 carries practical weight beyond this corpus. Because masked-but-present source tokens shift the target's positional encodings, masking inflated each mutual information term by roughly 5.5–5.9 bits/token. Although much of this offset cancels in the synergy computation, the residual distortion reversed the sign of synergy in 59% of samples, shifting mean synergy from n=2,3,45 to n=2,3,46. Had masking been used, the central qualitative finding of positive synergy would have been inverted. The authors argue the two methods are structurally appropriate to different settings—masking for multi-turn dialogue (where omission cascades), omission for independent texts—and demonstrate empirically that they are not interchangeable.
Limitations and open questions
The paper concedes several constraints candidly. Absolute information values are conditioned on the estimation model and are not comparable across models with different tokenisation or calibration; temperature scaling can alter surprisal estimates substantially, and miscalibration does not resolve with scale. Cross-model replication mitigates but does not eliminate this concern, and the authors recommend holding the model constant within studies.
Causal interpretation of STE inherits the assumptions of Granger causality: stationarity (which evolving natural dialogues may violate), absence of hidden confounders (shared topic, context, speaker background), and correct model specification. Where these are uncertain, STE should be read as directed predictive coupling rather than causal influence. Pointwise estimates can be negative even when expectations are positive, cautioning against small-sample applications. Computational cost scales with the number of conditioning sets, limiting very large corpora or real-time use without optimisation.
Experiment 1 remains the weakest evidential link: its stimulus material was LLM-generated, so surface regularities correlated with condition labels could drive the result independently of genuine cognitive flexibility. Validation against human speakers with independently assessed cognitive profiles is left open. Finally, the theoretical status of the "semantic" qualifier is clarified honestly: STE and SPID are formally structural quantities in Dretske's sense, measuring uncertainty reduction in bits; their semantic character derives from the LLM's context-sensitive probability landscape, not from a formal theory of meaning.
Conclusion
This paper contributes a principled, unified information-theoretic framework for quantifying directed semantic influence (STE) and multi-source contribution structure (SPID) in linguistic exchange, using LLMs as probabilistic estimators. Four experiments with distinct evidential roles—a controlled sensitivity test, two naturalistic validations against theoretically grounded constructs (persuasion role structure, MI quality), and a corpus-based validation against annotated argument structure—demonstrate that the measures capture phenomena inaccessible to lexical-feature or correlational approaches, with findings robust across architectures and scales. The masking-versus-omission analysis adds a methodological caution of general relevance to anyone computing counterfactual information quantities with transformers. The framework's portability to clinical, political, educational, and human-AI discourse is asserted rather than exhaustively demonstrated; whether STE generalises beyond the specific corpora tested, and whether its Granger-style assumptions hold in longer, messier interactions, remain open empirical questions.