Papers
Topics
Authors
Recent
Search
2000 character limit reached

Embedded Conditional Independence Tests for Large Language Model Generated Text with an Application to German Parliament Speeches

Published 1 Sep 2026 in stat.ML, cs.AI, cs.LG, math.ST, and stat.ME | (2609.00946v1)

Abstract: Conditional independence tests (CITs) test for conditional dependence between two random objects XX and YY given a third random object ZZ. Existing CITs have limited applicability to high-dimensional data, especially multimodal data like text. However, we show that such tests are of interest for LLM outputs, where we test whether an output XX generated from a source text ZZ carries information about an attribute YY beyond ZZ itself. For this purpose, we propose embedded CITs (eCITs), which embed XX and ZZ and apply an existing CIT to the resulting representations and to YY. We show that, provided the embedding of ZZ is sufficient, i.e. retains the information ZZ carries about either YY or the representation of XX, the null hypothesis transfers from XX and ZZ to their representations, so that a CIT valid for the embedded hypothesis is valid for the original one. We further give conditions for equivalence of the two hypotheses, and show that sufficiency weakens to mean sufficiency when the embedded test targets conditional mean independence. We propose a semi-synthetic simulation design to assess type I error (T1E) control and power of the eCITs for given embedding maps on a specific dataset and task, and use it to evaluate them on our application. Applying the eCITs to German Parliament speeches, we find for all combinations of embedding maps considered that the summaries of two LLMs contain information about the speaker's faction and gender beyond the speech they were generated from.

Summary

No one has generated a summary of this paper yet.

Paper to Video (Beta)

No one has generated a video about this paper yet.

Whiteboard

No one has generated a whiteboard explanation for this paper yet.

Continue Learning

We haven't generated follow-up questions for this paper yet.