Embedded Conditional Independence Tests for Large Language Model Generated Text with an Application to German Parliament Speeches
Abstract: Conditional independence tests (CITs) test for conditional dependence between two random objects and given a third random object . Existing CITs have limited applicability to high-dimensional data, especially multimodal data like text. However, we show that such tests are of interest for LLM outputs, where we test whether an output generated from a source text carries information about an attribute beyond itself. For this purpose, we propose embedded CITs (eCITs), which embed and and apply an existing CIT to the resulting representations and to . We show that, provided the embedding of is sufficient, i.e. retains the information carries about either or the representation of , the null hypothesis transfers from and to their representations, so that a CIT valid for the embedded hypothesis is valid for the original one. We further give conditions for equivalence of the two hypotheses, and show that sufficiency weakens to mean sufficiency when the embedded test targets conditional mean independence. We propose a semi-synthetic simulation design to assess type I error (T1E) control and power of the eCITs for given embedding maps on a specific dataset and task, and use it to evaluate them on our application. Applying the eCITs to German Parliament speeches, we find for all combinations of embedding maps considered that the summaries of two LLMs contain information about the speaker's faction and gender beyond the speech they were generated from.
Paper Prompts
Sign up for free to create and run prompts on this paper.