Determine whether language models use the learned impossibility distinction

Investigate whether the Gemma 3 4B IT model uses the observed representational distinction between necessary falsehood and contingent falsehood when generating answers and explanations.

Background

The study finds that Gemma 3 4B IT’s activation space distinguishes necessary falsehoods from contingent falsehoods along an impossibility direction that is nearly orthogonal to the truth direction. However, the probing results establish only that the distinction is linearly decodable from activations; they do not establish that the model causally uses this information in its judgments or generated responses.

The unresolved issue is therefore whether the observed activation geometry reflects an operational distinction employed by the model during inference, rather than merely a statistically accessible pattern in its representations. The paper also leaves open whether larger models would sharpen or dissolve this distinction.

References

Whether the model uses the observed distinction, and whether larger models sharpen or dissolve it, are questions this study leaves open.

Falsehood and Impossibility Are Different Directions in an AI's Representation of Language  (2608.12852 - Lee, 13 Aug 2026) in Section ‘A remark on sense and truth’, Section 4 of the Discussion