Papers
Topics
Authors
Recent
Search
2000 character limit reached

Propositional Interpretability in Logics & AI

Updated 4 June 2026
  • Propositional interpretability is the analysis of deducibility, meaning, and relative interpretability in formal systems using modal logics, order-theoretic frameworks, and categorical methods.
  • It integrates mathematical logic with AI techniques, employing empirical methods like sparse auto-encoders and causal tracing to map internal representations to propositional attitudes.
  • Recent developments, such as the slim and broad series axioms, extend interpretability logics to provide rigorous model-theoretic and categorical characterizations.

Propositional interpretability is the structural and semantic analysis of how deducibility, meaning, and relative interpretability between formal systems and agents can be rendered at the level of propositional logics. It encompasses: (1) the mathematical study of logics that capture provability, relative interpretability, and reductions among propositional deductive systems, (2) order-theoretic and categorical frameworks for formalizing mappings between such systems, and (3) contemporary applications in artificial intelligence, where internal representations of systems are interpreted as attitudes toward propositions. The subject integrates modal logic (especially interpretability logics), proof theory, algebraic semantics, and, more recently, empirical methods for mechanistic interpretability in large computational models.

A foundational approach to propositional interpretability uses bimodal logics with modalities for provability (□\Box) and relative interpretability (⊳\rhd, or ▹\triangleright). The language L={□,⊳}L = \{\Box, \rhd\} can formalize both "provable in TT" (□A\Box A) and "T + A interprets T + B" (A⊳BA \rhd B) for an arithmetic theory TT (Goris et al., 2015).

The interpretability logic IL(T)\mathrm{IL}(T) of a theory TT is the set of ⊳\rhd0-formulas valid under all arithmetical realizations in ⊳\rhd1. The "universal" interpretability logic ⊳\rhd2 is the intersection of all ⊳\rhd3 for all sufficiently strong ⊳\rhd4, yielding a minimal set of principles on relative interpretability valid across weak arithmetic bases.

In this framework, interpretability logics notably diverge from provability logics. Where provability logics for theories above minimal strength tend to stabilize (e.g., ⊳\rhd5, the modal logic of provability), interpretability logics can vary and display rich, theory-sensitive hierarchies.

2. Structure of Formal Interpretability Principles

Recent work delineates two infinite series of interpretability principles that strictly extend the known lower bounds for ⊳\rhd6 (Goris et al., 2015). The "slim" series ⊳\rhd7 and "broad" series ⊳\rhd8 are constructed recursively via auxiliary syntactic schemata:

  • ⊳\rhd9 schemata encapsulate the inductive passage between interpretability assertions, utilizing conjunction with provability and negations of interpretability, forming a proper cofinal hierarchy.
  • â–¹\triangleright0 extend â–¹\triangleright1 to incorporate disjunctive modal conditions (â–¹\triangleright2) and deeper recursive structure, forming an incomparable family.

For each series, explicit Kripke/Veltman frame conditions are derived in terms of ternary or quaternary relations, providing model-theoretic characterizations. For example, the slim series yields frame conditions â–¹\triangleright3 involving compositional constraints on the accessibility and interpretability relations in Veltman frames.

These results strictly extend the previously known lower bounds (e.g., â–¹\triangleright4) and approach the conjectural characterization of â–¹\triangleright5. Each family adds infinitely many new validities, sharpening the boundary between provability logic and interpretability logic.

3. Order-Theoretic and Categorical Formalization

The order-theoretic analysis of propositional interpretability models deductive systems as quantale-modules equipped with nuclei (closure operators) (Russo, 2012). In this framework, strong and weak forms of interpretability become algebraic properties:

Interpretation Type Structural requirement Categorical/morphisms
Strong interpretation Action-invariant, soundness under a quantale morphism â–¹\triangleright6 Module homomorphism (restriction of scalars)
Conservative interpretation Strong interpretation with injectivity Injective module map
Equivalence Bi-directional strong, conservative interpretations Module isomorphism
Weak interpretation Soundness only (no action-invariance) Sup-lattice homomorphism

This module-theoretic viewpoint unifies classical translation notions (Glivenko, Kolmogorov, Sheffer-stroke encodings) and enables transfer of syntactic, semantic, and computational properties along categorical functors (restriction/extension of scalars), providing powerful tools for analyzing logic equivalence, embedding, and reducibility.

4. Interpretability in Computational Models and AI

Propositional interpretability has been extended to analysis of AI systems by recasting their internal states as propositional attitudes—beliefs, credences, and desires—toward explicit propositions (Chalmers, 27 Jan 2025). The formal setting defines, for each agent ▹\triangleright7 and time ▹\triangleright8:

  • â–¹\triangleright9: binary belief/judgment function,
  • L={â–¡,⊳}L = \{\Box, \rhd\}0: credence/subjective probability,
  • L={â–¡,⊳}L = \{\Box, \rhd\}1: desire/utility function.

The "thought logging" problem is to construct a trace L={□,⊳}L = \{\Box, \rhd\}2 of these attitudes over a dynamically chosen or salient set of propositions L={□,⊳}L = \{\Box, \rhd\}3. Challenges include the infinity of L={□,⊳}L = \{\Box, \rhd\}4, saliency-based selection, and the detection of causal and mechanistic origins of attitudes (reason or mechanism logging).

Interpretability methods in this context include:

  • Linear and compositional probing (recovering truth values or confidence for L={â–¡,⊳}L = \{\Box, \rhd\}5 from activations);
  • Sparse auto-encoders for discovering monosemantic features mapping to concepts or propositions;
  • Causal tracing and activation patching for localizing information flow;
  • Chain-of-thought prompting for explicating intermediate reasoning steps;
  • Psychosemantic analyses to ground representational content in either information-theoretic or use-based conditions.

Empirical results show that propositional probes can robustly recover latent world state even when output behavior is unfaithful due to prompt injections or backdoor attacks (Feng et al., 2024). This evidences the feasibility of propositional interpretability for monitoring and analyzing complex AI systems.

5. Semantic Uniqueness and Categorical Fixity

In the context of intuitionistic logics, propositional interpretability has connections with categoricity. For intuitionistic propositional logic (IPL), any semantics agreeing with the consequence relation must necessarily interpret the logical connectives in the standard way; this is Carnap-categoricity (2207.14705). For each major semantics—Kripke, Beth, Dragalin, topological, algebraic—it can be demonstrated that any nonstandard interpretation consistent with IPL's deduction rules coincides with the standard one, underscoring the structural rigidity of IPL with respect to propositional interpretability.

6. Open Problems and Future Directions

Major open problems persist regarding the precise boundaries of interpretability logics for weak theories (e.g., is L={□,⊳}L = \{\Box, \rhd\}6 or properly smaller?) (Icard et al., 2020), as well as the possibility of a finite or recursively enumerable axiomatization of L={□,⊳}L = \{\Box, \rhd\}7 (Goris et al., 2015). In computational domains, scaling propositional interpretability and thought-logging to large, open-vocabulary AI systems remains an open research agenda—involving discovery of new unsupervised, compositional probes, psychosemantic metrics, and robust summarization methods (Chalmers, 27 Jan 2025).

Furthermore, the integration of propositional interpretability methods with rigorous philosophical metasemantics and advanced causal tracing continues to develop, with applications to AI alignment, safety, and the monitoring of latent world models (Feng et al., 2024). The extension of the modular, categorical, and psychosemantic formalisms to higher-order, temporal, or context-dependent propositions is another promising direction.

Topic to Video (Beta)

No one has generated a video about this topic yet.

Whiteboard

No one has generated a whiteboard explanation for this topic yet.

Follow Topic

Get notified by email when new papers are published related to Propositional Interpretability.