Propositional Interpretability in Logics & AI
- Propositional interpretability is the analysis of deducibility, meaning, and relative interpretability in formal systems using modal logics, order-theoretic frameworks, and categorical methods.
- It integrates mathematical logic with AI techniques, employing empirical methods like sparse auto-encoders and causal tracing to map internal representations to propositional attitudes.
- Recent developments, such as the slim and broad series axioms, extend interpretability logics to provide rigorous model-theoretic and categorical characterizations.
Propositional interpretability is the structural and semantic analysis of how deducibility, meaning, and relative interpretability between formal systems and agents can be rendered at the level of propositional logics. It encompasses: (1) the mathematical study of logics that capture provability, relative interpretability, and reductions among propositional deductive systems, (2) order-theoretic and categorical frameworks for formalizing mappings between such systems, and (3) contemporary applications in artificial intelligence, where internal representations of systems are interpreted as attitudes toward propositions. The subject integrates modal logic (especially interpretability logics), proof theory, algebraic semantics, and, more recently, empirical methods for mechanistic interpretability in large computational models.
1. Modal Logics of Propositional Interpretability
A foundational approach to propositional interpretability uses bimodal logics with modalities for provability () and relative interpretability (, or ). The language can formalize both "provable in " () and "T + A interprets T + B" () for an arithmetic theory (Goris et al., 2015).
The interpretability logic of a theory is the set of 0-formulas valid under all arithmetical realizations in 1. The "universal" interpretability logic 2 is the intersection of all 3 for all sufficiently strong 4, yielding a minimal set of principles on relative interpretability valid across weak arithmetic bases.
In this framework, interpretability logics notably diverge from provability logics. Where provability logics for theories above minimal strength tend to stabilize (e.g., 5, the modal logic of provability), interpretability logics can vary and display rich, theory-sensitive hierarchies.
2. Structure of Formal Interpretability Principles
Recent work delineates two infinite series of interpretability principles that strictly extend the known lower bounds for 6 (Goris et al., 2015). The "slim" series 7 and "broad" series 8 are constructed recursively via auxiliary syntactic schemata:
- 9 schemata encapsulate the inductive passage between interpretability assertions, utilizing conjunction with provability and negations of interpretability, forming a proper cofinal hierarchy.
- 0 extend 1 to incorporate disjunctive modal conditions (2) and deeper recursive structure, forming an incomparable family.
For each series, explicit Kripke/Veltman frame conditions are derived in terms of ternary or quaternary relations, providing model-theoretic characterizations. For example, the slim series yields frame conditions 3 involving compositional constraints on the accessibility and interpretability relations in Veltman frames.
These results strictly extend the previously known lower bounds (e.g., 4) and approach the conjectural characterization of 5. Each family adds infinitely many new validities, sharpening the boundary between provability logic and interpretability logic.
3. Order-Theoretic and Categorical Formalization
The order-theoretic analysis of propositional interpretability models deductive systems as quantale-modules equipped with nuclei (closure operators) (Russo, 2012). In this framework, strong and weak forms of interpretability become algebraic properties:
| Interpretation Type | Structural requirement | Categorical/morphisms |
|---|---|---|
| Strong interpretation | Action-invariant, soundness under a quantale morphism 6 | Module homomorphism (restriction of scalars) |
| Conservative interpretation | Strong interpretation with injectivity | Injective module map |
| Equivalence | Bi-directional strong, conservative interpretations | Module isomorphism |
| Weak interpretation | Soundness only (no action-invariance) | Sup-lattice homomorphism |
This module-theoretic viewpoint unifies classical translation notions (Glivenko, Kolmogorov, Sheffer-stroke encodings) and enables transfer of syntactic, semantic, and computational properties along categorical functors (restriction/extension of scalars), providing powerful tools for analyzing logic equivalence, embedding, and reducibility.
4. Interpretability in Computational Models and AI
Propositional interpretability has been extended to analysis of AI systems by recasting their internal states as propositional attitudes—beliefs, credences, and desires—toward explicit propositions (Chalmers, 27 Jan 2025). The formal setting defines, for each agent 7 and time 8:
- 9: binary belief/judgment function,
- 0: credence/subjective probability,
- 1: desire/utility function.
The "thought logging" problem is to construct a trace 2 of these attitudes over a dynamically chosen or salient set of propositions 3. Challenges include the infinity of 4, saliency-based selection, and the detection of causal and mechanistic origins of attitudes (reason or mechanism logging).
Interpretability methods in this context include:
- Linear and compositional probing (recovering truth values or confidence for 5 from activations);
- Sparse auto-encoders for discovering monosemantic features mapping to concepts or propositions;
- Causal tracing and activation patching for localizing information flow;
- Chain-of-thought prompting for explicating intermediate reasoning steps;
- Psychosemantic analyses to ground representational content in either information-theoretic or use-based conditions.
Empirical results show that propositional probes can robustly recover latent world state even when output behavior is unfaithful due to prompt injections or backdoor attacks (Feng et al., 2024). This evidences the feasibility of propositional interpretability for monitoring and analyzing complex AI systems.
5. Semantic Uniqueness and Categorical Fixity
In the context of intuitionistic logics, propositional interpretability has connections with categoricity. For intuitionistic propositional logic (IPL), any semantics agreeing with the consequence relation must necessarily interpret the logical connectives in the standard way; this is Carnap-categoricity (2207.14705). For each major semantics—Kripke, Beth, Dragalin, topological, algebraic—it can be demonstrated that any nonstandard interpretation consistent with IPL's deduction rules coincides with the standard one, underscoring the structural rigidity of IPL with respect to propositional interpretability.
6. Open Problems and Future Directions
Major open problems persist regarding the precise boundaries of interpretability logics for weak theories (e.g., is 6 or properly smaller?) (Icard et al., 2020), as well as the possibility of a finite or recursively enumerable axiomatization of 7 (Goris et al., 2015). In computational domains, scaling propositional interpretability and thought-logging to large, open-vocabulary AI systems remains an open research agenda—involving discovery of new unsupervised, compositional probes, psychosemantic metrics, and robust summarization methods (Chalmers, 27 Jan 2025).
Furthermore, the integration of propositional interpretability methods with rigorous philosophical metasemantics and advanced causal tracing continues to develop, with applications to AI alignment, safety, and the monitoring of latent world models (Feng et al., 2024). The extension of the modular, categorical, and psychosemantic formalisms to higher-order, temporal, or context-dependent propositions is another promising direction.