Papers
Topics
Authors
Recent
Search
2000 character limit reached

A Type Theory of Sense: Witnessed Choice in Stratified Semantic Spaces

Published 10 Jun 2026 in cs.LO | (2606.12504v1)

Abstract: We introduce TTS, a dependent type theory in which semantic composition is represented by horn filling and distinctions between possible completions are witnessed relative to explicit measurement regimes. TTS replaces globally canonical composition with regime-indexed indiscernibility and constructive apartness, allowing filler spaces to be classified as canonical when all completions are observationally connected and forked when two warranted completions are positively separated. Separation witnesses enter the calculus only through measurement contexts recording actual instrument outputs, yielding conservativity, provenance, and a no-fork-from-the-empty-record result. We prove that forks persist under refinement while canonicity may fail, and characterize exactly when an identification made by one regime can consistently coexist with a separation made by another. This framework supports a geometric account of Fregean sense as a choice of filler, reference as the boundary constraining that choice, and hyperintensional difference as measured apartness, while providing a falsifiable bridge to stratified representation spaces and branching behaviour in language-model generation.

Authors (1)

Summary

  • The paper introduces TTS, a dependent type theory that formalizes semantic composition as witnessed, instrument-relative choices in stratified spaces.
  • It demonstrates the inadequacy of classical type theories by formalizing semantic forks and singularities observed in LLM token spaces through empirical and theoretical analysis.
  • The theory provides constructive proofs for non-uniform identity, offering practical insights for improving LLM interpretability and evaluation of alignment protocols.

Type-Theoretic Foundations for Sense in Stratified Semantic Spaces

Overview

This work introduces TTS, a dependent type theory of sense, constructed in response to empirical findings that challenge the manifold hypothesis long assumed to underlie the geometry of LLM token spaces. Departing from classical smoothness assumptions, TTS formalizes semantic composition as instrument-relative, graded choices in a stratified space that exhibits locally disconnected filler spaces (forks) and regime-dependent verdicts. The system unifies a constructive proof-theoretic logic grounded in positive (witnessed) measurements with a Fregean semantic analysis, establishing sense as a witnessed choice of completion (filler), hyperintensional distinctness as apartness at a regime, and formalizing the persistence and interaction of these phenomena across instruments and resolutions.

Background and Motivation

Empirical studies [robinson2025manifold, bigelow2025forking] have decisively rejected the manifold hypothesis for LLM token subspaces, establishing nontrivial stratification and local singularities—positions where completion exists in non-homotopic, disconnected sets. Furthermore, forking events in neural generation expose semantic boundaries at which genuine decisions, not explainable by continuous deformation, are required. The conventional mathematical frameworks, such as MLTT or HoTT, are inadequate for describing such failures of composition: HoTT's Kan condition enforces contractibility of fillers, erasing the fork phenomenon central to actual semantic ambiguity.

TTS builds directly on these findings. It replaces global Kan assumptions with a locally-graded, regime-indexed analysis, producing a type theory sensitive to instrument resolution and strictly wedded to measurement-based evidence. The resulting system provides an account of meaning in LLMs rooted in the stratified and empirically anchored geometry now observed.

The Type Theory: Core Components

Semantic Primitives

At the core of TTS are semantic types—formalizations of points and continuations (hom-types) in the space of semantic states, governed by a signature generated from instrument-named data. Triangles encode compositional warrants: a triangle type witnesses coherent composition of a pair of continuations to a third.

Judgments in TTS are indexed over regimes, standing for specific instruments and resolutions (e.g., embedding distances at given thresholds). This explicit regime-indexation underpins all verdicts of identity, indiscernibility, and apartness.

Graded Relations

Indiscernibility (\simeq) and constructive apartness (#\#) are defined as regime-indexed binary relations, each valued in h-propositions, and subject to monotonicity in the regime poset. Apartness, following Brouwer-Heyting-Bishop, is strictly constructive: it can only be introduced via measurement context—explicit witnessing constants produced by actual instrument runs. This distinction enforces that the calculus "derives only what records force": no negative verdict exists without a constructive ground in observed data.

Grades of Compositional Demand

TTS characterizes three key local phenomena:

  1. Canonical (Canonicity): Single (up to indiscernibility) composition—no semantic decision; corresponds to classical Kan fillers.
  2. Forked: Disconnected space of fillers—genuine semantic choice is required. Forks are positively, witnessably established.
  3. Locked: Over-coherent regions where only degeneracy persists; all escape horns are uninhabited, corresponding to collapse/degeneration in LLMs.

The system derives—but does not stipulate—the mutual exclusion and persistence properties of these grades. Forks persist under regime refinement, canonicity does not, and the calculus reflects the empirical irreversibility of measured semantic distinctions.

Measurement Context and Provenance

Measurement contexts are finite lists of recorded (constant) verdicts—witnessed (in)discernibilities and separations at specific regimes. The type theory is parameterized by these records, ensuring all derivable judgments are conservative, of guaranteed provenance, and never asserted ex nihilo.

Formal Results

Exclusion, Persistence, and Cross-Regime Consistency

  • Exclusion: Fork and canonical verdicts at a horn are provably mutually exclusive at a given regime.
  • Persistence: Forks persist under restriction to finer regimes (i.e., stricter instruments), while canonicality does not (removable by refinement).
  • Cross-instrument Consistency: Disagreement is lawful precisely when the connecting instrument is strictly coarser or incomparable, never when it is a refinement of the separating one.

Semantic Interpretation

The intended models of TTS are regime-indexed simplicial apartness spaces. Key properties:

  • All nontrivial semantic identification is through regime-indexed indiscernibility; ambient identity is interpreted discretely.
  • Forks in the calculus are provably tied to witnessed multi-component structure in the measured data.
  • Basins—components where all inner horns are canonical—offer a mathematically precise formalization of semantic strata; forked horns demarcate singularities or boundaries between basins.

Empirical Justification

The theory is connected to actual ML/LLM measurement protocols:

  • Sampling and Embedding Experiments: Fork signatures—persistent cluster bimodality across resampling—appear precisely at semantically ambiguous boundaries; control conditions align tightly with unimodal basins.
  • Fill-in-the-Middle (FIM) Experiments: Conditioning on both sides exposes the effect of context in pruning/disambiguating completions, showing how right context collapses the fork distribution onto a unique (canonical) basin.
  • Multi-Instrument Analysis: Different generation/measurement regimes can yield divergent verdicts, exactly as permitted by TTS's formal cross-regime theorem. Alignment tuning induces mode collapse, functionally erasing observable forks—an effect TTS can formalize as non-refutable canonicity arising from instrumental impoverishment. Figure 1

    Figure 1: Sense-margin and cluster-structure analysis of sampled completions. The ambiguous boundary yields a bimodal (forked) distribution straddling two estimator basins; controls remain unimodal.

    Figure 2

    Figure 2: FIM horn experiment: the neutral horn (ambiguous completion) shows a distinctive fork signature, while suffix constraints collapse the sample distribution to a single basin—demonstrating the sufficiency of right-context for semantic disambiguation.

Theoretical and Practical Implications

TTS provides a rigorous, constructive logic for semantic verdicts in stratified spaces, with term-level witnesses tied directly to measurement. This enables:

  • Falsifiable semantic theory: All positive verdicts are grounded in, and limited by, actual data records.
  • A formal solution to the Fregean sense/reference distinction: Reference is the horn (boundary data); sense is the choice of filler; distinctness is witnessed apartness. Informativeness of semantic identity is explained via new instrument verdicts—consistent only when instruments are not comparable.
  • Instrument-relative, non-uniform meaning: No globally unique notion of identity or difference is asserted—verdicts are always local to instrument and context.

In AI, these findings have immediate applications to evaluation and interpretability protocols for LLMs. They offer a testable criterion for when systems confront genuine semantic ambiguity (forks) and when observed coherence may be an artifact of instrumental smoothing or model collapse. The relation between model alignment and erasure of semantic ambiguity emerges as a concrete, auditably formal phenomenon.

Conclusions

This paper presents TTS, a type theory of sense, as a mathematically precise response to the empirical failures of the manifold hypothesis in LLM embedding geometry. TTS accurately captures the witnessed, stratified, and instrument-relative character of semantic composition in modern AI systems. It replaces ungrounded negative verdicts with witnessed, measurement-based judgments, and frames semantic choices as positive, verifiable events. Through soundness, conservativity, and provenance theorems, it establishes a logic that is both philosophically robust and practically falsifiable.

Further work is warranted on the completeness and mechanization of the theory, on generalization to higher horns and complex compositional phenomena, and on large-scale empirical census of semantic forks in AI models. The connection to inferential closure properties and alignment-induced collapse suggests promising avenues for future research in the interpretability, auditing, and improvement of LLMs.

Whiteboard

Explain it Like I'm 14

Plain‑English Summary of “A Type Theory of Sense: Witnessed Choice in Stratified Semantic Spaces”

What is this paper about?

This paper builds a new logical system (called TTS) to describe meaning in language as something you do, not just something you have. It treats a “sense” as a way of completing a partially given message, and it treats “reference” as the parts that are already fixed. The key idea is that real language systems (like LLMs, LLMs) move through a bumpy, layered landscape of meanings, not a smooth highway. Because of that, sometimes there is exactly one natural way to continue a sentence, and sometimes there are several genuinely different ways — a real choice.

What questions does it ask?

The paper asks, in simple terms:

  • How can we formally talk about places where a sentence can be continued in more than one, truly different, meaningful way?
  • How can we make those differences depend on the measuring tool we use (like different zoom levels of a camera), since what looks the same at one level may look different at another?
  • How can we connect this logic to real measurements from LLMs, so “differences” are not just declared but witnessed by data?
  • Can we rebuild Frege’s old idea — sense vs reference — so that “sense” is the way you reach a meaning, not just the meaning itself?

How did the authors study it? (Methods and approach)

To explain the methods, imagine a trail map:

  • The map has points (places you can be in a conversation).
  • There are directed paths (ways to move from one point to another — like continuing a sentence).
  • A “triangle” is when you go from point A to B, then B to C, and ask: is there a single, smooth way to go from A to C that matches those edges? If yes, you have a “filler” — a way to read the whole passage smoothly.

Now add instruments with resolution:

  • Think of an “instrument” like a camera with a certain focus level (called a regime). At a coarse setting, two pictures can look the same; at a fine setting, details pop out and they look different.
  • The paper defines two relations, both tied to a chosen instrument:
    • Indiscernible: the instrument can’t tell two things apart (like two photos that look the same at this zoom).
    • Apart: the instrument can tell them apart — and you must show a witness (real data) that proves the difference.

Crucially, the paper only allows “apartness” to enter the system if you have actual measurement data. This is enforced using a “measurement context” — basically a lab notebook listing what an instrument run really observed (for example, distances between embeddings). The logic then reasons only from what the notebook actually records. No guessing.

Finally, the paper defines three “grades” you can encounter when trying to complete a sentence:

  • Canonical: as far as the chosen instrument can tell, there’s effectively one way to continue (no real decision).
  • Forked: there are two or more warranted ways to continue that the instrument can tell apart (a real choice).
  • Locked: you’re stuck in a tight loop (think of repetitive, collapsed text).

The formal system TTS adds:

  • Directed paths (one‑way continuations).
  • Triangle witnesses (proof a full reading really composes).
  • Regime‑indexed “same” vs “different.”
  • A measurement record so differences must be backed by data.

What did they find?

Here are the main results, stated plainly:

  • Sense as a way: A “sense” is a complete way to fill in a partially given bit of language (the “horn”). The “reference” is the fixed boundary you must respect (the edges of the horn). So, sense = how you get there; reference = where you must end up.
  • Forks vs canonicity can’t both happen at the same zoom: At one instrument setting, you can’t have “everything looks the same” and also “there is a witnessed difference” for the same case. That’s a theorem.
  • Differences persist when you zoom in: If at a coarse setting you can witness a difference (apartness), you’ll still be able to see that difference at finer settings. But if everything looked the same at a coarse setting, zooming in can reveal differences. So sameness does not necessarily persist.
  • Cross‑instrument peace treaty: Two instruments can disagree without contradiction, but only if the “connecting” instrument is not finer than the “separating” one. In everyday words: a coarser lens can merge what a finer lens splits; looking more closely won’t re‑merge what you already separated by looking closely.
  • No difference without data: The logic cannot invent “forks” (real choices) out of thin air. It only recognizes them if the measurement record actually shows them. This is called conservativity and provenance: the logic derives only what the data entitles you to claim.
  • Real‑world tie‑in:
    • Measurements show that token spaces in several LLMs are not smooth surfaces (“manifold hypothesis” fails). They’re layered/stratified: different parts have different shapes and dimensions.
    • Researchers have also found “forking tokens”: changing a single token can flip the model’s future behavior into a different mode — exactly the kind of fork this logic is built to recognize.
  • Frege revived: In this system, hyper‑fine differences in “sense” (ways of getting to a meaning) are real, witnessed, and can stay even if two senses end at the same reference. Identity can be informative because you must actually construct a bridge between two different ways and show they cohere.

A tiny example to make it concrete:

  • Take the sentence “We met them near the bank...” The word “bank” can mean “financial bank” or “river bank.”
  • At a coarse setting, some continuations might blur together. At a finer setting, the financial ones cluster apart from the river ones.
  • That split — supported by measured distances between embeddings — is a witnessed fork: the system must choose between genuinely different senses.

Why are these findings important?

  • Clearer picture of meaning: The paper gives a precise way to talk about meaning as action — as “how you get there” — and not just as a final destination.
  • Ambiguity with evidence: It turns fuzzy talk like “this is ambiguous” into testable, data‑backed statements. You can point to measurements that witness real choice points.
  • Instrument‑relative truth: It shows how different tools (zoom levels, metrics, probes) can legally disagree, and exactly when they must agree. That matches real scientific practice.
  • Safer claims: By design, the system won’t let you claim “difference” unless you show the data. That keeps the theory honest and testable.
  • Links to LLM behavior: The theory lines up with two observed facts about LLMs:
    • Their “meaning spaces” are not smooth everywhere (manifold hypothesis violations).
    • They sometimes branch into different futures based on one token (forking tokens).
  • Frege, modernized: The old philosophical idea — sense vs reference — gets a geometric, testable form. Sense lives in the space of ways to complete a demand; reference is its fixed boundary; “hyperintensional” differences are witnessed separations between ways.

What could this change or enable?

  • Better evaluation and control of LLMs: We can locate real choice points, check whether choices persist at finer analysis, and understand when seemingly identical outputs hide different “ways of getting there” (which could matter for safety, style, or bias).
  • Stronger semantics for AI: Designing systems that respect witnessed differences can prevent “over‑smoothing” and reduce collapse into repetitive or bland outputs (the “locked” phase).
  • Testable philosophy of language: The theory suggests concrete experiments (e.g., resampling, clustering under different thresholds) to study sense and ambiguity.
  • Practical rule of thumb: If a coarse probe merges two continuations that a finer probe separates, trust the finer probe about difference; but expect coarser probes to sometimes blur distinctions — and that’s okay, not a contradiction.

In short, the paper builds a clean, data‑first logic of meaning where:

  • Sense = a way to complete what is already fixed.
  • Differences must be witnessed by measurements.
  • Choices persist when you look more closely.
  • And the whole framework matches what we can actually observe in modern LLMs.

Open Problems

We haven't generated a list of open problems mentioned in this paper yet.

Collections

Sign up for free to add this paper to one or more collections.