Papers
Topics
Authors
Recent
Search
2000 character limit reached

Umwelt Representation Hypothesis (URH)

Updated 14 July 2026
  • URH is a framework defining how internal representations arise from agent-specific ecological and sensorimotor constraints rather than a universal world model.
  • It posits that alignment between biological systems and ANNs emerges from local optimal compressions driven by overlapping environmental factors.
  • The hypothesis guides using metrics like CKA, RSA, and SVCCA to map representational clusters and diagnose systematic divergences due to ecological variability.

Searching arXiv for the cited URH sources and closely related papers. The Umwelt Representation Hypothesis (URH) is a framework for explaining representational alignment and divergence across biological and artificial systems in terms of ecological and sensorimotor constraints rather than convergence to a single universal world model. In one formulation, URH states that alignment between brains and artificial neural networks (ANNs) arises from overlap in the ecological constraints under which systems develop and operate, so that systems learn locally optimal compressions of the world adaptive to their particular sensors, effectors, tasks, goals, resource limits, learning signals, social environments, and developmental histories (Bosch et al., 20 Apr 2026). In a more formal agent-theoretic lineage, URH identifies an agent’s internal world as the quotient of the external world induced by sensorimotor indistinguishability, so that the agent’s Umwelt is the minimal world model required to preserve its internal process (Ay et al., 2016). Across these usages, the common claim is that representation is fundamentally agent-relative: it is structured by what a system can sense, do, and exploit.

1. Concept and scope

URH takes its central term from biology and ecological psychology, where an Umwelt is the organism-specific “slice” of the world: the subset of sensory inputs, action affordances, goals, and constraints that matter for that organism (Bosch et al., 20 Apr 2026). On this view, the way a system perceives, values, and acts within its niche is determined by its sensors, effectors, energy budgets, developmental history, and local environment. URH extends this concept into comparative neuroscience, machine learning, embodied AI, and linguistic agents.

A central motivation for URH is the growing literature reporting striking representational alignment between ANNs and biological brains. Proposals of “universal representations,” including the Platonic Representation Hypothesis, treat such alignment as evidence that sufficiently capable systems trained on sufficiently rich data will converge to the same representational structure because all roads lead to a single globally optimal world model (Bosch et al., 20 Apr 2026). URH argues that this inference is premature. It permits the existence of universal representations, but treats them as special cases produced by unusually strong overlap in constraints rather than as the default explanation for alignment.

The hypothesis is therefore both explanatory and methodological. Explanatorily, it claims that partial alignment is expected: some systems share some representations because some aspects of their Umwelten overlap. Methodologically, it recommends replacing the search for a single optimal “world model” with the mapping of clusters of alignment in ecological constraint space (Bosch et al., 20 Apr 2026).

2. Historical and formal foundations

A mathematically explicit precursor appears in the measure-theoretic treatment of Umwelt in embodied agents. In that formulation, Uexküll’s function-circle is formalized as a stochastic sensorimotor loop over world states, sensors, controller states, and actions, each defined on Souslin spaces with Borel σ\sigma-algebras (Ay et al., 2016). The mechanisms are given by Markov kernels: the sensor mechanism β\beta, state update ϕ\phi, policy π\pi, and world update α\alpha. Together with an initial distribution, these define the joint process (Wt,St,Ct,At)tN(W_t,S_t,C_t,A_t)_{t\in\mathbb{N}} (Ay et al., 2016).

Within this framework, the external perspective is represented by the full σ\sigma-algebra of world events, whereas the intrinsic perspective is a smaller σ\sigma-algebra determined by what distinctions can be induced through the agent’s sensorimotor loop. The paper defines the agent’s Umwelt as

Σagent:=Wint:=σ{PSaN:aNAN},\Sigma_{\mathrm{agent}} := W_{\mathrm{int}} := \sigma\{P_S^{a_{\mathbb{N}}}: a_{\mathbb{N}}\in A^{\mathbb{N}}\},

where PSaN(w)P_S^{a_{\mathbb{N}}}(w) is the distribution of sensor sequences generated from world state β\beta0 under action sequence β\beta1 (Ay et al., 2016). Two world states are equivalent when they induce identical sensor-sequence distributions for all action sequences. The Umwelt is then the β\beta2-algebra of unions of these equivalence classes.

This yields a precise minimality result. Under compactness and continuity assumptions, there exists a simplified world model whose externally valid β\beta3-algebra equals the intrinsic one while preserving all internal processes relevant to the agent (Ay et al., 2016). This substantiates a strong version of URH: any externally richer description collapses, from the agent’s perspective, to the quotient induced by sensorimotor accessibility.

A complementary but distinct formalization appears in the comparative representation framework of the 2026 URH paper. There, ecological constraints are summarized by a structured descriptor

β\beta4

with dimensions including stimulus distributions, tasks, embodiment, goals, resources, learning signals, social/cultural environment, and developmental history (Bosch et al., 20 Apr 2026). For a system operating under constraints β\beta5 with input space β\beta6, its representation is modeled as

β\beta7

Alignment between two systems is measured by a function β\beta8, and ecological overlap is defined as

β\beta9

with the core prediction

ϕ\phi0

where ϕ\phi1 is monotone increasing and ϕ\phi2 bundles nuisance factors such as architecture class, capacity, optimization details, metric flexibility, and measurement noise (Bosch et al., 20 Apr 2026). This implies local alignment clusters rather than a single representational optimum.

3. Core claims and representational logic

The central claim of URH is that alignment arises from overlap in ecological constraints, not from a universal convergence law (Bosch et al., 20 Apr 2026). These constraints include stimulus distributions, task objectives, embodiment, goals and priors, resource limits, learning signals, social or cultural environment, and developmental history. Systems with overlapping constraints are predicted to learn overlapping representations; systems with divergent constraints are predicted to exhibit systematic, adaptive differences rather than arbitrary noise.

This logic directly opposes what the 2026 paper calls the “Anna Karenina” view that misalignment is merely idiosyncratic residual variance around a shared optimum (Bosch et al., 20 Apr 2026). Under URH, representational differences across species, individuals, and models are often informative because they track local ecological demands. Echolocation in bats and dolphins, magnetoreception in birds, stellar compass neurons in Bogong moths, and the mantis shrimp’s approximately 16 photoreceptor types versus the human approximately 3 are presented as adaptive specializations rather than deviations from a universal template (Bosch et al., 20 Apr 2026).

In this sense, URH treats internal representations as locally optimal compressions. A system does not encode the world “as such”; it encodes what is actionable, predictive, and useful given its niche. This theme also appears in information-theoretic work on evolved cognitive networks, which defines representation as information about relevant environmental features encoded in internal states beyond what is present in the sensors:

ϕ\phi3

Here ϕ\phi4 denotes environment states, ϕ\phi5 sensor states, and ϕ\phi6 internal states (Marstaller et al., 2012). This measure isolates internal modeling rather than direct sensor mirroring. The paper argues that such representations are “the expected consequence of an adaptive process” and that agents form them during their lifetime (Marstaller et al., 2012). This suggests a conceptual continuity between URH’s ecological framing and formal measures of sensor-beyond internal modeling.

A related computational instantiation appears in work on sensorimotor prediction under partial observability. There, compact representations emerge when an agent predicts future sensations from current sensations and motor commands, with memory integrating over time (Kulak et al., 2018). The learned codes capture local, action-conditioned affordances such as walls, corners, and ends of walls rather than a global metric map. This suggests a narrow but concrete reading of Umwelt: an internal model of regularities in the agent’s own sensorimotor stream.

4. Alignment metrics, ecological space, and model comparison

URH does not reject representational comparison; it reframes its purpose. Instead of using alignment metrics to search for a single best world model, it uses them to locate clusters of similarity in ecological constraint space (Bosch et al., 20 Apr 2026). The principal metrics discussed are Centered Kernel Alignment (CKA), Representational Similarity Analysis (RSA), and Singular Vector Canonical Correlation Analysis (SVCCA).

For representations over common stimuli ϕ\phi7 and ϕ\phi8, CKA is defined as

ϕ\phi9

and RSA is defined from dissimilarity matrices π\pi0 and π\pi1 by

π\pi2

SVCCA computes denoised shared linear subspaces and averages the resulting canonical correlations:

π\pi3

The 2026 paper emphasizes, however, that highly flexible mappings in encoding or decoding analyses can mask systematic differences, so less-flexible and standardized metrics are preferable when the goal is to expose ecological structure (Bosch et al., 20 Apr 2026).

The same paper proposes an “ecological constraint space” π\pi4 equipped with a metric π\pi5 over ecologies, together with a mapping from ecologies to learned representational manifolds. In this formalization,

π\pi6

so alignment should decay with ecological distance (Bosch et al., 20 Apr 2026). This predicts neighborhoods of high alignment corresponding to shared tasks, sensory channels, developmental exposures, or architectural constraints.

A plausible implication is that model comparison becomes inherently comparative and cartographic. Brains, species, individuals, and ANNs should be assigned ecological profiles recording stimulus regimes, embodiment, objectives, learning signals, resource limits, and developmental histories; these profiles can then be embedded into ecological space and compared against alignment maps (Bosch et al., 20 Apr 2026). Under this view, a failure to align is not necessarily evidence of model inadequacy. It may instead diagnose a mismatch in Umwelt.

5. Empirical evidence across species, individuals, and artificial systems

The empirical case for URH is assembled from several levels of comparison. At the cross-species level, the 2026 review emphasizes modality-specific adaptations such as echolocation, magnetoreception, stellar compass neurons, and expanded photoreceptor repertoires, alongside evidence that evolutionary hierarchy reflected in functional connectivity predicts representational similarity across species (Bosch et al., 20 Apr 2026). The central point is not merely that species differ, but that these differences are structured by niche and embodiment.

Within species, the same review notes that cultural and experiential factors shape perception and cortical organization. Examples include susceptibility to geometric illusions, the phenomenon of #thedress, reading acquisition and the visual word form area, category selectivity associated with Pokémon and car expertise, and language-specific representations spanning syntax, phonology, and semantics (Bosch et al., 20 Apr 2026). These findings are difficult to reconcile with a strong universality thesis in which higher-level misalignment is only noise.

For ANNs, the reviewed evidence includes texture bias in standard CNNs relative to the more global shape bias of human vision, divergent invariances exposed by model metamers, non-monotonicity between task accuracy and brain alignment, and systematic changes in learned representations induced by objective choice, data distribution, depth, dropout, initialization, and multimodal or topographic/recurrent architectures (Bosch et al., 20 Apr 2026). Recent brain-model comparisons are also noted as rarely reaching noise ceilings, which URH interprets as consistent with persistent ecological mismatch rather than simple underperformance.

Supportive evidence also comes from experimental settings designed around affordance structure. In deep reinforcement learning agents trained on a shortcut navigation task, environmental complexity was operationalized by shortcut openness frequency π\pi7, which determined both affordance exposure and cue frequency (Liu et al., 2024). Agents used PPO with fully connected layers and a GRU, received 12 sight lines yielding π\pi8, and acted via four discrete actions (Liu et al., 2024). Several findings align closely with URH. Spatial representations emerged early and stabilized before advanced strategies, while population-level codes for intended trajectory developed later and tracked shortcut usage. Frequent cue exposure supported early cue encoding, but stronger ultimate cue representations arose through integration into planning rather than exposure alone (Liu et al., 2024). The primary trajectory-separation score, based on normalized Wasserstein distance between pre-entrance and pre-shortcut activation sets, increased monotonically with training and correlated with shortcut usage, indicating that planned path rather than immediate location was encoded at the population level (Liu et al., 2024).

A different line of evidence comes from predictive sensorimotor learning in partially observable 2D environments. There, adding motor signals and memory strongly improved forward prediction, and latent clusters emerged for corners, walls, and context-dependent “nothing perceived” states (Kulak et al., 2018). In one reported comparison, the Recurrent-SM-encoder achieved lower π\pi9 prediction error than alternatives across Square, Rooms1, and Rooms2, with error reductions attributed first to motor information and then to memory (Kulak et al., 2018). The representation remained local, action-conditioned, and history-sensitive, which is consistent with URH’s emphasis on agent-centric regularities rather than externally fixed coordinates.

6. Variants and extensions: embodied, evolved, and linguistic Umwelten

URH is not a single formalism but a family resemblance across several research programs. In embodied agent theory, the key object is the intrinsic α\alpha0-algebra of sensorimotor accessibility and the quotient world induced by observational indistinguishability (Ay et al., 2016). In ecological-comparative neuroscience and ANN alignment, the key object is ecological constraint space and the overlap function α\alpha1 (Bosch et al., 20 Apr 2026). In information-theoretic evolutionary models, the key object is the representation measure α\alpha2 and its temporal extensions α\alpha3, α\alpha4, and α\alpha5 (Marstaller et al., 2012). In sensorimotor deep learning, the relevant object is a compact predictive code grounded in actions and memory (Kulak et al., 2018). These are not identical definitions, but they converge on the same explanatory motif: representation is shaped by what the agent can distinguish, remember, and use.

A more recent extension applies URH to linguistic agents. “Umwelt engineering” defines the linguistic cognitive environment as the vocabulary, grammar, conceptual primitives, and reasoning affordances available to an agent at inference time (Jehu-Appiah, 29 Mar 2026). For standard LLMs, the paper argues, the token stream does not merely report cognition; it is the cognition. The medium of reasoning α\alpha6 and a constraint set α\alpha7 therefore alter internal representations α\alpha8 and task performance α\alpha9, with the sensitivity claims (Wt,St,Ct,At)tN(W_t,S_t,C_t,A_t)_{t\in\mathbb{N}}0 and (Wt,St,Ct,At)tN(W_t,S_t,C_t,A_t)_{t\in\mathbb{N}}1 (Jehu-Appiah, 29 Mar 2026).

The reported experiments impose vocabulary constraints such as E-Prime and No-Have across seven task families and multiple models. The aggregate results show overall accuracy of 83.5% for control, 88.6% for No-Have, and 85.4% for E-Prime over scoreable (Wt,St,Ct,At)tN(W_t,S_t,C_t,A_t)_{t\in\mathbb{N}}2 trials, with No-Have compliance at 92.8% and E-Prime compliance at 48.1% (Jehu-Appiah, 29 Mar 2026). Task effects are heterogeneous rather than uniform: No-Have improves ethical reasoning by 19.1 percentage points and classification by 6.5 percentage points, while E-Prime improves causal reasoning by 14.1 percentage points but degrades syllogisms by 3.4 percentage points (Jehu-Appiah, 29 Mar 2026). Cross-model correlations of E-Prime task deltas reach (Wt,St,Ct,At)tN(W_t,S_t,C_t,A_t)_{t\in\mathbb{N}}3 between Haiku and GPT-4o-mini, indicating model-dependent restructuring (Jehu-Appiah, 29 Mar 2026).

In a second experiment, ensembles of linguistically constrained agents collectively achieved broader debugging coverage than any single agent. A minimal 3-agent ensemble reached 100% ground-truth coverage versus 88.2% for the control, and only 45 of the (Wt,St,Ct,At)tN(W_t,S_t,C_t,A_t)_{t\in\mathbb{N}}4 possible 3-agent subsets achieved full coverage (Jehu-Appiah, 29 Mar 2026). The paper interprets this in terms of “cognitive restructuring” and “cognitive diversification.” This suggests an important generalization of URH: the relevant Umwelt need not be sensory or embodied in the narrow biological sense. It can also be linguistic, provided the medium constrains what operations are available to cognition.

7. Predictions, implications, and limitations

URH yields a set of explicit empirical predictions. The 2026 comparative formulation predicts that alignment should increase monotonically and likely saturate with ecological overlap (Wt,St,Ct,At)tN(W_t,S_t,C_t,A_t)_{t\in\mathbb{N}}5, that misalignment should be systematic and adaptive where constraints diverge, and that shifting systems into new ecological regimes should move their representations predictably toward new local optima (Bosch et al., 20 Apr 2026). Proposed tests include controlled variation of stimulus statistics, task objectives, feedback structures, embodiment, energy and computation budgets, as well as cross-cultural and developmental sampling combined with matched or mismatched ANN ecologies (Bosch et al., 20 Apr 2026). A regression model of the form

(Wt,St,Ct,At)tN(W_t,S_t,C_t,A_t)_{t\in\mathbb{N}}6

is suggested for identifying high-impact constraints and testing (Wt,St,Ct,At)tN(W_t,S_t,C_t,A_t)_{t\in\mathbb{N}}7 (Bosch et al., 20 Apr 2026).

In neuroscience, URH implies that variability across individuals and species is expected, systematic, and often adaptive rather than a simple failure to reach a common optimum (Bosch et al., 20 Apr 2026). It recommends ethology-first experimental design, using tasks and stimuli that reflect natural behavior, active sensing, multimodal integration, and learning dynamics. In AI, it recommends aligning training distributions and objectives with deployment ecology, reporting ecological provenance in model cards, testing out-of-ecology generalization, and avoiding overgeneralization from narrow human-centric benchmarks (Bosch et al., 20 Apr 2026).

Several limitations and counterarguments are explicit. First, shared physics, geometry, and low-level sensory constraints may produce genuine commonalities. URH acknowledges this and predicts that such overlaps will be strongest at lower levels, with higher-level divergence growing as tasks, embodiment, and culture diverge (Bosch et al., 20 Apr 2026). Second, apparent alignment can be inflated by architecture, scale, optimization tricks, flexible mapping models, and measurement noise. URH therefore requires controlled experiments that vary ecology while holding architecture and optimization fixed, standardized metrics such as RSA, CKA, and SVCCA, explicit reporting of noise ceilings, and targeted stimuli that expose divergence (Bosch et al., 20 Apr 2026). Third, operationalizing ecological descriptors (Wt,St,Ct,At)tN(W_t,S_t,C_t,A_t)_{t\in\mathbb{N}}8, defining similarity functions (Wt,St,Ct,At)tN(W_t,S_t,C_t,A_t)_{t\in\mathbb{N}}9, selecting weights σ\sigma0, and learning or engineering the ecological metric σ\sigma1 remain open problems (Bosch et al., 20 Apr 2026).

The measure-theoretic formulation introduces further boundary conditions. Its strongest minimality result depends on compactness and continuity assumptions guaranteeing countable generation and joint measurability of the quotient model (Ay et al., 2016). Without such regularity, identifiability and factorization become more difficult. The paper also notes policy dependence in some constructions and the need for state augmentation in non-Markovian settings (Ay et al., 2016). In the linguistic extension, the main limitation is the absence of an active control matching constraint-prompt elaborateness, which leaves room for self-monitoring confounds (Jehu-Appiah, 29 Mar 2026).

Taken together, these lines of work present URH not as a denial of common structure, but as a shift in explanatory priority. Shared representations are expected where ecologies overlap strongly; divergent representations are expected where ecological constraints differ. The theoretical consequence is a move from universality as the default hypothesis to constraint-conditioned locality as the primary explanatory principle (Bosch et al., 20 Apr 2026).

Topic to Video (Beta)

No one has generated a video about this topic yet.

Whiteboard

No one has generated a whiteboard explanation for this topic yet.

Follow Topic

Get notified by email when new papers are published related to Umwelt Representation Hypothesis (URH).