Papers
Topics
Authors
Recent
Search
2000 character limit reached

Exploratory Experience Shapes the Geometry of Predictive Representations

Published 27 May 2026 in q-bio.NC and cs.LG | (2605.27929v1)

Abstract: Active sensing links behavior and learning through an action-perception loop: actions determine the observations used to update internal predictive models of perception, which subsequently guide the next actions. Predictive-coding frameworks provide a natural way to model this process, since internal representations are continuously updated to predict future observations. Here, we ask how exploratory and exploitative behavioral strategies shape these internal predictive representations. We build an online learning agent in a tree-like maze with a controllable parameter regulating the balance between exploratory and exploitative regimes. The agent updates a predictive-coding-based perception model from experience generated by its own behavior. The model predicts both future maze states and reward probability, allowing the agent to select actions either by expected information gain during exploration or by predicted reward during exploitation. We show that the resulting internal predictive representations depend strongly on the agent's behavioral regime. Exploratory agents develop representations that are more spatially organized and better preserve the structure of maze transitions in latent space. In contrast, exploitative agents learn less organized representations. We then train this predictive model on natural trajectories of water-deprived mice navigating the same maze and compare the resulting representations with those learned from agent trajectories. More exploratory mice show representational geometries that closely match those of exploratory agents, whereas mice with more restricted visitation patterns resemble reward-driven, exploitative agents. Together, these findings suggest that exploration enables predictive models to form generalized internal representations by organizing latent space around both spatial location and transition context in artificial agents and animals.

Summary

  • The paper demonstrates that varying only the exploration–exploitation policy in an online predictive-coding agent reorganizes latent representations, with exploratory agents showing stronger maze-depth alignment, transition consistency, and smoother trajectories than reward-driven agents.
  • The paper finds that navigation performance peaks at an intermediate reward-switching probability of approximately 0.025–0.05, because effective exploitation depends on sufficient exploration to build an accurate reward map.
  • The paper shows that mice with higher visitation entropy develop model-based latent geometries resembling exploratory agents, while noting that behavioral data alone establish an association rather than a causal neural mechanism.

Overview

This paper examines how the exploration–exploitation balance of a navigating agent shapes the geometry of the internal representations learned by a predictive-coding model. The authors, Shilova, Sharafeldin, Balakrishnan, and Choi, construct an online-learning agent in the binary-tree labyrinth introduced by Rosenberg et al., where a single action-conditioned predictive model supports both information-driven exploration (via expected information gain) and reward-driven navigation (via a learned value map). By varying only the probability of switching into reward-driven mode while holding architecture, objective, and environment fixed, they isolate behavioral sampling as the experimental variable. Their central finding is that exploratory experience produces latent spaces that are spatially organized and faithful to maze transition structure, whereas exploitative behavior yields less organized representations — and that mice with higher visitation entropy exhibit latent geometries resembling those of exploratory agents (2605.27929).

Task and modeling framework

The environment is a deterministic binary tree with N=127N = 127 nodes; each bout begins at the root, actions are drawn from {left, right, back}, and sparse reward is delivered only at a fixed water-port node. The agent's perception model is an action-conditioned variational predictive-coding network: a prior network pθ(zt∣ht,at,xt)p_\theta(z_t \mid h_t, a_t, x_t) predicts a 16-dimensional latent transition state before observing the outcome, a posterior network qϕ(zt∣ht,at,xt+1)q_\phi(z_t \mid h_t, a_t, x_{t+1}) infers it afterward, and a GRU updates a 64-dimensional recurrent history state. A shared decoder predicts both the next node and reward probability from (zt,ht+1)(z_t, h_{t+1}). Training minimizes next-node cross-entropy, reward binary cross-entropy, a KL term aligning prior and posterior latents, and an ℓ1\ell_1 sparsity regularizer on the posterior mean.

Action selection operates in two regimes. In the exploratory regime, actions are scored by expected information gain over the latent belief: the expected reduction in entropy of ztz_t after observing the next node, approximated by Monte Carlo sampling of hypothetical observations. In the reward-driven regime, predicted reward probabilities are averaged over recent visits to form a learned reward map R^(x)\hat{R}(x), and values are propagated via a soft backup over the empirically observed transition graph (γ=0.8\gamma = 0.8, T=0.1T = 0.1). A single switching probability pp governs entry into reward-driven mode at each exploratory step; after reward delivery, the agent is forced to explore for seven steps. This design ensures that all agents share identical learning machinery, so representational differences can be attributed to the sampling distribution alone.

Behavioral consequences of the exploration–exploitation balance

Varying pθ(zt∣ht,at,xt)p_\theta(z_t \mid h_t, a_t, x_t)0 induces a clear tradeoff between transition learning and reward seeking. Transition-balanced cross-entropy degrades monotonically as reward-driven behavior increases, consistent with restricted sampling providing fewer distinct transitions for model updating. Reward-seeking performance, however, is non-monotonic in pθ(zt∣ht,at,xt)p_\theta(z_t \mid h_t, a_t, x_t)1: purely exploratory agents learn accurate transition models but do not preferentially return to reward, while highly reward-driven agents exploit incomplete value maps built on insufficient coverage. Performance peaks at intermediate switching probabilities of approximately pθ(zt∣ht,at,xt)p_\theta(z_t \mid h_t, a_t, x_t)2–pθ(zt∣ht,at,xt)p_\theta(z_t \mid h_t, a_t, x_t)3, where agents retain enough exploration to form accurate reward maps yet exploit them effectively.

The reward-map analysis explains this tradeoff mechanistically. Predicted reward gradually localizes around the true water-port and then propagates through the observed transition graph, a process requiring sufficient terminal-node coverage. Exploratory and moderately reward-driven agents assign high predicted reward to the true water-port node, whereas strongly reward-driven agents frequently fail to do so across 30 random seeds. An important implication is that exploitation is beneficial only conditional on adequate exploration: reward-guided behavior improves efficiency precisely when the underlying world model contains enough structure to support it, and purely exploitative settings underperform balanced ones.

Exploration organizes predictive latent geometry

The authors analyze the geometry of the prior latent state pθ(zt∣ht,at,xt)p_\theta(z_t \mid h_t, a_t, x_t)4, which encodes the action-conditioned prediction of the upcoming transition, using three complementary metrics computed in the original high-dimensional space rather than on UMAP projections:

  • Spatial alignment (pθ(zt∣ht,at,xt)p_\theta(z_t \mid h_t, a_t, x_t)5): Spearman correlation between the first principal component of latent states along monotonic trajectory segments and maze depth.
  • Transition consistency (pθ(zt∣ht,at,xt)p_\theta(z_t \mid h_t, a_t, x_t)6): fraction of pθ(zt∣ht,at,xt)p_\theta(z_t \mid h_t, a_t, x_t)7-nearest latent neighbors sharing the same inward/outward transition direction.
  • Trajectory tortuosity (pθ(zt∣ht,at,xt)p_\theta(z_t \mid h_t, a_t, x_t)8): ratio of path length to endpoint distance in latent space, measuring smoothness.

Exploratory agents score better on all three measures. Their latent spaces form branched geometries aligned with the hierarchical structure of the tree, with clear separation between outward (root-to-leaf) and inward (leaf-to-root) transitions and smoother latent trajectories. More reward-driven agents produce compact, disorganized representations without directional separation. Because the latent geometry is never directly optimized, these differences arise indirectly from the training trajectories generated by each policy — broad sampling repeatedly exposes shared structural motifs (depth-matched transitions, direction of movement relative to the root), imposing pressure to organize the full transition structure, whereas value-map-driven sampling supports prediction only along frequently visited paths.

Correspondence with mouse behavior

To test whether this geometry–behavior relationship extends beyond synthetic policies, the authors train the same model on trajectory data from water-deprived mice navigating the same labyrinth task, one model per mouse with matched bout lengths and total experience. Each animal's sampling breadth is summarized by normalized visitation entropy over the 127 nodes. Mice with higher visitation entropy show stronger depth alignment, greater local consistency of transition direction, and lower tortuosity — quantitatively closer to exploratory agents — while low-entropy mice resemble mildly reward-driven agents. Representative animals B5 (exploratory) and C8 (reward-focused) illustrate the qualitative contrast: B5's latent space exhibits branching and depth organization, whereas C8's is compact with reduced spatial alignment.

The authors appropriately note that because mouse behavior is observational rather than experimentally manipulated, these results constitute a geometry-behavior association rather than causal evidence. Nevertheless, the association mirrors the causal pattern established in agents, and appendix analyses show that agent exploration entropies overlap the range observed in mice at low switching probabilities (pθ(zt∣ht,at,xt)p_\theta(z_t \mid h_t, a_t, x_t)9 to qϕ(zt∣ht,at,xt+1)q_\phi(z_t \mid h_t, a_t, x_{t+1})0), grounding the comparison. Supplementary unit-level analyses further show that individual units in both qϕ(zt∣ht,at,xt+1)q_\phi(z_t \mid h_t, a_t, x_{t+1})1 and qϕ(zt∣ht,at,xt+1)q_\phi(z_t \mid h_t, a_t, x_{t+1})2 develop spatially localized tuning at multiple scales — subtrees, branches, leaves, and specific paths — suggesting that population-level predictive geometry coexists with classical place-like tuning, though the authors explicitly decline to interpret these units as direct neural analogues.

Limitations and open questions

Several limitations constrain interpretation. The behavioral policy captures only a simplified exploration–exploitation tradeoff, omitting drives such as novelty seeking, memory, fatigue, and motivational shifts that shape natural behavior. The mouse analysis relies solely on behavioral trajectories without simultaneous neural recordings, so the claim that exploratory experience organizes biological predictive representations remains a model-based prediction rather than an empirical demonstration. The framework also operates on abstract node-level states; whether broad sampling organizes representations of visually grounded, sensory-rich observations is untested. Finally, the deterministic maze and fixed reward location simplify the credit-assignment problem relative to natural environments.

These constraints define concrete open questions: does the latent-geometry–behavior relationship hold in neural population activity recorded during navigation, and how does representational organization develop over learning relative to known hippocampal remapping dynamics?

Conclusion

This paper demonstrates that behavioral sampling regime, held against a fixed predictive objective, is sufficient to determine the organization of learned predictive representations. Exploration yields spatially aligned, transition-consistent, smoothly traversable latent maps; exploitation narrows experience and disrupts this organization, with reward-seeking performance peaking at an intermediate balance rather than under maximal exploitation. The replication of this pattern in models trained on mouse trajectories links artificial-agent findings to animal behavior and motivates future tests against neural recordings during navigation.

Paper to Video (Beta)

No one has generated a video about this paper yet.

Whiteboard

No one has generated a whiteboard explanation for this paper yet.

Open Problems

We haven't generated a list of open problems mentioned in this paper yet.