- The paper demonstrates that varying only the exploration–exploitation policy in an online predictive-coding agent reorganizes latent representations, with exploratory agents showing stronger maze-depth alignment, transition consistency, and smoother trajectories than reward-driven agents.
- The paper finds that navigation performance peaks at an intermediate reward-switching probability of approximately 0.025–0.05, because effective exploitation depends on sufficient exploration to build an accurate reward map.
- The paper shows that mice with higher visitation entropy develop model-based latent geometries resembling exploratory agents, while noting that behavioral data alone establish an association rather than a causal neural mechanism.
Overview
This paper examines how the exploration–exploitation balance of a navigating agent shapes the geometry of the internal representations learned by a predictive-coding model. The authors, Shilova, Sharafeldin, Balakrishnan, and Choi, construct an online-learning agent in the binary-tree labyrinth introduced by Rosenberg et al., where a single action-conditioned predictive model supports both information-driven exploration (via expected information gain) and reward-driven navigation (via a learned value map). By varying only the probability of switching into reward-driven mode while holding architecture, objective, and environment fixed, they isolate behavioral sampling as the experimental variable. Their central finding is that exploratory experience produces latent spaces that are spatially organized and faithful to maze transition structure, whereas exploitative behavior yields less organized representations — and that mice with higher visitation entropy exhibit latent geometries resembling those of exploratory agents (2605.27929).
Task and modeling framework
The environment is a deterministic binary tree with N=127 nodes; each bout begins at the root, actions are drawn from {left, right, back}, and sparse reward is delivered only at a fixed water-port node. The agent's perception model is an action-conditioned variational predictive-coding network: a prior network pθ(zt∣ht,at,xt) predicts a 16-dimensional latent transition state before observing the outcome, a posterior network qϕ(zt∣ht,at,xt+1) infers it afterward, and a GRU updates a 64-dimensional recurrent history state. A shared decoder predicts both the next node and reward probability from (zt,ht+1). Training minimizes next-node cross-entropy, reward binary cross-entropy, a KL term aligning prior and posterior latents, and an ℓ1 sparsity regularizer on the posterior mean.
Action selection operates in two regimes. In the exploratory regime, actions are scored by expected information gain over the latent belief: the expected reduction in entropy of zt after observing the next node, approximated by Monte Carlo sampling of hypothetical observations. In the reward-driven regime, predicted reward probabilities are averaged over recent visits to form a learned reward map R^(x), and values are propagated via a soft backup over the empirically observed transition graph (γ=0.8, T=0.1). A single switching probability p governs entry into reward-driven mode at each exploratory step; after reward delivery, the agent is forced to explore for seven steps. This design ensures that all agents share identical learning machinery, so representational differences can be attributed to the sampling distribution alone.
Behavioral consequences of the exploration–exploitation balance
Varying pθ(zt∣ht,at,xt)0 induces a clear tradeoff between transition learning and reward seeking. Transition-balanced cross-entropy degrades monotonically as reward-driven behavior increases, consistent with restricted sampling providing fewer distinct transitions for model updating. Reward-seeking performance, however, is non-monotonic in pθ(zt∣ht,at,xt)1: purely exploratory agents learn accurate transition models but do not preferentially return to reward, while highly reward-driven agents exploit incomplete value maps built on insufficient coverage. Performance peaks at intermediate switching probabilities of approximately pθ(zt∣ht,at,xt)2–pθ(zt∣ht,at,xt)3, where agents retain enough exploration to form accurate reward maps yet exploit them effectively.
The reward-map analysis explains this tradeoff mechanistically. Predicted reward gradually localizes around the true water-port and then propagates through the observed transition graph, a process requiring sufficient terminal-node coverage. Exploratory and moderately reward-driven agents assign high predicted reward to the true water-port node, whereas strongly reward-driven agents frequently fail to do so across 30 random seeds. An important implication is that exploitation is beneficial only conditional on adequate exploration: reward-guided behavior improves efficiency precisely when the underlying world model contains enough structure to support it, and purely exploitative settings underperform balanced ones.
Exploration organizes predictive latent geometry
The authors analyze the geometry of the prior latent state pθ(zt∣ht,at,xt)4, which encodes the action-conditioned prediction of the upcoming transition, using three complementary metrics computed in the original high-dimensional space rather than on UMAP projections:
- Spatial alignment (pθ(zt∣ht,at,xt)5): Spearman correlation between the first principal component of latent states along monotonic trajectory segments and maze depth.
- Transition consistency (pθ(zt∣ht,at,xt)6): fraction of pθ(zt∣ht,at,xt)7-nearest latent neighbors sharing the same inward/outward transition direction.
- Trajectory tortuosity (pθ(zt∣ht,at,xt)8): ratio of path length to endpoint distance in latent space, measuring smoothness.
Exploratory agents score better on all three measures. Their latent spaces form branched geometries aligned with the hierarchical structure of the tree, with clear separation between outward (root-to-leaf) and inward (leaf-to-root) transitions and smoother latent trajectories. More reward-driven agents produce compact, disorganized representations without directional separation. Because the latent geometry is never directly optimized, these differences arise indirectly from the training trajectories generated by each policy — broad sampling repeatedly exposes shared structural motifs (depth-matched transitions, direction of movement relative to the root), imposing pressure to organize the full transition structure, whereas value-map-driven sampling supports prediction only along frequently visited paths.
Correspondence with mouse behavior
To test whether this geometry–behavior relationship extends beyond synthetic policies, the authors train the same model on trajectory data from water-deprived mice navigating the same labyrinth task, one model per mouse with matched bout lengths and total experience. Each animal's sampling breadth is summarized by normalized visitation entropy over the 127 nodes. Mice with higher visitation entropy show stronger depth alignment, greater local consistency of transition direction, and lower tortuosity — quantitatively closer to exploratory agents — while low-entropy mice resemble mildly reward-driven agents. Representative animals B5 (exploratory) and C8 (reward-focused) illustrate the qualitative contrast: B5's latent space exhibits branching and depth organization, whereas C8's is compact with reduced spatial alignment.
The authors appropriately note that because mouse behavior is observational rather than experimentally manipulated, these results constitute a geometry-behavior association rather than causal evidence. Nevertheless, the association mirrors the causal pattern established in agents, and appendix analyses show that agent exploration entropies overlap the range observed in mice at low switching probabilities (pθ(zt∣ht,at,xt)9 to qϕ(zt∣ht,at,xt+1)0), grounding the comparison. Supplementary unit-level analyses further show that individual units in both qϕ(zt∣ht,at,xt+1)1 and qϕ(zt∣ht,at,xt+1)2 develop spatially localized tuning at multiple scales — subtrees, branches, leaves, and specific paths — suggesting that population-level predictive geometry coexists with classical place-like tuning, though the authors explicitly decline to interpret these units as direct neural analogues.
Limitations and open questions
Several limitations constrain interpretation. The behavioral policy captures only a simplified exploration–exploitation tradeoff, omitting drives such as novelty seeking, memory, fatigue, and motivational shifts that shape natural behavior. The mouse analysis relies solely on behavioral trajectories without simultaneous neural recordings, so the claim that exploratory experience organizes biological predictive representations remains a model-based prediction rather than an empirical demonstration. The framework also operates on abstract node-level states; whether broad sampling organizes representations of visually grounded, sensory-rich observations is untested. Finally, the deterministic maze and fixed reward location simplify the credit-assignment problem relative to natural environments.
These constraints define concrete open questions: does the latent-geometry–behavior relationship hold in neural population activity recorded during navigation, and how does representational organization develop over learning relative to known hippocampal remapping dynamics?
Conclusion
This paper demonstrates that behavioral sampling regime, held against a fixed predictive objective, is sufficient to determine the organization of learned predictive representations. Exploration yields spatially aligned, transition-consistent, smoothly traversable latent maps; exploitation narrows experience and disrupts this organization, with reward-seeking performance peaking at an intermediate balance rather than under maximal exploitation. The replication of this pattern in models trained on mouse trajectories links artificial-agent findings to animal behavior and motivates future tests against neural recordings during navigation.