- The paper introduces a rigorous multi-level probing framework that extends evaluation beyond mere neural activity prediction.
- It details how variations in linear decodability, single-unit tuning, and population geometry reveal distinct latent computations.
- The study highlights that similar predictive accuracies can mask significant functional and representational differences among models.
Multi-Level Probing of Latent Representations in Mouse V1 Digital Twins
Overview and Motivation
The paper "Beyond Neural Activity Prediction: Probing Latent Representations in Mouse V1 Digital Twins" (2605.23122) systematically investigates the latent representational properties of high-capacity neural-response predictors trained on freely moving mouse V1 recordings. While digital twins are conventionally evaluated by their predictive accuracy on held-out neural responses, such metrics do not fully characterize the functional computations these models encode or the structure of their internal states. The authors address this gap by introducing a rigorous multi-level probing framework, interrogating digital twins at the levels of linear decodability, emergent single-unit tuning, and population geometry. This approach reveals how representational features covary with neural prediction across diverse visual encoders and exposes persistent differences between models with similar predictive fidelity.
Figure 1: Multi-level probing of hidden representations in V1 digital twins: comparing architectures, neural prediction, and latent-level analyses.
Methodological Framework
Model Architecture and Training
A suite of eight convolutional-recurrent architectures—spanning CNNs, ResNets, MobileNet-V2, EfficientNet-B0, AlexNet, SqueezeNet, ShuffleNet-V2, and RegNet-Y-400MF—were trained on large-scale, publicly available mouse V1 data [parker_joint_2022], encompassing grayscaled movies paired with single-neuron activity. The pipeline employs a fixed recurrent backbone (single-layer GRU, 512 units) with visual encoders varied across the model family, enabling controlled analysis of architectural effects. Neural activity prediction is optimized via Poisson negative log-likelihood, with mouse-specific readouts serving as output heads.
Multi-Level Probes
The evaluation protocol comprises three principal axes:
- Linear Decodability (Functional Access): Linear probes are trained on frozen latent states to recover task-relevant information (orientation, contrast, and motion) using synthetic visual stimuli, providing quantitative readout of linearly accessible features.
- Emergent Single-Unit Tuning: Latent-unit responses to canonical parametric gratings are characterized, extracting metrics of global orientation selectivity (gOSI), contrast semisaturation (C50​), preferred spatial frequency, and phase modulation (F1/F0).
- Population Geometry: PCA eigenspectra of hidden-layer activity under naturalistic stimuli are computed, with power-law exponent α quantifying representation dimensionality and its correspondence to biological V1 [stringer2019high].
Empirical Results
Linear Probe Analyses: Decodability of Visual Features
All architectures exhibited non-trivial linear probe performance on orientation and contrast tasks, with substantial variability in accessible task information across models. For the orientation discrimination task, leading architectures approached $0.776$ accuracy (CNN) versus $0.411$ (ShuffleNet)—far above chance, yet with marked differences across architectures. Contrast detection and motion-direction discrimination similarly yielded psychometric and coherence-dependent performance curves with distinct saturations and slopes.
Figure 2: Controlled visual probes (orientation, contrast, motion) and architecture-level task performances versus neural prediction.
Notably, probe performance correlated with neural prediction (dynamic contrast: r=0.88, orientation: r=0.79), but similar predictive scores could mask divergent probe outcomes (e.g., SqueezeNet and CNN with equivalent correlation but $0.16$ difference in orientation accuracy). This underscores inadequacy of prediction metrics for fully capturing functional structure.
Latent Unit Tuning: Emergence of Canonical Selectivity
Single-unit analyses across architectures revealed structured tuning to orientation, spatial frequency, phase, and contrast. Distributions of gOSI, preferred SF, and C50​ varied by model, echoing differences in biological V1 populations. Moderate gOSI and low preferred SF predominated, mirroring classical mouse values [niell2008]. Tasks and tuning metrics correlated (contrast-probe vs. C50​: r=0.83, motion-probe vs. gOSI: α0), but no single tuning index fully predicted probe performance.
Figure 3: Latent-unit tuning metrics (orientation selectivity, contrast, spatial frequency, phase) across encoder families and their relationships to probe accuracy.
Population Geometry: Eigenspectrum Analysis
Population-level eigenspectra exposed pronounced variation in representational dimensionality. While all tested models exhibited steeper α1 than biological V1, better predictors consistently showed lower α2, corresponding to flatter and more distributed representations. Association between α3 and neural prediction was robust (α4), suggesting superior predictors are more "V1-like" in global hidden geometry.
Figure 4: Eigenspectra of hidden-layer activations, with power-law fits and correlations between exponent α5 and neural prediction.
Interpretative Discussion
Underdetermination and Functional Consequences
The results highlight underdetermination intrinsic to digital twin modeling: high-capacity architectures can achieve similar neural prediction via distinct latent structures. This is critical for downstream scientific uses (in silico hypothesis generation, experimental stimulus design), as models with matched scores may diverge in functional access or in population-level computations. The systematic probing framework transforms this challenge into an empirical investigation, revealing axes where representational differences persist among well-predicting twins.
Population Geometry as a Distinctive Axis
Population geometry, measured via eigenspectrum exponent, stratifies models more sharply than unit tuning or linear probe metrics. Models nearing biological V1 geometry (flatter spectra, lower α6) are superior predictors, yet none match in vivo population statistics without explicit regularization [liscai2025beyond]. The gap in hidden geometry versus output-layer geometry observed in prior studies is thus extended to latent spaces, hinting at the potential for targeted architectural or objective modifications for achieving richer representational alignment.
Limitations and Future Directions
The analysis is restricted to convolutional-recurrent models trained on freely moving mouse V1 and does not encompass transformers or radically divergent frameworks. Generalizing multi-level probing to additional architectures, datasets, and objectives will be needed to ascertain universal trends in representational organization versus prediction fidelity. Layer-wise application of probes and error-profile analysis could inform selection of digital twins for specific scientific applications, refining in silico experimentation.
Conclusion
This paper establishes a rigorous multi-level representational probing methodology for V1 digital twins, demonstrating that models with equivalent neural-prediction accuracy can diverge substantially in linear probe performance, unit tuning, and population geometry. Prediction correlation scores are informative but insufficient as sole evaluative criteria. Direct assessment of latent state properties is essential for understanding what computations and structures are encoded, guiding the use of digital twins as scientific instruments and informing architectural choices for neurobiological alignment.