- The paper shows that geometric compression (tight class clustering) and information-theoretic compression (low mutual information) are negatively and nonlinearly correlated.
- It introduces robust methodologies using Conditional Entropy Bottleneck and Gaussian dropout paradigms to measure MI and cluster tightness across various architectures.
- The study reveals that neither MI nor geometric compression alone reliably predicts generalization, urging cautious application of these metrics in model evaluation.
Introduction
In deep learning, the latent representations formed by DNNs are foundational for downstream task performance and generalization. Two principal perspectives for characterizing these representations are the geometric structure—mostly related to phenomena such as neural collapse and the clustering of class-specific encodings—and information theory, especially mutual information (MI) between network inputs and their latent representations. The central question addressed in "Geometric and Information Compression of Representations in Deep Learning" (2606.21593) is whether tight geometric compression (e.g., highly clustered, linearly separable representations) is necessarily linked with information-theoretical compression (low MI between input and representation), and whether this relationship is monotonic or robust to changes in training regimes.
Methodological Framework
The paper advances the empirical and theoretical study of the interplay between geometric and information-theoretic compression using a broad suite of architectures and regularization paradigms. Specifically, two DNN variants are considered: networks trained with the Conditional Entropy Bottleneck (CEB)—incorporating data-dependent additive noise in the latent space—and networks regularized with continuous multiplicative Gaussian dropout. For both paradigms, information-theoretic compression is measured via MI estimates: the variational bound intrinsic to the CEB loss for CEB models, and the Difference of Entropies (DoE) estimator for Gaussian dropout models.
Geometric compression is quantified by the neural collapse (NC1) criterion, which measures the ratio of within-class cluster variance to between-class centroid distances. By systematically sweeping regularization strengths (the β parameter for CEB, and explicit NC regularization for dropout models), the authors generate a diversity of models spanning ranges of MI and geometric tightness, permitting robust correlation analyses across settings.
A notable theoretical contribution is the extension of finite-MI guarantees, originally proven for ReLU-activated networks, to all deterministic DNNs with real-analytic activation functions under continuous dropout. This supports information-theoretic analyses for any modern architecture (e.g., those using GELU, sigmoid, or tanh activations).
Experimental Results
Across approximately 500 DNNs with varying depth, architecture, and training regimes—including fully connected MLPs, vision CNNs (LeNet, VGG, ResNet, DenseNet, WideResNets), and transformers (miniBERT) on image and text datasets—the central empirical finding is that geometric and information-theoretic compression are negatively and nonlinearly correlated at the end of training. Lower MI (i.e., more information-theoretic compression) is not reliably indicative of tighter geometric clustering (lower NC), and vice versa. The negative association deviates from the often-assumed direct linkage posited in prior studies.
The relationship is not invariant: changing regularization strength or the degree of over/underfitting can in some cases reverse the sign of the correlation between MI and NC. For example, in some hyperparameter regimes, models with superior generalization (lower train/test accuracy gap) exhibit nearly uncorrelated MI and NC. These findings are consistent across both MI estimation approaches (variational and DoE), affirming the robustness of the observation against estimator-induced artifacts.
Empirical evidence is further supported by a theoretical toy model: low MI between inputs and representations can emerge either from strong geometric compression (tight class clusters) or from high encoder noise that destroys input information without producing meaningful clusters. Only the former produces small NC values; thus, low MI and geometric compression can decorrelate.
Numerically, the rank correlations between MI and NC are consistently negative in CEB-trained models (train: -0.75, test: -0.83), and somewhat weaker but still negative in dropout-regularized models (train: -0.46, test: -0.24). Most importantly, the sign and nonlinearity of the relationship can flip under modified generalization behavior, revealing that neither MI nor NC is a universal marker for network generalization or representation quality.
Theoretical and Practical Implications
The paper directly contradicts the widespread assumption, motivated by both IB theory and empirical monitoring of training trajectories, that geometric and information-theoretic compressor are tightly aligned. Not only are the two measures not interchangeable proxies for latent representation quality, but their relationship is contingent on training, regularization, and architecture.
Practically, this undermines the frequent use of cluster-based or MI-based criteria as general indicators of good generalization, representation compactness, or overfitting avoidance. The theoretical extension of finite-MI guarantees to analytic activations further broadens the applicability of information-theoretic tools to any continuous-activation architecture.
From a methodological standpoint, the study highlights the importance of post-training analysis—focusing on the end-state properties of the learned representations—over dynamic tracking of training epochs, which can obscure subtle confounders, particularly generalization.
The findings imply that generalization might be the confounder driving the observed associations between geometry and information, rather than one being the consequence of the other. This resonates with Reichenbach’s principle of common cause and motivates future work on disentangling causal pathways via interventional or structural equation modeling analyses.
Conclusion
"Geometric and Information Compression of Representations in Deep Learning" provides a comprehensive, empirically and theoretically rigorous study refuting the straightforward equivalence between information and geometric compression in DNN representations. The observed negative and non-monotonic relationships call for caution in relying on either measure alone for diagnosing, optimizing, or claiming insight into network generalization or representation learning. Further, the work offers mathematical tools for future studies into the causal structure of generalization, geometry, and information in deep learning, anticipating advances in both representation analysis and DNN regularization design (2606.21593).