Grassmannian Dictionary Learning
- Grassmannian dictionary learning is a framework that leverages the geometry of subspaces to perform sparse coding and dictionary updates for data modeled as subspaces.
- It employs extrinsic embedding and Riemannian optimization to solve convex and non-convex problems, enabling robust performance in computer vision and dynamical system tasks.
- Kernelization extends this approach to nonlinear high-dimensional regimes, significantly improving accuracy in applications such as face recognition and dynamic texture analysis.
Grassmannian dictionary learning refers to a suite of algorithms and frameworks for structured, sparse representation and selection of subspace-based features, leveraging the geometry of the Grassmann manifold. This approach enables dictionary learning and sparse coding over spaces of linear subspaces, moving beyond vectorized representations to handle data elements best modeled by subspaces, such as image sets, video clips, or systems of observables in dynamical systems. The mathematical and algorithmic core exploits extrinsic embeddings, Riemannian geometry, and both convex and non-convex optimization over the Grassmannian, with kernelization and manifold optimization techniques for nonlinear or high-dimensional regimes (Harandi et al., 2014, Harandi et al., 2013, Schurig et al., 10 Nov 2025).
1. Grassmannian Geometry and Projection Embeddings
The Grassmann manifold is the set of all -dimensional linear subspaces of , formally realized as the quotient , where is the Stiefel manifold of orthonormal frames. Each point in the Grassmannian corresponds to a linear subspace, which can be represented extrinsically by any orthonormal basis or inherently by its orthogonal projection operator .
The projection map , , smoothly embeds the Grassmann manifold into the space of rank-0 symmetric, idempotent matrices, denoted 1. This embedding is isometric for the natural Riemannian metric, inducing the chordal distance
2
which is computationally tractable and coincides (up to constant scale) with projection distances derived from principal angles between subspaces (Harandi et al., 2014, Harandi et al., 2013). This extrinsic realization is central to enabling linear-algebraic manipulations and optimization in ambient matrix spaces.
2. Extrinsic Formulation of Sparse Coding and Dictionary Learning
Given a data set of Grassmann points (subspaces) 3, a dictionary 4 represented by atoms 5 is learned jointly with sparse coefficient vectors to best approximate each subspace in terms of dictionary atoms: 6 Coding is performed in the space of projections: each 7 is approximated as a linear combination of 8, penalized for sparsity. The Gram matrix 9 captures inner products of dictionary atoms, and data vectors 0 encode affinities between query and atoms. The coding (gSC) step minimizes, for each 1,
2
a convex Lasso problem (Harandi et al., 2014, Harandi et al., 2013).
Dictionary update leverages the following result: for each atom 3, form the aggregate matrix
4
and set 5 to the top 6 eigenvectors of 7, ensuring orthonormality and returning to the Grassmannian constraint. This closed-form update is a direct consequence of minimizing Frobenius distance to the convex set 8 and follows from the closest-point theorem (Harandi et al., 2014, Harandi et al., 2013).
3. Kernelization and Extensions to Nonlinear Regimes
To address data nonlinearity and high-dimensionality, the Grassmannian dictionary learning framework incorporates a kernelized extension. A feature map 9 and an associated kernel 0 are introduced. Subspaces are lifted into a reproducing kernel Hilbert space, with each lifted basis 1, where 2, 3 arise from the eigenstructure of the kernel Gram matrices 4. All coding and dictionary update operations become manipulations of kernel Gram matrices and small eigenproblems, effectively enabling the entire sparse coding and dictionary learning pipeline in a nonlinear ambient (Harandi et al., 2014, Harandi et al., 2013).
4. Riemannian Optimization and Grassmannian Dictionary Shaping
Recent developments include the use of full Riemannian optimization frameworks on the Grassmannian for dictionary selection, particularly where the representation problem is formulated as finding the best 5-dimensional subspace within a large candidate dictionary. The optimization is performed directly on the Grassmannian 6, using the quotient structure 7 with constraints imposed by orthogonality.
For example, in the context of Koopman operator theory and extended DMD, dictionary shaping is reduced to minimizing errors in approximate invariance and long-term prediction: 8 where 9 parametrizes the choice of observables, and 0 are predictions propagated through a Koopman-based reduced model. The Riemannian gradient is given analytically as 1, and optimization is executed via trust-region or other manifold-specific strategies, with tangent-projection, retraction, and vector transport defined accordingly. These techniques admit regularization and adaptation for invariance or additional constraints (Schurig et al., 10 Nov 2025).
5. Algorithm Summaries and Complexity
A canonical Grassmannian dictionary learning iteration alternates between sparse coding and atom update:
- Coding step (gSC): For each query subspace, solve the convex Lasso or quadratic program in coefficient space, using precomputed Gram matrices and inner products.
- Atom update: For each dictionary atom, aggregate contribution matrices as above, compute the symmetric sum, and set to the top 2 eigenvectors.
- Kernelized versions: Replace all projections and inner products with kernel-lifted analogues. Atom updates require generalized eigenproblems in kernel Gram spaces.
The computational complexity per outer iteration is dominated by Gram matrix computation (3), sparse coding solvers (4 total per dataset), and 5 eigen-decompositions of 6 matrices (7). Kernelization shifts the cost to kernel computation and SVDs in the reduced feature spaces (Harandi et al., 2014, Harandi et al., 2013).
6. Empirical Results and Comparative Performance
Empirical evaluations demonstrate substantial improvements of Grassmannian dictionary learning and their kernelized extensions over previous state-of-the-art approaches (e.g., canonical correlation–based, kernelized affine hull, and graph-embedding solvers). The canonical studies report:
| Dataset | Best Competing Method | GDL/gSC | KGDL/kgSC |
|---|---|---|---|
| YouTube Face Recognition | KAHM: 67.5% | 70.5% | 73.9% |
| Dynamic Textures (DynTex++) | GGDA: 84.1% | 90.3% | 92.8% |
| Ballet Actions | GGDA: 73.5% | 79.6% | 83.5% |
Across all tasks, extrinsic Grassmannian coding and dictionary learning yield consistent gains of 5–10 percentage points in discrimination accuracy compared to prior intrinsic or tangent-space methods. Kernelization further improves accuracy and enables modeling of complex, nonlinear relationships (Harandi et al., 2014, Harandi et al., 2013).
In dynamic systems application, shaping Koopman dictionaries via Grassmannian optimization leads to reduced projection errors (up to 30% reduction in invariance error), improved long-term prediction RMSE (by 20–50% over standard EDMD), and interpretable, sparse subspace selection with significant gains in forecasting efficiency (Schurig et al., 10 Nov 2025).
7. Applications and Theoretical Implications
Grassmannian dictionary learning is central to subspace modeling in computer vision, machine learning, and operator-theoretic system identification. Its deployment spans tasks such as image set classification, video-based recognition, dynamic texture modeling, and efficient compression of large observable libraries in dynamical systems. The mathematical foundation—reliance on projection geometry, isometric embedding, and Riemannian manifold optimization—makes these methods robust to non-Euclidean data structure and adaptable through kernelization.
A plausible implication is that future research will further integrate Grassmannian dictionary learning with deep architectures and scalable kernel frameworks, expanding to high-dimensional sensor data and online manifold-structured learning (Harandi et al., 2014, Harandi et al., 2013, Schurig et al., 10 Nov 2025).