---
title: Grassmannian Dictionary Learning
url: https://www.emergentmind.com/topics/grassmannian-dictionary-learning
type: topic
---

# Grassmannian Dictionary Learning

Grassmannian dictionary learning refers to a suite of algorithms and frameworks for structured, sparse representation and selection of subspace-based features, leveraging the geometry of the Grassmann manifold. This approach enables dictionary learning and sparse coding over spaces of linear subspaces, moving beyond vectorized representations to handle data elements best modeled by subspaces, such as image sets, video clips, or systems of observables in dynamical systems. The mathematical and algorithmic core exploits extrinsic embeddings, Riemannian geometry, and both convex and non-convex optimization over the Grassmannian, with kernelization and manifold optimization techniques for nonlinear or high-dimensional regimes [1401.8126][1310.4891][2511.07234].

## 1. Grassmannian Geometry and Projection Embeddings

The Grassmann manifold $\mathrm{Gr}(p,d)$ is the set of all $p$-dimensional linear subspaces of $\mathbb R^d$, formally realized as the quotient $\mathrm{St}(p,d)/O(p)$, where $\mathrm{St}(p,d)$ is the Stiefel manifold of orthonormal $d \times p$ frames. Each point in the Grassmannian corresponds to a linear subspace, which can be represented extrinsically by any orthonormal basis $X \in \mathrm{St}(p,d)$ or inherently by its orthogonal projection operator $P_X = X X^\top$.

The projection map $\Phi : \mathrm{Gr}(p,d) \to \mathrm{Sym}(d)$, $\Phi(\mathcal X) = X X^\top$, smoothly embeds the Grassmann manifold into the space of rank-$p$ symmetric, idempotent matrices, denoted $PG(p,d)$. This embedding is isometric for the natural Riemannian metric, inducing the chordal distance
\[
d_c(\mathcal X_i, \mathcal X_j) = \|\Phi(\mathcal X_i) - \Phi(\mathcal X_j)\|_F = \|X_i X_i^\top - X_j X_j^\top\|_F
\]
which is computationally tractable and coincides (up to constant scale) with projection distances derived from principal angles between subspaces [1401.8126][1310.4891]. This extrinsic realization is central to enabling linear-algebraic manipulations and optimization in ambient matrix spaces.

## 2. Extrinsic Formulation of Sparse Coding and Dictionary Learning

Given a data set of Grassmann points (subspaces) $\{\mathcal X_i\}_{i=1}^m$, a dictionary $\{\mathcal D_j\}_{j=1}^N$ represented by atoms $D_j \in \mathbb R^{d \times p}$ is learned jointly with sparse coefficient vectors to best approximate each subspace in terms of dictionary atoms:
\[
\min_{\{y_i\},\{D_j\}} \sum_{i=1}^m \left\| X_i X_i^\top - \sum_{j=1}^N y_{ij} D_j D_j^\top \right\|_F^2 + \lambda \sum_{i=1}^m \|y_i\|_1.
\]
Coding is performed in the space of projections: each $X_i X_i^\top$ is approximated as a linear combination of $D_j D_j^\top$, penalized for sparsity. The Gram matrix $\mathbb K_{jk} = \|D_j^\top D_k\|_F^2$ captures inner products of dictionary atoms, and data vectors $k_{i,j} = \|X_i^\top D_j\|_F^2$ encode affinities between query and atoms. The coding (gSC) step minimizes, for each $i$,
\[
y_i^\top \mathbb K y_i - 2 y_i^\top k_i + \lambda\|y_i\|_1,
\]
a convex Lasso problem [1401.8126][1310.4891].

Dictionary update leverages the following result: for each atom $D_r$, form the aggregate matrix
\[
S_r = \sum_{i=1}^m y_{ir} \left(X_i X_i^\top - \sum_{j \ne r} y_{ij} D_j D_j^\top\right)
\]
and set $D_r$ to the top $p$ eigenvectors of $S_r$, ensuring orthonormality and returning to the Grassmannian constraint. This closed-form update is a direct consequence of minimizing Frobenius distance to the convex set $PG(p,d)$ and follows from the closest-point theorem [1401.8126][1310.4891].

## 3. Kernelization and Extensions to Nonlinear Regimes

To address data nonlinearity and high-dimensionality, the Grassmannian dictionary learning framework incorporates a kernelized extension. A feature map $\phi: \mathbb R^d \to \mathcal H$ and an associated kernel $k(x, x') = \langle \phi(x), \phi(x') \rangle$ are introduced. Subspaces are lifted into a reproducing kernel Hilbert space, with each lifted basis $\Psi(X) = \Phi(X) U \Sigma^{-1/2}$, where $U$, $\Sigma$ arise from the eigenstructure of the kernel Gram matrices $K_{ij} = k(x_i, x_j)$. All coding and dictionary update operations become manipulations of kernel Gram matrices and small eigenproblems, effectively enabling the entire sparse coding and dictionary learning pipeline in a nonlinear ambient [1401.8126][1310.4891].

## 4. Riemannian Optimization and Grassmannian Dictionary Shaping

Recent developments include the use of full Riemannian optimization frameworks on the Grassmannian for dictionary selection, particularly where the representation problem is formulated as finding the best $p$-dimensional subspace within a large candidate dictionary. The optimization is performed directly on the Grassmannian $Gr(p, M)$, using the quotient structure $St(p, M)/O(p)$ with constraints imposed by orthogonality.

For example, in the context of Koopman operator theory and extended DMD, dictionary shaping is reduced to minimizing errors in approximate invariance and long-term prediction:
\[
g_N([U]) = \frac{1}{2JN} \sum_{i=1}^J \sum_{k=1}^N \| x(k; \bar{x}_i) - \hat{x}(k; \bar{x}_i; U) \|_2^2,
\]
where $U \in St(p, M)$ parametrizes the choice of observables, and $\hat{x}(k; \cdot)$ are predictions propagated through a Koopman-based reduced model. The Riemannian gradient is given analytically as $(I - U U^\top) \nabla f(U)$, and optimization is executed via trust-region or other manifold-specific strategies, with tangent-projection, retraction, and vector transport defined accordingly. These techniques admit regularization and adaptation for invariance or additional constraints [2511.07234].

## 5. Algorithm Summaries and Complexity

A canonical Grassmannian dictionary learning iteration alternates between sparse coding and atom update:

- **Coding step (gSC):** For each query subspace, solve the convex Lasso or quadratic program in coefficient space, using precomputed Gram matrices and inner products.
- **Atom update:** For each dictionary atom, aggregate contribution matrices as above, compute the symmetric sum, and set to the top $p$ eigenvectors.
- **Kernelized versions:** Replace all projections and inner products with kernel-lifted analogues. Atom updates require generalized eigenproblems in kernel Gram spaces.

The computational complexity per outer iteration is dominated by Gram matrix computation ($O(K^2 n p)$), sparse coding solvers ($O(N K^2)$ total per dataset), and $K$ eigen-decompositions of $n \times n$ matrices ($O(K n^3)$). Kernelization shifts the cost to kernel computation and SVDs in the reduced feature spaces [1401.8126][1310.4891].

## 6. Empirical Results and Comparative Performance

Empirical evaluations demonstrate substantial improvements of Grassmannian dictionary learning and their kernelized extensions over previous state-of-the-art approaches (e.g., canonical correlation–based, kernelized affine hull, and graph-embedding solvers). The canonical studies report:

| Dataset                    | Best Competing Method        | GDL/gSC | KGDL/kgSC |
|----------------------------|-----------------------------|---------|-----------|
| YouTube Face Recognition   | KAHM: 67.5%                 | 70.5%   | 73.9%     |
| Dynamic Textures (DynTex++)| GGDA: 84.1%                 | 90.3%   | 92.8%     |
| Ballet Actions             | GGDA: 73.5%                 | 79.6%   | 83.5%     |

Across all tasks, extrinsic Grassmannian coding and dictionary learning yield consistent gains of 5–10 percentage points in discrimination accuracy compared to prior intrinsic or tangent-space methods. Kernelization further improves accuracy and enables modeling of complex, nonlinear relationships [1401.8126][1310.4891].

In dynamic systems application, shaping Koopman dictionaries via Grassmannian optimization leads to reduced projection errors (up to 30% reduction in invariance error), improved long-term prediction RMSE (by 20–50% over standard EDMD), and interpretable, sparse subspace selection with significant gains in forecasting efficiency [2511.07234].

## 7. Applications and Theoretical Implications

Grassmannian dictionary learning is central to subspace modeling in computer vision, machine learning, and operator-theoretic system identification. Its deployment spans tasks such as image set classification, video-based recognition, dynamic texture modeling, and efficient compression of large observable libraries in dynamical systems. The mathematical foundation—reliance on projection geometry, isometric embedding, and Riemannian manifold optimization—makes these methods robust to non-Euclidean data structure and adaptable through kernelization.

A plausible implication is that future research will further integrate Grassmannian dictionary learning with deep architectures and scalable kernel frameworks, expanding to high-dimensional sensor data and online manifold-structured learning [1401.8126][1310.4891][2511.07234].

Source: https://www.emergentmind.com/topics/grassmannian-dictionary-learning