---
title: Low-Rank Hankel Tensor Completion
url: https://www.emergentmind.com/topics/low-rank-hankel-tensor-completion
type: topic
---

# Low-Rank Hankel Tensor Completion

Searching arXiv for recent and foundational papers on low-rank Hankel tensor completion.

Low-rank Hankel tensor completion denotes a family of recovery methods in which a partially observed signal, matrix, or tensor is lifted into a Hankel-structured higher-order tensor and reconstructed under a low-rank prior adapted to that lifted object. In the literature, this framework appears in at least three distinct but closely related forms: recovery of \(N\)-dimensional exponential signals via CP structure with Hankel-matrix nuclear-norm regularization on factor vectors; spatiotemporal traffic speed estimation via a fourth-order Hankel tensor and a truncated nuclear norm on a balanced unfolding; and multi-measurement spectral compressed sensing via a third-order Hankel tensor with low multilinear-rank Tucker structure and scaled gradient descent. Across these formulations, the central premise is that Hankelization exposes shift-invariant local structure, while completion exploits low-rankness in a model-specific sense [1604.02100], [2105.11335], [2507.04847].

## 1. Structural model and Hankel lifting

In the exponential-signal setting, the observed tensor \(\mathcal Y\in\mathbb C^{I_1\times I_2\times\cdots\times I_N}\) is modeled by
\[
y_{i_1,\ldots,i_N}=\sum_{r=1}^{R} d_r \prod_{n=1}^{N} z_{n,r}^{\,i_n-1},
\]
with only a subset of entries indexed by
\[
\Omega\subset\{1\ldots I_1\}\times\{1\ldots I_2\}\times\cdots\times\{1\ldots I_N\}
\]
available through the sampling operator \(\mathcal P_\Omega\). The reconstruction target is the full tensor \(\mathcal Y\), and the formulation explicitly exploits both low CP rank and exponential factor vectors [1604.02100].

In the traffic-speed setting, the starting point is an incomplete matrix \(\boldsymbol Y\in\mathbb R^{N\times T}\). A two-way delay embedding
\[
\mathcal H_{\tau_s,\tau_t}:\mathbb R^{N\times T}\longrightarrow \mathbb R^{\tau_s\times\tau_t\times (N-\tau_s+1)\times (T-\tau_t+1)}
\]
is defined by
\[
\bigl[\mathcal H_{\tau_s,\tau_t}(Z)\bigr]_{a,b,i,j}=Z_{i+a-1,\;j+b-1},
\]
thereby converting the original matrix into a fourth-order Hankel tensor \(X\). Its inverse \(\mathcal H_{\tau_s,\tau_t}^{-1}\) averages all copies of each original entry. This construction is used to approximate and characterize both global and local spatiotemporal patterns in a data-driven manner [2105.11335].

In the multi-measurement spectral compressed sensing setting, \(s\) equally sampled length-\(n\) time series are arranged as \(X_\star\in\mathbb C^{s\times n}\), and a third-order Hankel-tensor lifting
\[
\mathcal H:\mathbb C^{s\times n}\longrightarrow\mathbb C^{n_1\times n_2\times s},\qquad
[\mathcal H(X)](i_1,i_2,i_3)=X(i_3,\;i_1+i_2)
\]
produces a block-Hankel tensor. For spectral sparse signals, the lifted tensor admits a Tucker form with
\[
\mathrm{mulrank}\big(\mathcal H(X_\star)\big)=(r,r,r)\ll (n_1,n_2,s),
\]
so Hankel lifting converts spectral sparsity into low multilinear rank [2507.04847].

## 2. Rank models and optimization formulations

The phrase “low-rank” is not uniform across this literature. The formulations differ in the rank surrogate, in the object to which that surrogate is applied, and in the way the Hankel structure is enforced.

| Setting | Lifted object | Rank surrogate / structure |
|---|---|---|
| \(N\)-dimensional exponential signals | Factor-vector Hankel matrices and CP tensor | Nuclear norm of \(\mathcal R(Q_rU^{(n)})\) plus CP decomposition |
| Traffic speed estimation | Fourth-order Hankel tensor with balanced unfolding \(X_\square\) | Truncated nuclear norm of \(X_\square\) |
| Multi-measurement spectral compressed sensing | Third-order block-Hankel tensor \(\mathcal T\) | Low multilinear rank \((r,r,r)\) with Tucker decomposition |

For \(N\)-dimensional exponential signals, the reconstruction variable is written as
\[
\mathcal X=\sum_{r=1}^{\hat R}u_r^{(1)}\odot u_r^{(2)}\odot\cdots\odot u_r^{(N)},
\]
where \(\hat R\ge R\) is an overestimate of the true rank. The operator \(\mathcal R\) maps a vector into an \(S_1\times S_2\) Hankel matrix with \(S_1+S_2=I_n+1\), and the model minimizes
\[
\min_{\{U^{(n)}\}}
\sum_{r=1}^{\hat R}\sum_{n=1}^{N}\|\mathcal R(Q_rU^{(n)})\|_*
+\frac{\lambda}{2}\left\|\mathcal P_\Omega\!\left([[U^{(1)},\ldots,U^{(N)}]]-\mathcal Y\right)\right\|_F^2.
\]
This couples a CP decomposition with the nuclear norm of Hankel matrices so that low-CP-rank structure and exponential structure are enforced simultaneously [1604.02100].

For traffic speed estimation, the model solves
\[
\min_{X,Z}\;\|X_\square\|_{*,r}
\quad\text{s.t.}\quad
X=\mathcal H_{\tau_s,\tau_t}(Z),\qquad
Z_{i,j}=Y_{i,j}\ \ \forall (i,j)\in\Omega.
\]
Here \(X_\square\) is the balanced spatiotemporal unfolding of the fourth-order Hankel tensor, with
\[
p=\tau_s\tau_t,\qquad q=(N-\tau_s+1)(T-\tau_t+1),
\]
and the truncated nuclear norm is used to approximate the tensor rank. Each column of the unfolding represents the vectorization of a small patch in the original matrix [2105.11335].

For multi-measurement spectral compressed sensing, the constrained completion problem is
\[
\min_{\mathcal T\in\mathbb C^{n_1\times n_2\times s}}
\frac{1}{2}\left\|P_\Omega\bigl(G^*(\mathcal T)\bigr)-Y_\Omega\right\|_F^2
\]
subject to
\[
\mathrm{mulrank}(\mathcal T)=(r,r,r),\qquad (I-GG^*)\mathcal T=0.
\]
Equivalently, one may add a quadratic penalty for \((I-GG^*)\mathcal T\). The weighting operator \(D\) is chosen so that \(G\equiv \mathcal H\circ D^{-1}\) satisfies \(G^*G=I\), and the Hankel tensor is parametrized by a Tucker decomposition \((U_1,U_2,U_3)\bcdot S\) [2507.04847].

## 3. Numerical algorithms and computational structure

The 2016 HMRTC method solves a nonconvex and nonsmooth objective by ADMM after introducing auxiliary variables \(Z_r^{(n)}=\mathcal R(Q_rU^{(n)})\). The augmented Lagrangian uses multipliers \(D_r^{(n)}\) and penalty \(\beta\). The \(U\)-update fixes \(Z^{(n)}\) and \(D^{(n)}\), then minimizes the augmented Lagrangian; because the CP term is multilinear, an inner alternating minimization among the modes is applied, and the mode-\(n\) update reduces to a regularized least-squares system. The \(Z\)-update is a singular-value-shrinkage step,
\[
Z_{r;k+1}^{(n)}=S_{1/\beta}\!\left(\mathcal R(Q_rU_{k+1}^{(n)})+\frac{1}{\beta}D_{r;k}^{(n)}\right),
\]
followed by the multiplier update
\[
D_{r;k+1}^{(n)}=D_{r;k}^{(n)}+\beta\bigl[\mathcal R(Q_rU_{k+1}^{(n)})-Z_{r;k+1}^{(n)}\bigr].
\]
The penalty schedule starts at \(\beta_0=0.1\) and increases by \(\beta_{k+1}=\rho\beta_k\) with \(\rho\in(1,1.1]\), for example \(\rho=1.05\). The stopping criteria are relative change in reconstructed tensor \(<10^{-4}\) or \(k>10^3\). Per iteration, SVDs dominate the cost: one SVD of size \(\approx (I/2)\times (I/2)\) per \((r,n)\), i.e. \(O(I^3)\), for total \(O(N\hat RI^3)\); these SVDs are independent and can be parallelized, and storage is \(O(N\hat RI)\) [1604.02100].

The traffic-speed STH-LRTC framework also uses ADMM, but the variables are the lifted Hankel tensor \(X\), the completed matrix \(Z\), and a dual variable \(E\). The \(X\)-update applies the truncated singular-value shrinkage operator \(\mathcal D_{1/\rho^\ell}\) to the balanced unfolding \(A^\ell_\square\), where
\[
A^\ell=\mathcal H_{\tau_s,\tau_t}(Z^\ell)-\frac{1}{\rho^\ell}E^\ell.
\]
The \(Z\)-update uses \(\mathcal H_{\tau_s,\tau_t}^{-1}\) together with a hard constraint on \(\Omega\), the dual update is
\[
E^{\ell+1}=E^\ell+\rho^\ell\bigl[X^{\ell+1}-\mathcal H_{\tau_s,\tau_t}(Z^{\ell+1})\bigr],
\]
and the adaptive penalty rule is
\[
\rho^{\ell+1}=\min\{\beta\rho^\ell,\rho_{\max}\},\qquad \beta\in[1,1.2].
\]
The stopping criterion is \(\|Z^{\ell+1}-Z^\ell\|_F/\|Y_\Omega\|_F<\epsilon\). Each \(X\)-update requires an SVD of a \(p\times q\) unfolding with \(p=\tau_s\tau_t\ll q\), so a partial SVD costs \(\mathcal O(\min(p,q)\,p\,q)\), while the \(Z\)- and \(E\)-updates cost only \(\mathcal O(NT)\) [2105.11335].

The 2025 ScalHT method replaces ADMM by scaled gradient descent on Tucker factors and core. At iterate \(F^k=(U_1^k,U_2^k,U_3^k,S^k)\), each factor update is premultiplied by the inverse Gram matrix of the other factors, and the core update is scaled by \(\big((U_1^k)^HU_1^k\big)^{-1}\), \(\big((U_2^k)^HU_2^k\big)^{-1}\), and \(\big((U_3^k)^HU_3^k\big)^{-1}\). The paper states that this scaling dramatically reduces sensitivity to condition-number in practice. Fast identities exploiting the interaction between the low-rank Tucker structure and the Hankel lift reduce the overall per-iteration cost to
\[
O\bigl((s+n)r^2+n r^2\log n\bigr),
\]
with FFT-based fast convolution for single-mode multiplications. Initialization is “sequential spectral,” based on a back-projected estimate on a small batch \(\Omega_0\), top-\(r\) singular vectors of mode unfoldings, a projected core \(S^0\), and projection of \((U_1^0,U_2^0)\) to the incoherence ball [2507.04847].

## 4. Convergence theory and recovery guarantees

Theoretical guarantees differ sharply across formulations. In HMRTC, despite nonconvexity, two convergence statements are established: Theorem 1 states that the sequence \(\{U_k\}\) generated by the ADMM is Cauchy, hence convergent; Theorem 2 states that if the multiplier increments \(\|D_{k+1}-D_k\|_F\to 0\), then any limit point satisfies the Karush–Kuhn–Tucker conditions of the original problem. No global exact-recovery or sample-complexity bounds are given for the \(N\)-dimensional case, although the paper refers to existing \(1\)-D Hankel-matrix completion results predicting sample complexity \(O(R\log I)\) and to EMaC results [1604.02100].

For STH-LRTC, the theory reported is more generic: ADMM is known to converge, under mild conditions, at roughly a sublinear rate, and in practice an adaptive \(\rho\) often accelerates convergence. The principal methodological parameters are the spatial and temporal window lengths \(\tau_s,\tau_t\), which control the size of local patches. Larger values capture longer spatiotemporal dependencies, which is useful when data is very sparse, at the expense of higher SVD cost [2105.11335].

ScalHT provides the strongest formal guarantees among the three formulations. Under an incoherence assumption on the HOSVD factors and with step-size \(0<\eta\le 0.4\), Theorem 1 states that if
\[
m\gtrsim \varepsilon_0^{-2}\,\mu_0\,c_s\,s\,r^3\,\kappa^2\,\log(sn),
\]
then, with probability \(\ge 1-O((sn)^{-2})\), the iterates satisfy
\[
\|X^k-X_\star\|_F\le 3\,\varepsilon_0\,(1-0.5\eta)^k\,
\sigma_{\min}\!\bigl(\mathcal H(X_\star)\bigr).
\]
The paper further states exact recovery in \(k=O(\log(1/\epsilon_0))\) steps. Lemma 7 gives linear convergence once the iterate lies within a small constant relative neighborhood of the ground truth:
\[
\mathrm{dist}(F^{k+1},F_\star)\le (1-0.3\eta)\,\mathrm{dist}(F^k,F_\star).
\]
The proof relies on a new concentration bound for partial sampling through the Hankel adjoint \(G^*\), a sequential spectral initialization bound, and perturbation bounds for the Gram scaling. The authors describe these recovery and linear convergence guarantees as the first of their kind for low-rank Hankel tensor completion [2507.04847].

## 5. Empirical behavior and application domains

The empirical record in the 2016 paper is centered on simulated and real spectroscopy data. On simulated \(50\times 50\times 50\) three-dimensional signals, including both undamped and damped complex sinusoids, with rank \(R\) varying up to \(\sim 60\), sampling ratios \(SR\) from \(0.1\) to \(0.9\), and additive noise, HMRTC yields much lower RLNE than competing low-n-rank (ADM–TR) or weighted-CP (WCP) methods, especially at low \(SR\). Phase-transition plots of average RLNE show that HMRTC recovers accurately, with RLNE below the noise level, with as few as \(SR\approx 0.2\) for moderate \(R\), whereas the competing methods need \(SR>0.5\). HMRTC is also reported to be robust to over-estimation of rank: even if \(\hat R\) is \(2R\ldots 10R\), reconstruction remains good, whereas WCP fails when \(\hat R\gg R\). On real three-dimensional NMR spectroscopy (HNCO) data of size \(64\times 128\times 512\), nonuniformly Poisson-gap sampled at \(10\%\), HMRTC recovers peaks faithfully versus ADM–TR or WCP, enabling \(\sim 90\%\) time-saving. In a \(50\times 50\times 50\) benchmark, HMRTC runs in \(\sim 675\) s versus \(\sim 7395\) s for WCP and \(\sim 185\) s for ADM–TR, using \(12\) GB RAM [1604.02100].

The 2021 traffic-speed formulation is motivated by traffic state estimation using sparse observations from mobile sensors. The paper contrasts its approach with methods that rely on well-defined physical traffic flow models or require large amounts of simulation data to train machine learning models. The proposed framework is purely data-driven and model-free, involves only two hyperparameters, spatial and temporal window lengths, and numerical experiments on real-world high-resolution trajectory data demonstrate the effectiveness and superiority of the model in some challenging scenarios [2105.11335].

The 2025 ScalHT paper reports several numerical regimes. In phase-transition experiments with \(r=2\), \(s\in[16,48]\), and \(m\in[100,500]\), success occurs roughly when \(m\gtrsim 8s\); with \(s=32\) fixed, \(m\gtrsim 80r\) suffices. For \(n=63\), \(s=32\), \(r=8\), ScalHT, ScaledGD, and AM-FIHT achieve high success at \(p\gtrsim 0.5\), while ANM needs \(p\gtrsim 0.6\). For \(n=511\), \(s=512\), \(r=6\), ScalHT and ScaledGD reach relative error \(10^{-6}\) in \(\sim 60\) iterations in both \(p=0.17\) and \(p=0.22\) cases, whereas AM-FIHT takes \(\sim 90\) to \(140\) iterations. Across \(n=2^j-1\), ScalHT runs in \(O(n\log n)\) time and is reported at \(\approx O(1\!-\!2)\) sec for \(n=4095,s=256\). In DOA estimation using a \(17\)-element sparse linear array with \(n=64,s=32\), ScalHT+MUSIC achieves RMSE near the CRB at SNR \(>10\) dB, outperforming ANM+MUSIC and applying MUSIC directly to incomplete snapshots [2507.04847].

## 6. Conceptual distinctions, adjacent methods, and interpretation

A recurrent source of confusion is to treat low-rank Hankel tensor completion as a single model. The literature instead shows several structurally different formulations. One line exploits the CANDECOMP/PARAFAC structure together with the exponential structure of factor vectors by minimizing the nuclear norm of their Hankel matrices; another uses a balanced unfolding of a fourth-order spatiotemporal Hankel tensor and a truncated nuclear norm; a third imposes low multilinear rank on a third-order block-Hankel tensor through Tucker factors and a Hankel-consistency constraint or penalty [1604.02100], [2105.11335], [2507.04847].

The relation to adjacent methods is also explicit in the comparative baselines. In the spectroscopy setting, HMRTC is compared with ADM–TR and WCP. In traffic state estimation, the proposed method is positioned against physical traffic flow models and machine learning models trained on large amounts of simulation data. In multi-measurement spectral compressed sensing, ScalHT is compared with ANM, AM-FIHT, and ScaledGD. These comparisons indicate that the practical meaning of “Hankel tensor completion” depends on the ambient inverse problem, the lifted tensor order, the rank notion, and the optimization mechanism.

The theoretical trajectory across the cited works is equally nonuniform. Early HMRTC establishes convergence to a stationary point but does not provide global exact-recovery or \(N\)-dimensional sample-complexity results. STH-LRTC emphasizes tractable ADMM updates and patch-based spatiotemporal structure. ScalHT combines fast algebraic transforms with nonconvex Tucker factorization and provides recovery and linear convergence guarantees. This suggests a progression from convex or convex-surrogate regularization toward structured nonconvex parameterizations with sharper theory and lower computational cost, while preserving the defining role of Hankel lifting as the mechanism that exposes recoverable low-rank structure.

Source: https://www.emergentmind.com/topics/low-rank-hankel-tensor-completion