Papers
Topics
Authors
Recent
Search
2000 character limit reached

Tensor Nuclear Norm Overview

Updated 14 July 2026
  • Tensor nuclear norm is a generalization of the matrix nuclear norm defined via unit rank-one tensor decompositions and serves as the dual of the tensor spectral norm.
  • The t-product based tensor nuclear norm offers an exact convex envelope of tensor average rank, underpinning recovery theory in robust PCA and tensor completion.
  • Variants such as p-TNN, truncated TNN, and framelet-based TNN refine low-rank approximations, targeting specific data geometries and improving computational efficiency.

Tensor nuclear norm denotes several extensions of the matrix nuclear norm to multiway arrays. In the generic higher-order setting, it is the dual of the tensor spectral norm and can be written as the minimum total weight in a decomposition into unit rank-one tensors (Friedland et al., 2014). In the third-order tt-SVD literature, the term often refers to the norm induced by the tensor–tensor product, defined through a block-circulant embedding or, equivalently, by Fourier-domain frontal slices; within that framework it is the convex envelope of tensor average rank on the unit ball of the tensor spectral norm (Lu et al., 2018). The literature also contains several other tensor nuclear-norm constructions and surrogates, including truncated, transform-based, nonlinear, tensor-ring, and semidefinite-relaxation variants, reflecting the fact that low-rank structure in tensors is model-dependent (Xue et al., 2019).

1. General formulations and relation to matrices

For an order-dd tensor TRn1××ndT\in\mathbb R^{n_1\times\cdots\times n_d}, one standard definition is

T=min{i=1rλi:  T=i=1rλi(x1ixdi), xki2=1},\|T\|_*=\min\Bigl\{\sum_{i=1}^r|\lambda_i|:\;T=\sum_{i=1}^r\lambda_i\,(x_1^i\otimes\cdots\otimes x_d^i),\ \|x_k^i\|_2=1\Bigr\},

with dual spectral norm

Tσ=maxxk2=1T,x1xd.\|T\|_\sigma=\max_{\|x_k\|_2=1}\langle T,x_1\otimes\cdots\otimes x_d\rangle.

When d=2d=2, this reduces exactly to the matrix nuclear norm jσj(A)\sum_j \sigma_j(A) (Friedland et al., 2014). Qi et al. state the same decomposition-based definition for general tensors and emphasize the duality A=max{A,X:Xs1}\| \mathcal A\|_*=\max\{\langle \mathcal A,\mathcal X\rangle:\|\mathcal X\|_s\le 1\} (Qi et al., 2019).

This formulation preserves several matrix-like features, but not all. Every tensor has a nuclear norm attaining decomposition, and every symmetric tensor has a symmetric nuclear norm attaining decomposition; for symmetric tensors, the symmetric nuclear norm equals the nuclear norm (Friedland et al., 2014). Nie further develops moment-SOS and Lasserre-relaxation machinery for computing symmetric tensor nuclear norms and extracting symmetric nuclear decompositions in moderate-scale cases (Nie, 2016). At the same time, higher-order behavior departs sharply from matrices: for d3d\ge 3, there exist real tensors whose real and complex nuclear norms differ, and exact computation is NP-hard in several senses (Friedland et al., 2014).

A recurring source of ambiguity is that the literature contains several definitions of tensor nuclear norm for third-order tensors, especially in completion and recovery. The tt-SVD-based definition, the sum of unfolding nuclear norms, tensor-ring nuclear norms, and semidefinite dd0-norm relaxations all target low-rankness, but they act on different tensor models and induce different optimization geometries (Xue et al., 2017).

2. The dd1-product-based tensor nuclear norm

The construction introduced in tensor robust PCA is built on the tensor–tensor product. For dd2 and dd3,

dd4

where dd5 is the dd6 block-circulant matrix built from the frontal slices of dd7. Equivalently, one may FFT each tensor along the third dimension, multiply the resulting block-diagonal matrices slice-by-slice, and then inverse-FFT (Lu et al., 2018).

This algebra supports a dd8-SVD

dd9

with TRn1××ndT\in\mathbb R^{n_1\times\cdots\times n_d}0 orthogonal under the TRn1××ndT\in\mathbb R^{n_1\times\cdots\times n_d}1-product and TRn1××ndT\in\mathbb R^{n_1\times\cdots\times n_d}2 TRn1××ndT\in\mathbb R^{n_1\times\cdots\times n_d}3-diagonal. If TRn1××ndT\in\mathbb R^{n_1\times\cdots\times n_d}4, the tensor spectral norm and tensor nuclear norm are defined by

TRn1××ndT\in\mathbb R^{n_1\times\cdots\times n_d}5

In the Fourier-domain viewpoint, if TRn1××ndT\in\mathbb R^{n_1\times\cdots\times n_d}6 has frontal slices TRn1××ndT\in\mathbb R^{n_1\times\cdots\times n_d}7, then

TRn1××ndT\in\mathbb R^{n_1\times\cdots\times n_d}8

These definitions extend the matrix case exactly when TRn1××ndT\in\mathbb R^{n_1\times\cdots\times n_d}9 (Lu et al., 2018).

The same framework defines the tensor average rank

T=min{i=1rλi:  T=i=1rλi(x1ixdi), xki2=1},\|T\|_*=\min\Bigl\{\sum_{i=1}^r|\lambda_i|:\;T=\sum_{i=1}^r\lambda_i\,(x_1^i\otimes\cdots\otimes x_d^i),\ \|x_k^i\|_2=1\Bigr\},0

In related T=min{i=1rλi:  T=i=1rλi(x1ixdi), xki2=1},\|T\|_*=\min\Bigl\{\sum_{i=1}^r|\lambda_i|:\;T=\sum_{i=1}^r\lambda_i\,(x_1^i\otimes\cdots\otimes x_d^i),\ \|x_k^i\|_2=1\Bigr\},1-SVD formulations, one also encounters tubal rank, defined as the number of nonzero singular tubes or, equivalently, T=min{i=1rλi:  T=i=1rλi(x1ixdi), xki2=1},\|T\|_*=\min\Bigl\{\sum_{i=1}^r|\lambda_i|:\;T=\sum_{i=1}^r\lambda_i\,(x_1^i\otimes\cdots\otimes x_d^i),\ \|x_k^i\|_2=1\Bigr\},2 after FFT along the third mode (Zhang et al., 2022). The distinction matters: average-rank statements and tubal-rank statements are not interchangeable unless a paper states the relevant equivalence or reduction.

3. Convex-envelope geometry and recovery theory

A central theorem of the T=min{i=1rλi:  T=i=1rλi(x1ixdi), xki2=1},\|T\|_*=\min\Bigl\{\sum_{i=1}^r|\lambda_i|:\;T=\sum_{i=1}^r\lambda_i\,(x_1^i\otimes\cdots\otimes x_d^i),\ \|x_k^i\|_2=1\Bigr\},3-product-based construction is that on the unit ball

T=min{i=1rλi:  T=i=1rλi(x1ixdi), xki2=1},\|T\|_*=\min\Bigl\{\sum_{i=1}^r|\lambda_i|:\;T=\sum_{i=1}^r\lambda_i\,(x_1^i\otimes\cdots\otimes x_d^i),\ \|x_k^i\|_2=1\Bigr\},4

the convex envelope of T=min{i=1rλi:  T=i=1rλi(x1ixdi), xki2=1},\|T\|_*=\min\Bigl\{\sum_{i=1}^r|\lambda_i|:\;T=\sum_{i=1}^r\lambda_i\,(x_1^i\otimes\cdots\otimes x_d^i),\ \|x_k^i\|_2=1\Bigr\},5 is exactly T=min{i=1rλi:  T=i=1rλi(x1ixdi), xki2=1},\|T\|_*=\min\Bigl\{\sum_{i=1}^r|\lambda_i|:\;T=\sum_{i=1}^r\lambda_i\,(x_1^i\otimes\cdots\otimes x_d^i),\ \|x_k^i\|_2=1\Bigr\},6 (Lu et al., 2018). The proof proceeds through convex conjugates. Writing T=min{i=1rλi:  T=i=1rλi(x1ixdi), xki2=1},\|T\|_*=\min\Bigl\{\sum_{i=1}^r|\lambda_i|:\;T=\sum_{i=1}^r\lambda_i\,(x_1^i\otimes\cdots\otimes x_d^i),\ \|x_k^i\|_2=1\Bigr\},7, the conjugate T=min{i=1rλi:  T=i=1rλi(x1ixdi), xki2=1},\|T\|_*=\min\Bigl\{\sum_{i=1}^r|\lambda_i|:\;T=\sum_{i=1}^r\lambda_i\,(x_1^i\otimes\cdots\otimes x_d^i),\ \|x_k^i\|_2=1\Bigr\},8 can be expressed through the singular values of the block-circulant embedding, and von Neumann’s trace inequality yields the optimal alignment argument. The biconjugate T=min{i=1rλi:  T=i=1rλi(x1ixdi), xki2=1},\|T\|_*=\min\Bigl\{\sum_{i=1}^r|\lambda_i|:\;T=\sum_{i=1}^r\lambda_i\,(x_1^i\otimes\cdots\otimes x_d^i),\ \|x_k^i\|_2=1\Bigr\},9, restricted to Tσ=maxxk2=1T,x1xd.\|T\|_\sigma=\max_{\|x_k\|_2=1}\langle T,x_1\otimes\cdots\otimes x_d\rangle.0, then reduces to

Tσ=maxxk2=1T,x1xd.\|T\|_\sigma=\max_{\|x_k\|_2=1}\langle T,x_1\otimes\cdots\otimes x_d\rangle.1

This reproduces the matrix relation between nuclear norm, spectral norm, and rank in a tensor algebra induced by the Tσ=maxxk2=1T,x1xd.\|T\|_\sigma=\max_{\|x_k\|_2=1}\langle T,x_1\otimes\cdots\otimes x_d\rangle.2-product (Lu et al., 2018).

That convex-envelope statement is specific to the Tσ=maxxk2=1T,x1xd.\|T\|_\sigma=\max_{\|x_k\|_2=1}\langle T,x_1\otimes\cdots\otimes x_d\rangle.3-product model. The same paper stresses that, unlike just summing matrix nuclear norms of all slices, the new tensor nuclear norm averages over the block-circulant embedding, ensuring the convex-envelope property holds. In this sense, the Tσ=maxxk2=1T,x1xd.\|T\|_\sigma=\max_{\|x_k\|_2=1}\langle T,x_1\otimes\cdots\otimes x_d\rangle.4-SVD-based TNN is not merely a heuristic slice regularizer but a precise convex relaxation of tensor average rank in the associated non-commutative algebra (Lu et al., 2018).

Recovery theory in the same framework extends beyond robust PCA. By choosing an atomic set adapted to tubal rank, Zhang and Aeron show that TNN is a special atomic norm and that exact recovery from Gaussian measurements of a tensor of size Tσ=maxxk2=1T,x1xd.\|T\|_\sigma=\max_{\|x_k\|_2=1}\langle T,x_1\otimes\cdots\otimes x_d\rangle.5 and tubal rank Tσ=maxxk2=1T,x1xd.\|T\|_\sigma=\max_{\|x_k\|_2=1}\langle T,x_1\otimes\cdots\otimes x_d\rangle.6 requires

Tσ=maxxk2=1T,x1xd.\|T\|_\sigma=\max_{\|x_k\|_2=1}\langle T,x_1\otimes\cdots\otimes x_d\rangle.7

which is order optimal relative to the degrees of freedom Tσ=maxxk2=1T,x1xd.\|T\|_\sigma=\max_{\|x_k\|_2=1}\langle T,x_1\otimes\cdots\otimes x_d\rangle.8 (Lu et al., 2018). The same work gives tensor-completion guarantees under uniform random sampling and incoherence assumptions, with sample complexity Tσ=maxxk2=1T,x1xd.\|T\|_\sigma=\max_{\|x_k\|_2=1}\langle T,x_1\otimes\cdots\otimes x_d\rangle.9 (Lu et al., 2018).

In tensor robust PCA, the d=2d=20-product-based TNN leads to a convex program that exactly recovers low-rank and sparse components under incoherence and sparsity assumptions, paralleling matrix PCP theory. Matrix RPCA appears as a special case, and the reported applications include image recovery and background modeling (Lu et al., 2018).

4. Computation, proximal mappings, and algorithmic use

In the d=2d=21-SVD setting, computing the TNN and its proximal operator is FFT-centric. One performs an FFT along mode 3, computes matrix SVDs on the frontal slices in the transform domain, applies slice-wise singular-value thresholding, and then returns by inverse FFT. For the proximal map of d=2d=22, the tensor singular value thresholding (TSVT) procedure consists of: FFT, d=2d=23 slice SVDs, soft-thresholding of singular values, and inverse FFT (Zhang et al., 2022).

The per-iteration complexity reported for TSVT is one FFT and one inverse FFT of size d=2d=24, together with d=2d=25 independent SVDs of size d=2d=26 (Zhang et al., 2022). This decomposition is the basic computational reason why ADMM and proximal-gradient methods are practical for large third-order tensors in imaging and completion.

A direct application is dynamic cardiac MRI reconstruction. The TMNN model combines the d=2d=27-SVD-based TNN with the Casorati matrix nuclear norm: d=2d=28 The motivation is explicit: TNN exploits spatial structure of the dynamic MR data, while the Casorati matrix nuclear norm exploits temporal correlation. The resulting ADMM solver admits a fast Cartesian-sampling variant, and the reported experiments show up to d=2d=29 dB SNR gain over plain MNN together with an jσj(A)\sum_j \sigma_j(A)0 run-time reduction for the fast jσj(A)\sum_j \sigma_j(A)1-space implementation (Zhang et al., 2022).

The same study also states a useful structural caveat: in general,

jσj(A)\sum_j \sigma_j(A)2

Accordingly, TNN and Casorati nuclear norms should be regarded as complementary regularizers rather than interchangeable ones (Zhang et al., 2022).

5. Variants, tighter surrogates, and transform-induced extensions

Several later works keep the jσj(A)\sum_j \sigma_j(A)3-SVD backbone but modify the singular-value penalty or the transform domain to approximate tensor rank more tightly.

Tensor jσj(A)\sum_j \sigma_j(A)4-shrinkage nuclear norm. The jσj(A)\sum_j \sigma_j(A)5-TNN replaces soft-thresholding by the scalar jσj(A)\sum_j \sigma_j(A)6-shrinkage operator

jσj(A)\sum_j \sigma_j(A)7

applied to the diagonal entries of the jσj(A)\sum_j \sigma_j(A)8-SVD core in the Fourier domain. The resulting functional is positive, unitary-invariant, and non-convex for jσj(A)\sum_j \sigma_j(A)9, and it satisfies

A=max{A,X:Xs1}\| \mathcal A\|_*=\max\{\langle \mathcal A,\mathcal X\rangle:\|\mathcal X\|_s\le 1\}0

The paper states that A=max{A,X:Xs1}\| \mathcal A\|_*=\max\{\langle \mathcal A,\mathcal X\rangle:\|\mathcal X\|_s\le 1\}1-TNN is a better approximation of tensor average rank than TNN when A=max{A,X:Xs1}\| \mathcal A\|_*=\max\{\langle \mathcal A,\mathcal X\rangle:\|\mathcal X\|_s\le 1\}2, gives a statistical recovery-error bound for low-rank tensor completion, and analyzes an ADMM solver with adaptive momentum whose convergence rate is A=max{A,X:Xs1}\| \mathcal A\|_*=\max\{\langle \mathcal A,\mathcal X\rangle:\|\mathcal X\|_s\le 1\}3 under the smoothness assumption (Liu et al., 2019).

Tensor truncated nuclear norm. T-TNN generalizes matrix truncated nuclear norm to the A=max{A,X:Xs1}\| \mathcal A\|_*=\max\{\langle \mathcal A,\mathcal X\rangle:\|\mathcal X\|_s\le 1\}4-SVD setting by subtracting the contribution of the leading A=max{A,X:Xs1}\| \mathcal A\|_*=\max\{\langle \mathcal A,\mathcal X\rangle:\|\mathcal X\|_s\le 1\}5 singular values. In one formulation,

A=max{A,X:Xs1}\| \mathcal A\|_*=\max\{\langle \mathcal A,\mathcal X\rangle:\|\mathcal X\|_s\le 1\}6

so only the tail singular values of the first Fourier frontal slice are penalized. The stated motivation is that truncation avoids over-shrinkage of dominant components and more closely approximates tubal rank than the plain T-NN. The ADMM and APGL schemes require one matrix SVD of size A=max{A,X:Xs1}\| \mathcal A\|_*=\max\{\langle \mathcal A,\mathcal X\rangle:\|\mathcal X\|_s\le 1\}7 per iteration, and the reported image-completion experiments show a A=max{A,X:Xs1}\| \mathcal A\|_*=\max\{\langle \mathcal A,\mathcal X\rangle:\|\mathcal X\|_s\le 1\}8–A=max{A,X:Xs1}\| \mathcal A\|_*=\max\{\langle \mathcal A,\mathcal X\rangle:\|\mathcal X\|_s\le 1\}9 speedup over several baselines (Xue et al., 2017).

Framelet-based tensor nuclear norm. F-TNN replaces the DFT along each tube by a redundant framelet transform d3d\ge 30 satisfying d3d\ge 31. The norm is then defined as the sum of matrix nuclear norms of the framelet-transformed frontal slices. Because of framelet-basis redundancy, the representation of each tube is sparsely represented, and the paper reports that on MRI, video, and multispectral data the average truncated frontal-slice rank can drop by d3d\ge 32–d3d\ge 33. The completion model is convex, global minimizers can be obtained, and empirical gains of d3d\ge 34–d3d\ge 35 dB in PSNR over several baselines are reported (Jiang et al., 2019).

Nonlinear transform induced tensor nuclear norm. NTTNN augments a semi-orthogonal linear transform along mode 3 with an element-wise nonlinear activation d3d\ge 36, giving

d3d\ge 37

The paper argues that the nonlinearity can shrink small entries and emphasize dominant modes in the transformed frontal slices, and it solves the resulting nonlinear, nonconvex completion model with proximal alternating minimization under a Kurdyka–Łojasiewicz analysis. Reported gains are d3d\ge 38–d3d\ge 39 dB PSNR over TNN and tt0–tt1 dB over learned-transform TNNs on hyperspectral images, multispectral images, and videos (Li et al., 2021).

Multimode nonlinear transform-based TNN. MNT-TNN extends single-mode TTNN to a multimode setting using a face-wise mode-tt2 transform tt3, a mode-tt4 transform tt5, a mode-3 transform tt6, and an element-wise nonlinearity tt7. The paper formulates a PAM algorithm with sufficient-decrease and relative-error guarantees under the KŁ framework, and introduces ATTNN chains that first use linear TTNN variants and then nonlinear TTNN variants. At very high missing rates, ATTNNs are reported to deliver tt8–tt9 further MAPE/RMSE improvements over standalone MNT-TNN in spatiotemporal traffic imputation (Lu et al., 29 Mar 2025).

Tensor-ring nuclear norm. A different line of work defines

dd00

where dd01 are circular unfoldings. The stated theoretical link is dd02 when dd03 has TR ranks dd04. The resulting completion model is convex and is solved by ADMM; in stripe-missing image and video completion it is reported to outperform several conventional tensor-completion methods (Yu et al., 2019).

6. Higher-order structure, bounds, relaxations, and broader roles

For generic higher-order tensors, structural and computational questions remain difficult. Friedland and Lim show that nuclear norm depends on the base field for tensors of order at least three, that the nuclear norm unit ball and its weak-membership problem are computationally hard, and that computing spectral or nuclear norm is NP-hard for several restricted tensor classes (Friedland et al., 2014). Qi et al. add that the dd05-norm, Frobenius norm, and nuclear norm are tensor norms in the sense of vector-norm axioms plus submultiplicativity under outer products, whereas the infinity norm and spectral norm are not tensor norms (Qi et al., 2019).

Despite the intractability of exact computation, there are computable bounds. For an dd06-tensor dd07, the nuclear norm of every matrix flattening dd08 is a lower bound for dd09, and for 3-tensors one has

dd10

with both bounds sharp when dd11 (Hu, 2014). For third-order tensors, contraction to positive semidefinite biquadratic tensors gives additional lower bounds: the square roots of the nuclear norms of the three contracted biquadratic tensors are lower bounds of the tensor nuclear norm (Qi et al., 2019).

Semidefinite relaxations provide another route. The dd12-norms are defined through theta bodies of polynomial ideals generated by second-order minors of tensor matricizations, with the property that in the matrix case they reduce to the nuclear norm, while for order dd13 they give new norms. The unit-dd14-norm balls converge asymptotically to the unit tensor nuclear norm ball, and computing dd15-norms or minimizing them under affine constraints reduces to semidefinite programming (Rauhut et al., 2015).

Recent work on decomposability and subdifferentials shows that the tensor nuclear norm admits full decomposability over specific subspaces such as

dd16

and identifies the largest possible subspaces allowing exact additivity. The same work derives new inclusions for the subdifferential and uses them to establish the statistical performance of tensor robust PCA for tensors of arbitrary order, described there as the first such result in that generality (Guan et al., 6 Oct 2025).

The tensor nuclear norm also appears outside completion and denoising. In bilinear complexity, a bilinear operator dd17 can be identified with a 3-tensor, and its nuclear norm

dd18

equals the minimum growth factor of any bilinear algorithm for dd19. The forward-error bound proved in that setting ties numerical accuracy directly to dd20, yielding a tensorial notion of bilinear stability that is invariant under orthogonal change of coordinates (Dai et al., 2022).

Taken together, these results show that “tensor nuclear norm” is not a single universally adopted object but a family of norms and norm-like surrogates attached to distinct tensor models. The decomposition-based higher-order norm provides the broad convex analogue of matrix nuclear norm; the dd21-SVD-based norm provides an exact convex envelope of tensor average rank in the dd22-product algebra; and later transform-based or truncated variants modify the penalty to target specific data geometries or rank surrogates more tightly (Friedland et al., 2014).

Topic to Video (Beta)

No one has generated a video about this topic yet.

Whiteboard

No one has generated a whiteboard explanation for this topic yet.

Follow Topic

Get notified by email when new papers are published related to Tensor Nuclear Norm.