Papers
Topics
Authors
Recent
Search
2000 character limit reached

Sharp Convex Concentration for Symmetric Random Tensors with Subgaussian Coordinates

Published 20 Aug 2026 in math.PR | (2608.19832v1)

Abstract: Let X=(X1,…,Xn)X=(X_1,\ldots,X_n) have independent coordinates with mean zero, variance one, and ∣Xi∣<em>ψ2≤K|X_i|<em>{ψ_2}\le K, and let Hd=(R<sup>n)<sup>⊗2</sup></sup>dH_d=(\mathbb R<sup>n)<sup>{\otimes_2</sup></sup> d}. Let $L&gt;0$ and let f:Hd→Rf:H_d\to\mathbb R be convex and LL-Lipschitz. We prove that, for 0≤t≤cKLn<sup>d/20\le t\le c_KLn<sup>{d/2}, [ \textsf{P}\left{ \left\lvert f(X{\otimes d})-\textsf{E}f(X{\otimes d})\right\rvert >t \right} \le C\exp\left[-c_K\mathcal I{n,d}\left( \frac{t}{L n{(d-1)/2}} \right)\right], ] where [ \mathcal I_{n,d}(s)= \min\left{ \frac{s2}{d2}, \frac{s2}{d\log(e+nd/s2)} \right},\qquad s>0, \qquad \mathcal I_{n,d}(0)=0. ] The first rate is forced by changes in ∣X∣|X|. The second comes from changes of XX when its norm is nearly fixed. The proof constructs one coupling that controls both the coordinatewise conditional displacement and the mean squared Euclidean distance, and combines these bounds with a second-order estimate for x↦x<sup>⊗</sup>dx\mapsto x<sup>{\otimes</sup> d}. The rate is minimax sharp, scale by scale, even when the subgaussian norms are bounded by an absolute constant. For bounded coordinates the logarithm in the second rate disappears.

Authors (1)

Summary

  • The paper presents a minimally exponential in $d$ upper tail bound for convex functions of symmetric random tensors.
  • This concentration bound is achieved by summing fluctuations relative to the norm and to directions orthogonal to $x$ at typical radius $c=n^{1/2},$ rather than local approaches.
  • For bounded coordinates, the paper provides exact concentration bounds mirroring Huang-Tikhomirov.
  • Confirm that the upper bound of $mn$ holds when the expectation is divided by subgraph status is $m^2$
  • Confirm solutions $h(x)=0$ when $u$ is $c_Kn/d^2$

This paper by Xuanang Hu resolves a problem posed by Vershynin concerning concentration of convex Lipschitz functionals of a single symmetric random tensor X⊗dX^{\otimes d}, where X=(X1,…,Xn)X=(X_1,\ldots,X_n) has independent, centered, variance-one coordinates with ψ2\psi_2 norm bounded by an absolute constant KK (2608.19832). The main theorem gives a two-scale tail bound that the paper proves is minimax sharp, scale by scale, even within the full subgaussian class. The proof combines a new entropy-to-coupling machinery for product measures with a deterministic second-order estimate of the tensor map x↦x⊗dx\mapsto x^{\otimes d}.

The two-scale phenomenon

The central difficulty is geometric. Writing Φd(x)=x⊗d\Phi_d(x)=x^{\otimes d}, the derivative satisfies

∥DΦd(x)h∥2=d∥x∥2d−2∥h∥2+d(d−1)∥x∥2d−4⟨x,h⟩2,\|D\Phi_d(x)h\|^2 = d\|x\|^{2d-2}\|h\|^2 + d(d-1)\|x\|^{2d-4}\langle x,h\rangle^2,

so perturbations parallel to xx are amplified by a factor dd, while perturbations orthogonal to xx are amplified only by X=(X1,…,Xn)X=(X_1,\ldots,X_n)0. At the typical radius X=(X1,…,Xn)X=(X_1,\ldots,X_n)1 this predicts the deviation scale

X=(X1,…,Xn)X=(X_1,\ldots,X_n)2

with the first term driven by fluctuations of the norm and the second by fluctuations at nearly fixed norm. This separation is essential: Vershynin's results for simple tensors (independent factors) carry the smaller variance scale X=(X1,…,Xn)X=(X_1,\ldots,X_n)3 in the bounded case (2608.19832), but the symmetric model forces the larger scale X=(X1,…,Xn)X=(X_1,\ldots,X_n)4, as witnessed already by the Euclidean functional X=(X1,…,Xn)X=(X_1,\ldots,X_n)5, where X=(X1,…,Xn)X=(X_1,\ldots,X_n)6. Vershynin explicitly identified the symmetric tensor as an open case, noting that decoupling arguments are expected to lose factors exponential in X=(X1,…,Xn)X=(X_1,\ldots,X_n)7; this paper avoids such losses entirely.

The main result states that for every convex X=(X1,…,Xn)X=(X_1,\ldots,X_n)8-Lipschitz X=(X1,…,Xn)X=(X_1,\ldots,X_n)9, if ψ2\psi_20, then

ψ2\psi_21

with matching moment bounds. A companion formulation covers the entire natural range ψ2\psi_22 via the rate function

ψ2\psi_23

giving tails of order ψ2\psi_24. For bounded coordinates (ψ2\psi_25), a separate argument based on Talagrand's convex-distance inequality removes the logarithmic factor entirely, yielding subgaussian tails with variance scale ψ2\psi_26 — and the paper shows this scale is sharp even for one fixed bounded marginal law. This mirrors the Huang–Tikhomirov dichotomy between bounded product measures and the full subgaussian class.

The coupling machinery

The probabilistic core is a construction converting relative entropy into two distinct cost functionals simultaneously. For ψ2\psi_27 on a product of ψ2\psi_28-subgaussian marginals with ψ2\psi_29, the paper builds an explicit maximal coupling KK0 satisfying both

KK1

and

KK2

Both estimates must hold for a single coupling because both enter the deterministic tensor inequality. The one-dimensional building block is a maximal coupling that leaves common mass fixed and transports only the excess KK3 against the deficit KK4; pointwise entropy calculus converts KK5 into control of the conditional displacements in both directions. The product version chains these couplings coordinatewise using the chain rule KK6 and concavity of KK7.

A notable optimality result: the KK8 term in the mean squared distance is best possible uniformly over the subgaussian class. For biased measures on KK9 with bias x↦x⊗dx\mapsto x^{\otimes d}0, one has x↦x⊗dx\mapsto x^{\otimes d}1 while x↦x⊗dx\mapsto x^{\otimes d}2. Consequently, any improvement of the final concentration rates for special subclasses must exploit additional structure beyond general coupling bounds — which is exactly what the paper does for Euclidean functionals, whose polynomial structure permits cancellations invisible to the coupling argument.

The deterministic tensor estimate

On a sphere of radius x↦x⊗dx\mapsto x^{\otimes d}3, for x↦x⊗dx\mapsto x^{\otimes d}4,

x↦x⊗dx\mapsto x^{\otimes d}5

The proof integrates along geodesics: since the integrand's tangent direction is orthogonal to x↦x⊗dx\mapsto x^{\otimes d}6 after averaging, the parallel component of the mean displacement contributes only at second order (x↦x⊗dx\mapsto x^{\otimes d}7), while the orthogonal part enjoys the x↦x⊗dx\mapsto x^{\otimes d}8 factor from the derivative formula. A covering argument extends this to an annulus x↦x⊗dx\mapsto x^{\otimes d}9, adding a term Φd(x)=x⊗d\Phi_d(x)=x^{\otimes d}0 with Φd(x)=x⊗d\Phi_d(x)=x^{\otimes d}1, controlled by a Bernstein estimate for Φd(x)=x⊗d\Phi_d(x)=x^{\otimes d}2 with probability Φd(x)=x⊗d\Phi_d(x)=x^{\otimes d}3.

Proof architecture for concentration

The argument compares level sets around a median Φd(x)=x⊗d\Phi_d(x)=x^{\otimes d}4. Restricting to the shell Φd(x)=x⊗d\Phi_d(x)=x^{\otimes d}5, sets below and above Φd(x)=x⊗d\Phi_d(x)=x^{\otimes d}6 separated by a gap Φd(x)=x⊗d\Phi_d(x)=x^{\otimes d}7 are coupled via the two-set coupling theorem (which routes both conditional laws through one copy of Φd(x)=x⊗d\Phi_d(x)=x^{\otimes d}8, costing Φd(x)=x⊗d\Phi_d(x)=x^{\otimes d}9). For ∥DΦd(x)h∥2=d∥x∥2d−2∥h∥2+d(d−1)∥x∥2d−4⟨x,h⟩2,\|D\Phi_d(x)h\|^2 = d\|x\|^{2d-2}\|h\|^2 + d(d-1)\|x\|^{2d-4}\langle x,h\rangle^2,0 in the upper set, convexity gives

∥DΦd(x)h∥2=d∥x∥2d−2∥h∥2+d(d−1)∥x∥2d−4⟨x,h⟩2,\|D\Phi_d(x)h\|^2 = d\|x\|^{2d-2}\|h\|^2 + d(d-1)\|x\|^{2d-4}\langle x,h\rangle^2,1

so Lipschitzness forces ∥DΦd(x)h∥2=d∥x∥2d−2∥h∥2+d(d−1)∥x∥2d−4⟨x,h⟩2,\|D\Phi_d(x)h\|^2 = d\|x\|^{2d-2}\|h\|^2 + d(d-1)\|x\|^{2d-4}\langle x,h\rangle^2,2, and taking expectation against both coupling costs yields the contradiction when ∥DΦd(x)h∥2=d∥x∥2d−2∥h∥2+d(d−1)∥x∥2d−4⟨x,h⟩2,\|D\Phi_d(x)h\|^2 = d\|x\|^{2d-2}\|h\|^2 + d(d-1)\|x\|^{2d-4}\langle x,h\rangle^2,3 exceeds a constant multiple of the predicted scale. Conversion to moments uses a quantile representation with careful handling of the exponential factor ∥DΦd(x)h∥2=d∥x∥2d−2∥h∥2+d(d−1)∥x∥2d−4⟨x,h⟩2,\|D\Phi_d(x)h\|^2 = d\|x\|^{2d-2}\|h\|^2 + d(d-1)\|x\|^{2d-4}\langle x,h\rangle^2,4 arising near the edge of the moderate range, valid up to ∥DΦd(x)h∥2=d∥x∥2d−2∥h∥2+d(d−1)∥x∥2d−4⟨x,h⟩2,\|D\Phi_d(x)h\|^2 = d\|x\|^{2d-2}\|h\|^2 + d(d-1)\|x\|^{2d-4}\langle x,h\rangle^2,5.

For large degrees, ∥DΦd(x)h∥2=d∥x∥2d−2∥h∥2+d(d−1)∥x∥2d−4⟨x,h⟩2,\|D\Phi_d(x)h\|^2 = d\|x\|^{2d-2}\|h\|^2 + d(d-1)\|x\|^{2d-4}\langle x,h\rangle^2,6, the logarithmic term is dominated and the bound collapses to ∥DΦd(x)h∥2=d∥x∥2d−2∥h∥2+d(d−1)∥x∥2d−4⟨x,h⟩2,\|D\Phi_d(x)h\|^2 = d\|x\|^{2d-2}\|h\|^2 + d(d-1)\|x\|^{2d-4}\langle x,h\rangle^2,7 — shown sharp by the Gaussian radial example. A corollary for Hilbert-valued linear images ∥DΦd(x)h∥2=d∥x∥2d−2∥h∥2+d(d−1)∥x∥2d−4⟨x,h⟩2,\|D\Phi_d(x)h\|^2 = d\|x\|^{2d-2}\|h\|^2 + d(d-1)\|x\|^{2d-4}\langle x,h\rangle^2,8 yields the RMS shift bound ∥DΦd(x)h∥2=d∥x∥2d−2∥h∥2+d(d−1)∥x∥2d−4⟨x,h⟩2,\|D\Phi_d(x)h\|^2 = d\|x\|^{2d-2}\|h\|^2 + d(d-1)\|x\|^{2d-4}\langle x,h\rangle^2,9.

Sharpness

Two complementary lower bounds establish minimax sharpness under an absolute subgaussian constraint:

  • Radial lower bound: for xx0 and xx1, density comparisons near the mode xx2 give xx3, forcing the xx4 term.
  • Sparse-coordinate lower bound: a three-point law with rare large coordinates (xx5, probability xx6) combined with the convex functional xx7, where xx8 projects onto tensors with exactly one slot in xx9, produces deviations of order dd0 with probability dd1, forcing the logarithmic term.

These are assembled into a rate-function sharpness statement: whenever dd2 with dd3, some subgaussian law and convex one-Lipschitz functional achieve dd4. Minimax sharpness here is scale by scale: each admissible deviation level may use its own extremizing pair, and no single example is claimed to work across all scales.

Limitations and open questions

Several restrictions are explicit. The moment estimate requires dd5; outside this range the natural range contains no nontrivial fixed-exponent moderate regime, and the analysis of very high degrees relative to dd6 is not pursued. The logarithmic factor persists throughout for unbounded coordinates, consistent with the Huang–Tikhomirov lower bound, so removing it would require restricting beyond the full subgaussian class — and Proposition on dd7 optimality shows no uniform improvement of the coupling costs is possible. The paper notes but does not develop an extension to base vectors with multiplicities dd8, where parallel and orthogonal contributions scale as dd9 and xx0 respectively. Finally, stronger rates for structured subclasses (Euclidean functionals, bounded coordinates) rely on algebraic or isoperimetric structure beyond the coupling method; characterizing precisely which subclasses admit improvements remains open.

Conclusion

The paper settles the symmetric-tensor case of convex concentration with the correct dependence on dimension and degree, resolving an open question of Vershynin without exponential losses in xx1. Its two principal technical contributions — a maximal-coupling construction controlling conditional displacement and squared distance simultaneously under one entropy budget, and a norm-aware second-order tensor estimate separating radial from transverse motion — are of independent interest and appear tight in generality. Together with the matching lower bounds, the results delineate exactly where the xx2 versus xx3 variance scales and where the logarithmic correction apply across the bounded, subgaussian, and Gaussian hierarchies.

Paper to Video (Beta)

No one has generated a video about this paper yet.

Whiteboard

No one has generated a whiteboard explanation for this paper yet.

Tweets

Sign up for free to view the 1 tweet with 0 likes about this paper.