- The paper presents a minimally exponential in $d$ upper tail bound for convex functions of symmetric random tensors.
- This concentration bound is achieved by summing fluctuations relative to the norm and to directions orthogonal to $x$ at typical radius $c=n^{1/2},$ rather than local approaches.
- For bounded coordinates, the paper provides exact concentration bounds mirroring Huang-Tikhomirov.
- Confirm that the upper bound of $mn$ holds when the expectation is divided by subgraph status is $m^2$
- Confirm solutions $h(x)=0$ when $u$ is $c_Kn/d^2$
This paper by Xuanang Hu resolves a problem posed by Vershynin concerning concentration of convex Lipschitz functionals of a single symmetric random tensor X⊗d, where X=(X1​,…,Xn​) has independent, centered, variance-one coordinates with ψ2​ norm bounded by an absolute constant K (2608.19832). The main theorem gives a two-scale tail bound that the paper proves is minimax sharp, scale by scale, even within the full subgaussian class. The proof combines a new entropy-to-coupling machinery for product measures with a deterministic second-order estimate of the tensor map x↦x⊗d.
The two-scale phenomenon
The central difficulty is geometric. Writing Φd​(x)=x⊗d, the derivative satisfies
∥DΦd​(x)h∥2=d∥x∥2d−2∥h∥2+d(d−1)∥x∥2d−4⟨x,h⟩2,
so perturbations parallel to x are amplified by a factor d, while perturbations orthogonal to x are amplified only by X=(X1​,…,Xn​)0. At the typical radius X=(X1​,…,Xn​)1 this predicts the deviation scale
X=(X1​,…,Xn​)2
with the first term driven by fluctuations of the norm and the second by fluctuations at nearly fixed norm. This separation is essential: Vershynin's results for simple tensors (independent factors) carry the smaller variance scale X=(X1​,…,Xn​)3 in the bounded case (2608.19832), but the symmetric model forces the larger scale X=(X1​,…,Xn​)4, as witnessed already by the Euclidean functional X=(X1​,…,Xn​)5, where X=(X1​,…,Xn​)6. Vershynin explicitly identified the symmetric tensor as an open case, noting that decoupling arguments are expected to lose factors exponential in X=(X1​,…,Xn​)7; this paper avoids such losses entirely.
The main result states that for every convex X=(X1​,…,Xn​)8-Lipschitz X=(X1​,…,Xn​)9, if ψ2​0, then
ψ2​1
with matching moment bounds. A companion formulation covers the entire natural range ψ2​2 via the rate function
ψ2​3
giving tails of order ψ2​4. For bounded coordinates (ψ2​5), a separate argument based on Talagrand's convex-distance inequality removes the logarithmic factor entirely, yielding subgaussian tails with variance scale ψ2​6 — and the paper shows this scale is sharp even for one fixed bounded marginal law. This mirrors the Huang–Tikhomirov dichotomy between bounded product measures and the full subgaussian class.
The coupling machinery
The probabilistic core is a construction converting relative entropy into two distinct cost functionals simultaneously. For ψ2​7 on a product of ψ2​8-subgaussian marginals with ψ2​9, the paper builds an explicit maximal coupling K0 satisfying both
K1
and
K2
Both estimates must hold for a single coupling because both enter the deterministic tensor inequality. The one-dimensional building block is a maximal coupling that leaves common mass fixed and transports only the excess K3 against the deficit K4; pointwise entropy calculus converts K5 into control of the conditional displacements in both directions. The product version chains these couplings coordinatewise using the chain rule K6 and concavity of K7.
A notable optimality result: the K8 term in the mean squared distance is best possible uniformly over the subgaussian class. For biased measures on K9 with bias x↦x⊗d0, one has x↦x⊗d1 while x↦x⊗d2. Consequently, any improvement of the final concentration rates for special subclasses must exploit additional structure beyond general coupling bounds — which is exactly what the paper does for Euclidean functionals, whose polynomial structure permits cancellations invisible to the coupling argument.
The deterministic tensor estimate
On a sphere of radius x↦x⊗d3, for x↦x⊗d4,
x↦x⊗d5
The proof integrates along geodesics: since the integrand's tangent direction is orthogonal to x↦x⊗d6 after averaging, the parallel component of the mean displacement contributes only at second order (x↦x⊗d7), while the orthogonal part enjoys the x↦x⊗d8 factor from the derivative formula. A covering argument extends this to an annulus x↦x⊗d9, adding a term Φd​(x)=x⊗d0 with Φd​(x)=x⊗d1, controlled by a Bernstein estimate for Φd​(x)=x⊗d2 with probability Φd​(x)=x⊗d3.
Proof architecture for concentration
The argument compares level sets around a median Φd​(x)=x⊗d4. Restricting to the shell Φd​(x)=x⊗d5, sets below and above Φd​(x)=x⊗d6 separated by a gap Φd​(x)=x⊗d7 are coupled via the two-set coupling theorem (which routes both conditional laws through one copy of Φd​(x)=x⊗d8, costing Φd​(x)=x⊗d9). For ∥DΦd​(x)h∥2=d∥x∥2d−2∥h∥2+d(d−1)∥x∥2d−4⟨x,h⟩2,0 in the upper set, convexity gives
∥DΦd​(x)h∥2=d∥x∥2d−2∥h∥2+d(d−1)∥x∥2d−4⟨x,h⟩2,1
so Lipschitzness forces ∥DΦd​(x)h∥2=d∥x∥2d−2∥h∥2+d(d−1)∥x∥2d−4⟨x,h⟩2,2, and taking expectation against both coupling costs yields the contradiction when ∥DΦd​(x)h∥2=d∥x∥2d−2∥h∥2+d(d−1)∥x∥2d−4⟨x,h⟩2,3 exceeds a constant multiple of the predicted scale. Conversion to moments uses a quantile representation with careful handling of the exponential factor ∥DΦd​(x)h∥2=d∥x∥2d−2∥h∥2+d(d−1)∥x∥2d−4⟨x,h⟩2,4 arising near the edge of the moderate range, valid up to ∥DΦd​(x)h∥2=d∥x∥2d−2∥h∥2+d(d−1)∥x∥2d−4⟨x,h⟩2,5.
For large degrees, ∥DΦd​(x)h∥2=d∥x∥2d−2∥h∥2+d(d−1)∥x∥2d−4⟨x,h⟩2,6, the logarithmic term is dominated and the bound collapses to ∥DΦd​(x)h∥2=d∥x∥2d−2∥h∥2+d(d−1)∥x∥2d−4⟨x,h⟩2,7 — shown sharp by the Gaussian radial example. A corollary for Hilbert-valued linear images ∥DΦd​(x)h∥2=d∥x∥2d−2∥h∥2+d(d−1)∥x∥2d−4⟨x,h⟩2,8 yields the RMS shift bound ∥DΦd​(x)h∥2=d∥x∥2d−2∥h∥2+d(d−1)∥x∥2d−4⟨x,h⟩2,9.
Sharpness
Two complementary lower bounds establish minimax sharpness under an absolute subgaussian constraint:
- Radial lower bound: for x0 and x1, density comparisons near the mode x2 give x3, forcing the x4 term.
- Sparse-coordinate lower bound: a three-point law with rare large coordinates (x5, probability x6) combined with the convex functional x7, where x8 projects onto tensors with exactly one slot in x9, produces deviations of order d0 with probability d1, forcing the logarithmic term.
These are assembled into a rate-function sharpness statement: whenever d2 with d3, some subgaussian law and convex one-Lipschitz functional achieve d4. Minimax sharpness here is scale by scale: each admissible deviation level may use its own extremizing pair, and no single example is claimed to work across all scales.
Limitations and open questions
Several restrictions are explicit. The moment estimate requires d5; outside this range the natural range contains no nontrivial fixed-exponent moderate regime, and the analysis of very high degrees relative to d6 is not pursued. The logarithmic factor persists throughout for unbounded coordinates, consistent with the Huang–Tikhomirov lower bound, so removing it would require restricting beyond the full subgaussian class — and Proposition on d7 optimality shows no uniform improvement of the coupling costs is possible. The paper notes but does not develop an extension to base vectors with multiplicities d8, where parallel and orthogonal contributions scale as d9 and x0 respectively. Finally, stronger rates for structured subclasses (Euclidean functionals, bounded coordinates) rely on algebraic or isoperimetric structure beyond the coupling method; characterizing precisely which subclasses admit improvements remains open.
Conclusion
The paper settles the symmetric-tensor case of convex concentration with the correct dependence on dimension and degree, resolving an open question of Vershynin without exponential losses in x1. Its two principal technical contributions — a maximal-coupling construction controlling conditional displacement and squared distance simultaneously under one entropy budget, and a norm-aware second-order tensor estimate separating radial from transverse motion — are of independent interest and appear tight in generality. Together with the matching lower bounds, the results delineate exactly where the x2 versus x3 variance scales and where the logarithmic correction apply across the bounded, subgaussian, and Gaussian hierarchies.