---
title: Totally Convex Functionals
url: https://www.emergentmind.com/topics/totally-convex-functionals
type: topic
---

# Totally Convex Functionals

Totally convex functionals form a class of convex-analytic objects whose precise meaning depends on the ambient space. In the most direct recent arXiv treatment, the term refers to functionals \(\phi:\mathcal P_2(H)\to \mathbb R\cup\{+\infty\}\) on the quadratic Wasserstein space over a separable Hilbert space \(H\), where total convexity means convexity along every interpolation induced by every coupling, not merely along optimal transport geodesics; equivalently, the Lagrangian lift \(\hat\phi=\phi\circ\iota\) to \(L^2(Q,\mathbb M;H)\) is an ordinary convex functional on a Hilbert space [2509.01768]. In that setting total convexity supplies a full convex-analysis dictionary—Fenchel conjugacy, biconjugation, subdifferentials, cyclic monotonicity, Moreau–Yosida regularization, and Brenier-type optimal transport structure for laws of random measures [2509.01768].

## 1. Definition and conceptual position

Let \(H\) be a separable Hilbert space and \(\mathcal P_2(H)\) the quadratic Wasserstein space. The recent optimal-transport treatment defines a functional \(\phi:\mathcal P_2(H)\to\mathbb R\cup\{+\infty\}\) to be **totally convex** if for every coupling \(\boldsymbol\mu\in\mathcal P_2(H\times H)\) and every \(t\in[0,1]\),
\[
\phi(\mu_t)\le (1-t)\phi(\mu_0)+t\phi(\mu_1),
\qquad
\mu_t:=(\pi_t^{12})_\sharp \boldsymbol\mu,
\qquad
\pi_t^{12}(x_1,x_2):=(1-t)x_1+t x_2.
\]
The same source defines \(\phi\) to be **totally \(\lambda\)-convex** if
\[
\phi-\frac{\lambda}{2}m_2^2
\]
is totally convex, equivalently if the lifted functional \(\hat\phi\) is \(\lambda\)-convex on \(L^2(Q,\mathbb M;H)\) [2509.01768].

This notion is stronger than ordinary Wasserstein geodesic convexity. Geodesic convexity only requires convexity along at least one optimal interpolation between two measures, whereas total convexity requires convexity along **every interpolation induced by every coupling**. The same paper states that, if \(\dim H\ge 2\), every continuous geodesically convex functional on \(\mathcal P_2(H)\) is already totally convex [2509.01768]. This places total convexity at the point where Wasserstein geometry begins to recover Hilbert-space convex analysis.

A central structural ingredient is the maximal-correlation pairing
\[
[\mu_0,\mu_1]
:=
\max_{\gamma\in\Gamma(\mu_0,\mu_1)}
\int_{H\times H}\langle x_0,x_1\rangle\,d\gamma(x_0,x_1),
\]
which satisfies
\[
W_2^2(\mu_0,\mu_1)
=
m_2^2(\mu_0)+m_2^2(\mu_1)-2[\mu_0,\mu_1].
\]
This identity is the analogue of the Euclidean polarization formula and is the basis for the conjugacy theory of totally convex functionals [2509.01768].

## 2. Fenchel theory and total subdifferentials

The corresponding conjugate is the **Kantorovich–Legendre–Fenchel transform**
\[
\phi^\star(\nu)
:=
\sup_{\mu\in\mathcal P_2(H)}
\big([\nu,\mu]-\phi(\mu)\big).
\]
Under a mild linear lower bound, \(\phi^\star\) is proper, totally convex, and lower semicontinuous; moreover, \(\phi^{\star\star}\) is the largest totally convex lower semicontinuous functional below \(\phi\). The paper gives the equivalences
\[
\phi \text{ totally convex and l.s.c.}
\iff
\hat\phi \text{ convex and l.s.c. on } \mathbb H
\iff
\phi=\phi^{\star\star}
\]
and also the support-type representation
\[
\phi(\mu)=\sup_{(\nu,a)\in G}\big([\mu,\nu]-a\big)
\]
for a suitable \(G\subset \mathcal P_2(H)\times\mathbb R\) [2509.01768].

The **total subdifferential** is not merely a subset of \(\mathcal P_2(H)\times\mathcal P_2(H)\). It is a family of plans \(\gamma\in\mathcal P_2(H\times H)\). A plan \(\gamma\) belongs to \(\boldsymbol\partial\phi\) if \(\mu=\pi^1_\sharp\gamma\in D(\phi)\) and, for every \(\nu\in D(\phi)\) and every \(\eta\in\Gamma(\gamma,\nu)\),
\[
\phi(\nu)-\phi(\mu)
\ge
\int_{H^2\times H}\langle y_1,x_2-x_1\rangle\,d\eta(x_1,y_1;x_2).
\]
This refines the \([\cdot,\cdot]\)-subdifferential
\[
\partial^-\phi(\mu)
=
\left\{
\nu\in\mathcal P_2(H):
\phi(\mu')-\phi(\mu)\ge [\mu',\nu]-[\mu,\nu]
\ \forall \mu'
\right\},
\]
because
\[
\gamma\in \boldsymbol\partial\phi[\mu]
\iff
\nu:=\pi^2_\sharp\gamma\in\partial^-\phi(\mu)
\text{ and }
\gamma\in\Gamma_o(\mu,\nu).
\]
Thus the total subdifferential selects both a dual measure \(\nu\) and an optimal coupling between \(\mu\) and \(\nu\) [2509.01768].

The subdifferential theory inherits the usual monotone-operator structure. The lifted graph \(\widehat{\boldsymbol\partial\phi}\) coincides with the ordinary convex subdifferential \(\partial\hat\phi\) on \(L^2(Q,\mathbb M;H)\). Consequently, \(\boldsymbol\partial\phi\) is totally cyclically monotone; if \(\phi\) is totally convex and lower semicontinuous, then \(\boldsymbol\partial\phi\) is maximal totally monotone; and there is a unique deterministic minimal section \(^{\circ}\!\boldsymbol\partial\phi\), written
\[
^{\circ}\!\boldsymbol\partial\phi[\mu]
=
(\mathrm{Id}\times \nabla_W\phi(\cdot,\mu))_\sharp\mu.
\]
The paper defines \(\phi\) to be \(W\)-differentiable at \(\mu\) when \(\boldsymbol\partial\phi[\mu]\) is a singleton [2509.01768].

## 3. Lagrangian lifting to a Hilbert space

The decisive construction is the law map
\[
\iota:\mathbb H\to\mathcal P_2(H),
\qquad
\mathbb H:=L^2(Q,\mathbb M;H),
\qquad
\iota(X):=X_\sharp\mathbb M,
\]
with \((Q,\mathbb M)\) nonatomic. This map is surjective and \(1\)-Lipschitz. It also satisfies
\[
m_2(\iota(X))=\|X\|_{\mathbb H},
\qquad
W_2(\iota(X_1),\iota(X_2))\le \|X_1-X_2\|_{\mathbb H},
\]
together with the sharper identities
\[
[\iota(X_1),\iota(X_2)]
=
\sup_{g\in \mathrm{Aut}(Q,\mathbb M)}
\langle X_1,X_2\circ g\rangle_{\mathbb H},
\]
and
\[
W_2^2(\iota(X_1),\iota(X_2))
=
\inf_{g\in \mathrm{Aut}(Q,\mathbb M)}
\|X_1-X_2\circ g\|_{\mathbb H}^2.
\]
In this language, total convexity is exactly ordinary convexity of \(\hat\phi=\phi\circ\iota\) [2509.01768].

This lifting turns a non-Hilbertian Wasserstein problem into Hilbert-space convex analysis. Conjugation commutes with lifting:
\[
(\phi^\star)\circ\iota=(\hat\phi)^*.
\]
Moreau–Yosida regularization also commutes with lifting:
\[
\widehat{\phi_\tau}=(\hat\phi)_\tau.
\]
The same paper uses these identities to import Fenchel–Moreau theory, Rockafellar-type cyclic monotonicity, Yosida regularization, and differentiability theory from \(L^2(Q,\mathbb M;H)\) back to \(\mathcal P_2(H)\) [2509.01768].

This suggests that total convexity is the Wasserstein notion precisely calibrated to preserve linear convex analysis after passage through \(\iota\). It is not an auxiliary strengthening but the condition under which the Wasserstein problem becomes genuinely Hilbertian.

## 4. Optimal transport for laws of random measures

The main application is quadratic optimal transport on
\[
\mathcal P_2(\mathcal P_2(H)),
\]
the space of laws of random measures. For \(\mathcal M_0,\mathcal M_1\in\mathcal P_2(\mathcal P_2(H))\), the transport cost is
\[
\mathbb W_2^2(\mathcal M_0,\mathcal M_1)
=
\min_{\Pi\in\Gamma(\mathcal M_0,\mathcal M_1)}
\int W_2^2(\mu_0,\mu_1)\,d\Pi(\mu_0,\mu_1),
\]
and there is a lifted maximal-correlation identity
\[
\mathbb W_2^2(\mathcal M_0,\mathcal M_1)
=
\mathbb m_2^2(\mathcal M_0)+\mathbb m_2^2(\mathcal M_1)-2[\mathcal M_0,\mathcal M_1].
\]
Optimal Kantorovich potentials can then be chosen in the form of a totally convex proper lower semicontinuous functional \(\phi\) and its conjugate \(\phi^\star\), satisfying
\[
\phi(\mu)+\phi^\star(\nu)=[\mu,\nu]
\]
on the support of an optimal coupling [2509.01768].

Optimality admits the expected cyclic-monotonicity and subdifferential characterizations. In particular, optimal random couplings are concentrated on \(\boldsymbol\partial\phi\), and this is the random-measure analogue of the statement that optimal quadratic couplings in Euclidean space are concentrated on the graph of a convex subdifferential [2509.01768].

The Monge problem is solved under a regularity hypothesis on the source law. A measure \(\mathcal M\in\mathcal P_2(\mathcal P_2(H))\) is called **super-regular** when it vanishes on random exceptional sets and is concentrated on \(\mathcal P_2^r(H)\). For super-regular \(\mathcal M}_0\) and arbitrary \(\mathcal M}_1\), the paper proves existence and uniqueness of the strict Monge solution, induced by the minimal section of an optimal totally convex potential:
\[
\mathbf t=\nabla_W\phi.
\]
The source class is nontrivial: the paper shows that it includes laws with full support in \(\mathcal P_2(H)\) obtained as pushforwards of nondegenerate Gaussian measures on \(L^2(Q,\mathbb M;H)\), and in finite dimension the class of super-regular measures is dense [2509.01768].

## 5. Relation to ordinary convexity and stronger curvature notions

Ordinary convexity is much weaker. A useful background result states that if \(D\) is a convex subset of a real vector space, then a radially lower semicontinuous function
\[
f:D\to\mathbb R\cup\{+\infty\}
\]
is convex if and only if for all \(x,y\in D\) there exists \(\alpha=\alpha(x,y)\in(0,1)\) such that
\[
f(\alpha x+(1-\alpha)y)\le \alpha f(x)+(1-\alpha)f(y).
\]
Thus, under radial lower semicontinuity, one admissible interior interpolation point on each segment already forces full convexity [1709.08611]. By contrast, total convexity on \(\mathcal P_2(H)\) requires convexity along every interpolation induced by every coupling, and the corresponding calculus is correspondingly stronger [2509.01768].

A second, different usage of the phrase appears in Banach-space optimization. There, total convexity typically means that the functional bends away from its tangent planes in a quantitatively positive way, often formulated through positivity of
\[
\inf\{\, f(y)-f(x)-f'(x;y-x): \|y-x\|=t\,\}
\]
for \(t>0\) [1709.08611]. This is not the same definition as the Wasserstein-space one, even though both are stronger than plain convexity. This suggests that the term must always be read relative to its ambient geometry.

The distinction becomes operational in approximation theory. A recent reconstruction theorem for convex Lipschitz functionals on compact convex subsets of a separable Hilbert space produces explicit finite max-affine reconstructions
\[
f_{\varepsilon,d}(x)
=
\max_m \{\langle p_m,x\rangle+q_m\},
\]
which preserve convexity and the same Lipschitz constant, but the paper explicitly does not study strict, strong, uniform, or total convexity, and notes that max-affine functions typically have flat regions or affine faces [2605.08559]. This suggests that preserving ordinary convexity is compatible with finitely computable max-affine structure, whereas preserving stronger curvature properties generally requires additional structure.

## 6. Adjacent literatures and terminological boundaries

A recurrent source of confusion is that many papers study **functionals on convex objects** or **convex functionals** without studying total convexity. One geometric line studies **valuations** on the space
\[
\mathcal C^n
=
\{u:\mathbb R^n\to\mathbb R\cup\{\infty\}: u \text{ convex, lower semicontinuous, and } \lim_{|x|\to\infty}u(x)=\infty\},
\]
with max/min additivity, monotone decreasing order behavior, rigid-motion invariance, and \(m\)-continuity. Its main theorems classify simple and homogeneous valuations, but the paper explicitly states that it does not characterize convex, strictly convex, or totally convex functionals in the optimization sense [1502.06729].

Another line studies **increasing convex functionals** on lattices of measurable or continuous functions. On countable product spaces, increasing convex functionals under marginal constraints admit measure-valued dual representations that subsume transport and martingale-transport duality [1509.08988]. On \(C_b(X)\) with \(X\) Polish, upward continuity yields a \(\sigma\)-additive measure representation for convex increasing functionals, while downward continuity is characterized by weak\(^*\) compactness of dual sublevels and dual attainment [2209.08140]. These are duality and representation theorems, not theories of total convexity.

Related but distinct are robust convex integral functionals on \(L^\infty\), where the main issues are elimination of singular dual measures, Mackey continuity, Lebesgue properties, and exact \(L^1\)-conjugacy [1305.6023]; logarithmic means of two convex functionals, where the central objects are weighted arithmetic, harmonic, geometric, and logarithmic means in \(\Gamma_0(E)\) [2010.01259]; and convex integral functionals of càdlàg processes, where the emphasis is on interchange rules, conjugates, and subdifferentials on nondecomposable path spaces [1812.04086]. None of these papers develops total convexity as such.

The same caution applies in geometric analysis. For area-measure integral functionals on convex bodies, Brunn–Minkowski-type concavity implies monotonicity and, in the fully mixed setting, characterizes mixed volumes [1602.05994]. In shape optimization, different additions—Minkowski, Blaschke, or a functional-specific \(\mu\)-addition—lead to different concavity notions for shape functionals such as volume, capacity, torsional rigidity, and \(\lambda_1\) [1102.1887]. On metric spaces of convex bodies, continuity, affine invariance, and local Lipschitz behavior are developed through Hausdorff and Banach–Mazur metrics rather than through any theory of total convexity [1402.3684]. A further terminological variant appears in the study of measurable linear functionals under convex, i.e. logarithmically concave, measures on \(\mathbb R^\infty\), where “convex” refers to the measure rather than the functional [1602.06738].

This suggests a broad taxonomy. “Totally convex functional” is a precise term only after fixing the ambient structure: in \(\mathcal P_2(H)\) it means convexity along arbitrary couplings and Hilbertian lift-convexity [2509.01768]; in Banach-space optimization it commonly refers to quantitative separation from tangent planes [1709.08611]; and in several neighboring literatures the phrase is simply absent.

Source: https://www.emergentmind.com/topics/totally-convex-functionals