Papers
Topics
Authors
Recent
Search
2000 character limit reached

Deep Legendre Transform (DLT): Neural Convex Conjugates

Updated 13 March 2026
  • DLT is a neural framework for evaluating convex conjugates by leveraging the analytic relationship between a convex function and its Legendre transform.
  • The method trains neural networks to minimize an empirical squared residual, thereby sidestepping computationally expensive supremum optimizations.
  • DLT offers scalability and unbiased a posteriori error estimation, outperforming grid-based methods in high-dimensional convex optimization tasks.

The Deep Legendre Transform (DLT) is a neural framework for learning and evaluating convex conjugates (Legendre–Fenchel transforms) of differentiable convex functions in high dimensions. By exploiting an implicit analytic relation between a convex function and its conjugate, DLT circumvents direct supremum optimization or grid-based discretizations, facilitating scalable computation and a posteriori estimation of approximation error. DLT is applicable in convex optimization, variational analysis, Hamilton–Jacobi PDEs, and elsewhere in mathematics, physics, and economics (Minabutdinov et al., 22 Dec 2025).

1. Theoretical Foundations

For a differentiable convex function f:Rn→Rf:\mathbb{R}^n\to\mathbb{R}, the Legendre–Fenchel transform (convex conjugate) is

f∗(y)=sup⁡x∈Rn{⟨x,y⟩−f(x)}.f^*(y)=\sup_{x\in\mathbb{R}^n}\{\langle x,y\rangle - f(x)\}.

When ff is differentiable on an open domain C⊆RnC\subseteq\mathbb{R}^n, this relation simplifies on the dual domain D=∇f(C)D=\nabla f(C):

f∗(∇f(x))=⟨x,∇f(x)⟩−f(x),x∈C.f^*(\nabla f(x)) = \langle x,\nabla f(x)\rangle - f(x), \quad x\in C.

This identity follows from the Fenchel–Young theorem. Traditional grid-based approaches, even in optimized form (e.g., Lucet’s nested LLT algorithm), are computationally intractable for large dd, with costs scaling as O(N2d)O(N^{2d}) or O(dNd+1)O(dN^{d+1}) for uniform grid sizes NN per dimension (Minabutdinov et al., 22 Dec 2025). Neural approaches aiming to avoid grid enumeration often require solving max-min problems, which similarly scale poorly as f∗(y)=sup⁡x∈Rn{⟨x,y⟩−f(x)}.f^*(y)=\sup_{x\in\mathbb{R}^n}\{\langle x,y\rangle - f(x)\}.0 grows.

2. DLT Methodology and Implicit Optimization Objective

DLT leverages the analytical identity above to train a function f∗(y)=sup⁡x∈Rn{⟨x,y⟩−f(x)}.f^*(y)=\sup_{x\in\mathbb{R}^n}\{\langle x,y\rangle - f(x)\}.1 parameterized by neural networks to approximate f∗(y)=sup⁡x∈Rn{⟨x,y⟩−f(x)}.f^*(y)=\sup_{x\in\mathbb{R}^n}\{\langle x,y\rangle - f(x)\}.2. Instead of optimizing over f∗(y)=sup⁡x∈Rn{⟨x,y⟩−f(x)}.f^*(y)=\sup_{x\in\mathbb{R}^n}\{\langle x,y\rangle - f(x)\}.3 for each f∗(y)=sup⁡x∈Rn{⟨x,y⟩−f(x)}.f^*(y)=\sup_{x\in\mathbb{R}^n}\{\langle x,y\rangle - f(x)\}.4, DLT minimizes the empirical squared residual:

f∗(y)=sup⁡x∈Rn{⟨x,y⟩−f(x)}.f^*(y)=\sup_{x\in\mathbb{R}^n}\{\langle x,y\rangle - f(x)\}.5

where f∗(y)=sup⁡x∈Rn{⟨x,y⟩−f(x)}.f^*(y)=\sup_{x\in\mathbb{R}^n}\{\langle x,y\rangle - f(x)\}.6 is a training set sampled from f∗(y)=sup⁡x∈Rn{⟨x,y⟩−f(x)}.f^*(y)=\sup_{x\in\mathbb{R}^n}\{\langle x,y\rangle - f(x)\}.7, f∗(y)=sup⁡x∈Rn{⟨x,y⟩−f(x)}.f^*(y)=\sup_{x\in\mathbb{R}^n}\{\langle x,y\rangle - f(x)\}.8 is the weight decay, and gradient computations are performed via autodiff frameworks such as JAX or PyTorch. This loss equates to squared f∗(y)=sup⁡x∈Rn{⟨x,y⟩−f(x)}.f^*(y)=\sup_{x\in\mathbb{R}^n}\{\langle x,y\rangle - f(x)\}.9 error between ff0 and the true ff1 on the pushforward distribution ff2. This approach avoids the need to compute ff3 explicitly or solve a supremum for each ff4.

Regularization (e.g., weight decay) is applied in the standard way, and optimization uses Adam with a learning rate of ff5 and batch size determined empirically.

3. Neural Network Architectures and Training Protocol

DLT can be instantiated using various neural network architectures with specific design considerations:

Architecture Key Properties Convexity Guarantee
MLP 2×128 units, GELU None
ResNet 2 residual blocks (each Dense(128) → GELU → Dense(128)); skip connections None
ICNN Input-Convex Network; non-negative hidden weights, softplus nonlinearity; direct input skip Yes
MLP_ICNN MLP variant, non-negative weights, softplus, no skip Yes
KAN Kolmogorov–Arnold Network; basis expansion; for symbolic regression No

For KANs, each network layer sums over one-dimensional basis functions ff6 (including polynomials, exponential, logarithmic, and trigonometric forms); post-training, symbolic regression (e.g., sparse linear regression over ff7) enables closed-form recovery of learned conjugates.

Training follows a standard stochastic minibatch loop. For improved sampling in cases where ff8 distorts measure (notably in Neg-Log), DLT augments training with an inverse network ff9, reducing C⊆RnC\subseteq\mathbb{R}^n0 error by an order of magnitude in high dimensions with modest additional computation.

4. Error Estimation and A Posteriori Guarantees

DLT provides unbiased a posteriori error estimation without requiring knowledge of the exact C⊆RnC\subseteq\mathbb{R}^n1. Given C⊆RnC\subseteq\mathbb{R}^n2 i.i.d. samples C⊆RnC\subseteq\mathbb{R}^n3 from C⊆RnC\subseteq\mathbb{R}^n4 on C⊆RnC\subseteq\mathbb{R}^n5,

C⊆RnC\subseteq\mathbb{R}^n6

estimates the C⊆RnC\subseteq\mathbb{R}^n7 error C⊆RnC\subseteq\mathbb{R}^n8; the variance decreases as C⊆RnC\subseteq\mathbb{R}^n9, enabling confidence intervals via CLT. No ground-truth conjugate is necessary—an advantage in domains where D=∇f(C)D=\nabla f(C)0 lacks a closed analytic form.

5. Numerical Benchmarks and Comparative Performance

DLT achieves high-accuracy approximations across a spectrum of convex test functions, including high-dimensional quadratic, negative logarithm, negative entropy, and more complex cases without closed-form conjugates, such as quadratic-over-linear. Benchmark results display RMSE for DLT nearly identical to direct learning (when D=∇f(C)D=\nabla f(C)1 is available), with comparable training times. For the quadratic-over-linear function, small D=∇f(C)D=\nabla f(C)2 errors persist through D=∇f(C)D=\nabla f(C)3 dimensions.

In comparison to Lucet’s grid-based nested LLT, DLT remains computationally tractable as dimension increases, with active memory usage D=∇f(C)D=\nabla f(C)41 MB and second-scale runtimes, while grid-based methods become infeasible due to exponential cost in time (up to thousands of seconds) and memory (exceeding gigabytes) for D=∇f(C)D=\nabla f(C)5. Inverse-sampling for distorted gradients further reduces estimation error.

Method D=∇f(C)D=\nabla f(C)6 (s) Memory (MB) RMSE
Lucet (D=∇f(C)D=\nabla f(C)7) 1960 1530 29.3
DLT (D=∇f(C)D=\nabla f(C)8) 26.22 1.1 0.133

6. Advanced Variants: Symbolic Regression and Hamilton–Jacobi Extension

Utilizing Kolmogorov–Arnold Networks with symbolic regression, DLT can extract exact closed-form representations of D=∇f(C)D=\nabla f(C)9 for classes of separable convex functions, recovering, for example, f∗(∇f(x))=⟨x,∇f(x)⟩−f(x),x∈C.f^*(\nabla f(x)) = \langle x,\nabla f(x)\rangle - f(x), \quad x\in C.0 for a quadratic in f∗(∇f(x))=⟨x,∇f(x)⟩−f(x),x∈C.f^*(\nabla f(x)) = \langle x,\nabla f(x)\rangle - f(x), \quad x\in C.1 with residuals below f∗(∇f(x))=⟨x,∇f(x)⟩−f(x),x∈C.f^*(\nabla f(x)) = \langle x,\nabla f(x)\rangle - f(x), \quad x\in C.2. This regime permits both high-precision approximation and symbolic interpretability.

DLT generalizes to time-dependent convex-conjugate computations, such as those in Hamilton–Jacobi equations. For instance, the Hopf formula solution f∗(∇f(x))=⟨x,∇f(x)⟩−f(x),x∈C.f^*(\nabla f(x)) = \langle x,\nabla f(x)\rangle - f(x), \quad x\in C.3 of a Hamilton–Jacobi PDE is written in terms of conjugates and can be approximated by a time-augmented DLT (f∗(∇f(x))=⟨x,∇f(x)⟩−f(x),x∈C.f^*(\nabla f(x)) = \langle x,\nabla f(x)\rangle - f(x), \quad x\in C.4). In benchmarks for quadratic initial data/Hamiltonians at f∗(∇f(x))=⟨x,∇f(x)⟩−f(x),x∈C.f^*(\nabla f(x)) = \langle x,\nabla f(x)\rangle - f(x), \quad x\in C.5, time-parameterized DLT attains f∗(∇f(x))=⟨x,∇f(x)⟩−f(x),x∈C.f^*(\nabla f(x)) = \langle x,\nabla f(x)\rangle - f(x), \quad x\in C.6–f∗(∇f(x))=⟨x,∇f(x)⟩−f(x),x∈C.f^*(\nabla f(x)) = \langle x,\nabla f(x)\rangle - f(x), \quad x\in C.7 smaller f∗(∇f(x))=⟨x,∇f(x)⟩−f(x),x∈C.f^*(\nabla f(x)) = \langle x,\nabla f(x)\rangle - f(x), \quad x\in C.8 error than the Deep Galerkin Method (DGM) for f∗(∇f(x))=⟨x,∇f(x)⟩−f(x),x∈C.f^*(\nabla f(x)) = \langle x,\nabla f(x)\rangle - f(x), \quad x\in C.9, despite the latter’s lower PDE residual.

7. Limitations and Domain of Applicability

DLT requires dd0 to be convex and differentiable on an open set dd1; its learned approximation is valid on dd2, which may not coincide with the full effective domain of dd3. The framework’s predictive quality depends critically on neural network expressivity and sufficient computational resources. DLT does not address nonconvex or nondifferentiable targets directly.

DLT is particularly suited for situations prohibitive to grid-based or direct-supremum computation, including high-dimensional convex optimization duality, optimal transport, variational constructs (Moreau envelopes), indirect utility in economics, thermodynamic duality, Hamilton–Jacobi theory, and deep generative modeling (e.g., Wasserstein gradient flows) (Minabutdinov et al., 22 Dec 2025).

Definition Search Book Streamline Icon: https://streamlinehq.com
References (1)

Topic to Video (Beta)

No one has generated a video about this topic yet.

Whiteboard

No one has generated a whiteboard explanation for this topic yet.

Follow Topic

Get notified by email when new papers are published related to Deep Legendre Transform (DLT).