Moreau Envelope is a regularization technique that smooths nonsmooth objective functions by infimal convolution with a quadratic term, preserving key minimization structures.
It provides a differentiable, Lipschitz continuous approximation for both convex and weakly convex functions, ensuring gradient stability via the proximal mapping.
Its framework links proximal mapping, Fenchel duality, and variational convergence, enabling robust first- and second-order algorithm design in nonsmooth optimization.
The Moreau envelope is a fundamental regularization construct in variational analysis and optimization, which enables the smoothing of nonsmooth objective functions while preserving key minimization structures. Originally developed in the context of convex functionals, the Moreau envelope admits precise generalizations and retains substantial regularization power for the broader class of weakly convex objectives (functions whose deviation from convexity is controlled quadratically). Its links with the proximal mapping, infimal convolution, and Fenchel duality yield a toolkit central to contemporary nonsmooth optimization and variational convergence analysis. Below, the essential theory and properties of the Moreau envelope for both convex and weakly convex functions are outlined, following the framework in (Renaud et al., 17 Sep 2025).
1. Definition and Variational Representations
For a proper, lower semicontinuous function f:Rn→(−∞,+∞] and regularization parameter λ>0, the Moreau envelope Eλf and its associated proximal mapping are given by: Eλf(x):=y∈Rninf{f(y)+2λ1∥y−x∥2},proxλf(x):=argy∈Rnmin{f(y)+2λ1∥y−x∥2}.
This is equivalently formulated as the infimal convolution: Eλf=f□(2λ∥⋅∥2),
where (f□g)(x)=infy[f(y)+g(x−y)].
For convex f, Eλf admits a dual representation in terms of the Fenchel conjugate f∗: Eλf(x)=u∈Rnsup{⟨u,x⟩−f∗(u)−2λ∥u∥2}=[f∗+2λ∥⋅∥2]∗(x).
This structure links the Moreau envelope directly to convex analysis, duality, and infimal convolution machinery.
2. Classical Properties: The Convex Case
Suppose λ>00 is convex, proper, and lower semicontinuous, with λ>01.
Convexity and Differentiability:
The function λ>02 is convex as the infimum over convex mappings, and is everywhere finite and continuously differentiable on λ>03.
Gradient Formula:
λ>04
Lipschitz Gradient (Smoothness):
The proximal mapping λ>05 is nonexpansive (1-Lipschitz), so λ>06 is λ>07-Lipschitz: λ>08
Thus, λ>09 is Eλf0-smooth.
Minimizer Preservation:
The Moreau envelope preserves both minima and minimizers: Eλf1
If Eλf2, then Eλf3 and Eλf4.
3. Weakly Convex Case: Extensions and Delicate Properties
Let Eλf5 be Eλf6-weakly convex, i.e., Eλf7 is convex. For all Eλf8:
Well-Posedness, Continuity, Monotonicity:
For each Eλf9, Eλf(x):=y∈Rninf{f(y)+2λ1∥y−x∥2},proxλf(x):=argy∈Rnmin{f(y)+2λ1∥y−x∥2}.0 is nonincreasing on Eλf(x):=y∈Rninf{f(y)+2λ1∥y−x∥2},proxλf(x):=argy∈Rnmin{f(y)+2λ1∥y−x∥2}.1.
Eλf(x):=y∈Rninf{f(y)+2λ1∥y−x∥2},proxλf(x):=argy∈Rnmin{f(y)+2λ1∥y−x∥2}.2 and Eλf(x):=y∈Rninf{f(y)+2λ1∥y−x∥2},proxλf(x):=argy∈Rnmin{f(y)+2λ1∥y−x∥2}.3.
Differentiability and Gradient Formula:
Eλf(x):=y∈Rninf{f(y)+2λ1∥y−x∥2},proxλf(x):=argy∈Rnmin{f(y)+2λ1∥y−x∥2}.4 is differentiable and the gradient formula from the convex case extends: Eλf(x):=y∈Rninf{f(y)+2λ1∥y−x∥2},proxλf(x):=argy∈Rnmin{f(y)+2λ1∥y−x∥2}.5
The proof requires careful handling of upper and lower quadratic bounds and uses the one-sided directional envelope properties specific to weak convexity.
Proximal Inverse and Open Image:
If Eλf(x):=y∈Rninf{f(y)+2λ1∥y−x∥2},proxλf(x):=argy∈Rnmin{f(y)+2λ1∥y−x∥2}.6 is differentiable at Eλf(x):=y∈Rninf{f(y)+2λ1∥y−x∥2},proxλf(x):=argy∈Rnmin{f(y)+2λ1∥y−x∥2}.7,
If Eλf(x):=y∈Rninf{f(y)+2λ1∥y−x∥2},proxλf(x):=argy∈Rnmin{f(y)+2λ1∥y−x∥2}.9 is defined everywhere, then Eλf=f□(2λ∥⋅∥2),0, so Eλf=f□(2λ∥⋅∥2),1 is bijective with open image.
Cocoercivity and Nonexpansivity:
The generalized nonexpansivity becomes: Eλf=f□(2λ∥⋅∥2),2
yielding Eλf=f□(2λ∥⋅∥2),3-Lipschitz continuity for Eλf=f□(2λ∥⋅∥2),4.
Convexity/Strong Convexity of the Envelope:
Eλf=f□(2λ∥⋅∥2),5 is Eλf=f□(2λ∥⋅∥2),6-weakly convex.
If Eλf=f□(2λ∥⋅∥2),7 is Eλf=f□(2λ∥⋅∥2),8-strongly convex, Eλf=f□(2λ∥⋅∥2),9 is (f□g)(x)=infy[f(y)+g(x−y)]0-strongly convex.
Smoothness Regimes:
The gradient's Lipschitz constant is (f□g)(x)=infy[f(y)+g(x−y)]1. For (f□g)(x)=infy[f(y)+g(x−y)]2, (f□g)(x)=infy[f(y)+g(x−y)]3; for (f□g)(x)=infy[f(y)+g(x−y)]4, (f□g)(x)=infy[f(y)+g(x−y)]5.
Minima and Stationarity:
(f□g)(x)=infy[f(y)+g(x−y)]6. The characterizations
(f□g)(x)=infy[f(y)+g(x−y)]7
hold. For (f□g)(x)=infy[f(y)+g(x−y)]8 differentiable at (f□g)(x)=infy[f(y)+g(x−y)]9, f0.
4. Associated Subdifferential, Epi-Convergence, and Second-Order Structure
Subdifferential (Clarke) Connection:
For general (not necessarily convex) f1, critical points of f2 correspond to Clarke-critical points of f3: f4
where f5 denotes the Clarke subdifferential.
Epi-Convergence:
As f6, f7 pointwise; f8 epi-converges to f9 (in the Attouch–Wets sense), so Eλf0 and Eλf1 in the Painlevé–Kuratowski sense.
Second-Order Formula:
If Eλf2 locally and Eλf3 is Eλf4-Lipschitz with Eλf5, then: Eλf6
Eλf7
Geometry of Proximal Image:
The image of Eλf8 is "almost convex": the Lebesgue measure of points in Eλf9 not in f∗0 is zero, and similarly for the convex hull.
5. Overview Table: Main Properties by Function Class
Eλf(x)=u∈Rnsup{⟨u,x⟩−f∗(u)−2λ∥u∥2}=[f∗+2λ∥⋅∥2]∗(x).2 strongly convex if Eλf(x)=u∈Rnsup{⟨u,x⟩−f∗(u)−2λ∥u∥2}=[f∗+2λ∥⋅∥2]∗(x).3 is
Yes
Eλf(x)=u∈Rnsup{⟨u,x⟩−f∗(u)−2λ∥u∥2}=[f∗+2λ∥⋅∥2]∗(x).4-strongly convex if Eλf(x)=u∈Rnsup{⟨u,x⟩−f∗(u)−2λ∥u∥2}=[f∗+2λ∥⋅∥2]∗(x).5 is Eλf(x)=u∈Rnsup{⟨u,x⟩−f∗(u)−2λ∥u∥2}=[f∗+2λ∥⋅∥2]∗(x).6-strongly convex
This table condenses the quantitative and qualitative parallels and differences between convex and weakly convex regimes.
6. Significance, Generalization, and Applications
The Moreau envelope provides a universal smoothing mechanism:
For convex Eλf(x)=u∈Rnsup{⟨u,x⟩−f∗(u)−2λ∥u∥2}=[f∗+2λ∥⋅∥2]∗(x).7, it produces a Eλf(x)=u∈Rnsup{⟨u,x⟩−f∗(u)−2λ∥u∥2}=[f∗+2λ∥⋅∥2]∗(x).8, Eλf(x)=u∈Rnsup{⟨u,x⟩−f∗(u)−2λ∥u∥2}=[f∗+2λ∥⋅∥2]∗(x).9-smooth, strictly smaller-than-λ>000 function, preserving minimization structure and enabling first-order (and in strongly convex settings, higher-order) algorithm design.
For weakly convex λ>001, the envelope extends this regularization—preserving differentiability and smoothness up to the threshold λ>002, with all key stationarity and minimization identifications maintained.
This regularization is pivotal in:
Stochastic and deterministic first-order algorithms, where the Moreau envelope yields descent directions even for nonsmooth or weakly convex objectives.
Analysis of variational convergence (epi-convergence, Painlevé–Kuratowski minimizer limits).
Second-order theory, owing to the preservation or transfer of key curvature properties via explicit Hessian transformations.
Formulation and exact/approximate solution of bilevel optimization problems, composite minimization, and algorithmic frameworks reliant on the proximal point methodology.
7. Structural Limits and Advanced Topics
As λ>003, the Moreau envelope recovers λ>004 pointwise and in epigraphical topology, making it an essential tool in variational approximation.
For λ>005 (twice differentiable) weakly convex λ>006, the second derivative of the envelope encodes information on the local smoothability of λ>007 through the proximal transformation, central for higher-order algorithms.
The relationship between the image of the proximal mapping and the domain of λ>008 is "almost full measure," providing strong geometric guarantees for algorithmic coverage.
A plausible implication is that almost every point in λ>009 is represented as a prox point for some λ>010, facilitating analysis and justification of regularization or penalization procedures based on the Moreau envelope or related proximal maps.
The Moreau envelope thus constitutes a comprehensive smoothing and regularization theory—initially for convex but robustly extending to the weakly convex regime—endowing nonsmooth and nonconvex optimization with differentiability, smoothness, explicit first- and second-order structures, and robust links to variational analysis and algorithmic convergence (Renaud et al., 17 Sep 2025).