- The paper presents a new direct proof of the DKWM inequality using discrete-time martingales that recovers the optimal exponent 2nε².
- It employs minimax duality and Sion's theorem to simplify traditional, technical proofs, thereby reducing analytical complexity.
- The results offer sharp, non-asymptotic exponential concentration bounds crucial for empirical process theory and nonparametric inference.
An Elementary Proof of the Dvoretzky–Kiefer–Wolfowitz–Massart Inequality
Overview
The Dvoretzky–Kiefer–Wolfowitz–Massart (DKWM) inequality provides sharp, non-asymptotic exponential concentration bounds for the uniform deviation between the empirical distribution function (EDF) and the true cumulative distribution function (CDF) of independent and identically distributed (i.i.d.) samples. The DKWM inequality is foundational in probability theory and statistics, underpinning the finite-sample validity of nonparametric inference, uniform empirical process theory, and confidence bands.
This paper delivers a new, direct, and elementary proof of the DKWM inequality. In contrast to Massart's original, highly technical arguments and Reeve's recent martingale-centric refinements that utilize continuous time, this work employs a discrete-time martingale framework along with Sion's minimax theorem. The result is an explicit, concise, and accessible argument that recovers the optimal constants and generality of Massart's theorem without the analytical complexity typical of previous proofs (2607.04387).
Statement of Results
Given n i.i.d. real random variables (Xi)i=1n with distribution function F and empirical distribution function F^n(x), the DKWM bounds are:
P(x∈Rsup(F^n(x)−F(x))>ε)≤e−2nε2,∀ε>0
P(x∈Rsup∣F^n(x)−F(x)∣>ε)≤2e−2nε2,∀ε>0
The sharpness in the exponent 2nε2 is preserved.
Proof Strategy
Reduction to Canonical Setting
The proof begins by observing that the supremum deviation is preserved under probability integral transform. Thus, it suffices to establish the bound for n i.i.d. U[0,1] random variables. This reduction exploits the fact that the CDF transforms any sample into a uniform distribution, allowing the EDF–CDF deviation for general distributions to be controlled by the uniform case.
Discrete-Time Martingale Construction
Let (Ui)i=1n be i.i.d. (Xi)i=1n0. The supremum is translated into the event that the (Xi)i=1n1-th order statistic (Xi)i=1n2 drops beneath (Xi)i=1n3 for some (Xi)i=1n4. Introducing (Xi)i=1n5, the event (Xi)i=1n6 is equivalent to (Xi)i=1n7.
A reverse martingale sequence (Xi)i=1n8 is constructed, with (Xi)i=1n9 a function of F0 designed so that F1 upper bounds the desired probability via Doob’s reverse martingale maximal inequality. The explicit link to the binomial tail probability sets the stage for a sharp, exponential bound.
Minimax Optimization and Sion's Theorem
Optimizing the martingale-based upper bound over an exponential generating parameter F2 and using the explicit structure of the binomial likelihood, the argument proceeds by bounding the probability as:
F3
The function F4 is tailored to encapsulate the exponential moment method. Crucially, it is shown to possess quasi-concavity in F5 and quasi-convexity in F6, allowing Sion’s minimax theorem to be applied for saddle-point exchange.
Explicit Bound via Binary Relative Entropy
The minimax evaluation leads to an explicit representation in terms of the binary Kullback–Leibler divergence:
F7
Here, F8 denotes the binary KL divergence. Applying Pinsker's inequality, F9, recovers the canonical F^n(x)0 bound.
Numerical Sharpness and Claims
The proof precisely recovers the optimal DKWM constants as established by Massart and removes additional technical constraints on F^n(x)1. The approach is fully explicit and avoids reliance on measure-theoretic underpinnings or advanced empirical process theory.
Key strong claim: The exponent F^n(x)2 is achieved by a direct martingale and minimax argument, exposing the tightness of the DKWM inequality without the intricacy of previous treatments involving either symmetrization or continuous time limits.
Implications and Theoretical Impact
This proof substantially lowers the access barrier for the application and teaching of concentration inequalities for empirical processes. The explicit martingale construction demystifies the role of order statistics and binomial deviations, potentially facilitating generalizations to other settings such as Markov-dependent samples, sub-exponential tails, or high-dimensional function-indexed processes. The method may also be adapted to derive nonasymptotic bounds for other functionals of empirical measures with sharp constants.
The direct connection between martingale inequalities, saddle point duality, and information-theoretic divergence supports a unifying methodological template for future empirical process concentration results.
Prospects for Future Research
The discrete martingale and minimax approach advocated here may inspire several follow-ups:
- Extensions to non-i.i.d. settings: Adaptation to dependent data via coupling or block martingales.
- Sharper constants for small F^n(x)3: Fine tuning the argument for finite-sample improvements, possibly leveraging exact binomial tail behavior.
- High-dimensional analogues: Application to multivariate empirical distributions and processes on function spaces.
- Algorithmic certificate generation: Programmatic generation of nonasymptotic confidence bands and online estimation protocols with proved guarantees.
Conclusion
This paper provides an elementary, transparent, and optimal proof of the Dvoretzky–Kiefer–Wolfowitz–Massart inequality. Employing discrete-time martingale constructions and minimax duality, it offers both pedagogical clarity and practical power, yielding sharp deviation bounds fundamental to nonparametric statistics and empirical process theory (2607.04387). The methodology opens new directions for streamlined proofs in probability and statistics, and sets a foundation for further theoretical developments and practical algorithms in finite-sample statistical control.