Papers
Topics
Authors
Recent
Search
2000 character limit reached

Variation Brownian Kernel Ladders

Published 14 Aug 2026 in cs.LG | (2608.13882v1)

Abstract: Claims about the benefit of depth depend on the complexity assigned to a representation. We introduce the \emph{Variation Brownian Kernel Ladder} (VBKL), a path-atomic function-space framework that separates nonlinear recursive dictionary construction from linear variation superposition. Starting from linear projections, each atom recursively composes unit-ball profiles from the Brownian reproducing kernel Hilbert space; the full VBKL space is then the signed-measure variation hull of the completed dictionary. We identify each recursive dictionary as a union of Brownian pullback RKHS balls and establish variation-controlled Hölder regularity, compactness and attainment, and strict growth with depth under a local non-degeneracy condition whose trace lies in the support of the input measure. For associated finite lower-support architectures, we derive Rademacher and generalization bounds through Brownian quadratic chaos, signed threshold traces, and VC entropy. We also construct two-stage approximants by discretizing the outer measure and the selected outer Brownian profiles, obtaining an M<sup>1/2+m<sup>1/2M<sup>{-1/2}+m<sup>{-1/2} error bound, a sharp interpolation constant A/2\sqrt{A/2}, and at most $2M$ active outer-profile basis contributions per evaluation. Controlled experiments illustrate the approximation mechanisms and indicate a favorable limited-data accuracy--complexity trade-off.

Authors (1)

Summary

  • The paper develops a path-atomic variation framework that recursively composes normalized Brownian RKHS profiles before applying an outer signed-measure hull, distinguishing its geometry from layerwise-superposed variation spaces.
  • The paper proves quantitative Hölder regularity with exponent 2^{-(L-1)} and a strict depth hierarchy under explicit support and non-degeneracy conditions, showing that deeper norm-controlled spaces contain genuinely new functions.
  • The paper establishes finite-architecture Rademacher and generalization bounds and a two-stage approximation rate of O(N^{-1/2}), while experiments indicate a favorable accuracy–parameter trade-off mainly in limited-data settings.

Motivation and central question

Claims about the benefit of depth are meaningful only relative to a chosen notion of complexity. A function may be cheap when measured by width or parameter count and expensive when measured by weight magnitude, path variation, or an intrinsic function-space norm. The paper under review, "Variation Brownian Kernel Ladders" (VBKL) (2608.13882), addresses this issue by asking how the placement of linear superposition within a recursive function-space construction affects the geometry, statistical complexity, and constructive realizability of the resulting deep function classes.

The design choice is consequential. If convexification or measure superposition is applied at every layer—as in Deep Neural Variation Spaces (DNVS) [nakhleh2026deep]—then nonlinear feature generation and linear combination evolve together. VBKL instead postpones superposition: it first builds a nonlinear path dictionary recursively, then takes the signed-measure variation hull only at the outermost level. Because nonlinear composition and absolute convexification do not generally commute, the resulting space is not obtained by specializing DNVS to a Brownian activation family, and the author explicitly makes no general inclusion or equivalence claim between the two frameworks.

The construction

Starting from linear projections xωxx \mapsto \omega^\top x with ω\omega in an admissible direction set ΩSd1\Omega \subseteq S^{d-1}, each depth-LL atom composes exactly L1L-1 normalized Brownian RKHS profiles along one support path:

UL={xg(u(x)):uUL1,  gHk(B),  gHk(B)1}.U_L = \{x \mapsto g(u(x)) : u \in U_{L-1},\; g \in H_{k^{(B)}},\; \|g\|_{H_{k^{(B)}} \le 1}\}.

A key structural lemma identifies UlU_l exactly as a union of Brownian pullback RKHS unit balls uUl1{fHku:fHku1}\bigcup_{u \in U_{l-1}} \{f \in H_{k_u} : \|f\|_{H_{k_u}} \le 1\}, where ku=k(B)(u(),u())k_u = k^{(B)}(u(\cdot), u(\cdot')). The depth-LL VBKL space is then the Bochner-integral hull of finite signed measures on the completed dictionary, with intrinsic complexity given by the infimal total variation over all representing measures.

The Brownian kernel is not an arbitrary choice. Its RKHS is the anchored Sobolev–Cameron–Martin space of absolutely continuous functions with square-integrable derivative; pullbacks admit exact RKHS descriptions; the kernel metric propagates square-root regularity through composition; and the signed-threshold representation of the kernel connects finite variation classes to quadratic chaos and VC entropy. These properties allow analytical, statistical, and constructive theories to be developed within one model.

Analytical properties and strict depth hierarchy

The main analytical theorem establishes that every element of the full infinite-dimensional space admits a representative satisfying quantitative pointwise and Hölder bounds with exponent ω\omega0: ω\omega1 and ω\omega2, where ω\omega3 denotes the variation complexity. Under full support of the input measure this representative is unique, yielding a continuous linear injection into ω\omega4—an embedding into, not equality with, the Hölder space.

The strictness result deserves emphasis because it contrasts with DNVS's depth saturation for norm-controlled univariate ReLU classes. Under a local non-degeneracy condition—a Lipschitz curve through a point where a first-layer generator vanishes and grows linearly, with its trace lying in ω\omega5—the paper proves ω\omega6. The proof constructs a fractional-power profile ω\omega7 with exponent ω\omega8, which belongs to the Brownian RKHS but whose depth-ω\omega9 trace grows as ΩSd1\Omega \subseteq S^{d-1}0, violating the ΩSd1\Omega \subseteq S^{d-1}1-Hölder bound that any member of ΩSd1\Omega \subseteq S^{d-1}2 must satisfy. The implication is that adjacent intrinsic, potentially infinite-width, norm-controlled spaces contain genuinely different functions—a form of depth separation distinct from classical width- or parameter-based separations [telgarsky2016, eldan2016].

Statistical guarantees for finite architectures

Because the recursive dictionaries are infinite-dimensional, the statistical analysis concerns an explicitly parameterized finite lower-support architecture ΩSd1\Omega \subseteq S^{d-1}3 with piecewise-linear Brownian profiles, row-normalized mixing matrices, and an ΩSd1\Omega \subseteq S^{d-1}4-normalized readout, parameterized by at most ΩSd1\Omega \subseteq S^{d-1}5 real parameters. The associated outer dictionary ΩSd1\Omega \subseteq S^{d-1}6 is again a union of Brownian pullback RKHS unit balls, and its radius-ΩSd1\Omega \subseteq S^{d-1}7 variation hull defines the hypothesis class ΩSd1\Omega \subseteq S^{d-1}8.

The Rademacher bound separates three sources of complexity:

ΩSd1\Omega \subseteq S^{d-1}9

where LL0 is a sample-dependent lower-support envelope bounded deterministically by LL1, and LL2. The proof chain is notable: exact reduction from the variation hull to the outer dictionary via symmetry, reduction of the union-of-RKHS complexity to Brownian quadratic chaos, representation of that chaos through signed threshold traces of the lower-support architecture, and control of trace entropy through VC dimension and the Sauer–Shelah lemma. A high-probability generalization bound follows for Lipschitz losses in LL3.

Two scope restrictions should be stated plainly. First, the bound applies to LL4, not to the full infinite-dimensional ball LL5; the general finite architecture permits intermediate linear mixing and readout, so no inclusion in the path-atomic dictionary holds except in the path-only specialization. Second, attributing finite-parameter complexity to the unrestricted variation ball would be incorrect—the paper is careful to keep these objects distinct.

Constructive approximation with separated resources

For the full VBKL space, the approximation theory separates two resources. Stage one discretizes the outer signed measure: any LL6 admits an LL7-atomic approximant with equal weights LL8 and error at most LL9 in L1L-10, obtained by Monte Carlo sampling from the polar decomposition of an attaining measure (attainment itself follows from weak-* compactness under dictionary compactness). Stage two replaces each selected outer Brownian profile by its piecewise-linear interpolant on a uniform grid, adding error L1L-11 per atom, where L1L-12. Balanced refinement L1L-13 yields L1L-14 overall.

Two features strengthen the result. The interpolation constant L1L-15 is proved sharp: a normalized tent profile supported on a single grid interval attains the bound exactly, so neither the rate nor the constant can be improved. And evaluation is sparse: each hat function has support on two adjacent intervals, so evaluating the approximant requires at most L1L-16 active outer-profile basis contributions independently of L1L-17. Notably, the lower-level supports are never discretized—they remain elements of L1L-18—so the construction realizes every element of the full space while leaving recursive discretization open.

Experiments

The experiments are theory-directed rather than benchmark-driven. On the constructive side, fixed smooth profiles interpolate faster than the worst-case rate, while the normalized tent profile matches the theoretical bound at every tested resolution, empirically confirming sharpness. Signed-measure discretization follows the predicted L1L-19 slope, and balanced refinement exhibits the UL={xg(u(x)):uUL1,  gHk(B),  gHk(B)1}.U_L = \{x \mapsto g(u(x)) : u \in U_{L-1},\; g \in H_{k^{(B)}},\; \|g\|_{H_{k^{(B)}} \le 1}\}.0 behavior of the corollary.

On the supervised side, validation-selected VBKL realizations are compared against DNVS, kernel ridge regression, and random Fourier features on a matched recursive teacher and the Energy Efficiency benchmark. The results support a limited-data advantage rather than universal dominance:

UL={xg(u(x)):uUL1,  gHk(B),  gHk(B)1}.U_L = \{x \mapsto g(u(x)) : u \in U_{L-1},\; g \in H_{k^{(B)}},\; \|g\|_{H_{k^{(B)}} \le 1}\}.1 VBKL wins vs. DNVS Parameter ratio (DNVS/VBKL)
100 4/5 4.6×
250 2/5 13.7×
500 2/5 18.3×

At UL={xg(u(x)):uUL1,  gHk(B),  gHk(B)1}.U_L = \{x \mapsto g(u(x)) : u \in U_{L-1},\; g \in H_{k^{(B)}},\; \|g\|_{H_{k^{(B)}} \le 1}\}.2, VBKL achieves mean test MSE of 0.0517 versus 0.0927 for DNVS; at UL={xg(u(x)):uUL1,  gHk(B),  gHk(B)1}.U_L = \{x \mapsto g(u(x)) : u \in U_{L-1},\; g \in H_{k^{(B)}},\; \|g\|_{H_{k^{(B)}} \le 1}\}.3 and UL={xg(u(x)):uUL1,  gHk(B),  gHk(B)1}.U_L = \{x \mapsto g(u(x)) : u \in U_{L-1},\; g \in H_{k^{(B)}},\; \|g\|_{H_{k^{(B)}} \le 1}\}.4, DNVS, KRR, and RBF obtain lower mean errors. The honest reading—which the paper itself gives—is a favorable accuracy–parameter trade-off concentrated in small-sample regimes, not predictive dominance. Optimization studies verify the finite-difference directional estimator against automatic differentiation (with the expected accuracy degradation at step size UL={xg(u(x)):uUL1,  gHk(B),  gHk(B)1}.U_L = \{x \mapsto g(u(x)) : u \in U_{L-1},\; g \in H_{k^{(B)}},\; \|g\|_{H_{k^{(B)}} \le 1}\}.5) and variance reduction from Monte Carlo directional averaging; these validate implementation stability but do not constitute a global convergence theorem.

Limitations and open questions

Several limitations are conceded explicitly. The statistical guarantees cover only the finite-architecture class, leaving intrinsic full-space bounds open. The constructive theory leaves lower-level supports undiscretized, so a fully finite realization of the entire recursion remains to be developed. No optimization-convergence theory is provided for the explicit finite architectures. The strict-depth hierarchy depends on both the local non-degeneracy condition and the trace-support assumption; without them, strictness is not asserted. Finally, the empirical advantage is confined to limited-data regimes and two benchmarks, and the comparison protocol does not establish scalability claims beyond the tested sizes.

Conclusion

This paper contributes a path-atomic variation-space framework in which recursive Brownian dictionary construction and outer signed-measure superposition are mathematically distinct. It delivers an exact union-of-RKHS characterization of the recursive dictionaries, quantitative Hölder regularity with a strict depth hierarchy under explicit conditions, architecture-dependent Rademacher and generalization bounds whose proof passes through quadratic chaos and VC entropy, and a two-stage constructive approximation with a sharp interpolation constant and resolution-independent evaluation sparsity. The experiments corroborate the theoretical mechanisms and indicate a favorable limited-data accuracy–complexity trade-off, while the scope restrictions—finite-architecture statistics, undiscretized lower supports, and absent optimization theory—mark the precise boundaries of what has been established.

Paper to Video (Beta)

No one has generated a video about this paper yet.

Whiteboard

No one has generated a whiteboard explanation for this paper yet.

Open Problems

We found no open problems mentioned in this paper.