Activation-Integral Theory in Neural Networks
- Activation-Integral Theory is a mathematical framework that unifies neural network activation functions with integral representations to characterize function spaces.
- It establishes sharp approximation rates for finite-width networks through discretized integral representations and convex optimization methods.
- The theory extends to arbitrary activation functions and rule-based schemes, supporting principled design, improved trainability, and enhanced explainability.
Activation-Integral Theory (AIT) provides a rigorous mathematical and conceptual framework linking neural network activation functions and integral representations in both finite- and infinite-width regimes. Originating in the context of shallow neural networks with rectified polynomial (ReLU) activations, AIT unifies integral geometry, function space theory, and neural network expressivity. The theory yields a precise characterization of which functions can be represented by infinite-width neural architectures with bounded norms and provides sharp approximation rates for practical (finite-width) discretizations. Recent developments generalize AIT to arbitrary activation functions, integral transforms, and rule-based logic schemes, allowing principled design and analysis of novel nonlinearities, trainability, and explainability.
1. Integral Representation of Functions via Activation Functions
The foundational result of AIT is the exact integral representation of target functions as -weighted superpositions of ridge functions parameterized by the activation function. For ReLU activations, every in the Sobolev space (for a bounded Lipschitz domain ) admits the representation
where is the ReLU activation, is the unit sphere in 0, and 1 is a suitable coefficient function. Norm equivalence holds: 2 This mirrors the classical Barron integral representation (ReLU, 3) and identifies a fundamental correspondence between Sobolev regularity and neural network representations (Liu et al., 1 May 2025).
For the ordinary ReLU (ReLU4), the sphere-based and ridge-based formulations, and explicit inversion formulas involving Radon transforms, provide constructive recipes for infinite- and finite-width representers (Petrosyan et al., 2019).
2. Function Spaces Representable by Integral Neural Architectures
AIT precisely characterizes the admissible function class for each activation:
- ReLU/Barron space: ReLU-shallow networks with bounded 5 outer weights represent functions in Barron space, corresponding to 6 with finite first-moment Radon transform (Petrosyan et al., 2019).
- ReLU7/Sobolev spaces: For general 8, admissible 9 lie in 0, and the integral representation utilizes 1 coefficients (Liu et al., 1 May 2025).
- RePU polynomials: Using rectified power units 2, the representable class comprises piecewise polynomials (in 1D) and functions with bounded 3-th distributional Laplacian (multivariate) (Abdeljawad et al., 2021).
The minimal integral-norm (e.g., 4 or 5 norm of the coefficient function or measure) provides a natural complexity measure, controlling generalization and approximation error rates.
3. Approximation Rates and Linearized Network Constructions
A direct consequence of the integral representation is the derivation of optimal n-width (finite representation) approximation rates using "linearized" (fixed inner parameters, trainable linear coefficients) shallow networks. For ReLU6 networks, the rate is: 7 for 8 units, which is optimal for the corresponding Sobolev space. The construction uses random or deterministic quadrature rules to discretize the continuous integral, yielding explicit finite sets of parameters and convex optimization of the linear coefficients (Liu et al., 1 May 2025).
For RePU, the approximation error rate for uniform approximations on compacts is 9 for 0 atoms, reflecting the degree of the activation's homogeneity (Abdeljawad et al., 2021).
4. Generalizations Beyond ReLU: Activation-Integral Theory for Arbitrary Nonlinearities
AIT extends to arbitrary scalar or vector-valued activations, provided sufficient regularity and polynomial growth conditions. The theory systematically produces new activation functions by integrating desired gradient flows (e.g., choosing 1 and integrating to obtain 2). Piecewise or smooth gradient schemes generate a large family of novel nonlinearities, including those which interpolate between ReLU and exponential/sigmoid types (Huang et al., 2024).
The "Integral Signatures" framework formalizes the propagation and regularity statistics of arbitrary activations via a 9-dimensional vector of Gaussian moments, asymptotic slopes, and regularity measures, enabling principled taxonomy, stability classification, and kernel conditioning analysis (Mali et al., 9 Oct 2025).
5. Connection to Rule-Based and Sugeno-Integral Representations
AIT also admits a discrete, rule-based interpretation, especially for binarized neural networks (BNNs). Each BNN neuron's hard threshold can be written as a Sugeno integral over a suitable capacity (fuzzy measure), leading to an explicit, interpretable set-function and equivalent if–then rule set for the neuron's decision logic. The last-layer score in a BNN is similarly a Sugeno integral with a normalized measure. This connection facilitates symbolic reasoning, explainability, and extensions to multicriteria aggregation and verification (Baaj et al., 20 Apr 2026).
6. Extensions: Integral Activation Transform and Complex-Analytic Activations
The Integral Activation Transform (IAT) generalizes classical coordinate-wise nonlinearities to functional transforms based on lifting, nonlinear activation in functional space, and integration against output bases. Specializing IAT to smooth global bases and ReLU yields smoother and more expressive architectures, improving trainability and mitigating vanishing-gradient issues (Zhang et al., 2023).
Complex-analytic integral theorems (e.g., Cauchy’s integral formula) provide blueprints for new highly regular activation functions, such as the Cauchy activation, supporting universal approximation for analytic functions and yielding well-controlled gradient and locality properties. These structures can be efficiently integrated into modern deep architectures, providing theoretical guarantees for approximation and trainability (Li et al., 2024).
7. Implications for Neural Network Theory, Design, and Analysis
AIT recasts the design space of neural activation functions from heuristic experimentation to principled, mathematically grounded construction. It provides:
- Precise characterizations of representable function spaces for various activations and architectures.
- Constructive, sharp approximation rates for finite-width models.
- Taxonomic classification and stability guarantees based on integral signatures and Lyapunov analyses.
- Rule-based, interpretable decision models for binarized and discrete networks.
- Systematic pathways for activation function design via integration of tailored gradient flows.
- Unified perspectives linking classical functional analysis, convex geometry, and deep learning theory.
Recent advances in AIT have led to practical improvements in expressivity, trainability, and interpretability in both standard and novel neural architectures (Liu et al., 1 May 2025, Petrosyan et al., 2019, Abdeljawad et al., 2021, Huang et al., 2024, Mali et al., 9 Oct 2025, Baaj et al., 20 Apr 2026, Zhang et al., 2023, Li et al., 2024).