Papers
Topics
Authors
Recent
Search
2000 character limit reached

Rectified Linear Complexity in ReLU Networks

Updated 12 January 2026
  • Rectified Linear Complexity is a metric that quantifies how the interplay of depth and width in ReLU networks governs the creation of affine (piecewise linear) regions.
  • The analysis reveals that increased depth exponentially multiplies affine segments, thereby enhancing expressivity while imposing computational challenges.
  • Theoretical results, including depth–size gap theorems and zonotope-based lower bounds, underscore the need for deep architectures to efficiently approximate complex functions.

Rectified Linear Complexity denotes the interplay among depth, width, and the number of affine (piecewise linear) regions in functions realized by deep neural networks with rectified linear units (ReLU-DNNs). It quantifies the expressivity of ReLU networks by measuring how their architecture governs the partitioning of input space into regions where the computed function is affine. The notion synthesizes structural and functional complexity of ReLU-DNNs and establishes formal lower bounds relating network architecture to function representation and training complexity (Arora et al., 2016).

1. Function Class and Complexity Measures

A ReLU-DNN with input dimension w0w_0, output dimension wk+1w_{k+1}, and kk hidden layers of widths w1,…,wkw_1, \dots, w_k implements functions:

f(x)=Tk+1∘σ∘Tk∘⋯∘σ∘T1(x)f(x) = T_{k+1} \circ \sigma \circ T_k \circ \cdots \circ \sigma \circ T_1(x)

where Ti:Rwi−1→RwiT_i: \mathbb{R}^{w_{i-1}} \rightarrow \mathbb{R}^{w_i} is affine for i=1…ki=1\dots k, Tk+1T_{k+1} is linear, and σ\sigma applies coordinate-wise as σ(t)=max⁡{0,t}\sigma(t) = \max\{0, t\}.

Key structural and functional measures:

Term Definition Notation
Depth Total number of layers, including output wk+1w_{k+1}0
Width Max hidden layer width wk+1w_{k+1}1
Size Sum of hidden units across layers wk+1w_{k+1}2
Affine Pieces Maximal connected regions on which wk+1w_{k+1}3 is affine Number of PWL regions

Any ReLU-DNN computes a continuous piecewise linear (PWL) function. Conversely, every PWL function wk+1w_{k+1}4 can be represented by a ReLU-DNN of depth wk+1w_{k+1}5. The number of affine regions, i.e., the cardinality of maximal connected input regions mapped affinely, serves as a fundamental complexity metric.

2. Global Optimization for One Hidden Layer

Empirical risk minimization over ReLU networks with one hidden layer and convex loss wk+1w_{k+1}6 can be globally optimized as:

wk+1w_{k+1}7

A globally optimal algorithm proceeds via:

  1. Writing each hidden unit wk+1w_{k+1}8 as wk+1w_{k+1}9, kk0.
  2. Partitioning data by sign of kk1 for all kk2.
  3. Enumerating all sign and partition choices kk3 and all hyperplane partitions kk4, with possible count kk5.
  4. For each, solving the induced convex program in kk6.

Total runtime:

kk7

This is polynomial in sample size kk8 for fixed kk9 but exponential in w1,…,wkw_1, \dots, w_k0 and w1,…,wkw_1, \dots, w_k1, matching known computational hardness bounds.

3. Depth–Size Gap Theorems

The expressivity of ReLU-DNNs grows rapidly with increased depth compared to width or overall size. For integers w1,…,wkw_1, \dots, w_k2, w1,…,wkw_1, \dots, w_k3, there exists a 1D function w1,…,wkw_1, \dots, w_k4 such that:

  • A w1,…,wkw_1, \dots, w_k5-layer ReLU net of width w1,…,wkw_1, \dots, w_k6 represents w1,…,wkw_1, \dots, w_k7.
  • Any representation by a shallower w1,…,wkw_1, \dots, w_k8-layer net with w1,…,wkw_1, \dots, w_k9 incurs a lower bound on required size:

f(x)=Tk+1∘σ∘Tk∘⋯∘σ∘T1(x)f(x) = T_{k+1} \circ \sigma \circ T_k \circ \cdots \circ \sigma \circ T_1(x)0

Furthermore, for every f(x)=Tk+1∘σ∘Tk∘⋯∘σ∘T1(x)f(x) = T_{k+1} \circ \sigma \circ T_k \circ \cdots \circ \sigma \circ T_1(x)1, there exists a member of a smoothly-parameterized family (by f(x)=Tk+1∘σ∘Tk∘⋯∘σ∘T1(x)f(x) = T_{k+1} \circ \sigma \circ T_k \circ \cdots \circ \sigma \circ T_1(x)2):

  • f(x)=Tk+1∘σ∘Tk∘⋯∘σ∘T1(x)f(x) = T_{k+1} \circ \sigma \circ T_k \circ \cdots \circ \sigma \circ T_1(x)3 is realized by a depth f(x)=Tk+1∘σ∘Tk∘⋯∘σ∘T1(x)f(x) = T_{k+1} \circ \sigma \circ T_k \circ \cdots \circ \sigma \circ T_1(x)4 net of size f(x)=Tk+1∘σ∘Tk∘⋯∘σ∘T1(x)f(x) = T_{k+1} \circ \sigma \circ T_k \circ \cdots \circ \sigma \circ T_1(x)5.
  • Any depth f(x)=Tk+1∘σ∘Tk∘⋯∘σ∘T1(x)f(x) = T_{k+1} \circ \sigma \circ T_k \circ \cdots \circ \sigma \circ T_1(x)6 ReLU net computing f(x)=Tk+1∘σ∘Tk∘⋯∘σ∘T1(x)f(x) = T_{k+1} \circ \sigma \circ T_k \circ \cdots \circ \sigma \circ T_1(x)7 requires at least:

f(x)=Tk+1∘σ∘Tk∘⋯∘σ∘T1(x)f(x) = T_{k+1} \circ \sigma \circ T_k \circ \cdots \circ \sigma \circ T_1(x)8

The construction uses sawtooth-composed functions—composition amplifies the number of affine segments exponentially in depth.

4. Lower Bounds via Zonotope Constructions

A new lower bound for affine region count in ReLU-DNNs is established via the theory of zonotopes. For vectors f(x)=Tk+1∘σ∘Tk∘⋯∘σ∘T1(x)f(x) = T_{k+1} \circ \sigma \circ T_k \circ \cdots \circ \sigma \circ T_1(x)9, the zonotope Ti:Rwi−1→RwiT_i: \mathbb{R}^{w_{i-1}} \rightarrow \mathbb{R}^{w_i}0 is:

Ti:Rwi−1→RwiT_i: \mathbb{R}^{w_{i-1}} \rightarrow \mathbb{R}^{w_i}1

and its support function:

Ti:Rwi−1→RwiT_i: \mathbb{R}^{w_{i-1}} \rightarrow \mathbb{R}^{w_i}2

For general Ti:Rwi−1→RwiT_i: \mathbb{R}^{w_{i-1}} \rightarrow \mathbb{R}^{w_i}3, Ti:Rwi−1→RwiT_i: \mathbb{R}^{w_{i-1}} \rightarrow \mathbb{R}^{w_i}4 has:

Ti:Rwi−1→RwiT_i: \mathbb{R}^{w_{i-1}} \rightarrow \mathbb{R}^{w_i}5

distinct affine pieces. Ti:Rwi−1→RwiT_i: \mathbb{R}^{w_{i-1}} \rightarrow \mathbb{R}^{w_i}6 can be implemented by a two-layer ReLU net of size Ti:Rwi−1→RwiT_i: \mathbb{R}^{w_{i-1}} \rightarrow \mathbb{R}^{w_i}7.

Composition with a Ti:Rwi−1→RwiT_i: \mathbb{R}^{w_{i-1}} \rightarrow \mathbb{R}^{w_i}8-fold sawtooth map Ti:Rwi−1→RwiT_i: \mathbb{R}^{w_{i-1}} \rightarrow \mathbb{R}^{w_i}9 yields a ReLU net of depth i=1…ki=1\dots k0, size i=1…ki=1\dots k1, and number of segments:

i=1…ki=1\dots k2

Asymptotically, choosing i=1…ki=1\dots k3 or i=1…ki=1\dots k4,

i=1…ki=1\dots k5

For depth i=1…ki=1\dots k6, matching this piece count requires size at least i=1…ki=1\dots k7.

5. Synthesis and Implications of Rectified Linear Complexity

The composition of depth and width exponentially increases the count of affine regions:

  • Depth acts as the exponential composition resource; each layer can multiply the region count.
  • Width (or size) determines the parallel granularity per layer.
  • The total number of affine pieces grows as i=1…ki=1\dots k8 in 1D or i=1…ki=1\dots k9 in Tk+1T_{k+1}0.

The triplet Tk+1T_{k+1}1 forms a natural complexity measure—Rectified Linear Complexity—which encapsulates the expressive power of ReLU networks. Deeper networks attain exponential region growth with moderate width, while shallow networks require super-polynomial size for function approximation equivalence.

A plausible implication is that for function classes requiring exponentially many affine pieces, depth is indispensable for architectural efficiency. Furthermore, computational hardness in training aligns with the representational barriers: even for one hidden layer, the exponential increase in complexity with input dimensionality indicates that training algorithms are fundamentally limited by both representational and computational regimes (Arora et al., 2016).

Definition Search Book Streamline Icon: https://streamlinehq.com
References (1)

Topic to Video (Beta)

No one has generated a video about this topic yet.

Whiteboard

No one has generated a whiteboard explanation for this topic yet.

Follow Topic

Get notified by email when new papers are published related to Rectified Linear Complexity.