Papers
Topics
Authors
Recent
Search
2000 character limit reached

Limitations of Learning Tanh Neural Networks with Finite Precision

Published 9 Jun 2026 in cs.LG and stat.ML | (2606.11104v1)

Abstract: We investigate limitations of learning tanh\tanh neural networks from point evaluations under finite-precision computations and L<sup>pL<sup>p accuracy guarantees, building on Berner, Grohs, and Voigtländer (2023). Our approach is based on a novel construction of sharply localized bump functions via iterated tanh\tanh activations. Using this mechanism, we show that, in a finite-precision setting, no adaptive randomized algorithm based on mm samples can achieve a convergence rate higher than the Monte Carlo rate O(m<sup>1/p)O(m<sup>{-1/p}) in the L<sup>pL<sup>p norm, unless the sampling budget grows exponentially with the size of the network parameters and architecture. The results reveal fundamental limitations imposed by finite precision on the learnability of classes containing localized bump functions, extending previous results for ReLU networks to the tanh\tanh setting.

Authors (2)

Summary

  • The paper proves that any adaptive randomized learner for finite-precision tanh networks faces an L^p error lower bound of order m^{-1/p}, with exponentially large admissible sample budgets as network depth and width increase.
  • The authors construct effectively localized bumps by iterating scaled tanh functions toward nonzero fixed points, using quantization to make exponentially small tails indistinguishable from zero.
  • For L^∞ recovery, the lower bound can remain a positive constant regardless of sample count; in one example, even 10^71 samples leave worst-case error at least 0.49, demonstrating severe instability.

Context and motivation

The paper studies the sampling complexity of learning classes of feedforward neural networks with the hyperbolic tangent activation σ=tanh\sigma = \tanh from noiseless point evaluations, under the assumption that all computations are performed in finite-precision arithmetic. The work extends a line of research initiated by Berner, Grohs, and Voigtländer for ReLU networks (2606.11104), where it was shown that learning error decays only at the Monte Carlo rate m1/pm^{-1/p} in LpL^p norm while growing exponentially with network size — implying that even computationally unbounded algorithms cannot learn such classes efficiently from samples alone.

The extension to tanh\tanh is nontrivial for two structural reasons. First, unlike ReLU, tanh\tanh lacks positive homogeneity (ρ(λx)=λρ(x)\rho(\lambda x) = \lambda\rho(x)), which was a key tool in prior analyses. Second, tanh\tanh is real-analytic and strictly nonzero away from isolated points, so networks cannot vanish exactly on sets of positive measure; compactly supported bump constructions are unavailable in exact form. The authors overcome both obstacles by introducing finite precision as an essential modeling ingredient rather than a nuisance: values below a machine threshold ϵp>0\epsilon_p > 0 are quantized to zero, which permits "effective" compact support.

Main result

The central theorem states that if UU contains all tanh\tanh networks of input dimension m1/pm^{-1/p}0, depth m1/pm^{-1/p}1, width m1/pm^{-1/p}2, and coefficient bound m1/pm^{-1/p}3 in m1/pm^{-1/p}4 norm (with m1/pm^{-1/p}5, m1/pm^{-1/p}6), then under any quantizer mapping values below m1/pm^{-1/p}7 to zero, every adaptive randomized algorithm using m1/pm^{-1/p}8 samples incurs worst-case m1/pm^{-1/p}9 error at least

LpL^p0

The bound holds for arbitrary LpL^p1, with LpL^p2 an effective dimension parameter and LpL^p3 a depth budget allocated to maximizing the admissible sample count. Three qualitative features deserve emphasis:

  • Monte Carlo rate barrier: no algorithm can beat LpL^p4, regardless of adaptivity or randomization.
  • Exponential sample requirement: the admissible LpL^p5 grows exponentially in architecture parameters, so efficient learning is impossible as network size increases.
  • Intrinsic instability: for LpL^p6 the lower bound stays bounded below by a positive constant independent of LpL^p7. In a concrete instance (LpL^p8, LpL^p9, tanh\tanh0, tanh\tanh1, tanh\tanh2, tanh\tanh3), even tanh\tanh4 samples leave worst-case error tanh\tanh5, and there exist pairs of networks whose sampled values differ by only tanh\tanh6 while their tanh\tanh7 distance remains about tanh\tanh8. Unlike analytic interpolation, where node choice (e.g., Chebyshev points) can restore stability, here no placement of sample points yields simultaneously accurate and stable reconstruction within the stated regime.

The result applies to any finite-precision scheme with the threshold property, including IEEE 754 arithmetic; moreover, when tanh\tanh9 one can shift the reference level so that the relevant precision is relative (tanh\tanh0), making the infeasibility substantially more severe. An interactive application accompanies the paper for numerically verifying the full bound's assumptions.

Technical mechanism: fixed-point-driven bump construction

The proof rests on a novel construction of sharply localized bumps via iterated tanh\tanh1 activations, exploiting the fixed-point structure of tanh\tanh2 for tanh\tanh3. For tanh\tanh4, tanh\tanh5 has nontrivial fixed points tanh\tanh6 beyond tanh\tanh7, and its iterates converge pointwise to the step-like limit tanh\tanh8. The analysis quantifies this convergence: defining the critical point tanh\tanh9 and a jump scale ρ(λx)=λρ(x)\rho(\lambda x) = \lambda\rho(x)0 with ρ(λx)=λρ(x)\rho(\lambda x) = \lambda\rho(x)1, the authors establish exit-time bounds into the basin of attraction and exponential attraction estimates showing that after ρ(λx)=λρ(x)\rho(\lambda x) = \lambda\rho(x)2 iterations, all inputs above a shrinking threshold ρ(λx)=λρ(x)\rho(\lambda x) = \lambda\rho(x)3 land within ρ(λx)=λρ(x)\rho(\lambda x) = \lambda\rho(x)4 of the fixed point.

The bump network itself is built layer by layer: the first layer forms shifted sign-flipped ρ(λx)=λρ(x)\rho(\lambda x) = \lambda\rho(x)5 pairs centered at a grid point ρ(λx)=λρ(x)\rho(\lambda x) = \lambda\rho(x)6; the second averages these coordinate-wise primitives and thresholds via a bias, producing a coarse bump positive on ρ(λx)=λρ(x)\rho(\lambda x) = \lambda\rho(x)7 and negative outside ρ(λx)=λρ(x)\rho(\lambda x) = \lambda\rho(x)8; middle layers apply ρ(λx)=λρ(x)\rho(\lambda x) = \lambda\rho(x)9, sharpening the profile toward a step function; the final layer selects one coordinate and shifts by tanh\tanh0. The resulting network satisfies tanh\tanh1 on a cube of width tanh\tanh2 while decaying like tanh\tanh3 outside — fast enough that, given sufficient depth tanh\tanh4, the tail falls below tanh\tanh5 and becomes indistinguishable from zero.

The lower bound then follows a packing argument: choosing tanh\tanh6 and placing bump centers on a grid dense in tanh\tanh7 selected coordinates, a pigeonhole argument shows at least half the bumps avoid all sample points entirely; for those, the algorithm receives exactly the same quantized data as for the zero function despite positive tanh\tanh8 mass. A Markov-type reduction extends the argument from deterministic to adaptive randomized algorithms. A useful refinement is the effective dimension reduction: restricting the packing grid to an tanh\tanh9-dimensional active subspace improves the constant by a factor involving ϵp>0\epsilon_p > 00, so smaller ϵp>0\epsilon_p > 01 strengthens the bound.

Comparison with the ReLU setting

For ϵp>0\epsilon_p > 02, the analogous ReLU bound contains the factor ϵp>0\epsilon_p > 03. Two structural differences emerge. First, the characteristic factor ϵp>0\epsilon_p > 04 (here ϵp>0\epsilon_p > 05) present in the ReLU estimate does not appear in the ϵp>0\epsilon_p > 06 bound. Second, whereas the ReLU bound grows exponentially in architecture parameters — reflecting ReLU's value amplification through composition — the ϵp>0\epsilon_p > 07 bound remains uniformly bounded since ϵp>0\epsilon_p > 08; only the exponential dependence of the admissible sample budget on architecture persists. In particular, for ϵp>0\epsilon_p > 09 the ReLU bound still decays with UU0, while the UU1 bound does not, yielding the stronger instability phenomenon described above. Non-asymptotic comparison requires the full statement, as the remaining factors trade off in either direction depending on parameters.

Limitations and open questions

The paper concedes several restrictions. The choice of first-layer bias UU2 introduces an undesirable dependence on UU3 in the final bound; an alternative scaling UU4 could remove this dependence but leads to more involved estimates and is explicitly deferred to future work. Additionally, the intermediate slope UU5 may fall below one, increasing the threshold required to enter the basin of attraction of UU6, which weakens the constants. The quantitative guarantees also require technical inequalities among UU7, UU8, UU9, depth allocation tanh\tanh0, and tanh\tanh1 to hold; whether the bounds are sharp, or extend to other sigmoidal activations (GELU, logistic) or to noisy samples, remains open.

Conclusion

This paper establishes that finite-precision sampling complexity limitations previously known for ReLU networks persist for tanh\tanh2 networks: under finite-precision arithmetic, no adaptive randomized algorithm based on polynomially many samples can achieve better than Monte Carlo rate accuracy, and uniform recovery is unstable regardless of sample placement. The key methodological contribution is a fixed-point-based mechanism for constructing effectively localized bumps from real-analytic activations, demonstrating that finite precision can substitute exactly for the compact support and homogeneity properties unavailable in the sigmoidal setting.

Paper to Video (Beta)

No one has generated a video about this paper yet.

Whiteboard

No one has generated a whiteboard explanation for this paper yet.

Open Problems

We haven't generated a list of open problems mentioned in this paper yet.