Papers
Topics
Authors
Recent
Search
2000 character limit reached

Sample Complexity of Scientific Discovery: PAC Learnability of Compositional Function Trees

Published 28 Jun 2026 in cs.LG and stat.ML | (2606.29331v1)

Abstract: Scientific discovery via symbolic regression is often viewed as statistically and computationally intractable because the hypothesis space of expressions grows combinatorially with depth. This paper revisits the statistical side through the lens of PAC learning, focusing on compositional function trees built from a finite vocabulary of smooth operators (e.g., +,×,sin⁡,exp⁡{+,\times,\sin,\exp} and affine maps). We prove that the relevant generalization quantity, Rademacher complexity, hence the excess risk, does not necessarily blow up exponentially with the number of distinct symbolic structures, but is controlled by (i) the depth dd and (ii) the Lipschitz constants of the base operators along the composed computation graph. Concretely, under mild Lipschitz conditions on operators and bounded affine leaves, a finite-union bound over a vocabulary of size K=∣H<em>base∣K=|\mathcal{H}<em>{\mathrm{base}}| together with Maurer-type vector contraction yields Rn(H</em>comp<sup>d)</sup>≤(Kb2L)<sup>d−1R<em>n(H</em>comp<sup>1)\mathfrak{R}_n(\mathcal{H}</em>{\mathrm{comp}}<sup>{d})</sup> \leq (Kb\sqrt{2}L)<sup>{d-1}\mathfrak{R}<em>n(\mathcal{H}</em>{\mathrm{comp}}<sup>{1}) with arity bound bb; corresponding high-probability risk bounds scale as O(L<sup>d/n)\mathcal{O}(L<sup>{d}/\sqrt{n}) when K,b=O(1)K,b=O(1) and R<em>n(H</em>comp<sup>1)=O(n<sup>−1/2)\mathfrak{R}<em>n(\mathcal{H}</em>{\mathrm{comp}}<sup>{1})=O(n<sup>{-1/2}). We complement the theory with a modular codebase that trains differentiable operator trees (not MLPs) on synthetic "physics-like" targets of controlled depth and shows that the empirical generalization gap correlates positively with the predicted complexity term (L^<sup>d)/n(\widehat{L}<sup>{d})/\sqrt{n}.

Summary

  • The paper proves that depth-d compositional function trees have Rademacher complexity bounded by (Kb√2L)^(d−1) times the leaf complexity, separating statistical learnability from computational search difficulty.
  • The analysis shows excess risk scales approximately as L^d/√n for fixed operator vocabulary and arity, indicating that stable operators and modest depths can support reliable discovery from small datasets.
  • Controlled experiments confirm that error generally rises with depth and falls with sample size, but weak quantitative fit and loose high-depth bounds highlight the need for localized, agnostic, and multidimensional analyses.

Motivation and problem statement

Symbolic regression (SR) is frequently dismissed as statistically unlearnable because the number of candidate expressions grows combinatorially with expression depth, and structure search is NP-hard in many formulations. The paper argues that this objection conflates computational hardness with statistical learnability. In the PAC framework (2606.29331), a hypothesis class can have small sample complexity even when optimizing over it is computationally expensive. The authors formalize this separation for compositional function trees built from a finite vocabulary of smooth operators ({+,×,sin⁡,cos⁡,exp⁡,affine}\{+,\times,\sin,\cos,\exp,\mathrm{affine}\}), and show that generalization is governed by compositional stability—Lipschitz constants accumulated along the computation graph—rather than by the count of distinct symbolic structures.

Setup and hypothesis class

The learning setup is standard regression with squared loss over i.i.d. samples from an unknown distribution on bounded domains X⊂RpX \subset \mathbb{R}^p. A compositional hypothesis space is defined recursively: Hcomp1H_{\mathrm{comp}}^{1} contains base hypotheses (bounded affine leaves), and for d≥2d \ge 2,

Hcompd={ h∘(g1,g2):h∈Hbase, g1,g2∈Hcompd−1 },H_{\mathrm{comp}}^{d} = \{\, h \circ (g_1, g_2) : h \in H_{\mathrm{base}},\ g_1, g_2 \in H_{\mathrm{comp}}^{d-1} \,\},

with maximum arity bb (binary in the main analysis). Two assumptions carry the theory: bounded affine leaves (bounded covariates and leaf parameters, giving Rn(Hcomp1)=O(Wmax⁡Xmax⁡p/n)R_n(H_{\mathrm{comp}}^{1}) = \mathcal{O}(W_{\max}X_{\max}\sqrt{p/n})) and local Lipschitzness of each base operator on the data-induced range, with L=max⁡hLhL = \max_h L_h. The local treatment matters because operators like exp⁡\exp are not globally Lipschitz; the paper works on ranges induced by bounded measurement domains and learned parameters.

Main theoretical results

The technical core combines three ingredients into a tree-structured contraction argument:

  • Finite union over vocabulary: at each composition level, a union bound over the K=∣Hbase∣K = |H_{\mathrm{base}}| operators contributes a factor X⊂RpX \subset \mathbb{R}^p0, not the number of tree shapes.
  • Maurer's vector contraction: composing an X⊂RpX \subset \mathbb{R}^p1-Lipschitz map X⊂RpX \subset \mathbb{R}^p2 with a vector-valued class costs X⊂RpX \subset \mathbb{R}^p3 on the coordinatewise Rademacher complexity.
  • Product classes do not multiply complexity: pairing child subtrees adds complexities rather than multiplying them, costing a factor X⊂RpX \subset \mathbb{R}^p4.

Iterating these yields the depth bound

X⊂RpX \subset \mathbb{R}^p5

which translates via standard uniform convergence to a high-probability excess-risk bound of order X⊂RpX \subset \mathbb{R}^p6 plus a concentration term, holding X⊂RpX \subset \mathbb{R}^p7 and X⊂RpX \subset \mathbb{R}^p8 as constants. The corresponding PAC sample complexity for X⊂RpX \subset \mathbb{R}^p9-excess risk scales as Hcomp1H_{\mathrm{comp}}^{1}0.

The central claim is that depth enters sample complexity through multiplicative stability growth—Hcomp1H_{\mathrm{comp}}^{1}1 when tracked explicitly—rather than through the super-exponentially many tree shapes of depth Hcomp1H_{\mathrm{comp}}^{1}2. This implies that for stable vocabularies and modest depth, small-Hcomp1H_{\mathrm{comp}}^{1}3 scientific datasets can suffice for low excess risk; the bottleneck shifts to discrete search and inductive bias design, not statistics. The paper also sketches a PAC-Bayes extension combining MDL-style priors over tree shapes with the Lipschitz-depth multiplier, trading description length against stability accumulation.

Empirical validation

The experiments train differentiable operator trees (no MLPs) under deliberate realizability—ground-truth targets lie in the hypothesis family at matching depth and vocabulary—to isolate statistical scaling from misspecification. Targets range from affine (Hcomp1H_{\mathrm{comp}}^{1}4) to compositions involving Hcomp1H_{\mathrm{comp}}^{1}5 plus affine terms (Hcomp1H_{\mathrm{comp}}^{1}6), with Hcomp1H_{\mathrm{comp}}^{1}7, 20 seeds per configuration (400 runs total), and evaluation on Hcomp1H_{\mathrm{comp}}^{1}8 held-out samples. A data-dependent Lipschitz proxy Hcomp1H_{\mathrm{comp}}^{1}9 is estimated from gradient norms on a large batch.

The qualitative predictions hold: gaps increase with depth, decrease with d≥2d \ge 20, and align along an approximately linear upper envelope against d≥2d \ge 21. However, the quantitative fit is modest, and the paper reports it plainly: the log-log Pearson correlation between the predicted term and observed gap is d≥2d \ge 22, and a power-law fit through the origin yields exponent d≥2d \ge 23 with d≥2d \ge 24. The authors attribute the fitted exponent being far below 1 to nuisance variation—finite-batch noise in d≥2d \ge 25, imperfect convergence, and mismatch between local derivative norms and worst-case uniform Lipschitz constants—rather than falsification of the qualitative roles of d≥2d \ge 26 and depth amplification. Notably, at d≥2d \ge 27–d≥2d \ge 28 the predicted term spans orders of magnitude due to exponentiation in d≥2d \ge 29 while gaps remain small, indicating the bound is loose in these regimes.

Limitations and open questions

The paper is explicit about several constraints. The empirical study is deliberately controlled: one-dimensional uniformly bounded inputs, realizability-aligned architectures, no distribution shift, so the results say nothing about agnostic or misspecified settings beyond noting that an approximation-error term would be added. The estimator Hcompd={ h∘(g1,g2):h∈Hbase, g1,g2∈Hcompd−1 },H_{\mathrm{comp}}^{d} = \{\, h \circ (g_1, g_2) : h \in H_{\mathrm{base}},\ g_1, g_2 \in H_{\mathrm{comp}}^{d-1} \,\},0 measures stability of the trained map in-distribution rather than certifying a worst-case global constant; for Hcompd={ h∘(g1,g2):h∈Hbase, g1,g2∈Hcompd−1 },H_{\mathrm{comp}}^{d} = \{\, h \circ (g_1, g_2) : h \in H_{\mathrm{base}},\ g_1, g_2 \in H_{\mathrm{comp}}^{d-1} \,\},1 compositions, local slopes can grow rapidly with learned ranges, so Hcompd={ h∘(g1,g2):h∈Hbase, g1,g2∈Hcompd−1 },H_{\mathrm{comp}}^{d} = \{\, h \circ (g_1, g_2) : h \in H_{\mathrm{base}},\ g_1, g_2 \in H_{\mathrm{comp}}^{d-1} \,\},2 need not reflect uniform operator stability without envelope constraints or interval arithmetic. Uniform bounds are also potentially loose relative to localized or hypothesis-dependent alternatives (covering-number local Rademacher complexities, offset frameworks), which the paper identifies but does not develop. Open questions left by the work include multi-dimensional covariate bookkeeping, agnostic risk decompositions, and tightly integrating PAC-Bayes syntax priors with the contraction estimates in end-to-end SR systems that alternate discrete proposals with continuous refitting.

Conclusion

The paper establishes that the statistical complexity of depth-Hcompd={ h∘(g1,g2):h∈Hbase, g1,g2∈Hcompd−1 },H_{\mathrm{comp}}^{d} = \{\, h \circ (g_1, g_2) : h \in H_{\mathrm{base}},\ g_1, g_2 \in H_{\mathrm{comp}}^{d-1} \,\},3 compositional function trees scales as Hcompd={ h∘(g1,g2):h∈Hbase, g1,g2∈Hcompd−1 },H_{\mathrm{comp}}^{d} = \{\, h \circ (g_1, g_2) : h \in H_{\mathrm{base}},\ g_1, g_2 \in H_{\mathrm{comp}}^{d-1} \,\},4, yielding Hcompd={ h∘(g1,g2):h∈Hbase, g1,g2∈Hcompd−1 },H_{\mathrm{comp}}^{d} = \{\, h \circ (g_1, g_2) : h \in H_{\mathrm{base}},\ g_1, g_2 \in H_{\mathrm{comp}}^{d-1} \,\},5 excess risk under fixed vocabulary and arity, independent of the combinatorial count of symbolic structures. Controlled synthetic experiments qualitatively confirm the predicted scaling, though with substantial unexplained variance and loose constants at higher depths. The practical upshot is that for stable operator vocabularies and modest depth, the statistical side of symbolic discovery from small datasets is feasible, and the binding constraint lies in search algorithms and inductive bias rather than learnability.

Paper to Video (Beta)

No one has generated a video about this paper yet.

Whiteboard

No one has generated a whiteboard explanation for this paper yet.

Open Problems

We haven't generated a list of open problems mentioned in this paper yet.