Binary Perceptrons: Models, Capacity, & Learning
- Binary perceptrons are discrete threshold units with binary states that function as half-space indicators, forming the foundation of binary neural networks.
- They quantify storage capacity and representational power through rigorous methods such as replica theory, Gaussian integrals, and minimal hidden-width investigations.
- Algorithmic approaches, including CSP techniques and belief propagation, leverage solution space clustering and local entropy to enhance efficient learning.
Binary perceptrons are threshold units and networks whose synaptic states, activations, or both are restricted to discrete binary values, typically or . In the literature, the term spans several closely related objects: a single threshold neuron with binary weights, multilayer feedforward networks composed of binary units, and random constraint-satisfaction models in which one seeks a binary weight vector satisfying a set of random inequalities. Across these formulations, the central technical questions are the same: what functions binary perceptrons can represent, how many hidden units are minimally required, what storage capacity they achieve under random constraints, how their solution spaces are organized in Hamming space, and which algorithms can learn or find solutions efficiently (Crespin, 2013, Bi et al., 2019, Pastor et al., 2015, Stojnic, 2023).
1. Formal models and basic definitions
A standard threshold perceptron with real input and bias computes
In this form, a perceptron is exactly the characteristic function of a half-space, and a perceptron layer is a finite vector of such threshold units (Crespin, 2013). A closely related binary-neuron convention uses signed outputs,
with hidden and output units both binary-valued, typically in (0904.4587).
For storage problems, one often considers a binary-weight perceptron receiving binary patterns and producing
with . Introducing labels 0 and 1, the feasibility condition can be written as
2
where 3 is a stability margin and the storage load is 4 (Bi et al., 2019).
A second major formulation is the random half-space or Gaussian perceptron model. In the asymmetric binary perceptron, one seeks 5 such that
6
or equivalently 7 (Abbe et al., 2021). In the symmetric binary perceptron, the constraint is two-sided: 8 or, in Gaussian notation,
9
with constraint density 0 and SAT/UNSAT thresholds defined through Gaussian integrals (Barbier, 2024, Abbe et al., 2021).
Symmetric variants can also be expressed via even constraint indicators on Gaussian fields 1. Two canonical examples are the rectangle-binary-perceptron, 2, and the 3-function-binary-perceptron, 4. Their global 5 symmetry distinguishes them from the standard step perceptron (Aubin et al., 2019).
A different but related inference formulation is the teacher-student symmetric perceptron, where teacher weights 6 generate noiseless labels
7
while the student predicts
8
This turns a storage model into a planted learning problem parameterized by sample density 9, margin 0, and temperature 1 (Catania et al., 26 Mar 2026).
2. Representational structure and network construction
Because a single threshold perceptron is the indicator of a half-space, multilayer single-output perceptron networks can be characterized geometrically. The functional equivalence theorem states that, for a fixed list of half-spaces 2, the class of Boolean-valued functions 3 with 4 in the Boolean algebra generated by 5 coincides exactly with the class of single-output perceptron network functions over 6 (Crespin, 2013). Every such function can be written in DNF or CNF over cells and cocells induced by the half-spaces, and every single-output perceptron network is functionally equivalent to a 3-layer network over the same half-spaces (Crespin, 2013).
This geometric viewpoint has a direct constructive form. Given a DNF decomposition
7
one may build a 3-layer network whose first layer computes the half-spaces, whose second layer consists of AND-units recognizing the cells 8, and whose third layer is an OR-unit that unions them (Crespin, 2013). The same logic yields a CNF construction by duality. This establishes an exact representational correspondence rather than an approximation theorem.
Minimal architecture questions have also been studied directly for memorization. For three-layer feedforward networks with 9 binary inputs, one hidden layer of 0 sigmoid units, and one binary output, a local-neighbor complexity index is defined by
1
where 2 contains Hamming-distance-one neighbors of input 3 (Pastor et al., 2015). This complexity measure supports closed-form estimates of the minimal hidden-layer size 4 for regular, random, and intermediate binary patterns.
For regular patterns, including pseudo-parity structure, the minimal hidden width is
5
while random patterns obey
6
to first order, with a refined correction
7
Intermediate patterns are treated as perturbations of a nearest regular template, with Hamming-distance fraction 8, yielding the interpolation
9
This places hidden-layer size under an explicit complexity law rather than a generic “larger network for harder task” heuristic (Pastor et al., 2015).
Incremental constructive learning provides a second route to compact architecture. The NetLines algorithm grows a feedforward network with one hidden layer of binary units and a binary output unit. Its convergence theorem guarantees that, for any finite training set of 0 patterns with binary or real inputs, zero training error is reached with at most 1 hidden units (0904.4587). In this setting, representational sufficiency is tied to an explicit growth process rather than solely to VC-style counting arguments.
3. Capacity, thresholds, and phase structure
The storage capacity of the classical binary perceptron with threshold 2 is the critical load
3
Using fully lifted random duality theory, this capacity is characterized through a one-dimensional fixed-point equation at the second, first non-trivial lifting level: 4 For the zero-threshold case, the resulting scaled capacity is
5
matching the replica-symmetry-breaking prediction (Stojnic, 2023).
In the entropy-landscape analysis of the binary perceptron with random classifications, the satisfiable-to-unsatisfiable transition likewise occurs at
6
At this threshold the replica-symmetric entropy vanishes, while the typical inter-solution distance remains finite at approximately 7 (Huang et al., 2013). This is a geometric capacity statement: the total number of solutions goes to zero before the solution set collapses to a point.
For symmetric perceptrons, the annealed capacity can be exact over substantial parameter regions. In the rectangle-binary-perceptron, the critical density equals the annealed bound
8
for all 9, under the stated hypothesis controlling the second-moment saddle. In the 0-function-binary-perceptron, the same conclusion holds for narrow constraints 1, with
2
while for 3 the annealed bound is only an upper bound and one-step RSB lowers the threshold slightly; the paper concludes that full-RSB would be required to obtain the exact capacity in that regime (Aubin et al., 2019).
The teacher-student symmetric perceptron introduces a distinct phase diagram organized by overlap with the planted teacher. At zero temperature, the Bayes-optimal onset of teacher correlation occurs continuously at
4
while perfect teacher recovery appears through a first-order transition at 5, obtained by equating competing free energies (Catania et al., 26 Mar 2026). The resulting phases are paramagnetic (6), suboptimal correlated (7), and perfect-teacher (8), and the sequence of transitions depends on 9 and on whether the potential is piecewise-constant or linear (Catania et al., 26 Mar 2026).
Algorithmic thresholds can remain far below statistical capacity. For the asymmetric binary perceptron at 0, a discrepancy-based polynomial-time result proves
1
whereas the sharp SAT threshold remains near 2 (Li et al., 2024). In the large positive-margin regime, however, the same line of work shows
3
matching the information-theoretic capacity asymptotically as 4 (Li et al., 2024). The article literature therefore separates sharply between exact capacity, annealed bounds, and polynomial-time achievability.
4. Geometry of solution spaces
The solution-space geometry of binary perceptrons is one of the most developed aspects of the subject. For the classical binary perceptron, the entropy density of solutions at fixed Hamming distance from a reference configuration is obtained through a Legendre transform of a field-biased partition function,
5
At small constraint density, annealed and replica-symmetric descriptions agree well, but as 6 approaches capacity the allowed distance interval shrinks and an entropy-landscape gap develops (Huang et al., 2013). In the paired-solution landscape, the entropy curve becomes non-concave on the left wing, signaling a first-order transition in the overlap-conjugate field and supporting a clustered, glassy organization (Huang et al., 2013).
The same work concludes that the binary perceptron solution space near capacity consists of exponentially many isolated solutions: pure states have zero internal entropy, almost all spins are frozen within a state, and moving from one solution to another requires flipping 7 bits (Huang et al., 2013). This “clustering with freezing” picture is reinforced by symmetric perceptron analyses based on second moments and planted arguments, where the shape of
8
implies a forbidden interval near 9, so each reference solution is isolated inside a point-like frozen-1RSB cluster (Aubin et al., 2019).
At the same time, several works emphasize that typical geometry does not exhaust the algorithmically relevant structure. In the symmetric and asymmetric binary perceptrons at low load, there exists a subdominant but connected component of solutions—a “wide web”—with diameter 0 in the symmetric case and at least 1 in the asymmetric case (Abbe et al., 2021). An 2 randomized multiscale majority algorithm can find a solution in such a cluster with high probability when 3 (Abbe et al., 2021). This establishes formally that isolated typical solutions can coexist with rare linearly wide connected clusters.
Local-entropy analyses sharpen this contrast between typical and atypical states. In the binary negative-margin perceptron, typical solutions lie in exponentially many narrow frozen-1RSB clusters, but subdominant wide-flat minima can be selected by biasing toward high local entropy,
4
As the constraint density increases, the Franz-Parisi local entropy of maximally robust solutions ceases to be monotone at the local-entropy threshold 5, beyond which wide-flat minima fragment into disconnected finite-radius islands even though the SAT phase persists until 6 (Baldassi et al., 2023).
For zero-threshold asymmetric binary perceptrons, a large-deviation fully lifted random duality analysis locates local-entropy breakdown in the interval
7
with 8 for the underlying feasibility problem (Stojnic, 24 Jun 2025). The paper reports that this interval basically matches the range 9–0 that currently best solvers can handle, suggesting that the loss of positive local entropy at high overlap is a structural marker of the computational gap (Stojnic, 24 Jun 2025).
Connected atypical states in the symmetric binary perceptron have also been studied through chains of highly overlapping solutions. Under a no-memory ansatz with constant overlap 1, the overlap matrix is Markovian,
2
and the infinite-chain potential remains positive only above a second threshold
3
Below this threshold, decorrelated chains can still exist, but require a nested Markov chain ansatz with explicit memory effects (Barbier, 2024). A common misconception is therefore that isolated typical solutions preclude all long-range connectivity; the cited results show instead that connectivity survives in atypical, algorithmically significant sectors of the space.
5. Learning, search, and training algorithms
One classical family of algorithms studies the binary perceptron as a discrete CSP. In belief-propagation decimation, at step 4 one computes the marginal probability
5
chooses the most polarized unfixed variable
6
fixes 7 to its preferred value, and simplifies the factor graph (Bi et al., 2019). The associated message updates use cavity fields 8 and constraint-to-variable messages 9 derived from belief propagation on the factor graph (Bi et al., 2019).
The same paper analyzes two efficient solvers, SBPI and reinforced BP. SBPI maintains hidden odd-integer states 00 whose sign determines the weight, and updates them pattern by pattern, while reinforced BP adds a reinforcement term 01 to standard BP updates (Bi et al., 2019). A central empirical result is that most runtime is spent resolving late-decimation variables, whose values are strongly cross-correlated in the condensed residual subspace. Input sparseness reduces this bottleneck by weakening cross-correlations among late-fixed weights, thereby reducing the time used to assign them (Bi et al., 2019).
At low load, a different algorithmic mechanism operates. The multiscale majority algorithm partitions coordinates into blocks, repeatedly identifies the most troublesome rows, and assigns new coordinates by weighted majority votes
02
For both symmetric and asymmetric perceptrons, this yields an 03 randomized algorithm that finds a solution in a linearly wide cluster when 04 (Abbe et al., 2021). The proof combines concentration of partial row sums with tree-indexed interpolation paths in the solution graph (Abbe et al., 2021).
Discrepancy-minimization provides a third algorithmic paradigm. The Rothvoss-Eldan-Singh random-projection method solves a linear program over 05 and recursively rounds it; the Lovett-Meka edge-walk performs Gaussian steps in the subspace orthogonal to nearly tight constraints,
06
until at least half the coordinates are nearly frozen, then recurses (Li et al., 2024). These methods furnish the best-known polynomial-time guarantees across all 07, including asymptotic optimality for large positive 08 and an exponentially large algorithmic-statistical gap for 09 (Li et al., 2024).
Constructive supervised learning with binary neurons predates these CSP-style methods. NetLines incrementally adds hidden perceptrons, updates intermediate targets by
10
and trains each unit with the Minimerror cost
11
The algorithm has a finite-step convergence guarantee and was evaluated on parity, Monk’s problems, Wisconsin Breast Cancer, Pima Diabetes, Waveform, and Iris benchmarks (0904.4587).
The minimal-perceptron study validates theoretical hidden-width predictions by back-propagation on three-layer sigmoid networks. Training is declared successful when the mean-squared error falls below 12, and minimal 13 is located by sweeping hidden width until the success rate exceeds 14 over many random seeds (Pastor et al., 2015). The measured 15 falls on the theoretical curves 16 for regular patterns, 17 for random patterns, and the interpolation law for complex patterns (Pastor et al., 2015).
A recent development is fully binary-native multilayer training. A binary multilayer perceptron with 18 fully connected hidden layers can be trained using fixed random local classifiers 19, local 20–21 losses, binary activations 22, visible weights 23, and integer-valued hidden metaplastic weights 24 (Colombo et al., 2024). The forward pass and updates use only XNOR, Popcount, and increment/decrement operations, while the CP+R update rule is
25
plus a reinforcement step 26 applied with probability 27 (Colombo et al., 2024). On MNIST, FashionMNIST, and CIFAR-10 features, the method reports test-accuracy gains over the only existing fully binary single-layer state-of-the-art solution while using two to three orders of magnitude fewer Boolean gates than full-precision SGD under the same total memory demand (Colombo et al., 2024).
6. Modern binary architectures, distance transformation, and application domains
Binary perceptrons also appear as components of larger binary neural architectures. In BiMLP, vision MLP blocks are binarized using
28
with inference implemented by XNOR + POPCOUNT and gradients approximated by STE clipping (Xu et al., 2022). The paper argues that fully connected layers in vision MLPs behave like 29 convolutions, so binarization sharply restricts spatial and channel mixing capacity. To compensate, BiMLP introduces a multi-branch binary block and a universal shortcut
30
with branch outputs summed or concatenated depending on channel changes (Xu et al., 2022).
The capacity argument in BiMLP is explicitly perceptronic. A binarized 31 convolution aggregates 32 one-bit products, so each scalar output can take 33 levels, whereas a 34 binarized FC has only 35 levels. The proposed multi-branch binary block recovers part of this representational deficit with three branches rather than a 36-fold channel blowup for 37 (Xu et al., 2022). On ImageNet-1k with bit-width 38, BiMLP-S reaches Top-1 39 and BiMLP-M reaches Top-1 40 with lower OP counts than several prior binary CNN baselines (Xu et al., 2022).
In neuroscience-oriented analyses, networks of binary perceptrons have been studied as maps between Hamming spaces. For a perceptron with binary weight vector 41, threshold 42, and binary inputs 43 of equal Hamming weight 44 and mutual distance 45, the expected output disagreement probability 46 can be written exactly as a combinatorial sum over active-set occupancies 47 (Olypher et al., 2013). In the large-48 regime, a bivariate-normal approximation gives
49
where 50 and 51 is the number of active synapses (Olypher et al., 2013). Applied to a CA352CA1 hippocampal model with 53 and 54, the study finds maximal discriminability 55 at 56 for 57, while lowering the threshold to 58 reduces discrimination to 59 at the same 60 (Olypher et al., 2013).
Teacher-student and negative-margin studies connect geometry to generalization. In the binary negative-margin perceptron, the generalization error of a student 61 relative to a teacher 62 is
63
and the reported result is that solutions in wide-flat, local-entropy-selected clusters achieve higher teacher overlap and lower 64 than typical sharp minima, even in the highly underconstrained regime of very negative margins (Baldassi et al., 2023). This suggests that, within binary perceptron models, robustness and generalization are controlled less by the mere existence of solutions than by whether the accessible solutions belong to dense connected or high-local-entropy regions.
Taken together, these lines of work show that binary perceptrons are not a single model but a family of discrete threshold systems linking threshold logic, spin-glass theory, combinatorial optimization, constructive learning, and binary deep learning. Their mathematical interest lies in the unusually explicit relation between representation, capacity, geometry, and algorithmics: three-layer constructions can be exact (Crespin, 2013), minimal hidden widths can be expressed in terms of pattern complexity (Pastor et al., 2015), storage capacity can be computed at replica-symmetry-breaking level with 65 (Stojnic, 2023), and algorithmic success or failure can often be read directly from the existence or collapse of rare dense connected clusters (Abbe et al., 2021, Stojnic, 24 Jun 2025).