Papers
Topics
Authors
Recent
Search
2000 character limit reached

A Federated Many-to-One Hopfield model for associative Neural Networks

Published 20 Mar 2026 in cond-mat.dis-nn and stat.ML | (2603.19902v1)

Abstract: Federated learning enables collaborative training without sharing raw data, but struggles under client heterogeneity and streaming distribution shifts, where drift and novel data can impair convergence and cause forgetting. We propose a federated associative-memory framework that learns shared archetypes in heterogeneous, continual settings, where client data are independent but not necessarily balanced. Each client encodes its experience as a low-rank Hebbian operator, sent to a central server for aggregation and factorization into global archetypes. This approach preserves privacy, avoids centralized replay buffers, and is robust to small, noisy, or evolving datasets. We cast aggregation as a low-rank-plus-noise spectral inference problem, deriving theoretical thresholds for detectability and retrieval robustness. An entropy-based controller balances stability and plasticity in streaming regimes. Experiments with heterogeneous clients, drift, and novelty show improved global archetype reconstruction and associative retrieval, supporting the spectral view of federated consolidation.

Summary

  • The paper introduces a federated associative-memory method that exchanges low-rank Hebbian operators, enabling servers to recover shared archetypes without accessing raw client data or gradients.
  • The paper shows that archetype exposure controls spectral detectability through BBP thresholds, with simulations matching predicted eigenvalue locations at R² values of 0.955–0.977 and overlap errors as low as 0.0163.
  • The paper demonstrates that entropy-based client weighting balances stability and plasticity, detects newly emerging classes, and reduces a pure-noise attacker’s influence while maintaining mean retrieval above 0.8 under 20% contamination.

Overview and motivation

This paper proposes a federated associative-memory framework in which clients communicate low-rank Hebbian operators rather than raw data or gradients, and a central server reconstructs a shared set of latent archetypes via spectral factorization. The setting is deliberately non-i.i.d.: clients observe independent but heterogeneous, unbalanced, and time-varying subsets of the archetype population, with streaming drift and mid-training emergence of novel classes. The framework targets three persistent difficulties of federated learning (FL) simultaneously—statistical heterogeneity, catastrophic forgetting under distribution shift, and privacy constraints that preclude centralized replay buffers.

The core architectural idea is to cast federated consolidation as a spectral inference problem. Each client compresses its local experience into a Hebbian correlator; the server averages these operators into a global synaptic matrix and factorizes it into archetypes using an LL-layer Associative Memory (LAM) model with anti-imitative inter-layer couplings. Because the aggregated operator concentrates around a low-rank-plus-noise population matrix whose spike strengths are governed by class exposure, random-matrix theory yields explicit detectability thresholds for when an archetype becomes recoverable from federation-level statistics alone.

Problem setting

The generative model assumes KK binary Rademacher archetypes {ξμ}μ=1K∈{−1,+1}N\{\boldsymbol{\xi}^\mu\}_{\mu=1}^K \in \{-1,+1\}^N. Examples are generated by selecting an archetype μ\mu and flipping each bit independently with probability (1−r)/2(1-r)/2, where r∈(0,1]r\in(0,1] is the dataset quality. Client cc at round tt receives MctM_c^t examples drawn from a client-specific mixture πt,c(μ)\pi_{t,c}(\mu), which vanishes outside the client's support KK0, while the union over clients covers all KK1 archetypes. Two diagnostic quantities organize the analysis: the federated exposure KK2 (the empirical realization of the round-level global mixture KK3) and the coverage set KK4 of archetypes observed at least once up to round KK5.

The server-side reconstruction machinery builds on two prior results: the LAM architecture of Agliari et al., which can disentangle spurious mixture states through anti-imitative inter-layer interactions, and its extension operating directly on synaptic matrices. The paper's factorization pipeline is fully unsupervised: it approximates the pseudo-inverse (Kanter–Sompolinsky) kernel KK6 from the Hebbian matrix via iterative unlearning, extracts leading eigenvectors as orthogonalized pattern combinations, synthesizes nonlinear mixtures from these eigenvectors, runs the LAM dynamics on them, and prunes candidates using (i) a mutual-overlap deduplication threshold (KK7 flags duplicates) and (ii) a spectral acceptance test based on the quadratic form KK8 exceeding a threshold KK9 calibrated by Marchenko–Pastur or shuffle-null criteria.

The federated pipeline

The protocol alternates four steps over {ξμ}μ=1K∈{−1,+1}N\{\boldsymbol{\xi}^\mu\}_{\mu=1}^K \in \{-1,+1\}^N0 rounds. Clients initialize local correlators from their first batch; the server aggregates by equal-weight averaging into {ξμ}μ=1K∈{−1,+1}N\{\boldsymbol{\xi}^\mu\}_{\mu=1}^K \in \{-1,+1\}^N1, runs the LAM-based factorization to obtain {ξμ}μ=1K∈{−1,+1}N\{\boldsymbol{\xi}^\mu\}_{\mu=1}^K \in \{-1,+1\}^N2 reconstructed patterns, and re-encodes them as a broadcast operator {ξμ}μ=1K∈{−1,+1}N\{\boldsymbol{\xi}^\mu\}_{\mu=1}^K \in \{-1,+1\}^N3; each client then fuses this reconstruction with fresh local evidence through a convex combination controlled by a weight {ξμ}μ=1K∈{−1,+1}N\{\boldsymbol{\xi}^\mu\}_{\mu=1}^K \in \{-1,+1\}^N4. Small {ξμ}μ=1K∈{−1,+1}N\{\boldsymbol{\xi}^\mu\}_{\mu=1}^K \in \{-1,+1\}^N5 favors consolidation (low variance, bias toward stale information); large {ξμ}μ=1K∈{−1,+1}N\{\boldsymbol{\xi}^\mu\}_{\mu=1}^K \in \{-1,+1\}^N6 favors plasticity (low staleness bias, high batch-noise variance). The authors explicitly frame this knob as an operator-level analogue of the Complementary Learning Systems hypothesis, with fast episodic traces integrating into a slower consolidated store.

Reconstruction quality is measured by the magnetization—the maximum normalized overlap between any reconstructed pattern and the ground-truth archetype—and, when the teacher operator is available, by the normalized Frobenius error against the true Hebbian matrix.

Spectral theory: exposure governs detectability

The theoretical contribution centers on a spiked-Wishart decomposition of the rescaled round operator:

{ξμ}μ=1K∈{−1,+1}N\{\boldsymbol{\xi}^\mu\}_{\mu=1}^K \in \{-1,+1\}^N7

with per-coordinate noise variance {ξμ}μ=1K∈{−1,+1}N\{\boldsymbol{\xi}^\mu\}_{\mu=1}^K \in \{-1,+1\}^N8 in the Rademacher channel. Each archetype contributes a rank-one spike of strength {ξμ}μ=1K∈{−1,+1}N\{\boldsymbol{\xi}^\mu\}_{\mu=1}^K \in \{-1,+1\}^N9, so spike strength is directly proportional to exposure. This makes the organizing principle explicit: exposure determines spike strength, spike strength determines spectral separation, and separation determines whether an archetype is discoverable at that round.

Two theorems support the pipeline. The first establishes non-asymptotic concentration of μ\mu0 around its expectation in operator norm, with sub-Gaussian and bounded/Rademacher variants scaling exponentially in the effective round size μ\mu1. Notably, the variance proxy involves the spectral map μ\mu2, so low-noise regimes (large μ\mu3) yield tighter concentration—a favorable trade-off the authors highlight. The second theorem gives the full BBP phenomenology: bulk confinement within a band around the Marchenko–Pastur edges; a detection threshold at μ\mu4 (with aspect ratio μ\mu5); closed-form outlier locations μ\mu6; and eigenvector alignment μ\mu7. These results justify the operational detector μ\mu8 and the use of leading eigenvectors as reconstruction seeds.

Numerical validation at finite μ\mu9 shows empirical outlier locations matching the BBP prediction with (1−r)/2(1-r)/20 across 50 trials, and eigenvector overlaps tracking the alignment formula with median gap 0.0163. The residual gap concentrates on the weakest archetype (maximum 0.0486), consistent with the predicted (1−r)/2(1-r)/21 finite-size correction—an honest acknowledgment that the asymptotic formula degrades precisely in the low-exposure regime where detection matters most.

An important scope caveat is stated plainly: the guarantees assume a fixed effective noise level, whereas the experiments use fully adaptive client weights and pure-noise attacker clients. The theory predicts detectability provided the adaptive controller successfully suppresses high-variance clients; consistency between theory and practice in the adaptive regime is asserted rather than proven.

Entropy-based stability–plasticity control

The adaptive weight (1−r)/2(1-r)/22 is derived from sign-agreement statistics between the consolidated operator (1−r)/2(1-r)/23 and the current batch correlator (1−r)/2(1-r)/24. The empirical sign-agreement probability (1−r)/2(1-r)/25 is mapped through the binary entropy (1−r)/2(1-r)/26 and normalized against a noise floor (1−r)/2(1-r)/27, which accounts for the intrinsic corruption channel—so discrepancies explainable by noise do not trigger plasticity. An exponential moving average smooths the resulting schedule.

Two properties make this controller notable. First, it requires no labels or ground truth: each client computes the statistic locally. Second, it provides automatic robustness to corrupted clients. In a five-client experiment with one attacker receiving pure noise ((1−r)/2(1-r)/28), the attacker's weight collapses to near zero within roughly 10 rounds while good clients stabilize at (1−r)/2(1-r)/29–r∈(0,1]r\in(0,1]0; mean retrieval remains above 0.8 despite 20% contamination. This is a strong empirical claim: unsupervised sign-entropy alone suffices to discriminate signal from adversarial noise injection at the operator level.

Numerical results

Three experimental regimes validate the framework progressively.

Exposure drift. With r∈(0,1]r\in(0,1]1 clients and r∈(0,1]r\in(0,1]2 archetypes under a sequential dominance schedule, both fixed-weight extremes fail in complementary ways: r∈(0,1]r\in(0,1]3 prevents acquisition of later archetypes (r∈(0,1]r\in(0,1]4), while r∈(0,1]r\in(0,1]5 tracks the instantaneous mixture and forgets rapidly once exposures decline. The entropy controller spikes during transitions and relaxes during plateaus, sustaining high magnetization across all modes—acquiring newly dominant archetypes without overwriting consolidated ones.

Novelty emergence. Starting from r∈(0,1]r\in(0,1]6 archetypes with r∈(0,1]r\in(0,1]7 introduced at round 12 via a four-round ramp, the effective rank estimate r∈(0,1]r\in(0,1]8 transitions from 3 to 6 within a few rounds after introduction, indicating low-latency novelty detection driven by the BBP mechanism. Old archetypes show only a mild transient dip and recover, consistent with targeted rather than indiscriminate plasticity. A relative spectral gap at the r∈(0,1]r\in(0,1]9 boundary collapses during the ramp and restabilizes after expansion, providing an architecture-agnostic novelty signature.

Structured data. On Caltech-101 binary silhouettes (cc0, spatially correlated archetypes), the pipeline recovers all three patterns: magnetizations exceed 0.9 quickly and remain stable at cc1, while cc2 yields slower convergence with wider variability but still non-trivial reconstruction. The authors concede this experiment validates structured-but-binary data only and does not address high-dimensional image distributions; extending to richer representations is left open.

Limitations and open questions

Several limitations are acknowledged or evident. The theoretical guarantees cover fixed effective noise levels and balanced representative rounds (cc3); the fully adaptive, heterogeneous-budget, adversarial setting analyzed only empirically lacks formal treatment. The BBP alignment formula is asymptotic, and finite-cc4 deviations are largest exactly for weakly exposed archetypes. Privacy is preserved structurally (only operators are communicated) but no formal guarantee—such as differential privacy calibrated at the operator level—is provided, and the induced spectral degradation under DP noise is unanalyzed. The framework is restricted to linear Hebbian operators over binary patterns; kernelized or deep-feature extensions, asynchronous participation, and partial-client regimes connected to time-varying detectability phase transitions remain open questions posed by the authors.

Conclusion

This paper formulates federated archetype learning as a low-rank-plus-noise spectral inference problem over exchanged Hebbian operators, deriving BBP-type thresholds showing that exposure governs when archetypes become detectable and retrievable, and introducing an unsupervised entropy-based controller that resolves the stability–plasticity trade-off while automatically down-weighting corrupted clients. Experiments under drift, novelty, and structured data support the spectral view of federated consolidation, though the formal theory covers a narrower regime than the empirical setting, and privacy, scalability to rich data distributions, and asynchronous dynamics remain open.

Paper to Video (Beta)

No one has generated a video about this paper yet.

Whiteboard

No one has generated a whiteboard explanation for this paper yet.

Open Problems

We haven't generated a list of open problems mentioned in this paper yet.

Tweets

Sign up for free to view the 1 tweet with 0 likes about this paper.