- The paper develops a tilted large-deviations framework that produces exact, temperature-dependent free energy functionals for energy-based associative memories with arbitrary interaction functions.
- The analysis shows that polynomial DenseAMs with order k>2 retain a persistent nonretrieval minimum at zero overlap, making retrieval highly dependent on initialization even as successful retrieval becomes more accurate.
- For the Log-Sum-Exponential model, the paper derives the exact full-retrieval threshold α_c(λ)=λ(1−λ/2) for λ≤λ*, where exponentially many stored patterns can still be retrieved without error.
Overview
The paper develops a large-deviations framework for computing the free energy functional of a general class of energy-based associative memories, including dense associative memories (DenseAMs) with polynomial interactions of arbitrary order and the Log-Sum-Exponential (LSE) model. The central contribution is an exact, temperature-dependent free energy functional F(m) expressed in terms of the overlap vector m=(m1,…,mP) between a spin configuration and the P stored patterns. The framework reproduces classical replica results for the Hopfield model, extends them to k>2 interactions where no prior free energy expression was available, and yields the exact full-retrieval threshold αc(λ) for the LSE model with exponentially many stored patterns.
Method: tilted large deviations
Neurons are Ising spins si=±1; patterns ξμ are drawn i.i.d. from a binary distribution. For a general Hamiltonian H=−Nf(m), with mμ=N1∑iξiμsi, the authors first obtain the rate function I0(m) of the joint distribution of overlaps via a m=(m1,…,mP)0-dimensional Laplace transform and saddle-point evaluation, where the cumulant generating function is m=(m1,…,mP)1. The Gibbs-weighted distribution is then handled with the tilted large deviation principle, giving
m=(m1,…,mP)2
and the fixed-point condition reduces the saddle-point inversion to the self-consistency equation
m=(m1,…,mP)3
with the free energy functional
m=(m1,…,mP)4
This avoids the m=(m1,…,mP)5-dimensional inversion of m=(m1,…,mP)6 that becomes intractable at large m=(m1,…,mP)7, and—unlike the Hubbard–Stratonovich transformation—handles higher-order interactions directly.
Polynomial DenseAMs: exact free energy and ground-state landscape
For the m=(m1,…,mP)8-th order Hamiltonian m=(m1,…,mP)9, the framework yields closed-form fixed-point equations and free energy. For P0 the results match Amit, Gutfreund, and Sompolinsky's replica calculation, and the fixed-point equation agrees with the stochastic-dynamics treatment of Rooke et al. For P1 the free energy functional is new. For a single pattern, transitions are continuous for P2 and first order for P3, with transition temperatures P4 for P5 respectively.
In the extensive limit P6, with cross-overlaps treated as Gaussian noise of variance P7, the disorder-averaged ground-state energy (zero-temperature landscape of gradient descent) is
P8
with fixed points satisfying P9; for k>20 this reproduces Amit et al.'s result with k>21.
A key structural finding concerns higher-order networks: for k>22 the curvature at k>23 satisfies k>24 for all k>25, so the k>26 state remains a local minimum indefinitely, while a nonzero retrieval minimum k>27 appears only below a threshold k>28 at which it becomes global. Retrieval therefore depends strongly on the initial state: if the initial configuration lies in the k>29 basin, the pattern is never retrieved at zero noise, for any αc(λ)0. The authors note that escaping this basin requires a mechanism such as the stochastic noise considered by Rooke et al. As αc(λ)1 grows, the αc(λ)2 basin shrinks even though αc(λ)3 (e.g., αc(λ)4 for αc(λ)5), so higher order improves retrieval fidelity but restricts the set of successful initial states. The paper also distinguishes the transition thresholds αc(λ)6 from the conventional capacity threshold defined by an allowed error fraction—for polynomial models these do not coincide because αc(λ)7 at the transition (e.g., αc(λ)8 vs. αc(λ)9 for si=±10).
LSE model: exact full-retrieval threshold
For the LSE Hamiltonian with interaction strength si=±11, the authors introduce the auxiliary quantity si=±12. Using the tilted LDP with si=±13 and Gaussian cross-overlaps of variance si=±14, they obtain si=±15, with an extremum at si=±16. Retrieval is possible only in the phase si=±17; in the complementary phase the model reduces to the standard Hopfield model, whose si=±18 capacity is irrelevant against exponentially many patterns, so no retrieval occurs.
In the si=±19 phase and the ξμ0 limit, the ground-state landscape becomes ξμ1, with a unique fixed point at ξμ2: retrieval is error free. Equating ξμ3 gives the exact full-retrieval threshold
ξμ4
with ξμ5, reproducing the threshold recently derived via the random energy model by Lucibello and Mézard. Notably, unlike polynomial DenseAMs, the transition threshold and the retrieval threshold coincide for LSE because ξμ6 exactly at and below the transition.
Limitations and open questions
The formalism is derived for binary neurons; extension to continuous or other discrete variables is asserted but not demonstrated. The extensive-limit analysis of the polynomial case treats cross-overlap contributions as Gaussian with variance ξμ7, an assumption valid by the central limit theorem in the stated scaling but not verified beyond it. The LSE treatment assumes an exponential number of stored patterns, though the method is claimed to apply to arbitrary ξμ8. The initial-state dependence of gradient descent at ξμ9 raises the open question of what dynamical mechanisms (e.g., finite-temperature stochasticity) suffice to escape the persistent H=−Nf(m)0 basin, and how the basin geometry scales with H=−Nf(m)1 and H=−Nf(m)2.
Conclusion
The paper provides a systematic large-deviations procedure for computing exact free energy functionals of associative memory models with arbitrary cost functions. It recovers classical Hopfield results, supplies the previously unknown free energy landscape for polynomial DenseAMs of any order, and establishes that higher-order networks exhibit persistent initial-state dependence in retrieval, while for the LSE model it derives the exact full-retrieval capacity threshold H=−Nf(m)3. The framework extends naturally to other architectures within the class of energy-based memories defined by H=−Nf(m)4.