Papers
Topics
Authors
Recent
Search
2000 character limit reached

Free energy landscape of Dense Associative Memory

Published 21 Jul 2026 in cond-mat.dis-nn, cond-mat.stat-mech, and cs.AI | (2607.19195v1)

Abstract: Using large deviations theory, we solve and obtain a general expression for the free energy functional for a broad class of associative memories, including dense associative memories. We illustrate the method by reproducing classical results for the Hopfield model. For a finite number of patterns, we derive the temperature-dependent free energy functional for dense associative memories featuring polynomial interactions and Log-Sum-Exponential (LSE) activation. We also evaluate the disorder-averaged ground-state energy of these systems in the extensive limit. Our analytical framework reveals how memory retrieval depends on the initial state in higher-order dense networks, and gives the exact full-retrieval threshold for the LSE model. This method provides a systematic procedure for analyzing diverse, complex architectures in associative memory.

Authors (2)

Summary

  • The paper develops a tilted large-deviations framework that produces exact, temperature-dependent free energy functionals for energy-based associative memories with arbitrary interaction functions.
  • The analysis shows that polynomial DenseAMs with order k>2 retain a persistent nonretrieval minimum at zero overlap, making retrieval highly dependent on initialization even as successful retrieval becomes more accurate.
  • For the Log-Sum-Exponential model, the paper derives the exact full-retrieval threshold α_c(λ)=λ(1−λ/2) for λ≤λ*, where exponentially many stored patterns can still be retrieved without error.​​​​

Overview

The paper develops a large-deviations framework for computing the free energy functional of a general class of energy-based associative memories, including dense associative memories (DenseAMs) with polynomial interactions of arbitrary order and the Log-Sum-Exponential (LSE) model. The central contribution is an exact, temperature-dependent free energy functional F(m)\mathcal{F}(\mathbf{m}) expressed in terms of the overlap vector m=(m1,,mP)\mathbf{m} = (m^1,\dots,m^P) between a spin configuration and the PP stored patterns. The framework reproduces classical replica results for the Hopfield model, extends them to k>2k>2 interactions where no prior free energy expression was available, and yields the exact full-retrieval threshold αc(λ)\alpha_c(\lambda) for the LSE model with exponentially many stored patterns.

Method: tilted large deviations

Neurons are Ising spins si=±1s_i = \pm 1; patterns ξμ\boldsymbol{\xi}^\mu are drawn i.i.d. from a binary distribution. For a general Hamiltonian H=Nf(m)H = -N f(\mathbf{m}), with mμ=1Niξiμsim^\mu = \frac{1}{N}\sum_i \xi_i^\mu s_i, the authors first obtain the rate function I0(m)I_0(\mathbf{m}) of the joint distribution of overlaps via a m=(m1,,mP)\mathbf{m} = (m^1,\dots,m^P)0-dimensional Laplace transform and saddle-point evaluation, where the cumulant generating function is m=(m1,,mP)\mathbf{m} = (m^1,\dots,m^P)1. The Gibbs-weighted distribution is then handled with the tilted large deviation principle, giving

m=(m1,,mP)\mathbf{m} = (m^1,\dots,m^P)2

and the fixed-point condition reduces the saddle-point inversion to the self-consistency equation

m=(m1,,mP)\mathbf{m} = (m^1,\dots,m^P)3

with the free energy functional

m=(m1,,mP)\mathbf{m} = (m^1,\dots,m^P)4

This avoids the m=(m1,,mP)\mathbf{m} = (m^1,\dots,m^P)5-dimensional inversion of m=(m1,,mP)\mathbf{m} = (m^1,\dots,m^P)6 that becomes intractable at large m=(m1,,mP)\mathbf{m} = (m^1,\dots,m^P)7, and—unlike the Hubbard–Stratonovich transformation—handles higher-order interactions directly.

Polynomial DenseAMs: exact free energy and ground-state landscape

For the m=(m1,,mP)\mathbf{m} = (m^1,\dots,m^P)8-th order Hamiltonian m=(m1,,mP)\mathbf{m} = (m^1,\dots,m^P)9, the framework yields closed-form fixed-point equations and free energy. For PP0 the results match Amit, Gutfreund, and Sompolinsky's replica calculation, and the fixed-point equation agrees with the stochastic-dynamics treatment of Rooke et al. For PP1 the free energy functional is new. For a single pattern, transitions are continuous for PP2 and first order for PP3, with transition temperatures PP4 for PP5 respectively.

In the extensive limit PP6, with cross-overlaps treated as Gaussian noise of variance PP7, the disorder-averaged ground-state energy (zero-temperature landscape of gradient descent) is

PP8

with fixed points satisfying PP9; for k>2k>20 this reproduces Amit et al.'s result with k>2k>21.

A key structural finding concerns higher-order networks: for k>2k>22 the curvature at k>2k>23 satisfies k>2k>24 for all k>2k>25, so the k>2k>26 state remains a local minimum indefinitely, while a nonzero retrieval minimum k>2k>27 appears only below a threshold k>2k>28 at which it becomes global. Retrieval therefore depends strongly on the initial state: if the initial configuration lies in the k>2k>29 basin, the pattern is never retrieved at zero noise, for any αc(λ)\alpha_c(\lambda)0. The authors note that escaping this basin requires a mechanism such as the stochastic noise considered by Rooke et al. As αc(λ)\alpha_c(\lambda)1 grows, the αc(λ)\alpha_c(\lambda)2 basin shrinks even though αc(λ)\alpha_c(\lambda)3 (e.g., αc(λ)\alpha_c(\lambda)4 for αc(λ)\alpha_c(\lambda)5), so higher order improves retrieval fidelity but restricts the set of successful initial states. The paper also distinguishes the transition thresholds αc(λ)\alpha_c(\lambda)6 from the conventional capacity threshold defined by an allowed error fraction—for polynomial models these do not coincide because αc(λ)\alpha_c(\lambda)7 at the transition (e.g., αc(λ)\alpha_c(\lambda)8 vs. αc(λ)\alpha_c(\lambda)9 for si=±1s_i = \pm 10).

LSE model: exact full-retrieval threshold

For the LSE Hamiltonian with interaction strength si=±1s_i = \pm 11, the authors introduce the auxiliary quantity si=±1s_i = \pm 12. Using the tilted LDP with si=±1s_i = \pm 13 and Gaussian cross-overlaps of variance si=±1s_i = \pm 14, they obtain si=±1s_i = \pm 15, with an extremum at si=±1s_i = \pm 16. Retrieval is possible only in the phase si=±1s_i = \pm 17; in the complementary phase the model reduces to the standard Hopfield model, whose si=±1s_i = \pm 18 capacity is irrelevant against exponentially many patterns, so no retrieval occurs.

In the si=±1s_i = \pm 19 phase and the ξμ\boldsymbol{\xi}^\mu0 limit, the ground-state landscape becomes ξμ\boldsymbol{\xi}^\mu1, with a unique fixed point at ξμ\boldsymbol{\xi}^\mu2: retrieval is error free. Equating ξμ\boldsymbol{\xi}^\mu3 gives the exact full-retrieval threshold

ξμ\boldsymbol{\xi}^\mu4

with ξμ\boldsymbol{\xi}^\mu5, reproducing the threshold recently derived via the random energy model by Lucibello and Mézard. Notably, unlike polynomial DenseAMs, the transition threshold and the retrieval threshold coincide for LSE because ξμ\boldsymbol{\xi}^\mu6 exactly at and below the transition.

Limitations and open questions

The formalism is derived for binary neurons; extension to continuous or other discrete variables is asserted but not demonstrated. The extensive-limit analysis of the polynomial case treats cross-overlap contributions as Gaussian with variance ξμ\boldsymbol{\xi}^\mu7, an assumption valid by the central limit theorem in the stated scaling but not verified beyond it. The LSE treatment assumes an exponential number of stored patterns, though the method is claimed to apply to arbitrary ξμ\boldsymbol{\xi}^\mu8. The initial-state dependence of gradient descent at ξμ\boldsymbol{\xi}^\mu9 raises the open question of what dynamical mechanisms (e.g., finite-temperature stochasticity) suffice to escape the persistent H=Nf(m)H = -N f(\mathbf{m})0 basin, and how the basin geometry scales with H=Nf(m)H = -N f(\mathbf{m})1 and H=Nf(m)H = -N f(\mathbf{m})2.

Conclusion

The paper provides a systematic large-deviations procedure for computing exact free energy functionals of associative memory models with arbitrary cost functions. It recovers classical Hopfield results, supplies the previously unknown free energy landscape for polynomial DenseAMs of any order, and establishes that higher-order networks exhibit persistent initial-state dependence in retrieval, while for the LSE model it derives the exact full-retrieval capacity threshold H=Nf(m)H = -N f(\mathbf{m})3. The framework extends naturally to other architectures within the class of energy-based memories defined by H=Nf(m)H = -N f(\mathbf{m})4.

Paper to Video (Beta)

No one has generated a video about this paper yet.

Whiteboard

No one has generated a whiteboard explanation for this paper yet.

Open Problems

We haven't generated a list of open problems mentioned in this paper yet.