---
title: Generalized Elephant Random Walk
url: https://www.emergentmind.com/topics/generalized-elephant-random-walk
type: topic
---

# Generalized Elephant Random Walk

The generalized elephant random walk (GERW) denotes a family of non-Markovian random walks in which the distribution of the next increment depends on a retained, sampled, transformed, or shared record of the past. The classical one-dimensional elephant random walk (ERW), introduced by Schütz and Trimper, already exhibits a memory-driven phase transition: depending on the memory parameter, the walk is diffusive, critical, or superdiffusive. Subsequent work has generalized this mechanism along several axes, including higher-dimensional state spaces, nonlinear reinforcement maps, arbitrary step laws, random step sizes, graph-based shared memory, multiple extractions from the past, varying or random memory windows, and interacting particle systems with exclusion. Across these variants, a recurrent theme is that asymptotics are governed by spectral data, fixed-point derivatives, or stochastic-approximation drifts rather than by independence or Markovianity [1608.01305], [2406.19383], [2004.02004].

## 1. Classical model and principal directions of generalization

In the classical ERW, one considers a discrete-time process on \(\mathbb Z\) with increments \(\eta_n\in\{+1,-1\}\), initial bias \(q\in[0,1]\), and memory parameter \(p\in[0,1]\). The first step satisfies
\[
P\{\eta_1=+1\}=q,\qquad P\{\eta_1=-1\}=1-q,
\]
and for \(n\ge 2\), one chooses \(k\) uniformly from \(\{1,\dots,n-1\}\), then sets \(\eta_n=\eta_k\) with probability \(p\) and \(\eta_n=-\eta_k\) with probability \(1-p\). The position is \(X_n=\sum_{i=1}^n \eta_i\). In the original higher-dimensional extension, one replaces \(\pm1\) by the \(2d\) nearest-neighbor vectors \(\{\pm e_1,\dots,\pm e_d\}\), repeats the sampled direction with probability \(p\), and otherwise chooses uniformly among the other \(2d-1\) directions [1608.01305].

The classical model already has exact moment formulas and a three-regime asymptotic theory. In one dimension, the threshold is \(p=3/4\): \(p<3/4\) gives \(\sqrt n\)-scale Gaussian fluctuations, \(p=3/4\) gives a \(\sqrt{n\log n}\) correction, and \(p>3/4\) yields almost-sure convergence of \(n^{-(2p-1)}S_{\lfloor tn\rfloor}\) to \(Yt^{2p-1}\) with non-Gaussian \(Y\) [1608.01305].

The term “generalized elephant random walk” is used in the literature for several non-equivalent extensions. Some retain the linear memory law but change the geometry or the step law; others replace the linear reinforcement by a nonlinear memory map, or replace single-step recall by sampling several past steps, or allow a population of elephants to share memory through a graph, or let memory be restricted, gradually increasing, or randomly thinned [2406.19383], [2410.22969].

| Variant | Defining mechanism | Representative control parameter |
|---|---|---|
| Multi-dimensional ERW | Complete memory on \(\mathbb Z^d\) with \(2d\) directions | \(p_c^d=\frac{2d+1}{4d}\) |
| Graph-based shared memory | Each elephant samples from in-neighbours in a directed graph | spectrum of \(B\) |
| Nonlinear GERW | \(P(X_{n+1}=+1\mid\mathcal F_n)=h(V_n/n)\) with \(h(x)=(1-p)+(2p-1)f(x)\) | \(\alpha=h'(y^*)\) |
| Multiple extractions | \(k\) sampled past times with majority or general \(f\) | \(p^*(k),\,p_c(k)\) or \(\tau\) |
| Random/varying memory | Memory set is truncated, growing, or randomly selected | \(m_n/n\), or two-stage sampling |
| Random step sizes / general step law | Memory acts on signs or past increments, but step magnitudes are random | \(a_n,\ v_n,\ \alpha\) |

## 2. Analytical representations

A central representation is the embedding into generalized Pólya urns. In the one-dimensional ERW, the two-color urn has mean replacement matrix
\[
A=\begin{pmatrix} p&1-p\\[2pt] 1-p&p\end{pmatrix},
\]
with eigenvalues \(\lambda_1=1\) and \(\lambda_2=2p-1\), and the walk satisfies
\[
(X_n)_{n\ge0}\stackrel d= (B_n-R_n)_{n\ge0}.
\]
This identifies the position with the projection of the urn composition onto the second eigenspace and makes Janson’s urn limit theorems applicable. The same idea extends to \(\mathbb Z^d\): the \(2d\)-color urn has mean replacement matrix
\[
A=\frac{1-p}{2d-1}J_{2d}+\frac{2dp-1}{2d-1}I_{2d},
\]
with simple eigenvalue \(\lambda_1=1\) and subleading eigenvalue
\[
\lambda_2=\frac{2dp-1}{2d-1}
\]
of multiplicity \(2d-1\) [2004.02004].

A second dominant framework is stochastic approximation. For graph-based shared memory, if \(S_n=(S_1(n),\dots,S_k(n))\) and \(B\) is the memory matrix defined by
\[
b_{i,j}=\frac{2p_j-1}{d_j^{in}}\quad \text{if } (i,j)\in E,
\]
then
\[
E[X_{n+1}\mid \mathcal F_n]=(S_n/n)\,B,
\]
and with \(Z_n=S_n/n\),
\[
Z_{n+1}=Z_n+\frac1{n+1}\bigl[Z_n(B-I)+\Delta M_{n+1}\bigr].
\]
This places the model in standard stochastic-approximation form with drift \(h(\theta)=\theta(I-B)\) [2410.22969].

Nonlinear GERW models use the same paradigm with a nonlinear drift. In one dimension, if \(V_n\) is the number of \(+1\) steps up to time \(n\) and \(p_n=V_n/n\), then
\[
P(X_{n+1}=+1\mid \mathcal F_n)=h(p_n),\qquad h(x)=(1-p)+(2p-1)f(x),
\]
where \(f:[0,1]\to[0,1]\) is a memory map. In the multidimensional version, the normalized auxiliary chain \(I_n\) satisfies
\[
I_{n+1}=I_n-\frac1{n+1}\bigl(I_n-H(I_n)\bigr)+\frac1{n+1}\varepsilon_{n+1},
\]
so asymptotics are determined by the stable zeros of \(x-H(x)\) and the Jacobian \(JH(x^*)\) [2406.19383].

Several other techniques coexist with urn and SA methods. Martingale normalization underlies the analysis of random step sizes and of some generalized Bernoulli formulations; strong invariance and Brownian embedding yield approximation rates; a Fokker–Planck equation is used for coupled-memory multidimensional models; and the Hill–Lane–Sudderth nonlinear urn is the appropriate analogue for majority-type multiple-extraction models [2302.06311], [1806.04173], [2507.06478].

## 3. Regimes, thresholds, and scaling laws

The most persistent structural feature of elephant random walks is the existence of asymptotic regimes separated by a critical quantity. In the classical linear one-dimensional model, the critical parameter is \(p=3/4\). For \(0\le p<3/4\),
\[
n^{-1/2}S_{\lfloor tn\rfloor}\Longrightarrow W(t)
\]
in \(D([0,\infty))\), where \(W\) is a centered Gaussian process with covariance
\[
\mathbb E[W(s)W(t)] = \frac{s}{3-4p}\Bigl((t/s)^{2p-1}-1\Bigr),\qquad 0<s\le t.
\]
At \(p=3/4\),
\[
(n\log n)^{-1/2}S_{\lfloor tn\rfloor}\Longrightarrow B(t),
\]
and for \(p>3/4\),
\[
n^{-(2p-1)}S_{\lfloor tn\rfloor}\to Yt^{2p-1}
\quad \text{a.s.}
\]
with nondegenerate non-Gaussian \(Y\) [1608.01305].

In the multi-dimensional ERW on \(\mathbb Z^d\), the role of \(2p-1\) is taken by
\[
\alpha=\frac{2dp-1}{2d-1},
\]
and the critical memory becomes
\[
p_c^d=\frac{2d+1}{4d}.
\]
For \(p<p_c^d\), one has
\[
\bigl(n^{-1/2}S_{\lfloor tn\rfloor}\bigr)_{t\ge0}\Rightarrow (W_t)_{t\ge0},
\]
with covariance kernel
\[
\mathbb E[W_sW_t']
=\frac{2d-1}{d(1+2d-4dp)}\,s\Bigl(\frac ts\Bigr)^\alpha I_d,\qquad 0<s\le t.
\]
At criticality, the correct scaling is \(n^t\) in time and \(\sqrt{\log n}\,n^{t/2}\) in space, leading to \((1/\sqrt d)B_t\), while in the superdiffusive regime,
\[
n^{-\alpha}S_{\lfloor tn\rfloor}\to t^\alpha Y
\quad \text{a.s.},
\]
where \(Y\in\mathbb R^d\setminus\{0\}\) is non-Gaussian [2004.02004].

In nonlinear GERW, the transition is not tied to a universal \(p\)-value. If \(y^*\) is the unique stable fixed point of \(h\) and
\[
\alpha=h'(y^*)=(2p-1)f'(y^*),
\]
then the regimes are determined by \(\alpha\): \(\alpha<1/2\) yields a CLT with variance
\[
\sigma^2=\frac{1-(s^*)^2}{1-2\alpha},\qquad s^*=2y^*-1,
\]
\(\alpha=1/2\) gives a \(\sqrt{n/\ln n}\) normalization, and \(1/2<\alpha<1\) gives almost-sure convergence of
\[
n^{1-\alpha}\Bigl(\frac{S_n}{n}-s^*\Bigr)
\]
to a nondegenerate random variable \(L\) [2406.19383].

Graph-based and multiple-extraction models generate analogous but differently parameterized trichotomies. For graph-based shared memory, if \(\eta=\max \Re(\lambda_j)\) over the eigenvalues of \(B\), then \(\eta<1/2\) is globally diffusive, \(\eta=1/2\) is globally critical, and \(\eta>1/2\) produces non-Gaussian leading modes. For fixed-\(k\) general reinforcement, the parameter is \(\tau=1-H'(x^*)\); for growing \(k(n)\), it is \(\tau=1-g'(x^*)\). In both cases, \(\tau>1/2\) is diffusive, \(\tau=1/2\) is critical, and \(\tau<1/2\) is superdiffusive [2410.22969], [2507.14626].

A recurrent misconception is that every generalized ERW inherits the classical threshold \(p=3/4\). The literature shows instead that the threshold depends on the mechanism: it is \(p_c^d=(2d+1)/(4d)\) in the MERW, \(\alpha=1/2\) in nonlinear map-based models, \(\eta=1/2\) in graph-based shared memory, and \(\tau=1/2\) in multiple-extraction models. Some extensions remove the phase transition entirely: finitely restricted memory produces linear drift plus \(\sqrt n\)-Gaussian fluctuations for all \(p\in(0,1)\) [1812.01915].

## 4. Principal model families

The multi-dimensional elephant random walk (MERW) is the direct \(d\)-dimensional analogue of the classical ERW. It is a non-Markovian walk on \(\mathbb Z^d\) with complete memory of previous directions. Its urn embedding is exact, and in the diffusive regime each coordinate of the scaling limit is an independent noise-reinforced Brownian motion. This construction differs from the coupled-memory model of Marquioni, where each coordinate may copy or invert a past step taken in another coordinate. There the coupling matrix
\[
A_{ik}=\alpha_k^i\,\gamma_k^i
\]
controls both first moments and second moments, and two relative-motion regimes appear: a “following” regime and an “opposite” regime. When eigenvalues coincide at or above \(1/2\), logarithmically corrected superdiffusion such as \(t(\ln t)^3\) or \(t^{2\lambda}(\ln t)^2\) can occur, phenomena explicitly described as not found in the classical one-dimensional ERW [2004.02004], [1806.04173].

Shared-memory systems replace a single elephant by a population. In the ERWG model, \(k\) elephants move on \(\mathbb Z\), and elephant \(j\) samples uniformly from the past steps of an in-neighbour \(i\) in a directed graph \(G=(V,E)\). The mean dynamics are encoded by the graph-based matrix \(B\). If \(B\) has a real eigenvalue \(1\) with left-eigenvector \(v\), then
\[
Z_n\to \theta_\infty:=v/\|v\|_1
\]
with \(\theta_\infty B=\theta_\infty\) and \(\theta_\infty\cdot 1=1\). The paper states that in strongly connected graphs with \(\lambda=1\), all elephants coalesce almost surely into a common random limit, and that second-order fluctuations reflect the full spectrum of \(B\) [2410.22969].

Restricted and varying memory lead to another major branch. In the restricted-memory models of Gut and Stadtmüller, the memory set may be the first \(m\) steps, the last \(k\) steps, or a mixture such as \(\{1,n\}\) or \(\{1,2,n\}\). The distant-past-only case becomes independent after the initial stage once the early steps are fixed; recent-past and mixed cases become finite-state Markov chains. Consequently, these models have no phase transition and exhibit standard \(\sqrt n\)-scale CLTs for all \(p\in(0,1)\) [1812.01915]. In the gradually increasing memory model, the memory window \(m_n\to\infty\) but need not equal \(n\). If \(m_n/n\to 0\), the phase transition at \(p=3/4\) reappears, yet \(S_n/n\to 0\) in probability and in \(L^2\). If \(m_n/n\to \alpha\in(0,1]\), the authors derive asymptotics of the mean and variance, including an extra variance term \(\alpha(1-\alpha)\), while stating that the full CLT remains open [2110.13497].

Random-memory ERW introduces an additional averaging layer. At time \(n\), one first chooses \(Y(n)\sim \mathrm{Unif}\{1,\dots,n\}\), then chooses \(K_n\sim \mathrm{Unif}\{1,\dots,Y(n)\}\), and finally sets \(X_{n+1}=a_nX_{K_n}\) with \(a_n\) Rademacher. The conditional mean increment becomes
\[
E[X_{n+1}\mid \mathcal F_n]
=\frac{2p-1}{n}\sum_{r=1}^n \frac{S_r}{r}.
\]
The paper emphasizes that the replacement of the single sum \(\frac{a}{n}S_n\) by the double sum \(\frac{a}{n}\sum_{r=1}^n S_r/r\) slows the growth of the drift and moderates the variance growth relative to the classical full-memory ERW [2501.12866].

Interacting elephant random walks introduce spatial interaction on top of temporal memory. In the exclusion model of Arita and Ragoucy, \(N\) elephants move on a ring of \(L\) sites with simple exclusion and complete memory of attempted displacements \(\sigma_n^{(k)}\in\{-1,0,+1\}\). Monte Carlo simulations and mean-field arguments reveal a condensation phenomenon with a phase transition that is manifestly of first order, and the transition point depends on the initial configuration [1808.04477].

## 5. Generalized increments, weighted observables, and refined limit theory

One important generalization preserves the memory rule but changes the increment law. In the superdiffusive ERW with arbitrary step distribution, the initial jump is \(X_1=\xi_1\), where \(\{\xi_n\}\) are i.i.d. with moments
\[
m_k=E[\xi_1^k],\qquad M_k=E[(\xi_1-m_1)^k].
\]
For \(\alpha\in(1/2,1]\), if \(E[\xi_1^2]<\infty\),
\[
\frac{S_n-nm_1}{n^\alpha}\xrightarrow{\text{a.s.}} Q,
\]
where \(Q\) is non-degenerate, and if \(E[|\xi_1|^p]<\infty\) for some even \(p\), the convergence also holds in \(L^p\). Assuming \(E[\xi_1^4]<\infty\), the first four moments of \(Q\) satisfy
\[
E[Q]=0,
\]
\[
E[Q^2]=\frac{M_2}{(2\alpha-1)\Gamma(2\alpha)},
\]
\[
E[Q^3]=\frac{4M_3}{(3\alpha-1)\Gamma(3\alpha)},
\]
\[
E[Q^4]
=
\frac{6\left[3(2\alpha-1)^2M_4+2(1-\alpha)(5\alpha-2)M_2^2\right]}
{(2\alpha-1)^2(4\alpha-1)\Gamma(4\alpha)}.
\]
These formulas recover Bercu’s symmetric \(\pm1\) case and make the role of skewness and kurtosis explicit for non-symmetric step laws [2112.00066].

Random step-size models separate direction from magnitude. If \(X_n\in\{\pm1\}\) evolves by the usual elephant rule and \(Y_n=X_nZ_n\) with i.i.d. positive step sizes \(Z_n\), then
\[
S_n=\sum_{i=1}^n Y_i,\qquad
a_n=\frac{\Gamma(n)}{\Gamma(n+2p-1)},\qquad
v_n=\sum_{i=1}^n a_i^2.
\]
Under \(E[Z_1^2]<\infty\) and \(0<p\le 3/4\),
\[
\frac{S_n-a_n n(2q-1)}{\sqrt{v_n+a_n^2n\sigma^2}}
\xrightarrow{\mathcal D} N(0,1),
\qquad \sigma^2=\Var(Z_1).
\]
The same paper proves a law of the iterated logarithm and rates of normal approximation in Kolmogorov, Zolotarev, and Wasserstein distances. In particular, under \(E[Z_1^{2+\rho}]<\infty\),
\[
K(W_n)\le
C\Bigl(n^{-\rho/2}+(v_n+a_n^2n\sigma^2)^{-1/2}\Bigr),
\]
with analogous bounds for \(\zeta_r\) and \(W_r\). The case \(Z_i\equiv\) constant recovers the usual ERW, and the paper states that even there the CLT rates are new [2302.06311].

A broader stochastic-algorithm formulation incorporates both varying memory and random step sizes in multiple dimensions. In Zhang’s notation,
\[
S_n=\sum_{k=1}^n \sigma_k,\qquad
T_n=\sum_{k=1}^n \sigma_k Z_k,\qquad
O_n=\Bigl(\frac{S_n}{n},\frac{T_n}{n}\Bigr),
\]
and \(O_n\) obeys a two-component recursive algorithm with drift
\[
h(x,y)=((1-p)x,\; y-\mu_Z x).
\]
This yields Gaussian approximation, CLT, precise law of the iterated logarithm, almost sure central limit theorem, and Chung-type law of the iterated logarithm for the multi-dimensional ERW, the multi-dimensional ERW with random step sizes, and their centers of mass. The paper also defines
\[
C_n=\frac1n\sum_{k=1}^n T_k
\]
and derives joint asymptotics for \((T_n,C_n)\) in all three regimes [2405.12495].

## 6. Structural phenomena, interpretation, and open problems

Generalized elephant random walks occupy a boundary zone between reinforced random walks, generalized urns, stochastic approximation, and interacting particle systems. The classical Pólya-urn representation explains why spectral gaps and eigenspace projections determine functional limits; the stochastic-approximation representation explains why stable fixed points, Jacobians, and linearizations organize fluctuation regimes; and interacting or multiple-extraction extensions show that memory can generate synchronization, attractor selection, or first-order condensation rather than only anomalous diffusion [1608.01305], [2410.22969].

Several qualitative phenomena recur across distinct models. One is non-Gaussian superdiffusion: it appears in the classical ERW, the MERW, graph-based shared memory, nonlinear GERW, arbitrary-step superdiffusive models, and multiple-extraction models. Another is dependence on initial conditions. In the multiple-extraction model with odd \(k\), if \(p>p_c(k)\), the urn equation has three fixed points and the process converges to one of two stable attractors according to the finite-time seed. In the interacting exclusion model, the transition point depends on the initial configuration. These examples show that reinforcement can preserve a persistent memory of early fluctuations even after macroscopic rescaling [2507.06478], [1808.04477].

The literature also corrects a second common oversimplification: more memory does not uniformly imply stronger superdiffusion. Finitely restricted memory removes the classical phase transition; gradually increasing memory with \(m_n/n\to 0\) preserves the threshold but forces \(S_n=o(n)\); and random memory slows the drift relative to full memory because the conditional mean uses the averaged sum \(\sum_{r\le n}S_r/r\) rather than the instantaneous state \(S_n\) alone [1812.01915], [2110.13497], [2501.12866].

Open problems are explicit. For graph-based shared memory, the listed directions include non-uniform sampling from neighbours, time-varying graphs, nonlinear reinforcement leading to ballistic regimes, and replacing the integer line by higher-dimensional targets [2410.22969]. For nonlinear GERW, the stated open questions include recurrence versus transience in the symmetric superdiffusive case, the law of the non-Gaussian limit variable, the critical exponent \(\alpha=1\), and higher-order corrections and large deviations [2406.19383]. For gradually increasing memory with \(m_n/n\to\alpha\in(0,1]\), the full CLT is conjectured in the subcritical regime but not proved [2110.13497].

Taken together, these results show that the generalized elephant random walk is not a single model but a research program: a class of reinforced, history-dependent walks whose asymptotic behavior is shaped by how memory is sampled, aggregated, coupled, or transformed. The unifying mathematical content is the conversion of path dependence into analyzable drift structures—spectral in urn models, dynamical in stochastic approximation, and combinatorial in multiple-extraction or random-memory schemes—while the diversity of outcomes ranges from Brownian limits to non-Gaussian attractors, synchronization, first-order condensation, and sub-linear entropy growth [2507.14626], [2507.06478].

Source: https://www.emergentmind.com/topics/generalized-elephant-random-walk