---
title: 'Acoustic Memory: Mechanisms and Applications'
url: https://www.emergentmind.com/topics/acoustic-memory
type: topic
---

# Acoustic Memory: Mechanisms and Applications

Searching arXiv for recent and foundational papers on acoustic memory across physical, mathematical, and machine-learning contexts.
arXiv search query: "acoustic memory"
Acoustic memory denotes a set of mechanisms by which acoustic degrees of freedom, acoustic propagation histories, or acoustic representations are retained and later re-accessed. In integrated photonics it refers to coherent mapping of optical information onto traveling acoustic phonons and its subsequent retrieval; in dispersive wave theory it refers to after-effect terms represented by temporal convolutions and memory-type boundary conditions; in contemporary audio-language modeling it refers to retention of non-linguistic audio cues such as environmental sounds across multiple conversational turns [1608.08767], [2505.07405], [2605.27039]. Across these literatures, the term names a common functional role—history retention—rather than a single physical implementation.

## 1. Conceptual scope

The literature uses “acoustic memory” in several technically distinct senses. In dispersive and viscoelastic media, memory is constitutive: the present wave field depends on the entire history of deformation, and this dependence is written explicitly through convolution kernels in the governing equations and in acoustic boundary conditions. In transcranial optoacoustics, memory is a propagation invariant: nearby intracranial sources experience nearly the same skull-induced spatio-temporal distortion, so the medium “memorizes” a local distortion kernel. In large audio language models, acoustic memory is representational: the model must retain non-linguistic audio cues heard earlier and later answer a probe about them [2505.07405], [2108.03958], [2605.27039].

A common source of confusion is the conflation of acoustic memory with semantic memory for speech. EnvMem separates these explicitly: semantic memory concerns spoken linguistic content such as words and numeric facts, whereas acoustic memory concerns later recognition or classification of an earlier environmental sound. This distinction is operational rather than philosophical, because the benchmark evaluates the two with matched multi-turn dialogue structure and multiple context lengths [2605.27039].

Another common simplification is to treat acoustic memory as necessarily a resonant storage cavity. The broader record includes traveling-wave Brillouin buffers, multimode bulk-acoustic systems, inverse problems for memory kernels, acoustically trained suspensions, and temporally modulated electroacoustic media. *This suggests that acoustic memory is best understood as a family resemblance across storage, constitutive history dependence, and invariant propagation effects rather than as a single device class.*

## 2. Coherent phononic storage of optical information

In chip-integrated stimulated Brillouin scattering (SBS), acoustic memory is realized by transferring optical information to a coherent hypersound wave and recovering it by the reverse process. The waveguide platform of sputtered chalcogenide glass As\(_2\)S\(_3\) embedded in silica confines both light and GHz-frequency longitudinal acoustic modes through refractive-index and acoustic-impedance mismatch. In the demonstrated rib waveguide, the cross-section is \(2.2~\mu\mathrm{m}\times0.8~\mu\mathrm{m}\), spiral lengths are \(9\)–\(46~\mathrm{cm}\) on a \(20~\mathrm{mm}\times0.7~\mathrm{mm}\) footprint, the effective area is \(1.5~\mu\mathrm{m}^2\), the optical loss is \(0.2~\mathrm{dB/cm}\), the Brillouin shift is \(\Omega_B/2\pi \simeq 7.7~\mathrm{GHz}\), and the acoustic lifetime is \(\tau_{\mathrm{ph}}\simeq 10.5~\mathrm{ns}\). The coupled-mode dynamics are written as
\[
\frac{\partial A_p}{\partial z} + \frac{n}{c}\frac{\partial A_p}{\partial t}
= -\frac{g_0}{2A_{\mathrm{eff}}}QA_s - \frac{\alpha}{2}A_p,
\]
\[
-\frac{\partial A_s}{\partial z} + \frac{n}{c}\frac{\partial A_s}{\partial t}
= +\frac{g_0}{2A_{\mathrm{eff}}}Q^*A_p - \frac{\alpha}{2}A_s,
\]
\[
2\tau_{\mathrm{ph}}\frac{\partial Q}{\partial t}+Q=A_pA_s^*.
\]
The memory cycle is correspondingly write, decay, and read: the optical data and write pulses excite \(Q\), the acoustic envelope decays as \(Q(T)=Q(0^+)\exp(-T/\tau_{\mathrm{ph}})\), and a read pulse generates the retrieved field \(A_{\mathrm{out}}(t)\simeq \kappa Q(T)A_{\mathrm{read}}(t)\) [1608.08767].

The experimentally demonstrated performance establishes a distinctive regime: pulses as short as \(500~\mathrm{ps}\) were stored and retrieved, corresponding to \(\gtrsim 1.5~\mathrm{GHz}\) instantaneous bandwidth, even though the intrinsic Brillouin linewidth is \(\Delta\nu_B\simeq 30~\mathrm{MHz}\). The storage time was continuously tunable from \(0\) to \(\approx 21\times\) the pulse width or up to \(\approx 10.5~\mathrm{ns}\); retrieval efficiency decayed exponentially, with \(32\%\) at \(T=3.5~\mathrm{ns}\) and \(\approx 15\%\) at \(T=10~\mathrm{ns}\). Multi-wavelength operation was demonstrated on two channels separated by \(100~\mathrm{GHz}\), with cross-talk suppression \(<1\%\) and no measurable depletion of channel 2 when writing channel 1. Because the phase-matching condition \(k_{\mathrm{data}}-k_{\mathrm{write}}=K_{\mathrm{acoustic}}\) assigns a unique acoustic wavevector to each wavelength pair, the retrieved photon emerges at the original \(\omega_{\mathrm{data}}\), enabling frequency-multiplexed buffering [1608.08767].

The short native phonon lifetime motivated refreshed-phonon storage. In the refreshed scheme, synchronized optical pulses resonantly reinforce the acoustic wave and counteract intrinsic decay. The minimal model writes the acoustic amplitude \(a(t)\) as
\[
\frac{da}{dt}=-\Gamma_a a(t)+g_{\mathrm{SBS}}E_p^*(t)E_s(t)e^{i\Delta kz-i\Delta\omega t}+\xi(t),
\]
and under periodic refresh the recursion
\[
a_n=a_{n-1}e^{-\Gamma_a\tau}+\delta a
\]
yields a steady-state amplitude \(a_{\mathrm{ss}}=\delta a/[1-e^{-\Gamma_a\tau}]\). If \(\delta a=a_0(1-e^{-\Gamma_a\tau})\), each refresh step exactly compensates the preceding decay. Experimentally, a chalcogenide rib waveguide of length \(\approx 40~\mathrm{cm}\) with \(\Omega/2\pi=7.8~\mathrm{GHz}\), group velocity \(\approx 2.5~\mathrm{km/s}\), and intrinsic phonon lifetime \(\tau_0\approx10~\mathrm{ns}\) was driven with \(500~\mathrm{ps}\) data/write/read pulses and \(300~\mathrm{ps}\) refresh pulses. At \(8~\mathrm{ns}\), unrefreshed readout efficiency was \(\approx 4\%\); with \(4\) or \(7\) refresh pulses it rose to \(\approx 10\%\) and \(\approx 20\%\), respectively. With \(39\) refresh pulses spaced by \(1~\mathrm{ns}\), clear readout was achieved at \(40~\mathrm{ns}\), four times the intrinsic lifetime, and homodyne detection verified phase preservation after \(40~\mathrm{ns}\). The same analysis estimates storage times on the order of \(350~\mathrm{ns}\) without changing the apparatus and anticipates microsecond storage under improved extinction, reduced spontaneous Brillouin noise, and chirped refresh pulses [1904.13167].

Taken together, these results show that traveling-wave acoustic memory can combine GHz-class bandwidth with coherent phase retention. *A plausible implication is that the usual bandwidth–delay constraint is not intrinsic to phonon-based storage, but depends on how the phonon population is generated, refreshed, and read out.*

## 3. Quantum acoustic memories and phonon-mode control

In superconducting and optomechanical platforms, acoustic memory is often a genuine quantum memory: a long-lived phonon mode stores a qubit or bosonic state and is accessed through a nonlinear ancilla. A representative example is the multimode bulk-acoustic system comprising an Xmon-type transmon coupled to equally spaced HBAR modes on sapphire. Under low-frequency longitudinal modulation,
\[
H(t)=\frac{1}{2}\bigl[\omega_0+A\cos(\omega_{\mathrm{mod}}t)\bigr]\sigma_z+\sum_i \omega_i a_i^\dagger a_i + g_m\sum_i(\sigma_++\sigma_-)(a_i+a_i^\dagger),
\]
and the effective sideband interaction for the \(k\)-th branch is \(g_{\mathrm{eff}}(k)=g_mJ_{|k|}(A/\omega_{\mathrm{mod}})\). This permits selective access to individual modes despite their uniform spacing. The measured parameters are \(\omega_{\mathrm{FSR}}/2\pi\simeq39~\mathrm{MHz}\), \(g_m/2\pi\simeq3.1~\mathrm{MHz}\), qubit \(T_1\simeq1.09~\mu\mathrm{s}\), acoustic-mode lifetime \(T_{1m}\simeq210~\mathrm{ns}\), maximum \(g_{\mathrm{eff}}/2\pi\simeq1.8~\mathrm{MHz}\), and \(\pi\)-swap time \(\simeq240~\mathrm{ns}\). The initialize–write–store–read protocol uses a qubit \(\pi\)-pulse, a sideband-mediated phonon swap, optional switch-off modulation, a storage interval \(\tau_{\mathrm{store}}\lesssim T_{1m}\), and a reverse swap followed by dispersive readout. The demonstrated storage window extends to \(\sim200~\mathrm{ns}\) [2006.05446].

At the opposite lifetime extreme, a crystalline-silicon optomechanical crystal nanobeam cavity with a phononic bandgap shield localizes a \(5~\mathrm{GHz}\) breathing mode and suppresses radiation loss exponentially with shield length. At \(N=8\) shield periods, all excitation protocols yielded the same intrinsic damping rate \(\gamma_0/2\pi\approx0.11~\mathrm{Hz}\), corresponding to \(\tau=1/\gamma_0\approx1.47~\mathrm{s}\), \(Q\approx4.9\times10^{10}\), and \(f\cdot Q\approx2.6\times10^{20}~\mathrm{Hz}\); the implied effective phonon propagation length is \(\approx7.5~\mathrm{km}\). The measured damping is consistent with non-resonant two-level systems on etched silicon surfaces rather than three-phonon scattering. Because red and blue sideband pulses support write/read operations and \(G/2\pi\) can reach \(\gtrsim1~\mathrm{MHz}\), the corresponding cooperativity can reach \(C\approx10^{13}\), implying near-unity write/read fidelity and a storage time of \(\tau\approx1~\mathrm{s}\) in the idealized quantum-memory picture described in the study [1901.04129].

Multimode acoustic storage also enables memory architectures beyond simple delay. In the proposed hybrid QRAM, a single transmon is piezoelectrically coupled to many high-\(Q\) acoustic modes, with off-resonant drives engineering effective beamsplitter and three-mode interactions. The platform admits BAW resonators with \(T_1\sim10~\mathrm{ms}\), SAW resonators with \(T_1\sim10~\mu\mathrm{s}\), and phononic-crystal resonators with \(T_1\sim1~\mathrm{s}\); engineered couplings satisfy \(g_v^{(1)}/2\pi\sim10\)–\(100~\mathrm{kHz}\) and \(g_v^{(2)}/2\pi\sim5\)–\(25~\mathrm{kHz}\). For \(\kappa/2\pi\lesssim1~\mathrm{kHz}\), \(\gamma/2\pi\sim50~\mathrm{kHz}\), \(\nu/2\pi=10~\mathrm{MHz}\), and \(\Delta\nu/2\pi=1~\mathrm{MHz}\), the virtual-gate fidelity exceeds \(99\%\). The same work uses designated address modes and a bucket-brigade routing scheme to realize QRAM on a single chip, with quantum information stored directly in phonon Fock states [1906.11340].

A more specialized proposal stores photonic orbital angular momentum in a mechanical shear mode on a cavity mirror. The optoacoustic interaction
\[
\hat H_{\mathrm{int}}=-i\hbar g \chi_{ll'pp'} \hat a^\dagger \hat a(\hat b+\hat b^\dagger)
\]
enforces the selection rule \(|l'|=2|l|\), since the optical intensity profile carries \(\cos^2(l\theta)\). Under linearization and resolved-sideband conditions, a \(\pi\)-pulse state swap has duration \(\tau=\pi/[2g\sqrt{n_c\chi_{ll'pp'}}]\). With \(\omega_m/2\pi\sim1\)–\(5~\mathrm{MHz}\), \(Q_m\sim10^5\), \(\kappa/2\pi\sim50~\mathrm{kHz}\), \(g/2\pi\sim0.2~\mathrm{Hz}\), and \(n_c\sim10^{18}\), the analysis yields fidelities \(F>0.99\) for \(l\le2\), and \(F>0.95\) up to \(l\sim6\) or higher with optimized radial indices [1306.5813].

These platforms span traveling hypersound, multimode HBARs, phononic-crystal cavities, and optomechanical shear modes. *Taken together, they indicate that acoustic memory can be engineered either for large bandwidth and modest delay or for extreme coherence and discrete quantum access, depending on whether the relevant phonons are traveling waves or highly shielded cavity modes.*

## 4. Wave-propagation memory, imaging, and inverse problems

Acoustic memory also appears when propagation distortions persist locally across nearby source positions. In transcranial optoacoustic imaging, broadband ultrasonic pulses generated at neighboring intracranial locations traverse the skull with nearly identical mode conversions, reverberations, and attenuation profiles; the measured waveforms differ primarily by a delay. This “optoacoustic memory effect” is modeled through the wave equation
\[
\frac{\partial^2 p(\mathbf r,t)}{\partial t^2}
-c(\mathbf r)^2\rho(\mathbf r)\nabla\cdot\Bigl(\frac{1}{\rho(\mathbf r)}\nabla p(\mathbf r,t)\Bigr)
=\Gamma(\mathbf r)H(\mathbf r)\frac{\partial \delta(t)}{\partial t},
\]
discretized as \(d=A p_0\). The local memory effect is quantified by cross-correlation \(C(\Delta r,\tau)\) and by the normalized peak correlation \(\rho(\Delta r)\). Experimentally, \(\rho(\Delta r)>0.9\) for lateral shifts up to \(\approx3~\mathrm{mm}\), and a memory-based inversion that approximates each forward-model column as a delayed version of a measured reference sinogram resolves \(100~\mu\mathrm{m}\) spheres separated by \(\ge300~\mu\mathrm{m}\), matching the \(\approx300~\mu\mathrm{m}\) skull-free resolution. In three-dimensional random microsphere phantoms, the same method restores point-like foci with peak-to-sidelobe ratio \(>20~\mathrm{dB}\), whereas homogeneous-acoustics reconstructions remain grossly distorted [2108.03958].

A different use of the term concerns permanent shifts generated by nonlinear sound waves. For a one-dimensional barotropic perfect fluid, the exact Riemann-wave equation
\[
\partial_t v + (c_{s0}+\beta v)\partial_x v = 0
\]
implies that if the initial profile has a constant tail \(v(x,0)=B\), then the density after the wave has passed is permanently shifted to
\[
\rho_\infty=\rho_0\Bigl[1+(\beta-1)\frac{B}{c_{s0}}\Bigr]^{1/(\beta-1)},
\]
with acoustic memory
\[
\Delta\rho=\rho_0\Bigl[\Bigl(1+(\beta-1)\frac{B}{c_{s0}}\Bigr)^{1/(\beta-1)}-1\Bigr].
\]
For weak disturbances, \(\Delta\rho\approx \rho_0 B/c_{s0}\). The proposed experimental realization uses a box-trapped Bose–Einstein condensate, a phase-imprinted pulse with an oscillatory front and constant tail, and in situ absorption imaging before the shock time \(t_{\mathrm{shock}}=1/(\beta A k)\). The effect is presented as an acoustic analogue of gravitational-wave memory whose nonlinearity comes from the perfect-fluid equations rather than from Einstein’s equations [2011.05837].

In PDE analysis, acoustic memory is formalized through integral terms and acoustic boundary conditions. The one-dimensional inverse problem for a dispersive bar uses
\[
\partial_{tt}u-\partial_{xx}u-\beta \partial_{xxtt}u+\int_0^t k(t-s)\partial_{xx}u(x,s)\,ds=0
\]
on \(I=(0,\ell)\), with \(u(0,t)=0\) and boundary conditions at \(x=\ell\),
\[
\bigl[u_x-\int_0^t k(s)u_x(x,t-s)\,ds\bigr]_{x=\ell}=y'(t),\qquad
u_t(\ell,t)=-p y'(t)-q y(t),
\]
where \(p,q>0\). The inverse task is to recover the kernel \(k(t)\) from the overdetermination
\[
\int_0^\ell [\phi(x)-\beta\phi''(x)]u_x(x,t)\,dx=f(t).
\]
After reduction to homogeneous boundary conditions via \(z(t)=p y'(t)+q y(t)\) and \(v(x,t)=u_t(x,t)+z(t)x/\ell\), contraction arguments in Sobolev spaces yield global existence and uniqueness:
\[
v\in H^2(0,T;H_0^1(I)\cap H^2(I)),\quad
k\in H^1(0,T),\quad
y\in H^3(0,T).
\]
The non-degeneracy condition \(\int_0^\ell \phi'(x)u_0''(x)\,dx\neq0\) ensures invertibility of the Volterra equation for \(k\) [2505.07405].

These studies show that acoustic memory can be a local propagation invariance, a permanent nonlinear remnant, or a constitutive kernel to be identified from data. *A plausible implication is that “memory” in wave acoustics is often less about storing a signal in a separate register than about preserving a history-dependent transformation law.*

## 5. Material and biological embodiments

In dense suspensions under shear, acoustic memory can be embedded directly into the rheological microstructure. The training protocol uses a stress-controlled rheometer with a piezoelectric disk applying a \(1.16~\mathrm{MHz}\) sine wave at nominal acoustic power \(P\). At volume fraction \(\phi_v=0.565\), a constant shear stress \(\sigma_{\mathrm{app}}=380~\mathrm{Pa}\) is applied while an acoustic power \(P_{\mathrm{train}}\) is maintained for \(50~\mathrm{s}\); when \(P_{\mathrm{train}}\to0\) at \(t=50~\mathrm{s}\), every trained sample shear-jams within \(<1~\mathrm{s}\). The stress is decomposed as
\[
\sigma_{\mathrm{app}}=\sigma_P+\sigma_S,
\]
with primary force chains aligned with the maximum compressive axis and secondary chains aligned with the orthogonal extensional axis. During shear cessation, the normalized stress is fit by
\[
\frac{\sigma(t)}{\sigma_{\mathrm{app}}}=Ae^{-t/\tau_1}+Be^{-t/\tau_2}-Ce^{-t/\tau_3},
\]
where \(\tau_1\approx11~\mathrm{s}\), \(\tau_2\approx1000~\mathrm{s}\), and \(\tau_3\approx67~\mathrm{s}\), and \(\sigma_P=(A+B)\sigma_{\mathrm{app}}\), \(\sigma_S=C\sigma_{\mathrm{app}}\). Different training powers produce markedly different responses: after shear reversal, the minimum viscosity rises from \(\eta_{\min}\approx10~\mathrm{Pa\cdot s}\) at \(P_{\mathrm{train}}=0~\mathrm{W}\) to \(\approx10^4~\mathrm{Pa\cdot s}\) at \(P_{\mathrm{train}}=16~\mathrm{W}\), while the reverse strain needed to re-jam decreases from \(\gtrsim1\) to \(\lesssim0.2\). The same training can induce shear jamming below the conventional threshold, both by lowering \(\phi\) to \(\phi_v^-\approx0.56\) and by lowering the applied stress to \(\sigma_{\mathrm{app}}^-=38~\mathrm{Pa}\) [2404.15850].

The proposed electroacoustic model of the neuron relocates acoustic memory into a temporally modulated biological medium. In that framework, Coulomb forces associated with the action potential deform the membrane and modulate the compressibility of the axoplasm-plus-shell system. The effective stiffness satisfies
\[
\frac{1}{K(t)}=\frac{1}{K_M}+\frac{1}{K_{\mathrm{ion}}(t)},
\]
and with \(\rho\approx10^3~\mathrm{kg\,m^{-3}}\) the acoustic speed becomes
\[
c(t)=\frac{1}{\sqrt{\rho K(t)}}\approx5~\mathrm{m/s}
\]
at half-depolarization. The resulting one-dimensional acoustic equation is
\[
\frac{\partial^2 p(z,t)}{\partial z^2}-\frac{\partial}{\partial t}\Bigl[\frac{1}{K(t)}\frac{\partial p}{\partial t}(z,t)\Bigr]=0,
\]
or equivalently
\[
\frac{\partial^2 p}{\partial z^2}-\frac{1}{c^2(t)}\frac{\partial^2 p}{\partial t^2}
+\frac{\dot c(t)}{c^3(t)}\frac{\partial p}{\partial t}=0.
\]
For a periodic train of action potentials with period \(T\), \(K'(t)=K'(t+T)\) admits a Floquet expansion and yields an infinite-dimensional system for the harmonic amplitudes \(q_m\); numerical treatment produces allowed bands and forbidden temporal gaps. The study proposes that repeated pulse trains may open or close these dynamic band gaps and thereby modulate pathway selectivity, with possible relevance to plasticity, memory consolidation, and brain-field resonances, but presents this as a theoretical framework rather than an established neurophysiological mechanism [2312.13745].

These material and biological examples emphasize that acoustic memory need not be a localized resonator. It may instead be encoded in anisotropic contact networks or in a time-periodic medium whose transmission properties depend on prior excitation history.

## 6. Acoustic memory in machine learning and audio-language models

In machine learning, “acoustic memory” has two distinct uses. In streaming ASR it usually denotes mechanisms for carrying long-range acoustic context with bounded latency and bounded memory complexity; in large audio language models it denotes the ability to remember non-speech sounds across turns. The augmented-memory transformer for acoustic modeling segments the frame sequence into blocks \(C_n\), augments each with left and right context, and appends a bank of summary vectors \(M_n=\{m_1,\dots,m_{n-1}\}\). Self-attention is then performed jointly over current frames and memory slots, with the final summarization query producing the new memory \(m_n\). With \(80\)-dimensional log-Mel features, a VGG front-end, \(D=512\), \(8\) heads, segment length \(B=128\) frames, left context \(L=64\), right context \(R=32\), and memory-bank sizes up to \(\infty\), the \(40\)M-parameter AMTrf attains \(3.3\%\) test-clean and \(7.6\%\) test-other WER at \(0.32~\mathrm{s}\) look-ahead, outperforming the \(40\)M LC-BLSTM baseline at \(3.8\%\) and \(9.9\%\), while retaining strict streaming behavior [2005.08042].

Emformer modifies this idea by distilling long-range context into a fixed-size augmented memory bank and caching left-context key/value projections. For a sequence of length \(T\), full self-attention costs \(O(T^2 d)\) time and \(O(T^2)\) space, whereas blockwise attention with memory of size \(M\) and block length \(N=|L|+|C|+|R|\) costs approximately \(O((M+N)Nd)\) per block and becomes quasi-linear in \(T\) when \(M\ll N\). The memory summary is
\[
s_i^n=\frac{1}{|C|}\sum_{t\in C_i}x_t^n,\qquad
m_i^n=\mathrm{Attn}(W_q s_i^n; K_i^n,V_i^n),
\]
with a FIFO update \(M_{i+1}^n\leftarrow[\mathrm{tail}(M_i^n),m_i^n]\). In LibriSpeech hybrid modeling at medium latency, Emformer 24L reaches \(2.72\%\) / \(6.01\%\) WER on clean/other with RTF \(0.13\) and training \(0.25~\mathrm{h/epoch}\), compared with AM-TRF at \(3.27\%\) / \(6.66\%\), RTF \(0.16\), and \(1.14~\mathrm{h/epoch}\); the reported gains are a \(4.6\times\) training speedup and an \(18\%\) RTF reduction [2010.10759].

Earlier acoustic models treated memory as recurrent state or as a computational bottleneck. The TC-DNN-BLSTM-DNN architecture uses BLSTM cell state to “remember” features across tens to hundreds of frames and reduces WSJ eval92 WER from \(3.79\%\) for a \(4\times2048\) ReLU DNN baseline to \(3.47\%\) for the full model. Self-attentional acoustic models highlight a different issue: raw self-attention memory grows as \(O(T^2 D)\) for utterances up to \(2{,}026\) frames, which is mitigated through downsampling and Gaussian biasing, yielding a best TEDLIUM WER of \(14.90\%\) / \(15.89\%\) for dev/test while training at \(2.4\)k characters/s compared with \(1.1\)k for LSTM baselines. A separate line of work pursues memory efficiency of the model itself rather than memory of context: binary hidden-layer weights reduce weight storage by roughly \(32\times\), shrinking a WSJ \(6\times1024\)-unit network from \(\approx38~\mathrm{MB}\) to \(\approx1.2~\mathrm{MB}\), at the cost of WER degradation from \(6.8\%\) / \(3.8\%\) to \(7.7\%\) / \(4.8\%\) on dev93/eval92 [1504.01482], [1803.09519], [1706.09453].

The recent multi-turn literature returns to acoustic memory in the literal sense of recalling non-linguistic sounds. EnvMem constructs dialogues of length \(N\in\{2,4,8,16\}\) in which only the first user turn contains a sound mixed with speech at fixed \(10~\mathrm{dB}\) SNR, and the final turn asks either an acoustic question or a matched semantic question. The benchmark contains \(2{,}000\) acoustic-probe cases and \(2{,}000\) semantic-probe cases. Latent analyses define drift \(\Delta_t=\|z_t-z_0\|_2\), use \(10\)-way linear probes, and compute linear CKA across layers. The central finding is that representational trajectory drift, rather than retrieval failure, explains most degradation: failed long-context trials align with short-context trajectories in mid-layers, and attention interventions change accuracy by only \(\pm3\%\) with confidence intervals spanning zero. By contrast, activation patching at the best-probe layer can raise accuracy on failed \(N=16\) cases from near chance \((\approx0.13)\) to up to \(0.75\) for Qwen2.5-Omni, provided the donor representation is format-compatible and from the same acoustic class [2605.27039].

A more speculative computational proposal is Phonetic Trajectory Memory, which replaces growing key–value caches with a fixed-size state \(S_t\in\mathbb T^{16}\) evolving under irrational rotations and phonetic injections. The architecture reports bridge-token updates of \(64\) bytes versus \(\approx192~\mathrm{KB}\) for dense FP16 key–value storage, compression \(>256\times\) for bridges, overall compression \(>3{,}000\times\) when \(80\%\) of tokens are bridged, and retrieval latency approximately \(34~\mathrm{ms}\) independent of context depth. The work frames retrieval as “Signal Consensus” between semantic prior and geometric likelihood, and reports up to approximately \(92\%\) factual accuracy, but it should be read as a proposed biomimetic memory architecture rather than as an established acoustic-memory benchmark [2512.20245].

Across these computational literatures, acoustic memory ranges from explicit memory banks for frame sequences, to recurrent state for acoustic context, to the retention of environmental sound cues in long-context audio dialogue. *This suggests that the computational question has shifted from whether a model can carry acoustic history at all to which internal representation format remains decodable after long temporal separation.*

Source: https://www.emergentmind.com/topics/acoustic-memory