---
title: Marginally Absorbing Manifold (MAM)
url: https://www.emergentmind.com/topics/marginally-absorbing-manifold-mam
type: topic
---

# Marginally Absorbing Manifold (MAM)

to=arxiv_search.search  天天中彩票双色球json
{"query":"arXiv:2509.01296 OR \"Learning by training: emergent return-point memory from cyclically tuning disordered sphere packings\"", "max_results": 5}
to=arxiv_search.search  天天中彩票提现json
{"query":"\"return-point memory\" disordered systems cyclically driven", "max_results": 10}
to=arxiv_search.search  สำนักเลขานุการองค์กร  聚利json
{"query":"\"marginally absorbing manifold\" arXiv", "max_results": 10}
A marginally absorbing manifold (MAM) is a subset of parameter space that is absorbing for gradient-based training dynamics over a prescribed target interval, yet loses that absorbing character immediately when the target is extended beyond the interval’s endpoints. In the formulation introduced for cyclic inverse design of athermal disordered sphere packings, the MAM is the geometric object that encodes memory of the training range and reproduces a two-sided form of return-point memory under cyclic tuning of elastic properties [2509.01296].

## 1. Formal definition in parameter space

The construction is posed in an $n$-dimensional parameter space $\Theta \subset \mathbb{R}^n$ with coordinates
$$
\theta = (\theta_1,\theta_2,\ldots,\theta_n)\in\Theta.
$$
A scalar observable or training function
$$
F:\Theta\to\mathbb{R}
$$
maps parameters to a mechanical response, such as the Poisson ratio $\nu(\theta)$ or an elastic-modulus component $c_{ijkl}(\theta)$. Training is defined relative to a target value $F^\*$ through the objective
$$
\ell(\theta;F^\*) = [F(\theta)-F^\*]^2,
$$
with idealized continuous-time steepest-descent dynamics
$$
\frac{d\theta}{dt}=-\nabla_\theta \ell(\theta;F^\*).
$$

For a closed interval of training targets $[F_{\min},F_{\max}]$, an absorbing manifold $\mathcal{M}\subset\Theta$ is defined by the requirement that for every $F^\*\in[F_{\min},F_{\max}]$ and every $\theta_0\in\mathcal{M}$, the training dynamics beginning at $\theta_0$ with target $F^\*$ returns to the same point $\theta_0$ up to machine precision. Equivalently, the solution with $\theta(0)=\theta_0$ satisfies $\theta(T)=\theta_0$ for some finite $T>0$ when $F(\theta(T))=F(\theta(0))$. In the discretized setting with step size $\eta$, the update rule is
$$
\theta_{k+1}=\theta_k-\eta\nabla \ell,
$$
and absorption is operationally defined by exact return to $\theta_0$ after one up-and-down target cycle within the interval [2509.01296].

A manifold is *marginally* absorbing when it remains absorbing for all $F^\*\in[F_{\min},F_{\max}]$ but ceases to be absorbing as soon as training targets satisfy $F^\*<F_{\min}$ or $F^\*>F_{\max}$. The term “marginally” is therefore literal: the manifold just barely traps the dynamics, and the two endpoints $F_{\min}$ and $F_{\max}$ are the encoded memories.

The stability condition on $\mathcal{M}$ is stated in terms of tangent displacements $n(\theta)$:
$$
\nabla_\theta \ell(\theta;F^\*)\cdot n(\theta)=0
\qquad
\text{for all }F^\*\in[F_{\min},F_{\max}].
$$
Within the interval, motion tangent to $\mathcal{M}$ is compatible with absorption, whereas normal displacements become unstable once the target exits the interval. A common misconception is to identify any absorbing set with a MAM. In this formulation, absorption alone is insufficient; marginality requires the immediate loss of absorption outside the trained target range.

## 2. Gradient discontinuities as the origin of the manifold

The physical mechanism proposed for MAM formation is the presence of gradient discontinuities in the training function. In the sphere-packing model, $F(\theta)$ is continuous in $\theta$, but $\nabla_\theta F$ is discontinuous whenever a particle pair $\langle i,j\rangle$ goes into or out of contact. The system energy is given by the 2D Hertzian form
$$
U(\{r_i\};\theta)=\sum_{i<j}u(r_{ij};\sigma_{ij}),
\qquad
\sigma_{ij}=(D_i+D_j)/2,
$$
with pair potential
$$
u(r;\sigma)=\frac{k}{2.5}\cdot \max\{0,1-(r/\sigma)\}^{2.5}.
$$
Linear response is governed by the Hessian
$$
H_{ab}=\frac{\partial^2 U}{\partial x_a\partial x_b},
$$
and an elastic-modulus component is written as
$$
c_{pqrs}(\theta)=
\frac{1}{V}\frac{\partial^2 U}{\partial \epsilon_{pq}\partial \epsilon_{rs}}
-\frac{1}{V}\,\Xi^T\cdot H^{-1}\cdot \Xi,
$$
where $\Xi=(\partial^2 U)/(\partial \epsilon \partial x)$ and $x=(x_1,\ldots,x_{2N})$ [2509.01296].

Because $u(r;\sigma)\in C^2$ but not $C^3$, contact formation or breaking at $r_{ij}=\sigma_{ij}$ causes a jump in $\partial^3 U/\partial \sigma \partial x^2$, and thus a jump in $\nabla_\theta c_{pqrs}(\theta)$. Each contact event therefore defines a codimension-1 surface $\Sigma\subset\Theta$ across which the gradient has a finite discontinuity.

At a point $\theta\in\Sigma$, the jump is decomposed relative to the local unit normal $\hat{u}(\theta)$ and tangent directions $\hat{v}$:
$$
\Delta \nabla_\theta F(\theta)=\nabla_\theta F(\theta^+)-\nabla_\theta F(\theta^-)
=[\Delta s]\hat{v}+[\Delta p]\hat{u}.
$$
Since $F$ itself is continuous, the tangential contribution $\Delta s\equiv \hat{v}\cdot \Delta\nabla F$ must vanish, and the relevant discontinuity is
$$
\Delta p\equiv \hat{u}\cdot\big(\nabla F(\theta^+)-\nabla F(\theta^-)\big).
$$

Two classes of gradient-discontinuity surfaces are then distinguished. In a Type 1 GD, trajectories cross the surface smoothly. In a Type 2 GD, the normal component of the gradient flips sign across the surface so that both sides push toward $\Sigma$; a steepest-descent or ascent trajectory that reaches such a surface becomes bound to it. The paper’s central claim is that one or more Type 2 GD surfaces at $F=F_{\min}$ and $F=F_{\max}$ convert an otherwise reversible optimization path into a marginally absorbing one.

## 3. Cyclic inverse design and convergence to a MAM

The simulation protocol is organized as a cyclic inverse-design procedure on mechanically stable jammed packings. Initialization begins from a jammed packing of $N$ particles at volume fraction $\phi_0$. A set of $n_{sp}$ species diameters $D_\alpha$ is chosen, the trainable parameters are set as
$$
\theta=(D_1,\ldots,D_{n_{sp}}),
$$
and the initial configuration is quenched so that $U(\{r_i\};\theta_0)$ reaches a local minimum.

Single-target training then fixes a target such as $F^\*=F_{\max}$, for example $\nu^\*=0.5$, and minimizes
$$
\ell(\theta)=[F(\theta)-F^\*]^2
$$
using automatic differentiation plus RMSProp, or steepest descent, with energy re-minimization with respect to particle positions at each step. This establishes the local optimization dynamics that will subsequently be driven cyclically.

The cyclic sweep defines a target sequence
$$
F^\*=F_{\max},\,F_{\max}-\delta,\,F_{\max}-2\delta,\ldots,F_{\min},\,F_{\min}+\delta,\ldots,F_{\max},
$$
using the final parameter state from one target as the initial condition for the next. One full return to $F_{\max}$ constitutes a cycle. Repetition continues until three operational signatures of convergence are satisfied: the cycle-to-cycle parameter trajectory becomes indistinguishable with $|\Delta\theta_{\text{cycle}}|\to 0$, the training trajectories $F(t)$ become smooth with no spikes, and no contact changes occur for intermediate targets $F\in(F_{\min},F_{\max})$ [2509.01296].

Read-out is performed by sweeping a fine grid of $F_{\text{read}}$ values, both inside and outside the training interval, from a point on $\mathcal{M}$ at $F=F_{\max}$. The recorded observables are the net displacement $|\theta_{\text{out}}-\theta_{\text{in}}|$, the number of iteration steps, the contact-change counts, and the perpendicular distance of $\theta$ to the principal-axis direction on $\mathcal{M}$. The text further states that each individual training step is costly, of order $\sim 10^3$–$10^4$ gradient steps times FIRE relaxations, and that one typically needs only $O(10$–$30)$ cycles before convergence. Within the article’s internal logic, these operational criteria define the empirical identification of a MAM.

## 4. Memory encoding and return-point phenomena

Once the dynamics is trapped on $\mathcal{M}$ by two Type 2 GD surfaces, one at $F=F_{\min}$ and one at $F=F_{\max}$, cyclic training within the interval $[F_{\min},F_{\max}]$ leaves $\theta$ on the same loop in $\mathcal{M}$. By contrast, training to $F^\*>F_{\max}$ or $F^\*<F_{\min}$ crosses a bounding GD surface and forces motion into a new region of parameter space. In this sense, the MAM stores the two training extremes as a bounded memory.

The structure is presented as reproducing the classic hallmarks of return-point memory. First, the read-out observables display sharp kinks at $F_{\min}$ and $F_{\max}$. The listed observables include the net parameter change $\Delta\theta$, the number of gradient steps needed, and the distance of the final $\theta$ to the line joining $\theta(F_{\min})\to \theta(F_{\max})$. Second, there is complete reversibility, or absorption, for any cycle contained within the training range, but irreversibility appears once the cycle amplitude exceeds either endpoint [2509.01296].

A useful clarification is that the memory is not formulated as symbolic storage or explicit state labeling. Rather, it is encoded geometrically in the accessibility structure of parameter space under the specified training dynamics. The article’s formalism therefore treats memory as a property of constrained trajectory recurrence.

This also helps distinguish the MAM from an ordinary reversible path. The reversibility is not generic; it is conditional on the path being confined by the two endpoint GD surfaces. A plausible implication is that the “memory” resides less in any single configuration than in the manifold-plus-boundary structure generated by repeated cyclic training.

## 5. Gradient Discontinuity Learning as the general mechanism

The paper abstracts the sphere-packing results into a broader framework called Gradient Discontinuity Learning (GDL). In that formulation, one assumes a continuous but piecewise-smooth function $F:\Theta\to\mathbb{R}$ with a collection of GD surfaces $\Sigma_i$ where $\nabla F$ jumps. Each $\Sigma_i$ is classified as Type 1 or Type 2 according to the sign change of the normal component of $\nabla F$ across the surface.

Under gradient descent, or descent in $\ell$, trajectories cross Type 1 surfaces smoothly but become bound to Type 2 surfaces. While bound, they slide along the surface according to the tangential component $\nabla_{\parallel}F$. If a closed cycle of target values encounters two distinct Type 2 surfaces, for example at $F=F_{\min}$ and $F=F_{\max}$, then the dynamics is driven into the codimension-2 intersection of those surfaces. Once there, cyclic variation of $F^\*$ can move points only back and forth along that intersection, and the resulting set is marginally absorbing [2509.01296].

Crossing beyond either bounding surface requires leaving the intersection, which produces an irreversible jump and thereby implements two-sided memory. The stated requirements for GDL are a continuous $F$ with non-trivial GD surfaces, cyclic variation of the target, and a gradient-based update rule. The text also notes that non-infinitesimal step sizes or momentum-based optimizers may produce “effective GDs” even if $F$ were $C^\infty$; the stated key requirement is that the path is not perfectly reversible unless pinned at special surfaces.

Within this framework, the MAM is not introduced as an idiosyncrasy of jammed solids. It is instead presented as a generic geometric outcome of cyclic training in piecewise-smooth landscapes. The paper identifies potential applications extending beyond jammed solids to sheared suspensions, biological evolution, phenotypic plasticity, machine-learning models with constrained parameterizations, and more. Because these are listed as potential applications rather than demonstrated cases, they are best read as scope conditions for the proposed mechanism rather than as established empirical generalizations.

## 6. Conceptual significance, limitations, and interpretation

The principal result is that athermal disordered sphere packings, when inverse-designed by cyclic tuning of elastic constants between two endpoints, evolve toward a MAM that traps the parameters. The manifold is bounded by exactly two Type 2 gradient-discontinuity surfaces associated with the endpoint training values, is absorbing for cycles within $[F_{\min},F_{\max}]$, and becomes non-absorbing immediately outside that range. The physical origin of the gradient discontinuities is contact change, and the diagnostic for Type 2 behavior is the sign flip of the normal component of $\nabla_\theta F$ [2509.01296].

Several interpretive points follow directly from this formulation. First, the MAM is a dynamical object defined by training trajectories, not merely a level set of $F$. Second, its memory content is explicitly bounded and two-sided, tied to $F_{\min}$ and $F_{\max}$ rather than to arbitrary interior targets. Third, the mechanism depends on non-smooth structure in the optimization landscape, here supplied by contact changes in Hertzian packings.

A potential misconception is that the manifold stores all details of prior training history. The formulation given here supports a narrower statement: it encodes the two extreme memories $F_{\min}$ and $F_{\max}$ and yields return-point-like signatures under read-out. Another possible misconception is that gradient discontinuities must always arise from non-differentiable microscopic physics. The text instead suggests a broader view in which “effective GDs” may be induced by optimizer details such as non-infinitesimal step sizes or momentum.

The broader significance of the MAM concept lies in its provision of a precise geometric description of how cyclic environmental variation can generate both adaptation and memory. In the paper’s terms, this furnishes a simple and broadly applicable physical framework for understanding how adaptive systems learn under environmental change and retain memory of past experiences. A plausible implication is that the relevant unit of analysis in such systems is not solely the optimized configuration, but the manifold of recurrently accessible states generated by the training protocol.

Source: https://www.emergentmind.com/topics/marginally-absorbing-manifold-mam