---
title: Structural Susceptibility Matrix
url: https://www.emergentmind.com/topics/structural-susceptibility-matrix
type: topic
---

# Structural Susceptibility Matrix

Searching arXiv for the cited papers and closely related work to ground the article.
Searching arXiv for 1112.5418 and 2605.07980.
The **structural susceptibility matrix** is a matrix-valued object used to quantify how structural perturbations propagate through a modeled system, but its precise definition is context-dependent. In nonlinear dynamical systems, it is the Hessian of a least-squares trajectory-mismatch cost, equivalently \(S=J^\top J\), and measures how infinitesimal perturbations of the dynamical law affect an observed trajectory [1112.5418]. In Bayesian learning, it is a covariance-based linear-response matrix whose entries pair model components with data perturbations, and it is, up to a factor of \(n\beta\), the Jacobian of the map from data distributions to structural coordinates [2605.07980]. In a susceptibility-stratified contagion model, the same phrase is used for a block matrix that encodes how susceptibility classes contribute to new infections and whose infection block yields \(R_0\) through its spectral radius [2009.09150]. This diversity of usage indicates that the term denotes a family of sensitivity operators rather than a single universally standardized construction.

## 1. Dynamical-systems definition as a Hessian

In the dynamical-systems formulation, one begins from a least-squares cost
\[
C(\theta)=\tfrac12\sum_k [y_k(\theta)-d_k]^2
\]
or, in the continuous-time setting,
\[
C(\theta)=\tfrac12\int_0^N \|z(t;\theta)-z_0(t)\|^2\,dt,
\]
where \(\theta=(\theta_1,\dots,\theta_P)\) are parameters, \(z_0(t)\) is the unperturbed trajectory, and \(z(t;\theta)\) the perturbed one. The structural susceptibility matrix is then defined by
\[
S_{ij}\equiv \frac{\partial^2 C}{\partial \theta_i\,\partial \theta_j}\Big|_{\theta=0}.
\]
At the best fit, if \(J\) is the Jacobian of the model outputs with respect to parameters,
\[
J_{ki}=\frac{\partial z_k}{\partial \theta_i},
\]
then
\[
S=J^\top J.
\]
This construction is explicitly identified with the Hessian matrix of the least-squares cost function [1112.5418].

The associated interpretation is geometric. The eigenvalues of \(S\) quantify sensitivity along linear combinations of parameters, and diagonalization
\[
S\,v_k=\lambda_k\,v_k,\qquad \lambda_1\ge \lambda_2\ge \cdots \ge \lambda_P
\]
defines eigenparameters through the eigenvectors \(v_k\). Large \(\lambda_k\) correspond to stiff directions in parameter space, whereas small \(\lambda_k\) correspond to sloppy directions. The data summarize this as an apparently inherent insensitivity to large magnitude variations in certain linear combinations of parameters, with sloppiness quantified by Hessian eigenvalues that typically span many orders of magnitude [1112.5418].

A sensitivity-based representation follows from the variational equation. Defining
\[
s_i(t)\equiv \frac{\partial z(t)}{\partial \theta_i},
\]
one has
\[
\frac{d}{dt}s_i = \left(\frac{\partial f}{\partial z}\right)s_i + \frac{\partial f}{\partial \theta_i},
\]
with \(s_i(0)=0\) or, as stated in the source, with appropriate phase-matching and period corrections. The Hessian entries can then be written as
\[
S_{ij}=\int_0^T [s_i(t)]^\top s_j(t)\,dt.
\]
In the special case where the cost compares directly the vector field perturbations \(g_i(z)=\partial f/\partial \theta_i\), one finds the formally simpler approximation
\[
S_{ij}\simeq \int_0^T g_i(z(t))^\top g_j(z(t))\,dt,
\]
neglecting state-sensitivity propagation.

## 2. Polynomial perturbations in the van der Pol oscillator

The canonical van der Pol equations are written in slow-fast form by setting \(\epsilon=\mu^{-2}\):
\[
\epsilon \dot x = x-\frac{x^3}{3}-y,\qquad \dot y = x,
\]
or, equivalently,
\[
\dot x = y,\qquad \dot y = \mu(1-x^2)y-x.
\]
To study structural susceptibility, a polynomial perturbation is added to the right-hand side of the \(\dot x\)-equation:
\[
\epsilon \dot x = x-\frac{x^3}{3}-y
+\sum_{m+n\le N} a_{m,n}\left(x-\frac{x^3}{3}-y\right)^m x^n,\qquad \dot y=x.
\]
The perturbation amplitudes \(\{a_{m,n}\}\) are the parameters, with
\[
P=\frac{(N+1)(N+2)}{2}.
\]

This polynomial basis induces a natural decomposition. For \(m\neq 0\), the perturbation terms vanish on the critical manifold \(y=x-x^3/3\), and these are termed **fast parameters**. For \(m=0\), the perturbations act on the slow manifold and are termed **slow parameters** [1112.5418].

The significance of this construction is that the structural susceptibility matrix is not merely a local curvature object but also a classifier of perturbation types. The data state that perturbations in the van der Pol dynamics show that most directions in parameter space weakly affect the limit cycle, whereas only a few directions are stiff. This provides an explicit realization of sloppiness in a nonlinear oscillator with multiple time scales.

## 3. Time-scale separation, eigenvalue clustering, and the singular limit

The van der Pol example is used to connect structural susceptibility to separation of time scales. The summary states that the quantity \(\log \lambda_k\) versus \(k\) typically shows an almost straight line over many decades, which is identified as the hallmark of sloppiness. In the single time-scale limit \(\mu\sim 1\), the eigenvalues exhibit a broad but moderate spread. As \(\mu\to\infty\) and \(\epsilon\to 0\), the eigenvalues separate into two clusters [1112.5418].

The asymptotic structure is explicit. The top \(N+1\) eigenvalues, corresponding to stiff modes, approach \(O(1)\) constants as \(\mu\) increases. The remaining \(P-(N+1)\) eigenvalues decay as power laws in \(\mu\), with two modes numerically scaling as \(\mu^{-2\,\dots\,-3}\) and the rest as \(\mu^{-5\,\dots\,-6}\). The source therefore states that, as \(\epsilon\to 0\), slow-manifold perturbations remain costly, with stiff \(\lambda\approx \mathrm{const}\), whereas jump-only perturbations become increasingly negligible in the least-squares cost, with sloppy \(\lambda\to 0\) [1112.5418].

The eigenvectors sharpen this interpretation. The stiff eigenvectors lie primarily in the subspace of slow parameters \(a_{0,n}\), and these combinations deform the slow manifold on which the system spends most of its period. The sloppy eigenvectors lie in the subspace of fast parameters with \(m>0\); they only deform the short jumps, and as \(\epsilon\to 0\) these jumps occupy vanishing time. In the singular limit, the matrix is approximated analytically by
\[
S_{\text{slow,slow}}\simeq \int_{\text{slow manifold}} (\cdots)\,dt = O(1),\qquad
S_{\text{fast,fast}}\simeq \int_{\text{jumps}} (\cdots)\,dt \sim O(\epsilon^p)\to 0,
\]
so the Hessian acquires a block-diagonal form with stiff and sloppy blocks.

A plausible implication is that structural susceptibility provides a quantitative route from slow-fast geometry to parameter-space anisotropy: the temporal occupancy of different regions of phase space controls which perturbations remain visible to a trajectory-based cost.

## 4. Bayesian learning: covariance, linear response, and structural coordinates

In Bayesian learning, the structural susceptibility matrix is introduced in a different but formally related way. One first chooses a family of \(H\) component observables \(\{\phi_{C_j}(w)\}_{j=1}^H\). If \(W\cong U\times C_j\) is a product decomposition isolating component \(C_j\) and \(w^*=(u^*,v^*)\) is the trained parameter, then
\[
\phi_{C_j}(w)=\delta(u-u^*)\,[L(w)-L(w^*)],
\]
where
\[
L(w)=\mathbb E_{(x,y)\sim q}[-\ln p(y\mid x,w)]
\]
is the population loss. The corresponding structural coordinate is its posterior expectation
\[
\mu_j=\langle \phi_{C_j}\rangle
=\int_W \phi_{C_j}(w)\,\Pi^{\mathrm{pop}_\beta}(dw),
\]
with population Gibbs posterior
\[
\Pi^{\mathrm{pop}_\beta}(dw)\propto e^{-n\beta L(w)}\,\pi(w)\,dw.
\]

Data perturbations are introduced by
\[
q_h=(1-h)\,q+h\,q',\qquad h\in[0,1].
\]
Under \(q_h\), the loss becomes \(L^h=\mathbb E_{q_h}[\ell_z]\). By the fluctuation-dissipation theorem, for any observable \(\phi\),
\[
\frac{\partial}{\partial h}\langle \phi\rangle_h\Big|_{h=0}
=-\,n\beta\,\mathrm{Cov}_{\Pi^{\mathrm{pop}_\beta}}[\phi,\Delta L],
\qquad
\Delta L(w)=L^{q'}(w)-L(w).
\]
For the basis of perturbations obtained by upweighting a single data point \(z_k=(x_k,y_k)\), with \(q'=\delta_{z_k}\), one has \(\Delta L(w)=\ell_{z_k}(w)-L(w)\), and the structural susceptibility matrix \(S\in\mathbb R^{H\times D}\) is defined by
\[
S_{j,k}
=\chi_{z_k}(\phi_{C_j})
=\frac{1}{n\beta}\frac{\partial}{\partial h}\langle \phi_{C_j}\rangle_h\Big|_{h=0}
=-\,\mathrm{Cov}\!\bigl[\phi_{C_j},\,\ell_{z_k}-L\bigr].
\]
Equivalently,
\[
S_{j,k}=\frac{\partial}{\partial q_k}\mu_j
=\frac{\partial}{\partial q_k}\langle \phi_{C_j}\rangle.
\]

The matrix is therefore, up to the factor \(n\beta\), the Jacobian of the map
\[
\mu:\{q:\text{data dists}\}\longrightarrow \mathbb R^H,\qquad
q\mapsto (\mu_1,\dots,\mu_H),
\]
and its differential form is
\[
d\mu = n\beta\,S\,dq.
\]
In this setting, the structural susceptibility matrix pairs model components with data patterns and supplies a linearized map from distributional perturbations to structural change [2605.07980].

## 5. Pseudo-inverse, empirical estimation, and patterning

The Bayesian formulation leads directly to an inverse problem. Given a target change \(d\mu^*\in\mathbb R^H\) in structural coordinates, linearization gives
\[
d\mu^* \approx n\beta\,S\,dq
\qquad\Longrightarrow\qquad
S\,dq=\frac{1}{n\beta}d\mu^*.
\]
When \(S\) is not square or not full-rank, the minimal-norm least-squares solution is given by the Moore-Penrose pseudo-inverse:
\[
dq_{\mathrm{opt}}=\frac{1}{n\beta}S^+d\mu^*,
\qquad
S^+=V\Sigma^+U^\top\quad\text{when }S=U\Sigma V^\top.
\]
Equivalently, if \(S=\sum_\alpha \sigma_\alpha u_\alpha v_\alpha^\top\), then
\[
dq_{\mathrm{opt}}
=\frac{1}{n\beta}\sum_{\alpha:\sigma_\alpha>0}
\frac{u_\alpha^\top d\mu^*}{\sigma_\alpha}\,v_\alpha.
\]
The source identifies this as the **patterning** prescription: the smallest-norm distributional perturbation that achieves the desired first-order change in \(\mu\) [2605.07980].

Empirically, one works with a finite dataset \(\{z_i\}_{i=1}^n\), the empirical posterior
\[
\Pi^{\mathrm{emp}_\beta}(dw)\propto e^{-n\beta L_n(w)}\pi(w)\,dw,
\qquad
L_n(w)=\tfrac1n\sum_{i=1}^n \ell_{z_i}(w),
\]
and the estimator
\[
\widehat S_{j,k}
=-\,\mathrm{Cov}^{\mathrm{emp}}\!\bigl[\phi_{C_j}(w),\,\ell_{z_k}(w)-L_n(w)\bigr]
=\frac{\partial}{\partial \rho_k}\langle \phi_{C_j}\rangle_\rho\Big|_{\rho=1},
\]
where \(\rho_k\) is the weight on sample \(z_k\). A posterior sampler such as SGLD can be used to draw \(M\) samples \(w^{(t)}\) and estimate empirical covariances by sample averages. Because \(\phi_{C_j}\) contains a delta-slice, the implementation uses a **weight-restricted** sampler on \(C_j\), estimates a renormalized susceptibility \(\tilde\chi^j_{z_k}\), and then **standardizes** columns to remove unknown renormalization constants before SVD or patterning. To stabilize inversion, the source gives the ridge-regularized inverse
\[
R_\lambda(\widehat S)
=\widehat S^\top(\widehat S\,\widehat S^\top+\lambda I)^{-1},
\]
which replaces \(1/\sigma_\alpha\) by \(\sigma_\alpha/(\sigma_\alpha^2+\lambda)\) in the SVD representation [2605.07980].

The computational profile is also explicit: \(S\) has size \(H\times D\); typically \(\mathrm{rank}(S)\ll \min(H,D)\); low-rank modes can be extracted by randomized SVD or Lanczos with cost \(O((H+D)r^2)\) rather than \(O(H^2D)\); and if \(D\) is very large, one may subsample a representative batch of data points or cluster data points a priori and treat each cluster as one column. If the posterior is sharply peaked and roughly Gaussian, covariances may be approximated via the Hessian inverse \(H^{-1}\) at \(w^*\), whereas in deep nets the posterior is typically non-Gaussian, so SGLD is used.

## 6. Susceptibility-class matrices in contagion dynamics and terminological variation

In a susceptibility-stratified contagion model, \(\sigma\) denotes individual susceptibility, \(p(\sigma)\) its population density, and \(S(t,\sigma)\) the susceptible mass at time \(t\) among those with susceptibility \(\sigma\). The governing assumption is that individuals of susceptibility \(\sigma\) are removed from \(S\) at rate \(\beta \sigma S(t,\sigma)I(t)\), yielding
\[
\partial_t S(t,\sigma)=-\,\beta\,\sigma\,S(t,\sigma)\,I(t).
\]
With
\[
m_1(t)=\int_0^\infty \sigma S(t,\sigma)\,d\sigma,
\]
the aggregated equations are
\[
\frac{d}{dt}S(t)=-\,\beta\,I(t)\int_0^\infty \sigma S(t,\sigma)\,d\sigma,
\qquad
\frac{d}{dt}I(t)=\beta I(t)m_1(t)-\gamma I(t),
\qquad
\frac{d}{dt}R(t)=\gamma I(t).
\]
More generally, the moments
\[
m_k(t)=\int_0^\infty \sigma^k S(t,\sigma)\,d\sigma
\]
satisfy
\[
\dot m_k=-\,\beta\,I\,m_{k+1},
\]
so closure requires carrying infinitely many moments [2009.09150].

For discrete susceptibility classes \(\sigma_1,\dots,\sigma_n\), with \(S_i(t)=S(t,\sigma_i)\), the state vector is
\[
\mathbf x(t)=\bigl[S_1(t),\dots,S_n(t),\,I(t),\,R(t)\bigr]^T,
\]
and the dynamics are written in block form as
\[
\dot{\mathbf x}
=
\underbrace{
\begin{pmatrix}
-\beta I\,\mathrm{diag}(\sigma_i) & \mathbf 0 & \mathbf 0\\
\beta \sum_i \sigma_i & -\gamma & 0\\
0 & +\gamma & 0
\end{pmatrix}
}_{\mathbf S}
\mathbf x(t).
\]
This large matrix \(\mathbf S\) is called the structural susceptibility matrix. More compactly,
\[
\mathbf S=
\begin{pmatrix}
S_{SS} & \mathbf 0 & \mathbf 0\\
s_{IS} & -\gamma & 0\\
0 & +\gamma & 0
\end{pmatrix},
\qquad
S_{SS}=-\beta I\,\mathrm{diag}(\sigma_i),
\qquad
s_{IS}=\bigl[\beta \sigma_1,\dots,\beta \sigma_n\bigr].
\]

The infection block then gives a next-generation matrix
\[
K=\frac{1}{\gamma}\bigl[\beta \sigma_i p_i \delta_{j,i}\bigr]_{i=1\ldots n,\,j=1\ldots n}
=\frac{\beta}{\gamma}\,\mathrm{diag}(\sigma_1 p_1,\dots,\sigma_n p_n),
\]
with spectral radius
\[
R_0=\rho(K)=\max_i \frac{\beta \sigma_i p_i}{\gamma}.
\]
For a continuous distribution, the analogous quantity is
\[
R_0=\int_0^\infty \frac{\beta \sigma p(\sigma)}{\gamma}\,d\sigma
=\frac{\beta}{\gamma}E[\sigma].
\]
The herd-immunity condition is expressed through
\[
R_{\mathrm{eff}}(t)=\frac{\beta}{\gamma}\langle \sigma\rangle_t\,S(t)
\le \frac{\beta}{\gamma}\langle \sigma\rangle_0\,S(t)
=R_0\,S(t),
\]
and the herd-immunity fraction is
\[
1-S^*=1-\frac{1}{R_0}.
\]

A common misconception would be to treat these three uses of the term as interchangeable. The summarized literature does not support that. In the van der Pol analysis, the matrix is a Hessian of a trajectory cost; in Bayesian learning, it is a covariance-defined response matrix and Jacobian of structural coordinates; in contagion modeling, it is a state-transition block matrix or infection block tied to \(R_0\). The common thread is susceptibility of structure to perturbation, but the underlying state spaces, derivatives, and inferential roles are distinct.

Source: https://www.emergentmind.com/topics/structural-susceptibility-matrix