---
title: 'Input-to-Input Gain: Systems Perspective'
url: https://www.emergentmind.com/topics/input-to-input-gain-iig
type: topic
---

# Input-to-Input Gain: Systems Perspective

Searching arXiv for recent papers on “Input-to-Input Gain” and closely related formulations.
Input-to-Input Gain (IIG) is not a uniformly standardized term across the arXiv literature. In the most explicit formulation presently represented in the cited corpus, IIG is introduced as an input-side analogue of output-to-output gain (OOG) for linear time-invariant fault-detection settings, where it measures the maximum energy of undetectable faults for a given disturbance intensity [2509.11194]. In adjacent literatures, however, closely related ideas appear under other names: componentwise external-input-to-output gain in nonlinear small-gain theory [1409.6991], nonlinear input-output amplification under small-signal finite-gain $\mathcal{L}_p$ stability in transitional shear flows [2506.08129], forward scattering gain in microwave SQUID amplifiers [1206.4706], region-dependent input-to-state gain composition in interconnected nonlinear systems [1503.02884], and Optimal Input Gain (OIG) in feed-forward neural-network training [2303.17732]. Taken together, these works indicate that “IIG” functions less as a single universal definition than as a family of gain-level constructions relating input-side perturbations, channels, or transformations to downstream system behavior.

## 1. Formal definition and scope

The clearest named definition appears in "Fundamental limitations of sensitivity metrics for anomaly impact analysis in LTI systems" [2509.11194]. There, IIG is proposed as a new measure of robust fault sensitivity and is defined as the maximum energy of undetectable faults for a given disturbance intensity. The fault/detection model is
$$
\tilde{\Sigma}:\left\{ \begin{array}{l} \dot{x}_2(t)=\tilde{A}x_2(t)+B_d d(t)+B_f f(t),\\[2mm] r(t)=Cx_2(t)+D_d d(t)+D_f f(t), \end{array} \right.
$$
with disturbance $d$, fault $f$, and residual $r$, where
$$
r = \mathds{T}_{dr}[d]+\mathds{T}_{fr}[f].
$$
A fault $f$ is called $\mathcal D$-undetectable if there exists $d\in\mathcal D$ such that
$$
\|r\|_{\mathcal L_2}^2 = \|\mathds{T}_{dr}[d]+\mathds{T}_{fr}[f]\|_{\mathcal L_2}^2 \le \|\mathds{T}_{dr}[d]\|_{\mathcal L_2}^2.
$$
The corresponding optimization problem is
$$
\|\tilde{\Sigma}\|^2_{\mathcal{L}_{2e},f \leftarrow d} \triangleq \sup_{f,\,d\in \mathcal L_{2e}} \|f\|_{\mathcal L_2}^2
$$
subject to the system dynamics, $x_2(0)=0$, $\|d\|_{\mathcal L_2}^2\le 1$, and the above undetectability condition [2509.11194].

In this formal sense, larger IIG means larger faults can be masked by disturbances, while smaller IIG means the detector is more robust against disturbance masking [2509.11194]. This is the most precise current arXiv definition in the supplied record.

A broader reading is required elsewhere. Several papers do not define IIG as a named property, but do derive gain maps that are naturally interpreted as input-side amplification or masking relations. This suggests that IIG is better understood as a cross-domain gain concept whose exact semantics depend on the modeling framework.

## 2. Nonlinear systems and small-gain interpretations

In "An Extended Small-Gain Theorem" [1409.6991], the term Input-to-Input Gain is not used explicitly, but the paper derives exactly the kind of gain-level interconnection analysis that can be read as an IIG-like framework. The setting is a two-subsystem interconnection
$$
\begin{array}{cc}
\dot{x_1}=f_1(x_1,y_2,u_1), & y_1=h_1(x_1,y_2,u_1)\\
\dot{x_2}=f_2(x_2,y_1,u_2), & y_2=h_2(x_2,y_1,u_2)
\end{array}
$$
with an auxiliary smooth mapping $h$ such that $(y_1,y_2)=h(x_1,x_2,u_1,u_2)$ solves the algebraic output coupling equations [1409.6991].

The paper defines boundedness observability (UO) by requiring a class $K$ function $\alpha^0$ and a nonnegative constant $D^0$ such that
$$
|x(t)|\le \alpha^0\bigl(|x(0)|+\|(u_t^T,y_t^T)^T\|\bigr)+D^0,\qquad \forall t\in[0,T').
$$
It also defines input-to-output practical stability (IOpS) by
$$
|y(t)|\le \beta(|x(0)|,t)+\gamma(\|u\|)+d,
$$
with $\beta$ of class $KL$, $\gamma$ of class $K$, and $d\ge 0$; when $d=0$, the system is input-to-output stable (IOS) [1409.6991].

For the interconnected case, each subsystem is assumed to satisfy
$$
|y_1(t)|\le \beta_1(|x_1(0)|,t)+\gamma_1^y(\|y_{2t}\|)+\gamma_1^u(\|u_1\|)+d_1,
$$
$$
|y_2(t)|\le \beta_2(|x_2(0)|,t)+\gamma_2^y(\|y_{1t}\|)+\gamma_2^u(\|u_2\|)+d_2.
$$
Here $\gamma_i^y$ is the gain from the interconnection signal to $y_i$, while $\gamma_i^u$ is the gain from the external input $u_i$ to $y_i$ [1409.6991]. If an IIG interpretation is sought, the exact objects closest to it are $\gamma_1^u$, $\gamma_2^u$, and the derived interconnected gains $r_1,r_2$.

The extended small-gain condition is
$$
\left. \begin{array}{c}
(Id+\rho_2)\circ\gamma_2^y\circ(Id+\rho_1)\circ\gamma_1^y(s)\le s,\\
(Id+\rho_1)\circ\gamma_1^y\circ(Id+\rho_2)\circ\gamma_2^y(s)\le s,
\end{array}\right\}
\qquad \forall s\ge s_l,
$$
with class $K_\infty$ functions $\rho_1,\rho_2$ and threshold $s_l\ge 0$ [1409.6991]. Under the subsystem IOpS and UO assumptions, the full interconnection is IOpS and has the UO property. More specifically, the theorem yields separate output bounds
$$
|y_1(t)|\le \beta_1'(|x(0)|,t)+(r_1+r_3^1)(\|u\|)+d_1',
$$
$$
|y_2(t)|\le \beta_2'(|x(0)|,t)+(r_2+r_3^2)(\|u\|)+d_2',
$$
where
$$
\left\{ \begin{array}{c}
r_1(s)=(Id+\rho_1^{-1})\circ(Id+\rho_3)^2\circ[\gamma_1^u+\gamma_1^y\circ(Id+\rho_2^{-1}\circ(Id+\rho_3)^2\circ\gamma_2^u)](s),\\
r_2(s)=(Id+\rho_2^{-1})\circ(Id+\rho_3)^2\circ[\gamma_2^u+\gamma_2^y\circ(Id+\rho_1^{-1}\circ(Id+\rho_3)^2\circ\gamma_1^u)](s).
\end{array} \right.
$$
These formulas are the paper’s most direct support for an IIG-style reading: they provide explicit interconnection-level gain bounds from the total external input to each output component [1409.6991].

A related but distinct nonlinear perspective appears in "Interconnecting a System Having a Single Input-to-State Gain With a System Having a Region-Dependent Input-to-State Gain" [1503.02884]. That paper does not define IIG; instead, it studies ISS input-to-state gains and their composition. The interconnected system is
$$
\left\{\begin{array}{rcl}
\dot{x} &=& f(x,z),\\
\dot{z} &=& g(x,z).
\end{array}\right.
$$
The key gain objects are $\gamma$ and $\delta$, and, in the region-dependent formulation, the local and non-local gains $\gamma_\ell$ and $\gamma_g$ together with the compositions $\gamma_\ell\circ\delta$ and $\gamma_g\circ\delta$ [1503.02884]. The local small-gain condition is
$$
\forall s\in (0,M_\ell],\qquad \gamma_\ell\circ\delta(s)<s,
$$
the non-local one is
$$
\forall s\in [M_g,\infty),\qquad \gamma_g\circ\delta(s)<s,
$$
and when $M_g<M_\ell$, the origin is globally asymptotically stable [1503.02884]. The paper therefore treats gain-to-gain composition across an interconnection rather than a distinct named input-to-input gain.

## 3. Input-output amplification in transitional shear flows

"Nonlinear input-output analysis of transitional shear flows using small-signal finite-gain $\mathcal{L}_p$ stability" [2506.08129] addresses the same general issue through nonlinear input-output amplification. The system is written as
$$
\dot{\boldsymbol{a}} = \boldsymbol{L}\boldsymbol{a} + \boldsymbol{\Upsilon}(\boldsymbol{a}) + \boldsymbol{f}, \qquad \boldsymbol{y}=\boldsymbol{a},
$$
where $\boldsymbol{f}$ is the external disturbance input and $\boldsymbol{y}$ is the output [2506.08129]. The operational gain bound is
$$
\left\|\boldsymbol{y}_\tau\right\|_{\mathcal{L}_p} \le \gamma \left\|\boldsymbol{f}_\tau\right\|_{\mathcal{L}_p} + \beta.
$$
This is presented as the nonlinear analog of a gain bound, and the paper explicitly states that the nonlinear gain is guaranteed only when the input forcing is below a permissible threshold [2506.08129].

The underlying theorem is the Small-Signal Finite-Gain $\mathcal{L}_p$ stability theorem. It assumes exponential stability of the unforced equilibrium and a Lyapunov function $V(t,\boldsymbol{a})$ satisfying
$$
c_1 \|\boldsymbol{a}\|^2 \le V(t,\boldsymbol{a}) \le c_2 \|\boldsymbol{a}\|^2,
$$
$$
\frac{\partial V}{\partial t}+\frac{\partial V}{\partial \boldsymbol{a}}\mathcal{N}(t,\boldsymbol{a},\boldsymbol{0}) \le -c_3 \|\boldsymbol{a}\|^2,
$$
$$
\left\|\frac{\partial V}{\partial \boldsymbol{a}}\right\| \le c_4 \|\boldsymbol{a}\|.
$$
Then, for sufficiently small forcing amplitude,
$$
\sup_{0\le t\le \tau}\|\boldsymbol{f}(t)\| \le \min\left\{r_u,\frac{c_1c_3\delta}{c_2c_4K}\right\},
$$
the output satisfies the finite-gain bound [2506.08129].

For the nine-mode shear-flow system, the paper sets $K=1$, $\eta_1=1$, and $\eta_2=0$, yielding
$$
\gamma=\frac{\lambda_{\max}(\boldsymbol{P})\|2\boldsymbol{P}\|}{\lambda_{\min}(\boldsymbol{P})\varepsilon}, \qquad
\beta=\|\boldsymbol{a}_0\|\sqrt{\frac{\lambda_{\max}(\boldsymbol{P})}{\lambda_{\min}(\boldsymbol{P})}\rho},
$$
and the forcing threshold
$$
\sup_{0\le t\le \tau}\|\boldsymbol{f}(t)\| \le f_{\text{LMI}} := \frac{\lambda_{\min}(\boldsymbol{P})\varepsilon\delta}{\lambda_{\max}(\boldsymbol{P})\|2\boldsymbol{P}\|}.
$$
The paper computes these bounds via Linear Matrix Inequalities (LMI) and Sum-of-Squares (SOS), using a quadratic Lyapunov function $V=\boldsymbol{a}^T\boldsymbol{P}\boldsymbol{a}$ [2506.08129].

Several reported conclusions are directly relevant to IIG-like amplification. The nonlinear $\mathcal{L}_p$ gain from SSFG analysis is several orders of magnitude higher than the linear $\mathcal{L}_p$ gain; both nonlinear and linear $\mathcal{L}_p$ gains are much larger than the linear $\mathcal{L}_2$ gain; and finite gain is only guaranteed below a permissible forcing amplitude, which the paper describes as an inherently nonlinear property that cannot be predicted by linear input-output analysis [2506.08129]. The reported scaling laws are
- nonlinear $\mathcal{L}_p$ gain via LMI: $\gamma_{\max}\sim \mathrm{Re}^{4.04}$,
- nonlinear $\mathcal{L}_p$ gain via SOS: $\gamma_{\max,\text{SOS}}\sim \mathrm{Re}^{3.61}$,
- linear $\mathcal{L}_p$ gain: $\gamma_{cor}\sim \mathrm{Re}^{5.00}$,
- linear $\mathcal{L}_2$ gain: $\gamma_{\mathcal{L}_2}\sim \mathrm{Re}^{2.00}$ [2506.08129].

This use of IIG is therefore not a masking metric, but a nonlinear amplification bound from disturbance input to flow response.

## 4. Scattering, directionality, and gain asymmetry in microwave SQUID amplifiers

In "Gain, directionality and noise in microwave SQUID amplifiers: Input-output approach" [1206.4706], IIG is again not formally defined, but the concept is embodied in the small-signal transmission from the input mode to the output mode. The dc SQUID is modeled as a running-state Josephson device whose shunt resistors are replaced by semi-infinite transmission lines of impedance $Z_C = R$, and the wave amplitudes are expressed in input-output form
$$
A_{i}^{\rm in/out}(t) =\frac{V^{i}\pm Z_C I^{i}}{2\sqrt{Z_C}}, \qquad [i \in \{L, R\}]
$$
with common and differential coordinates
$$
\varphi^{C} = \frac{\varphi_L+\varphi_R}{2},\qquad \varphi^{D} = \frac{\varphi_L-\varphi_R}{2}.
$$
The system is analyzed as a linear scattering problem among the participating modes, leading to an admittance matrix
$$
\mathbb{Y} = \widehat{\mathbb{M}}\widecheck{\mathbb{M}}^{-1},
$$
and scattering matrix
$$
\mathbb{S} = (\mathbb{U}+\mathbb{Y})^{-1}(\mathbb{U}-\mathbb{Y})
$$
[1206.4706].

The relevant off-diagonal coefficients are the forward and reverse conversion channels:
- forward: $s^{CD}$ or $z^{CD}$, describing differential $\to$ common conversion,
- reverse: $s^{DC}$ or $z^{DC}$, describing common $\to$ differential conversion [1206.4706].

The paper explicitly interprets the forward channel as the amplification path relevant to input-to-output gain. In that sense, the IIG-like quantity is encoded in $|s^{CD}|^2$ or, in power-gain form, $|z^{CD}|^2$ [1206.4706]. The power gain and reverse gain are
$$
G_P[\omega_m] = |z^{CD}[\omega_m]|^2 \,{\rm Re}[z^{CC}[\omega_m]]\,{\rm Re}[z^{DD}[\omega_m]],
$$
$$
G_P^{\rm rev} = |z^{DC}[\omega_m]|^2 \,{\rm Re}[z^{CC}[\omega_m]]\,{\rm Re}[z^{DD}[\omega_m]],
$$
and directionality is defined as $G_P - G_P^{\rm rev}$ [1206.4706].

A central result is that including only the fundamental Josephson frequency gives no forward/backward asymmetry, while higher harmonics produce $|s^{CD}|^2 \neq |s^{DC}|^2$, yielding nonreciprocal gain [1206.4706]. The physical mechanism is multiharmonic Josephson mixing: the running Josephson phase generates multiple harmonics, these act like a multitone pump, signal conversion proceeds through multiple interfering pathways, and the phases of the harmonics are not symmetric under $t\to -t$, so frequency conversion becomes asymmetric [1206.4706].

The gain depends strongly on bias and frequency. With
$$
\varepsilon \equiv \frac{I_0}{I_B}=\frac{\omega_0}{\omega_B},
$$
the paper finds an optimal bias around
$$
\varepsilon \approx 0.455
$$
for several performance measures. At low signal frequency, the quasistatic result is
$$
G_P^{dc} \approx \rho_g \left(\frac{\omega_0}{\omega_m}\right)^2,
$$
so the power gain scales roughly as $1/\omega_m^2$, and the paper states that no power gain is obtained for signal frequencies close to the plasma frequency of the junctions [1206.4706]. Here IIG is best read as forward mode-conversion gain with inherent nonreciprocity.

## 5. Neural-network training and Optimal Input Gain

"Optimal Input Gain: All You Need to Supercharge a Feed-Forward Neural Network" [2303.17732] uses the term Optimal Input Gain (OIG), which is the most explicit input-side use of “gain” outside control and systems theory. The paper considers two equivalent MLPs related by a linear preprocessing matrix:
$$
\mathbf{x}'_p = \mathbf{A}\mathbf{x}_p,
$$
with weight equivalence
$$
\mathbf{W}' \mathbf{A} = \mathbf{W}.
$$
The central claim is that training equivalent networks is not dynamically equivalent under gradient-based methods: linear preprocessing changes the effective learning direction [2303.17732].

Let
$$
\mathbf{G} = \frac{1}{N_v}\sum_{p=1}^{N_v}\boldsymbol{\delta}_p \mathbf{x}_p^T
$$
denote the negative gradient matrix for the input weights of the original network. For the transformed network,
$$
\mathbf{G}' = \mathbf{G}\mathbf{A}^T,
$$
and when mapped back to the original network,
$$
\mathbf{G}'' = \mathbf{G}\mathbf{R}_i, \qquad \mathbf{R}_i = \mathbf{A}^T\mathbf{A}.
$$
The paper states that linear preprocessing of the inputs is equivalent to multiplying the original negative gradient matrix by an autocorrelation matrix $\mathbf{R}_i$ at each iteration [2303.17732]. In this formulation, input gain is a trainable preconditioning mechanism in input space.

The most important special case is diagonal:
$$
\mathbf{R}= \begin{bmatrix}
r(1) & 0 & \cdots & 0 \\
0 & r(2) & \cdots & 0 \\
\vdots & \vdots & \ddots & \vdots \\
0 & 0 & \cdots & r(N+1)
\end{bmatrix},
$$
where the diagonal entries $r(n)$ are the input gains [2303.17732]. The update becomes
$$
\mathbf{W} \leftarrow \mathbf{W} + \mathbf{r}\cdot \mathbf{G},
$$
so each input coordinate receives its own learned gain rather than a single scalar learning rate.

The gains are obtained via a second-order method. The derivative of the error with respect to a gain $r(m)$ is
$$
d_r(m) \equiv \frac{\partial E}{\partial r(m)} = -\frac{2}{N_v}\sum_{p=1}^{N_v} x_p(m) \sum_{i=1}^{M}\big[t_p(i)-y_p(i)\big]\,v(i,m),
$$
with
$$
v(i,m)=\sum_{k=1}^{N_h} w_{oh}(i,k)o'_p(k)g(k,m),
$$
and the Gauss-Newton approximation to the Hessian is
$$
h_{ig}(m,u) \equiv \frac{\partial^2 E}{\partial r(m)\partial r(u)} = \frac{2}{N_v}\sum_{p=1}^{N_v} x_p(m)x_p(u)\sum_{i=1}^{M} v(i,m)v(i,u).
$$
The gain vector is then found from
$$
\mathbf{H}_{ig}\mathbf{r}=\mathbf{d}_r
$$
[2303.17732].

The paper also states that Hidden Weight Optimization (HWO) is equivalent to BP with whitening applied to the inputs. With
$$
\mathbf{R}_i = \frac{1}{N_v}\sum_{p=1}^{N_v}\mathbf{x}_p\mathbf{x}_p^T,
$$
HWO solves
$$
\mathbf{G}_{hwo}\mathbf{R}_i = \mathbf{G}, \qquad \mathbf{G}_{hwo}=\mathbf{G}\mathbf{R}_i^{-1}.
$$
Using the SVD
$$
\mathbf{R}_i = \mathbf{U}\boldsymbol{\Sigma}\mathbf{U}^T,
$$
the whitening matrix is identified in the paper as
$$
\mathbf{A}=\boldsymbol{\Sigma}^{1/2}\mathbf{U}^T
$$
[2303.17732]. Empirically, the paper reports that OIG-BP improves over OWO-BP on all datasets, OIG-HWO improves over OIG-BP, and OIG-HWO often performs close to or sometimes matching Levenberg-Marquardt with far lower computational cost [2303.17732]. In this literature, IIG corresponds to adaptive input-space gain selection rather than disturbance masking or dynamical amplification.

## 6. Structural themes, limitations, and recurrent misconceptions

A recurring misconception is that IIG denotes a single theorem or universally accepted systems property. The supplied literature does not support that claim. Only [2509.11194] introduces IIG as a named metric. The other papers either do not use the term explicitly or deploy closely related but domain-specific notions: external-input-to-output gains in nonlinear interconnections [1409.6991], input-to-state gain compositions [1503.02884], nonlinear input-output gain under small-signal forcing [2506.08129], forward scattering gain and reverse-gain asymmetry [1206.4706], and optimal input scaling/preconditioning in neural-network training [2303.17732].

A second misconception is that gain is always reciprocal or symmetric. The SQUID amplifier analysis directly contradicts this: with higher harmonics included, forward and reverse gains differ, and directionality is the difference between those gains [1206.4706]. Likewise, the anomaly-detection formulation is intrinsically asymmetric because it asks how disturbance can conceal fault energy, not how the channels interchange roles [2509.11194].

A third misconception is that gain bounds are always global. The nonlinear control and flow papers repeatedly introduce restricted regimes. In the extended small-gain theorem, the practical threshold $s_l$ yields a condition that need only hold for all $s\ge s_l$ [1409.6991]. In the region-dependent ISS setting, local and non-local small-gain inequalities hold on different intervals and are then “glued” via the overlap condition $M_g<M_\ell$ [1503.02884]. In transitional shear flows, finite gain is guaranteed only below a permissible forcing amplitude [2506.08129].

The explicit limitation theory is most developed in [2509.11194]. Using left coprime factorization,
$$
\mathds{T}_{fr}=M_I^{-1}N_f,\qquad \mathds{T}_{dr}=M_I^{-1}N_d,
$$
the paper defines
$$
S_I \triangleq \frac{N_d}{N_f}
$$
and derives the lower bound
$$
\tilde\gamma^* \ge 4\|S_I\|_{\mathcal H_\infty}^2.
$$
It then applies the Poisson integral relation and Blaschke-product factorization to show that non-minimum-phase zeros impose fundamental lower bounds on IIG [2509.11194]. The corresponding bound is
$$
\|S_I\|_{\mathcal H_\infty} \ge \max_{\mu_h\in\mathcal Z_{S_I},\nu_k\in\mathcal Z_{P_I}}
\left\{ |\mathcal B^{-1}_{S_I}(\nu_k)|,\, |\mathcal B^{-1}_{P_I}(\mu_h)|-1 \right\}.
$$
This establishes that IIG cannot in general be made arbitrarily small; it is constrained by transmission-zero geometry [2509.11194].

## 7. Comparative view across domains

The following comparison summarizes the principal meanings of IIG-like constructions in the cited literature.

| Domain | IIG or closest object | Main role |
|---|---|---|
| LTI anomaly impact analysis | IIG as maximum energy of undetectable faults for a given disturbance intensity | Robust fault sensitivity under disturbance masking |
| Nonlinear small-gain interconnections | $\gamma_i^u$, $r_1$, $r_2$ | Componentwise external-input-to-output gain bounds |
| Transitional shear flows | Nonlinear $\mathcal{L}_p$ gain $\gamma$ with permissible forcing amplitude | Nonlinear disturbance amplification |
| Microwave SQUID amplifiers | Forward scattering gain $|s^{CD}|^2$ or $|z^{CD}|^2$ | Differential-to-common mode conversion and directionality |
| Feed-forward neural networks | Optimal Input Gain (OIG) | Learned input-space preconditioning |
| Region-dependent ISS interconnections | $\gamma_\ell\circ\delta$, $\gamma_g\circ\delta$ | Gain composition across interconnected subsystems |

Across these settings, the common structure is a gain relation defined on the input side of a system, interconnection, or learning rule. What changes is the object being bounded: undetectable fault energy, subsystem output response, flow amplification, scattering transmission, or gradient preconditioning. This suggests that “Input-to-Input Gain” functions as a unifying interpretive label only at a high level. At the formal level, the literature remains plural: the term is explicit and optimization-based in [2509.11194], implicit and componentwise in [1409.6991], region-dependent and ISS-based in [1503.02884], nonlinear and Lyapunov-certified in [2506.08129], scattering-theoretic and nonreciprocal in [1206.4706], and optimization/preconditioning-based in [2303.17732].

Source: https://www.emergentmind.com/topics/input-to-input-gain-iig