Input-to-Input Gain: Systems Perspective
- Input-to-Input Gain (IIG) is defined in fault detection as the maximum energy of undetectable faults for a given disturbance intensity in LTI systems.
- In nonlinear and ISS frameworks, IIG-like gains characterize how input perturbations amplify downstream outputs, influencing stability and robustness.
- In neural network training, optimal input gain adapts input-space preconditioning, enhancing gradient-based learning efficiency.
Searching arXiv for papers on “Input-to-Input Gain” and closely related formulations. Input-to-Input Gain (IIG) is not a uniformly standardized term across the arXiv literature. In the most explicit formulation presently represented in the cited corpus, IIG is introduced as an input-side analogue of output-to-output gain (OOG) for linear time-invariant fault-detection settings, where it measures the maximum energy of undetectable faults for a given disturbance intensity (Dong et al., 14 Sep 2025). In adjacent literatures, however, closely related ideas appear under other names: componentwise external-input-to-output gain in nonlinear small-gain theory (Li et al., 2014), nonlinear input-output amplification under small-signal finite-gain stability in transitional shear flows (Wei et al., 9 Jun 2025), forward scattering gain in microwave SQUID amplifiers (Kamal et al., 2012), region-dependent input-to-state gain composition in interconnected nonlinear systems (Shiromoto et al., 2015), and Optimal Input Gain (OIG) in feed-forward neural-network training (Rane et al., 2023). Taken together, these works indicate that “IIG” functions less as a single universal definition than as a family of gain-level constructions relating input-side perturbations, channels, or transformations to downstream system behavior.
1. Formal definition and scope
The clearest named definition appears in "Fundamental limitations of sensitivity metrics for anomaly impact analysis in LTI systems" (Dong et al., 14 Sep 2025). There, IIG is proposed as a new measure of robust fault sensitivity and is defined as the maximum energy of undetectable faults for a given disturbance intensity. The fault/detection model is
$\tilde{\Sigma}:\left\{ \begin{array}{l} \dot{x}_2(t)=\tilde{A}x_2(t)+B_d d(t)+B_f f(t),\[2mm] r(t)=Cx_2(t)+D_d d(t)+D_f f(t), \end{array} \right.$
with disturbance , fault , and residual , where
$r = \mathds{T}_{dr}[d]+\mathds{T}_{fr}[f].$
A fault is called -undetectable if there exists such that
$\|r\|_{\mathcal L_2}^2 = \|\mathds{T}_{dr}[d]+\mathds{T}_{fr}[f]\|_{\mathcal L_2}^2 \le \|\mathds{T}_{dr}[d]\|_{\mathcal L_2}^2.$
The corresponding optimization problem is
$\tilde{\Sigma}:\left\{ \begin{array}{l} \dot{x}_2(t)=\tilde{A}x_2(t)+B_d d(t)+B_f f(t),\[2mm] r(t)=Cx_2(t)+D_d d(t)+D_f f(t), \end{array} \right.$0
subject to the system dynamics, $\tilde{\Sigma}:\left\{ \begin{array}{l} \dot{x}_2(t)=\tilde{A}x_2(t)+B_d d(t)+B_f f(t),\[2mm] r(t)=Cx_2(t)+D_d d(t)+D_f f(t), \end{array} \right.$1, $\tilde{\Sigma}:\left\{ \begin{array}{l} \dot{x}_2(t)=\tilde{A}x_2(t)+B_d d(t)+B_f f(t),\[2mm] r(t)=Cx_2(t)+D_d d(t)+D_f f(t), \end{array} \right.$2, and the above undetectability condition (Dong et al., 14 Sep 2025).
In this formal sense, larger IIG means larger faults can be masked by disturbances, while smaller IIG means the detector is more robust against disturbance masking (Dong et al., 14 Sep 2025). This is the most precise current arXiv definition in the supplied record.
A broader reading is required elsewhere. Several papers do not define IIG as a named property, but do derive gain maps that are naturally interpreted as input-side amplification or masking relations. This suggests that IIG is better understood as a cross-domain gain concept whose exact semantics depend on the modeling framework.
2. Nonlinear systems and small-gain interpretations
In "An Extended Small-Gain Theorem" (Li et al., 2014), the term Input-to-Input Gain is not used explicitly, but the paper derives exactly the kind of gain-level interconnection analysis that can be read as an IIG-like framework. The setting is a two-subsystem interconnection
$\tilde{\Sigma}:\left\{ \begin{array}{l} \dot{x}_2(t)=\tilde{A}x_2(t)+B_d d(t)+B_f f(t),\[2mm] r(t)=Cx_2(t)+D_d d(t)+D_f f(t), \end{array} \right.$3
with an auxiliary smooth mapping $\tilde{\Sigma}:\left\{ \begin{array}{l} \dot{x}_2(t)=\tilde{A}x_2(t)+B_d d(t)+B_f f(t),\[2mm] r(t)=Cx_2(t)+D_d d(t)+D_f f(t), \end{array} \right.$4 such that $\tilde{\Sigma}:\left\{ \begin{array}{l} \dot{x}_2(t)=\tilde{A}x_2(t)+B_d d(t)+B_f f(t),\[2mm] r(t)=Cx_2(t)+D_d d(t)+D_f f(t), \end{array} \right.$5 solves the algebraic output coupling equations (Li et al., 2014).
The paper defines boundedness observability (UO) by requiring a class $\tilde{\Sigma}:\left\{ \begin{array}{l} \dot{x}_2(t)=\tilde{A}x_2(t)+B_d d(t)+B_f f(t),\[2mm] r(t)=Cx_2(t)+D_d d(t)+D_f f(t), \end{array} \right.$6 function $\tilde{\Sigma}:\left\{ \begin{array}{l} \dot{x}_2(t)=\tilde{A}x_2(t)+B_d d(t)+B_f f(t),\[2mm] r(t)=Cx_2(t)+D_d d(t)+D_f f(t), \end{array} \right.$7 and a nonnegative constant $\tilde{\Sigma}:\left\{ \begin{array}{l} \dot{x}_2(t)=\tilde{A}x_2(t)+B_d d(t)+B_f f(t),\[2mm] r(t)=Cx_2(t)+D_d d(t)+D_f f(t), \end{array} \right.$8 such that
$\tilde{\Sigma}:\left\{ \begin{array}{l} \dot{x}_2(t)=\tilde{A}x_2(t)+B_d d(t)+B_f f(t),\[2mm] r(t)=Cx_2(t)+D_d d(t)+D_f f(t), \end{array} \right.$9
It also defines input-to-output practical stability (IOpS) by
0
with 1 of class 2, 3 of class 4, and 5; when 6, the system is input-to-output stable (IOS) (Li et al., 2014).
For the interconnected case, each subsystem is assumed to satisfy
7
8
Here 9 is the gain from the interconnection signal to 0, while 1 is the gain from the external input 2 to 3 (Li et al., 2014). If an IIG interpretation is sought, the exact objects closest to it are 4, 5, and the derived interconnected gains 6.
The extended small-gain condition is
7
with class 8 functions 9 and threshold 0 (Li et al., 2014). Under the subsystem IOpS and UO assumptions, the full interconnection is IOpS and has the UO property. More specifically, the theorem yields separate output bounds
1
2
where
3
These formulas are the paper’s most direct support for an IIG-style reading: they provide explicit interconnection-level gain bounds from the total external input to each output component (Li et al., 2014).
A related but distinct nonlinear perspective appears in "Interconnecting a System Having a Single Input-to-State Gain With a System Having a Region-Dependent Input-to-State Gain" (Shiromoto et al., 2015). That paper does not define IIG; instead, it studies ISS input-to-state gains and their composition. The interconnected system is
4
The key gain objects are 5 and 6, and, in the region-dependent formulation, the local and non-local gains 7 and 8 together with the compositions 9 and $r = \mathds{T}_{dr}[d]+\mathds{T}_{fr}[f].$0 (Shiromoto et al., 2015). The local small-gain condition is
$r = \mathds{T}_{dr}[d]+\mathds{T}_{fr}[f].$1
the non-local one is
$r = \mathds{T}_{dr}[d]+\mathds{T}_{fr}[f].$2
and when $r = \mathds{T}_{dr}[d]+\mathds{T}_{fr}[f].$3, the origin is globally asymptotically stable (Shiromoto et al., 2015). The paper therefore treats gain-to-gain composition across an interconnection rather than a distinct named input-to-input gain.
3. Input-output amplification in transitional shear flows
"Nonlinear input-output analysis of transitional shear flows using small-signal finite-gain $r = \mathds{T}_{dr}[d]+\mathds{T}_{fr}[f].$4 stability" (Wei et al., 9 Jun 2025) addresses the same general issue through nonlinear input-output amplification. The system is written as
$r = \mathds{T}_{dr}[d]+\mathds{T}_{fr}[f].$5
where $r = \mathds{T}_{dr}[d]+\mathds{T}_{fr}[f].$6 is the external disturbance input and $r = \mathds{T}_{dr}[d]+\mathds{T}_{fr}[f].$7 is the output (Wei et al., 9 Jun 2025). The operational gain bound is
$r = \mathds{T}_{dr}[d]+\mathds{T}_{fr}[f].$8
This is presented as the nonlinear analog of a gain bound, and the paper explicitly states that the nonlinear gain is guaranteed only when the input forcing is below a permissible threshold (Wei et al., 9 Jun 2025).
The underlying theorem is the Small-Signal Finite-Gain $r = \mathds{T}_{dr}[d]+\mathds{T}_{fr}[f].$9 stability theorem. It assumes exponential stability of the unforced equilibrium and a Lyapunov function 0 satisfying
1
2
3
Then, for sufficiently small forcing amplitude,
4
the output satisfies the finite-gain bound (Wei et al., 9 Jun 2025).
For the nine-mode shear-flow system, the paper sets 5, 6, and 7, yielding
8
and the forcing threshold
9
The paper computes these bounds via Linear Matrix Inequalities (LMI) and Sum-of-Squares (SOS), using a quadratic Lyapunov function 0 (Wei et al., 9 Jun 2025).
Several reported conclusions are directly relevant to IIG-like amplification. The nonlinear 1 gain from SSFG analysis is several orders of magnitude higher than the linear 2 gain; both nonlinear and linear 3 gains are much larger than the linear 4 gain; and finite gain is only guaranteed below a permissible forcing amplitude, which the paper describes as an inherently nonlinear property that cannot be predicted by linear input-output analysis (Wei et al., 9 Jun 2025). The reported scaling laws are
- nonlinear 5 gain via LMI: 6,
- nonlinear 7 gain via SOS: 8,
- linear 9 gain: 0,
- linear 1 gain: 2 (Wei et al., 9 Jun 2025).
This use of IIG is therefore not a masking metric, but a nonlinear amplification bound from disturbance input to flow response.
4. Scattering, directionality, and gain asymmetry in microwave SQUID amplifiers
In "Gain, directionality and noise in microwave SQUID amplifiers: Input-output approach" (Kamal et al., 2012), IIG is again not formally defined, but the concept is embodied in the small-signal transmission from the input mode to the output mode. The dc SQUID is modeled as a running-state Josephson device whose shunt resistors are replaced by semi-infinite transmission lines of impedance 3, and the wave amplitudes are expressed in input-output form
4
with common and differential coordinates
5
The system is analyzed as a linear scattering problem among the participating modes, leading to an admittance matrix
6
and scattering matrix
7
The relevant off-diagonal coefficients are the forward and reverse conversion channels:
- forward: 8 or 9, describing differential $\|r\|_{\mathcal L_2}^2 = \|\mathds{T}_{dr}[d]+\mathds{T}_{fr}[f]\|_{\mathcal L_2}^2 \le \|\mathds{T}_{dr}[d]\|_{\mathcal L_2}^2.$0 common conversion,
- reverse: $\|r\|_{\mathcal L_2}^2 = \|\mathds{T}_{dr}[d]+\mathds{T}_{fr}[f]\|_{\mathcal L_2}^2 \le \|\mathds{T}_{dr}[d]\|_{\mathcal L_2}^2.$1 or $\|r\|_{\mathcal L_2}^2 = \|\mathds{T}_{dr}[d]+\mathds{T}_{fr}[f]\|_{\mathcal L_2}^2 \le \|\mathds{T}_{dr}[d]\|_{\mathcal L_2}^2.$2, describing common $\|r\|_{\mathcal L_2}^2 = \|\mathds{T}_{dr}[d]+\mathds{T}_{fr}[f]\|_{\mathcal L_2}^2 \le \|\mathds{T}_{dr}[d]\|_{\mathcal L_2}^2.$3 differential conversion (Kamal et al., 2012).
The paper explicitly interprets the forward channel as the amplification path relevant to input-to-output gain. In that sense, the IIG-like quantity is encoded in $\|r\|_{\mathcal L_2}^2 = \|\mathds{T}_{dr}[d]+\mathds{T}_{fr}[f]\|_{\mathcal L_2}^2 \le \|\mathds{T}_{dr}[d]\|_{\mathcal L_2}^2.$4 or, in power-gain form, $\|r\|_{\mathcal L_2}^2 = \|\mathds{T}_{dr}[d]+\mathds{T}_{fr}[f]\|_{\mathcal L_2}^2 \le \|\mathds{T}_{dr}[d]\|_{\mathcal L_2}^2.$5 (Kamal et al., 2012). The power gain and reverse gain are
$\|r\|_{\mathcal L_2}^2 = \|\mathds{T}_{dr}[d]+\mathds{T}_{fr}[f]\|_{\mathcal L_2}^2 \le \|\mathds{T}_{dr}[d]\|_{\mathcal L_2}^2.$6
$\|r\|_{\mathcal L_2}^2 = \|\mathds{T}_{dr}[d]+\mathds{T}_{fr}[f]\|_{\mathcal L_2}^2 \le \|\mathds{T}_{dr}[d]\|_{\mathcal L_2}^2.$7
and directionality is defined as $\|r\|_{\mathcal L_2}^2 = \|\mathds{T}_{dr}[d]+\mathds{T}_{fr}[f]\|_{\mathcal L_2}^2 \le \|\mathds{T}_{dr}[d]\|_{\mathcal L_2}^2.$8 (Kamal et al., 2012).
A central result is that including only the fundamental Josephson frequency gives no forward/backward asymmetry, while higher harmonics produce $\|r\|_{\mathcal L_2}^2 = \|\mathds{T}_{dr}[d]+\mathds{T}_{fr}[f]\|_{\mathcal L_2}^2 \le \|\mathds{T}_{dr}[d]\|_{\mathcal L_2}^2.$9, yielding nonreciprocal gain (Kamal et al., 2012). The physical mechanism is multiharmonic Josephson mixing: the running Josephson phase generates multiple harmonics, these act like a multitone pump, signal conversion proceeds through multiple interfering pathways, and the phases of the harmonics are not symmetric under $\tilde{\Sigma}:\left\{ \begin{array}{l} \dot{x}_2(t)=\tilde{A}x_2(t)+B_d d(t)+B_f f(t),\[2mm] r(t)=Cx_2(t)+D_d d(t)+D_f f(t), \end{array} \right.$00, so frequency conversion becomes asymmetric (Kamal et al., 2012).
The gain depends strongly on bias and frequency. With
$\tilde{\Sigma}:\left\{ \begin{array}{l} \dot{x}_2(t)=\tilde{A}x_2(t)+B_d d(t)+B_f f(t),\[2mm] r(t)=Cx_2(t)+D_d d(t)+D_f f(t), \end{array} \right.$01
the paper finds an optimal bias around
$\tilde{\Sigma}:\left\{ \begin{array}{l} \dot{x}_2(t)=\tilde{A}x_2(t)+B_d d(t)+B_f f(t),\[2mm] r(t)=Cx_2(t)+D_d d(t)+D_f f(t), \end{array} \right.$02
for several performance measures. At low signal frequency, the quasistatic result is
$\tilde{\Sigma}:\left\{ \begin{array}{l} \dot{x}_2(t)=\tilde{A}x_2(t)+B_d d(t)+B_f f(t),\[2mm] r(t)=Cx_2(t)+D_d d(t)+D_f f(t), \end{array} \right.$03
so the power gain scales roughly as $\tilde{\Sigma}:\left\{ \begin{array}{l} \dot{x}_2(t)=\tilde{A}x_2(t)+B_d d(t)+B_f f(t),\[2mm] r(t)=Cx_2(t)+D_d d(t)+D_f f(t), \end{array} \right.$04, and the paper states that no power gain is obtained for signal frequencies close to the plasma frequency of the junctions (Kamal et al., 2012). Here IIG is best read as forward mode-conversion gain with inherent nonreciprocity.
5. Neural-network training and Optimal Input Gain
"Optimal Input Gain: All You Need to Supercharge a Feed-Forward Neural Network" (Rane et al., 2023) uses the term Optimal Input Gain (OIG), which is the most explicit input-side use of “gain” outside control and systems theory. The paper considers two equivalent MLPs related by a linear preprocessing matrix:
$\tilde{\Sigma}:\left\{ \begin{array}{l} \dot{x}_2(t)=\tilde{A}x_2(t)+B_d d(t)+B_f f(t),\[2mm] r(t)=Cx_2(t)+D_d d(t)+D_f f(t), \end{array} \right.$05
with weight equivalence
$\tilde{\Sigma}:\left\{ \begin{array}{l} \dot{x}_2(t)=\tilde{A}x_2(t)+B_d d(t)+B_f f(t),\[2mm] r(t)=Cx_2(t)+D_d d(t)+D_f f(t), \end{array} \right.$06
The central claim is that training equivalent networks is not dynamically equivalent under gradient-based methods: linear preprocessing changes the effective learning direction (Rane et al., 2023).
Let
$\tilde{\Sigma}:\left\{ \begin{array}{l} \dot{x}_2(t)=\tilde{A}x_2(t)+B_d d(t)+B_f f(t),\[2mm] r(t)=Cx_2(t)+D_d d(t)+D_f f(t), \end{array} \right.$07
denote the negative gradient matrix for the input weights of the original network. For the transformed network,
$\tilde{\Sigma}:\left\{ \begin{array}{l} \dot{x}_2(t)=\tilde{A}x_2(t)+B_d d(t)+B_f f(t),\[2mm] r(t)=Cx_2(t)+D_d d(t)+D_f f(t), \end{array} \right.$08
and when mapped back to the original network,
$\tilde{\Sigma}:\left\{ \begin{array}{l} \dot{x}_2(t)=\tilde{A}x_2(t)+B_d d(t)+B_f f(t),\[2mm] r(t)=Cx_2(t)+D_d d(t)+D_f f(t), \end{array} \right.$09
The paper states that linear preprocessing of the inputs is equivalent to multiplying the original negative gradient matrix by an autocorrelation matrix $\tilde{\Sigma}:\left\{ \begin{array}{l} \dot{x}_2(t)=\tilde{A}x_2(t)+B_d d(t)+B_f f(t),\[2mm] r(t)=Cx_2(t)+D_d d(t)+D_f f(t), \end{array} \right.$10 at each iteration (Rane et al., 2023). In this formulation, input gain is a trainable preconditioning mechanism in input space.
The most important special case is diagonal:
$\tilde{\Sigma}:\left\{ \begin{array}{l} \dot{x}_2(t)=\tilde{A}x_2(t)+B_d d(t)+B_f f(t),\[2mm] r(t)=Cx_2(t)+D_d d(t)+D_f f(t), \end{array} \right.$11
where the diagonal entries $\tilde{\Sigma}:\left\{ \begin{array}{l} \dot{x}_2(t)=\tilde{A}x_2(t)+B_d d(t)+B_f f(t),\[2mm] r(t)=Cx_2(t)+D_d d(t)+D_f f(t), \end{array} \right.$12 are the input gains (Rane et al., 2023). The update becomes
$\tilde{\Sigma}:\left\{ \begin{array}{l} \dot{x}_2(t)=\tilde{A}x_2(t)+B_d d(t)+B_f f(t),\[2mm] r(t)=Cx_2(t)+D_d d(t)+D_f f(t), \end{array} \right.$13
so each input coordinate receives its own learned gain rather than a single scalar learning rate.
The gains are obtained via a second-order method. The derivative of the error with respect to a gain $\tilde{\Sigma}:\left\{ \begin{array}{l} \dot{x}_2(t)=\tilde{A}x_2(t)+B_d d(t)+B_f f(t),\[2mm] r(t)=Cx_2(t)+D_d d(t)+D_f f(t), \end{array} \right.$14 is
$\tilde{\Sigma}:\left\{ \begin{array}{l} \dot{x}_2(t)=\tilde{A}x_2(t)+B_d d(t)+B_f f(t),\[2mm] r(t)=Cx_2(t)+D_d d(t)+D_f f(t), \end{array} \right.$15
with
$\tilde{\Sigma}:\left\{ \begin{array}{l} \dot{x}_2(t)=\tilde{A}x_2(t)+B_d d(t)+B_f f(t),\[2mm] r(t)=Cx_2(t)+D_d d(t)+D_f f(t), \end{array} \right.$16
and the Gauss-Newton approximation to the Hessian is
$\tilde{\Sigma}:\left\{ \begin{array}{l} \dot{x}_2(t)=\tilde{A}x_2(t)+B_d d(t)+B_f f(t),\[2mm] r(t)=Cx_2(t)+D_d d(t)+D_f f(t), \end{array} \right.$17
The gain vector is then found from
$\tilde{\Sigma}:\left\{ \begin{array}{l} \dot{x}_2(t)=\tilde{A}x_2(t)+B_d d(t)+B_f f(t),\[2mm] r(t)=Cx_2(t)+D_d d(t)+D_f f(t), \end{array} \right.$18
The paper also states that Hidden Weight Optimization (HWO) is equivalent to BP with whitening applied to the inputs. With
$\tilde{\Sigma}:\left\{ \begin{array}{l} \dot{x}_2(t)=\tilde{A}x_2(t)+B_d d(t)+B_f f(t),\[2mm] r(t)=Cx_2(t)+D_d d(t)+D_f f(t), \end{array} \right.$19
HWO solves
$\tilde{\Sigma}:\left\{ \begin{array}{l} \dot{x}_2(t)=\tilde{A}x_2(t)+B_d d(t)+B_f f(t),\[2mm] r(t)=Cx_2(t)+D_d d(t)+D_f f(t), \end{array} \right.$20
Using the SVD
$\tilde{\Sigma}:\left\{ \begin{array}{l} \dot{x}_2(t)=\tilde{A}x_2(t)+B_d d(t)+B_f f(t),\[2mm] r(t)=Cx_2(t)+D_d d(t)+D_f f(t), \end{array} \right.$21
the whitening matrix is identified in the paper as
$\tilde{\Sigma}:\left\{ \begin{array}{l} \dot{x}_2(t)=\tilde{A}x_2(t)+B_d d(t)+B_f f(t),\[2mm] r(t)=Cx_2(t)+D_d d(t)+D_f f(t), \end{array} \right.$22
(Rane et al., 2023). Empirically, the paper reports that OIG-BP improves over OWO-BP on all datasets, OIG-HWO improves over OIG-BP, and OIG-HWO often performs close to or sometimes matching Levenberg-Marquardt with far lower computational cost (Rane et al., 2023). In this literature, IIG corresponds to adaptive input-space gain selection rather than disturbance masking or dynamical amplification.
6. Structural themes, limitations, and recurrent misconceptions
A recurring misconception is that IIG denotes a single theorem or universally accepted systems property. The supplied literature does not support that claim. Only (Dong et al., 14 Sep 2025) introduces IIG as a named metric. The other papers either do not use the term explicitly or deploy closely related but domain-specific notions: external-input-to-output gains in nonlinear interconnections (Li et al., 2014), input-to-state gain compositions (Shiromoto et al., 2015), nonlinear input-output gain under small-signal forcing (Wei et al., 9 Jun 2025), forward scattering gain and reverse-gain asymmetry (Kamal et al., 2012), and optimal input scaling/preconditioning in neural-network training (Rane et al., 2023).
A second misconception is that gain is always reciprocal or symmetric. The SQUID amplifier analysis directly contradicts this: with higher harmonics included, forward and reverse gains differ, and directionality is the difference between those gains (Kamal et al., 2012). Likewise, the anomaly-detection formulation is intrinsically asymmetric because it asks how disturbance can conceal fault energy, not how the channels interchange roles (Dong et al., 14 Sep 2025).
A third misconception is that gain bounds are always global. The nonlinear control and flow papers repeatedly introduce restricted regimes. In the extended small-gain theorem, the practical threshold $\tilde{\Sigma}:\left\{ \begin{array}{l} \dot{x}_2(t)=\tilde{A}x_2(t)+B_d d(t)+B_f f(t),\[2mm] r(t)=Cx_2(t)+D_d d(t)+D_f f(t), \end{array} \right.$23 yields a condition that need only hold for all $\tilde{\Sigma}:\left\{ \begin{array}{l} \dot{x}_2(t)=\tilde{A}x_2(t)+B_d d(t)+B_f f(t),\[2mm] r(t)=Cx_2(t)+D_d d(t)+D_f f(t), \end{array} \right.$24 (Li et al., 2014). In the region-dependent ISS setting, local and non-local small-gain inequalities hold on different intervals and are then “glued” via the overlap condition $\tilde{\Sigma}:\left\{ \begin{array}{l} \dot{x}_2(t)=\tilde{A}x_2(t)+B_d d(t)+B_f f(t),\[2mm] r(t)=Cx_2(t)+D_d d(t)+D_f f(t), \end{array} \right.$25 (Shiromoto et al., 2015). In transitional shear flows, finite gain is guaranteed only below a permissible forcing amplitude (Wei et al., 9 Jun 2025).
The explicit limitation theory is most developed in (Dong et al., 14 Sep 2025). Using left coprime factorization,
$\tilde{\Sigma}:\left\{ \begin{array}{l} \dot{x}_2(t)=\tilde{A}x_2(t)+B_d d(t)+B_f f(t),\[2mm] r(t)=Cx_2(t)+D_d d(t)+D_f f(t), \end{array} \right.$26
the paper defines
$\tilde{\Sigma}:\left\{ \begin{array}{l} \dot{x}_2(t)=\tilde{A}x_2(t)+B_d d(t)+B_f f(t),\[2mm] r(t)=Cx_2(t)+D_d d(t)+D_f f(t), \end{array} \right.$27
and derives the lower bound
$\tilde{\Sigma}:\left\{ \begin{array}{l} \dot{x}_2(t)=\tilde{A}x_2(t)+B_d d(t)+B_f f(t),\[2mm] r(t)=Cx_2(t)+D_d d(t)+D_f f(t), \end{array} \right.$28
It then applies the Poisson integral relation and Blaschke-product factorization to show that non-minimum-phase zeros impose fundamental lower bounds on IIG (Dong et al., 14 Sep 2025). The corresponding bound is
$\tilde{\Sigma}:\left\{ \begin{array}{l} \dot{x}_2(t)=\tilde{A}x_2(t)+B_d d(t)+B_f f(t),\[2mm] r(t)=Cx_2(t)+D_d d(t)+D_f f(t), \end{array} \right.$29
This establishes that IIG cannot in general be made arbitrarily small; it is constrained by transmission-zero geometry (Dong et al., 14 Sep 2025).
7. Comparative view across domains
The following comparison summarizes the principal meanings of IIG-like constructions in the cited literature.
| Domain | IIG or closest object | Main role |
|---|---|---|
| LTI anomaly impact analysis | IIG as maximum energy of undetectable faults for a given disturbance intensity | Robust fault sensitivity under disturbance masking |
| Nonlinear small-gain interconnections | $\tilde{\Sigma}:\left\{ \begin{array}{l} \dot{x}_2(t)=\tilde{A}x_2(t)+B_d d(t)+B_f f(t),\[2mm] r(t)=Cx_2(t)+D_d d(t)+D_f f(t), \end{array} \right.$30, $\tilde{\Sigma}:\left\{ \begin{array}{l} \dot{x}_2(t)=\tilde{A}x_2(t)+B_d d(t)+B_f f(t),\[2mm] r(t)=Cx_2(t)+D_d d(t)+D_f f(t), \end{array} \right.$31, $\tilde{\Sigma}:\left\{ \begin{array}{l} \dot{x}_2(t)=\tilde{A}x_2(t)+B_d d(t)+B_f f(t),\[2mm] r(t)=Cx_2(t)+D_d d(t)+D_f f(t), \end{array} \right.$32 | Componentwise external-input-to-output gain bounds |
| Transitional shear flows | Nonlinear $\tilde{\Sigma}:\left\{ \begin{array}{l} \dot{x}_2(t)=\tilde{A}x_2(t)+B_d d(t)+B_f f(t),\[2mm] r(t)=Cx_2(t)+D_d d(t)+D_f f(t), \end{array} \right.$33 gain $\tilde{\Sigma}:\left\{ \begin{array}{l} \dot{x}_2(t)=\tilde{A}x_2(t)+B_d d(t)+B_f f(t),\[2mm] r(t)=Cx_2(t)+D_d d(t)+D_f f(t), \end{array} \right.$34 with permissible forcing amplitude | Nonlinear disturbance amplification |
| Microwave SQUID amplifiers | Forward scattering gain $\tilde{\Sigma}:\left\{ \begin{array}{l} \dot{x}_2(t)=\tilde{A}x_2(t)+B_d d(t)+B_f f(t),\[2mm] r(t)=Cx_2(t)+D_d d(t)+D_f f(t), \end{array} \right.$35 or $\tilde{\Sigma}:\left\{ \begin{array}{l} \dot{x}_2(t)=\tilde{A}x_2(t)+B_d d(t)+B_f f(t),\[2mm] r(t)=Cx_2(t)+D_d d(t)+D_f f(t), \end{array} \right.$36 | Differential-to-common mode conversion and directionality |
| Feed-forward neural networks | Optimal Input Gain (OIG) | Learned input-space preconditioning |
| Region-dependent ISS interconnections | $\tilde{\Sigma}:\left\{ \begin{array}{l} \dot{x}_2(t)=\tilde{A}x_2(t)+B_d d(t)+B_f f(t),\[2mm] r(t)=Cx_2(t)+D_d d(t)+D_f f(t), \end{array} \right.$37, $\tilde{\Sigma}:\left\{ \begin{array}{l} \dot{x}_2(t)=\tilde{A}x_2(t)+B_d d(t)+B_f f(t),\[2mm] r(t)=Cx_2(t)+D_d d(t)+D_f f(t), \end{array} \right.$38 | Gain composition across interconnected subsystems |
Across these settings, the common structure is a gain relation defined on the input side of a system, interconnection, or learning rule. What changes is the object being bounded: undetectable fault energy, subsystem output response, flow amplification, scattering transmission, or gradient preconditioning. This suggests that “Input-to-Input Gain” functions as a unifying interpretive label only at a high level. At the formal level, the literature remains plural: the term is explicit and optimization-based in (Dong et al., 14 Sep 2025), implicit and componentwise in (Li et al., 2014), region-dependent and ISS-based in (Shiromoto et al., 2015), nonlinear and Lyapunov-certified in (Wei et al., 9 Jun 2025), scattering-theoretic and nonreciprocal in (Kamal et al., 2012), and optimization/preconditioning-based in (Rane et al., 2023).