- The paper proves that Equilibrium Propagation, Coupled Learning, and Adjoint Coupled Learning converge locally and exponentially near solutions when an active-incidence rank condition ensures coercivity, with guarantees for both continuous and discrete updates.
- The analysis shows that Equilibrium Propagation and Adjoint Coupled Learning are gradient flows, while Coupled Learning has a controlled cubic correction that remains convergent sufficiently close to the solution manifold.
- The paper identifies symmetry-driven failures in a kite circuit, including spurious fixed points, but uses Sard’s theorem to show that coercivity holds on the entire solution manifold for almost every target output.
Overview
This paper provides the first local convergence analysis of three physical learning methods—Equilibrium Propagation (EP), Coupled Learning (CL), and a newly introduced method, Adjoint Coupled Learning (AL)—for linear resistor networks. Physical learning trains a physical system to perform a computational task using only local update rules at each network element, with the physics itself performing the global transfer of information that backpropagation would otherwise require of a central processor (2606.15443). The paper's central contribution is a rigorous characterization of when the training loss decays exponentially near a solution: a coercivity condition, expressible as a rank condition on a matrix built from the circuit's active-incidence structure and the constrained Laplacian.
The analysis covers both the continuous-time limit (small nudge η, small learning rate τ) and the discrete update rule at finite learning rate and finite nudge. The main results establish that if the initial parameters lie sufficiently close to a solution—a parameter configuration achieving the desired output—and coercivity holds, then the loss decays exponentially and the parameters converge to the solution manifold. The paper also demonstrates that coercivity can fail, via an explicit "kite" circuit exhibiting a symmetry-induced degeneracy, but proves using Sard's theorem that such failures are non-generic: for almost every choice of desired output, coercivity holds at every point of the solution manifold.
Framework and the three learning rules
The setting is a smooth energy G(x;k), convex in the state x∈RN for each admissible parameter k∈RM. Inputs v are imposed through constraints P⊤x=v; the free state x0(k) minimizes G under these constraints. The state error is r:=Q⊤x0−w, where τ0 selects output nodes and τ1 is the target. The three methods differ in how they construct a nudged state τ2:
- EP adds a force proportional to the residual to the energy: it minimizes τ3 subject to the input constraints.
- CL clamps the output partway toward the target: it minimizes τ4 subject to both input constraints and τ5. Clamping is physically simpler than forcing, requiring no active source.
- AL, introduced here, solves a fully constrained reference problem with τ6 enforced, extracts the Lagrange multiplier τ7 (the "adjoint error," interpretable as a current), and then nudges by clamping the output to τ8.
All methods share the contrastive update τ9, whose small-G(x;k)0 limit yields the fundamental evolution equation G(x;k)1.
Gradient flow structure
A key structural result distinguishes the three methods. EP performs exact gradient descent on the natural loss G(x;k)2: differentiating the KKT system gives G(x;k)3, so the loss strictly decreases away from fixed points. AL likewise performs gradient flow, but on the adjoint loss G(x;k)4, where G(x;k)5 is the Lagrange multiplier enforcing the target constraint at the reference state; the same identity G(x;k)6 holds.
CL does not perform gradient flow of any apparent loss. Its dynamics follow those of a weighted loss G(x;k)7, where G(x;k)8 is a positive-definite matrix identified as a discrete Dirichlet-to-Neumann map converting voltage boundary conditions into equivalent currents. Because G(x;k)9 depends on x∈RN0, differentiating along the CL flow produces an additional cubic correction term x∈RN1—cubic in the residual x∈RN2—on top of the dissipative gradient term. For sufficiently small x∈RN3 the dissipative term dominates, which is what makes local convergence possible despite CL not being gradient flow.
The paper distills a general design principle from this comparison: gradient flow is recovered precisely when the error and the nudge are of dual type across the voltage–current duality. EP crosses the duality (voltage error applied as current); CL matches types (voltage error, voltage nudge) and deviates; AL matches types but uses the adjoint (current) error with a voltage-type constraint nudge, restoring gradient flow. The authors use this principle to predict a fourth method with current inputs, outputs, and errors, nudged by voltage, that should also be gradient flow.
Coercivity estimates
Because x∈RN4 generically exceeds the output dimension x∈RN5, solutions form a manifold x∈RN6 of dimension x∈RN7. The analysis asks when the dynamics are coercive near regular points of x∈RN8, i.e., points where x∈RN9 has full row rank. Specializing to linear circuits with trainable conductances, where k∈RM0 and k∈RM1 is the graph Laplacian, the paper derives a necessary and sufficient condition: coercivity holds if and only if the matrix k∈RM2 has full column rank, where k∈RM3 is the active-incidence matrix built from edges carrying nonzero voltage drop in the free state, k∈RM4 is the input-constrained Laplacian, and k∈RM5 encodes the output nodes.
Under this condition, EP satisfies k∈RM6 locally, with k∈RM7 the smallest nonzero squared voltage drop among active edges and k∈RM8 the smallest singular value above—hence exponential decay. A simple sufficient condition is that every edge carries a nonzero free-state voltage drop, since the full incidence matrix of a connected graph has full rank. For CL, the same rank condition applies (the Dirichlet-to-Neumann map only rescales the forcing without changing its support), yielding k∈RM9, where the cubic term arises from the v0-dependence of v1 and vanishes on the solution manifold. AL admits the identical coercivity argument with the active edge set defined by the fully constrained reference state rather than the free state.
Failure of coercivity: the kite circuit
For single-output circuits (v2), the coercivity condition reduces to a transparent physical statement: at least one edge carrying current in the free state must also carry a nonzero voltage drop in the clamped state. Coercivity fails exactly when the free and clamped voltage drops have disjoint support. In this case the loss decay rate factors explicitly as v3, depending only on conductances.
The kite circuit realizes this failure. With a symmetry condition v4 forcing equal voltages at two symmetric nodes, and v5, the free and clamped states become invisible to each other on all active edges. Two regimes result:
- On the solution manifold (v6): the loss is zero but the exponential convergence guarantee degenerates; convergence is slow near these codimension-2 points.
- Off the manifold (v7): the update rule vanishes identically at nonzero loss, producing genuine spurious fixed points of the dynamics. Notably, these spurious fixed points are shared by EP and CL, since rescaling by v8 cannot change the support of the forcing.
AL behaves differently off the manifold: clamping the output to v9 forces voltage drops across the middle branch, so the degenerate points are no longer fixed points. Instead, the dynamics are confined to the invariant codimension-2 manifold and drive P⊤x=v0 toward zero—a boundary minimum of the current-error loss rather than a contradiction with AL's gradient-flow structure.
Genericity via Sard's theorem
The degeneracies above depend on special choices of the target. Since the output map P⊤x=v1 is real-analytic (rational) in P⊤x=v2, Sard's theorem implies that its set of critical values has Lebesgue measure zero. Consequently, for almost every target P⊤x=v3, the solution manifold is a smooth P⊤x=v4-dimensional submanifold and the coercivity condition holds at every point of it. This result is non-vacuous whenever the circuit admits at least one coercive configuration, since the image of P⊤x=v5 then contains an open set while the critical values have measure zero.
Numerical experiments on the kite circuit corroborate this picture. On a two-parameter slice measuring deviation from the two symmetry conditions, the coercivity constant degenerates only at the origin, with roughly elliptical level sets elongated in the direction of the ratio-condition parameter—the circuit is more sensitive to breaking P⊤x=v6 than to breaking the ratio condition. Semi-log plots confirm straight-line (exponential) loss decay for EP at every nonzero perturbation, with decay rates tracking the coercivity constant; EP, CL, and AL all inherit the same coercivity-controlled exponential decay, with rates rescaled by factors involving P⊤x=v7.
Continuous-time local convergence
The main theorem combines the Implicit Function Theorem with coercivity in a bootstrap argument. If P⊤x=v8 lies on the solution manifold with P⊤x=v9 of full row rank, then there exists a neighborhood such that all trajectories starting within it remain inside, satisfy x0(k)0 (with x0(k)1, x0(k)2 for EP; x0(k)3, x0(k)4 for CL), converge to a point of x0(k)5, and do so with Cauchy convergence of the parameter trajectory. The proof controls total displacement via x0(k)6 together with the exponential decay, ensuring the trajectory never exits the region where the local estimates hold.
An important corollary concerns spurious fixed points: because EP and AL are gradient flows, any fixed point with nonzero loss requires x0(k)7 to be rank-deficient there. Thus EP can get stuck only at non-coercive configurations like the kite's degenerate set—which is codimension-2, so generic trajectories should miss it.
Discrete-time convergence
At finite learning rate and finite nudge, the linear structure of the circuit yields an exact edgewise identity separating the gradient step from an x0(k)8 correction:
x0(k)9
Combining the descent lemma with the Polyak–Łojasiewicz inequality supplied by coercivity, the paper proves geometric decay G0 with G1, provided G2, the initial iterate is close to G3, and the initial loss satisfies a smallness bound scaling like G4.
The CL case differs substantively: its non-gradient correction survives the G5 limit, being G6 independent of G7. The corresponding smallness hypothesis involves G8 instead of G9, meaning that starting near the solution manifold—not merely shrinking the nudge—is required to control CL's non-gradient part. This is the discrete analogue of the cubic r:=Q⊤x0−w0 remainder in the continuous theory. AL inherits the discrete EP theorem verbatim with constants computed against the reference-state active edge set.
Limitations and open questions
Several restrictions are stated plainly. All quantitative results are proved for linear circuits with trainable conductances; the authors expect linearity may be dropped when r:=Q⊤x0−w1 is sufficiently smooth, but this is conjectural and unproven. The convergence guarantees are local: the basin must exclude the spurious fixed points that exist off the solution manifold at critical targets, and shifting the target r:=Q⊤x0−w2 alone may not suffice globally if initialization lands near such a point. The genericity result ensures coercivity on the entire solution manifold for almost every r:=Q⊤x0−w3, but says nothing about the size of the convergence ball or about global behavior. Whether the predicted fourth method (current inputs, outputs, and errors, with voltage nudging) admits the same analysis is asserted by symmetry arguments but not carried out. Finally, the algorithms are studied as idealized mathematical procedures; their implementation in stochastic or noisy physical substrates, and extension beyond quadratic energies, remain outside the scope of the proofs.
Conclusion
This paper supplies the missing rigorous foundation for local convergence of contrastive physical learning in linear circuits. It establishes that EP and AL are true gradient flows (of the state-error and adjoint-error losses respectively), that CL follows modified dynamics with a cubic correction controlled near solutions, and that a concrete rank condition on the active-incidence structure governs exponential loss decay in both continuous and discrete time. The kite counterexample shows coercivity can fail through symmetry, producing both slow convergence on the solution manifold and genuine spurious fixed points off it, while the Sard-theorem genericity result shows such failures are measure-zero events in the target. The work converts the empirical success of physical learning devices into provable guarantees under explicit, physically interpretable conditions, and identifies the voltage–current duality between error and nudge as the structural feature determining whether a contrastive rule is gradient descent.