---
title: Spike Agreement-Dependent Plasticity
url: https://www.emergentmind.com/topics/spike-agreement-dependent-plasticity-sadp-d3360c42-e02e-454e-9b7f-3b378efbe752
type: topic
---

# Spike Agreement-Dependent Plasticity

Searching arXiv for the cited SADP papers to ground the article in current arXiv records.
arxiv_search(query="Spike Agreement Dependent Plasticity", max_results=10)
Spike Agreement Dependent Plasticity (SADP) denotes a class of spike-driven local learning rules in which plastic change is governed by some notion of agreement between neural events rather than by precise ordered spike-pair timing alone. Across the cited arXiv literature, that agreement takes three distinct forms: near-synchronous pre–post spiking independent of order in recurrent STDP analyses, convergence of causally effective presynaptic arrival times in delay plasticity, and chance-corrected agreement between binary spike trains over a finite window using Cohen’s $\kappa$ in explicit SADP formulations. This suggests that SADP is best understood as an umbrella for agreement-driven plasticity mechanisms rather than a single canonical update rule [2501.09296] [2211.08397] [2508.16216] [2601.08526].

## 1. Conceptual scope and terminology

In classical pairwise STDP, synaptic updates depend on the timing difference $\Delta t=t_{\mathrm{post}}-t_{\mathrm{pre}}$, typically through an asymmetric kernel with potentiation for pre-before-post events and depression for post-before-pre events. A canonical form given in the recurrent-network analysis is
$$
K(\Delta t)=A_{+}e^{-\Delta t/\tau_{+}}\Theta(\Delta t)-A_{-}e^{\Delta t/\tau_{-}}\Theta(-\Delta t),
$$
with $\Theta$ the Heaviside function, $\tau_{\pm}$ decay constants, and $A_{\pm}$ amplitudes. SADP departs from this template by rewarding some form of coincidence, agreement, or alignment that is not exhausted by millisecond-scale causal order [2501.09296].

The term is not used uniformly. In the recurrent-assembly work, SADP corresponds to the acausal or symmetric excitatory STDP rule $L^{\mathrm{ac}}(s)$, which is approximately even and positive near $s=0$, with depression for larger $|s|$; the rule therefore rewards spike coincidence independent of order. In the delay-learning work, the paper does not use the term “SADP”; here SADP denotes the principle that plastic changes depend on the agreement among presynaptic spikes contributing to a postsynaptic event. In the 2025 and 2026 SADP papers, the term is explicit and refers to synaptic weight updates driven by population-level agreement metrics such as Cohen’s $\kappa$ rather than pairwise $\Delta t$ comparisons [2501.09296] [2211.08397] [2508.16216] [2601.08526].

| Instantiation | Plastic quantity | Operational notion of agreement |
|---|---|---|
| Acausal/symmetric STDP | Synaptic weight | Near-synchronous pre–post spiking independent of order |
| Activity-dependent delay learning | Propagation delay | Convergence of causal presynaptic arrival times toward a local mean |
| $\kappa$-based SADP | Synaptic weight | Chance-corrected agreement between pre- and post-synaptic spike trains |

A common misconception is that SADP is simply “STDP without timing.” The cited work does not support that reduction. In one line of work, SADP is still a timing-kernel rule, but an approximately even one; in another, it is a delay-retiming mechanism; in a third, it is a discrete-time spike-train agreement rule defined over whole windows rather than spike pairs. The shared principle is agreement dependence, not a single universal mathematical form [2501.09296] [2211.08397] [2508.16216].

## 2. SADP as acausal spike-timing plasticity in recurrent networks

The most mechanistic treatment of SADP-like dynamics in recurrent assemblies is given by the analysis of overlapping neuronal assemblies in recurrently coupled networks of spiking neuron models. There, excitatory synapses follow
$$
L(s)=\kappa L^{\mathrm{c}}(s)+(1-\kappa)L^{\mathrm{ac}}(s),
$$
where $\kappa\in[0,1]$ controls causality, $\kappa=1$ is strictly causal, and $\kappa=0$ is the SADP or acausal case. The causal kernel is antisymmetric:
$$
L^{\mathrm{c}}(s)=
\begin{cases}
f^{\mathrm{c}}_{+}e^{-s/\tau^{\mathrm{c}}_{+}}, & s>0,\\
-f^{\mathrm{c}}_{-}e^{s/\tau^{\mathrm{c}}_{-}}, & s\le 0,
\end{cases}
$$
whereas the agreement-based kernel is approximately even and positive near synchrony:
$$
L^{\mathrm{ac}}(s)=
\begin{cases}
f^{\mathrm{ac}}_{+}e^{-s/\tau^{\mathrm{ac}}_{+}}(T_{+}-s), & s>0,\\
f^{\mathrm{ac}}_{-}e^{s/\tau^{\mathrm{ac}}_{-}}(T_{-}+s), & s\le 0.
\end{cases}
$$
Both are used in a balanced regime with $\int L(s)\,ds\sim O(\epsilon)\|L\|_1$ and $\epsilon=1/N$, to avoid trivial drift driven by chance coincidences [2501.09296].

The slow synaptic dynamics are written as
$$
\frac{dw_{ij}}{dt}=\int_{-\infty}^{\infty}L(\tau)\,[r_i r_j+C_{ij}(\tau)]\,d\tau,
$$
with $C_{ij}(\tau)=\langle (y_i(t)-r_i)(y_j(t+\tau)-r_j)\rangle$. Under weak coupling, cross-covariances are expanded by linear response, and the resulting mean-field ODEs for population-averaged weights $W_{\alpha\beta}$ are organized by connectivity motifs: first-order direct feedforward and backward motifs, and second-order disynaptic motifs including common input and two-link chains. A central structural fact is that common input produces symmetric correlations, so the motif kernels $M^{cc}(s)$ and $(M^{fc}+M^{bc})(s)$ are even functions of $s$ [2501.09296].

That symmetry yields the core SADP result. With strictly causal, approximately odd $L(s)$, integrating against symmetric correlations nearly cancels:
$$
S^{cc}=\int L(s)M^{cc}(s)\,ds\approx 0,\qquad
S^{fb}=\int L(s)(M^{fc}+M^{bc})(s)\,ds\approx 0.
$$
Dynamics are then approximately linear and dominated by first-order forward LTP and backward LTD. By contrast, with SADP or acausal $L(s)$, the same symmetric correlations yield positive second-order contributions,
$$
S^{cc}\gg 0,\qquad S^{fb}\gg 0,
$$
and these terms are quadratic in weights, so overlap-driven common input promotes both within-assembly and cross-assembly growth [2501.09296].

The network is a recurrent E–I network of exponential integrate-and-fire neurons with $N=1000$, $N_E=800$, $N_I=200$, sparse random adjacency with $p_0=0.1$, and synaptic filter $J_{\mathrm{syn}}(t)=(1/\tau_{\mathrm{syn}})e^{-t/\tau_{\mathrm{syn}}}H(t)$ with $\tau_{\mathrm{syn}}=5~\mathrm{ms}$. Training uses two independent white-noise signals $s_A(t)$ and $s_B(t)$ delivered to excitatory subsets defining assemblies $A$ and $B$, while an overlap population $O$ receives both; the overlap ratio is $\lambda=N_O/N$, varying from $0$ to $\gamma_E=0.8$ [2501.09296].

For $\kappa=1$, $S^{cc}$ and $S^{fb}$ are negligible, $\lambda$ enters only through quadratic terms, and simulations show that for both $\lambda=0.1$ and $\lambda=0.6$, $W_{AA}$ continues to grow after training while $W_{AB}$ decays. Assemblies remain segregated, and trial-averaged simulations agree with the 5D mean-field prediction. For $\kappa=0$, the reduced dynamics of $W_{AB}$ take the form
$$
\frac{dW_{AB}}{dt}=k_0+k_1W_{AB}+k_2W_{AB}^2,
$$
with $k_2>0$, yielding an unstable threshold $W_{AB}^t$ that decreases with increasing overlap $\lambda$. Small overlap can still permit segregation, but large overlap drives growth of $W_{AB}$ and fusion of assemblies. The reported fusion boundary in $(\kappa,\lambda)$ space therefore separates sufficiently causal rules, which prevent fusion even at high overlap, from sufficiently SADP-like rules, which fuse beyond a critical $\lambda$ [2501.09296].

## 3. SADP as activity-dependent delay alignment

A second usage of SADP arises in local delay learning, where the plastic variable is not synaptic strength but axonal propagation delay. In this formulation, the relevant quantity is the arrival time
$$
a_{ij}=t_i+d_{ij},
$$
for presynaptic neuron $i$ and postsynaptic neuron $j$. A presynaptic spike is eligible only if it arrives before the postsynaptic spike and within a causal window,
$$
\Delta t_{\mathrm{lag}(i\!\to\! j)}\equiv t_j-(t_i+d_{ij}),\qquad 0\le \Delta t_{\mathrm{lag}}<10~\mathrm{ms}.
$$
For each postsynaptic spike at time $t_j$, the average arrival time over the causal set $S_j$ is
$$
\bar a_j=\frac{1}{|S_j|}\sum_{i\in S_j}(t_i+d_{ij}),
$$
and the alignment error is
$$
E_j=\sum_{i\in S_j}\left[(t_i+d_{ij})-\bar a_j\right]^2.
$$
SADP, in this interpretation, pushes arrivals toward $\bar a_j$ so that causally effective inputs become more temporally coherent [2211.08397].

The local update rule is
$$
\Delta d_{ij}=-\,3\,\tanh\!\left(\frac{(t_i+d_{ij})-\bar a_j}{3}\right),
\quad \text{if } 0\le t_j-(t_i+d_{ij})<10~\mathrm{ms},
$$
and $\Delta d_{ij}=0$ otherwise. Earlier-than-average arrivals receive positive delay updates and are slowed; later-than-average arrivals receive negative updates and are sped up. The tanh saturates step magnitude at $\pm 3~\mathrm{ms}$ per update. The paper presents the rule heuristically, but the provided interpretation is that it acts like a bounded step toward gradient descent on $E_j$ while maintaining locality and bounded updates [2211.08397].

The neuron model is Izhikevich regular spiking, with
$$
\frac{dv}{dt}=0.04v^2+5v+140-u+I(t),\qquad
\frac{du}{dt}=a(bv-u),
$$
and spike reset $v\ge 30~\mathrm{mV}\Rightarrow v\leftarrow c,\;u\leftarrow u+d$, using the typical regular-spiking parameters $a=0.02$, $b=0.2$, $c=-65~\mathrm{mV}$, $d=8$. Inputs are latency-coded over a fixed window $T_{\max}=40~\mathrm{ms}$ according to
$$
t_i=T_{\max}x_i,
$$
with normalized input channel values $x_i\in[0,1]$. A three-layer feedforward SNN with 100 neurons per layer, layer-to-layer connection probability $0.1$, fixed homogeneous weight magnitude $6$, and delays initialized uniformly in $(0,40)~\mathrm{ms}$ is trained on downscaled $10\times 10$ MNIST digits [2211.08397].

Outputs are decoded as polychronous group patterns (PGPs). The similarity between two PGPs $A$ and $B$ is
$$
s(A,B)=\frac{M(A,B)}{\frac{|A|+|B|}{2}},
$$
where $M(A,B)$ counts spikes that match in identity and order within a tolerance. Hierarchical clustering uses thresholds of $80\%$ or $90\%$, and classification assigns the most common cluster label. Training presents 20 instances per digit class, with digits 0 and 1 used for training and digit 2 used only during testing for unseen-class generalization. Accuracy improved after delay learning in nearly all cases where classes were separable. Some networks failed to separate classes: $2.4\%$ at $90\%$ threshold and $45\%$ at $80\%$ threshold. Generalization to the unseen class was observed at the $80\%$ threshold, with max accuracy $64\%$ and mean $32\%$ across 38 networks that could separate the unseen class [2211.08397].

This formulation broadens the SADP concept. Agreement is not coincidence between a presynaptic and a postsynaptic spike train as whole objects, but reduced dispersion of causally effective presynaptic arrivals. The stated goal is to stabilize reproducible time-locked output patterns and increase separability through polychronization rather than through weight modulation alone [2211.08397].

## 4. SADP as chance-corrected spike-train agreement

The explicit formalization of SADP as a synaptic learning paradigm appears in the 2025 work that defines it as a biologically inspired learning rule for SNNs that replaces pairwise spike-timing updates with population-level agreement between pre- and post-synaptic spike trains. For batch index $b$, presynaptic neuron $i$, postsynaptic neuron $j$, and discrete time bins $t=1,\dots,T$, binary spike trains are
$$
\mathbf{X}_{b,i,:}\in\{0,1\}^T,\qquad \mathbf{S}_{b,j,:}\in\{0,1\}^T.
$$
The SNN uses leaky integrate-and-fire units:
$$
\mathbf{V}_t=\lambda \mathbf{V}_{t-1}+\mathbf{I}_t,\qquad \mathbf{I}_t=\mathbf{X}_t\mathbf{W},
$$
followed by normalization,
$$
\tilde{\mathbf{V}}_t=\frac{\mathbf{V}_t-\min(\mathbf{V}_t)}{\max(\mathbf{V}_t)-\min(\mathbf{V}_t)+\epsilon},
$$
and thresholding with reset,
$$
\mathbf{S}_t=H(\tilde{\mathbf{V}}_t-\theta),\qquad
\mathbf{V}_t\leftarrow \mathbf{V}_t\cdot (1-\mathbf{S}_t).
$$
Agreement is measured by Cohen’s $\kappa$ rather than $\Delta t$ [2508.16216].

Using contingency counts $n_{11}$, $n_{10}$, $n_{01}$, and $n_{00}$ over $T$ bins, observed and expected agreement are
$$
p_o=\frac{n_{11}+n_{00}}{T},
$$
$$
p_e=
\left(\frac{n_{11}+n_{10}}{T}\right)\left(\frac{n_{11}+n_{01}}{T}\right)+
\left(\frac{n_{01}+n_{00}}{T}\right)\left(\frac{n_{10}+n_{00}}{T}\right),
$$
and
$$
\kappa_{ij}^{(b)}=\frac{p_o^{(b)}-p_e^{(b)}}{\max(1-p_e^{(b)},\epsilon)}.
$$
The weight update is
$$
\Delta w_{ij}^{(t)}=\frac{\eta_t}{B}\sum_{b=1}^B\mathcal{L}\!\left(\kappa_{ij}^{(b)}\right),
$$
$$
w_{ij}^{(t+1)}=
\operatorname{clip}\!\left(
\operatorname{sign}\!\big(w_{ij}^{(t)}+\Delta w_{ij}^{(t)}\big)\cdot
\max\!\left(\left|w_{ij}^{(t)}+\Delta w_{ij}^{(t)}\right|,\epsilon\right),
-1,1
\right).
$$
The learning function $\mathcal{L}$ is either piecewise linear,
$$
\mathcal{L}_{\mathrm{linear}}(\kappa)=
\begin{cases}
\alpha_+\kappa, & \kappa\ge 0,\\
\alpha_-\kappa, & \kappa<0,
\end{cases}
$$
or spline-based, using device-calibrated or ideal reference curves [2508.16216].

A defining feature of this SADP is computational complexity. Per-synapse agreement is computed in one pass over $T$ bins, giving
$$
\mathcal{O}(N_{\mathrm{pre}}N_{\mathrm{post}}T),
$$
whereas classical pairwise STDP is given as
$$
\mathcal{O}(N_{\mathrm{pre}}N_{\mathrm{post}}S^2),
$$
with $S$ the number of spikes. The algorithm admits efficient implementation using bitwise AND, NOT, XOR, and POPCOUNT. The paper explicitly connects this to hardware efficiency and to iontronic organic memtransistor data, where conductance modulation under $1000$ excitatory write pulses $(-3.0~\mathrm{V}, 50~\mathrm{ms})$ and inhibitory pulses $(+1.0~\mathrm{V}, 50~\mathrm{ms})$ with reads $(-0.5~\mathrm{V}, 50~\mathrm{ms})$ yields gradual, nearly linear potentiation and depression. Device-measured conductance changes $\Delta G/G_0(t)$ are fitted with smoothing splines to derive the SADP kernels $f_+(\delta)$ and $f_-(\delta)$ [2508.16216].

The experiments use MNIST and Fashion-MNIST with either rate coding over 10 time steps or time-to-first-spike coding, and either a 784-to-400 fully connected LIF layer (“1layer”) or a 784-to-64 version (“1layer_small”). Unsupervised SADP is trained for 10 epochs with batch size 64, and a downstream classifier is then trained for 50 supervised epochs on extracted features. Reported MNIST results include: linear $\mid$ rate $\mid$ 1layer, accuracy $0.9068$, F1 $0.9055$, runtime/epoch $720.89~\mathrm{s}$; spline\_ideal $\mid$ rate $\mid$ 1layer, accuracy $0.9106$, F1 $0.9094$, runtime/epoch $758.83~\mathrm{s}$; STDP $\mid$ rate $\mid$ 1layer, accuracy $0.1294$, F1 $0.0518$, runtime/epoch $2611.38~\mathrm{s}$; and Hebbian $\mid$ rate $\mid$ 1layer, accuracy $0.1135$, F1 $0.0204$, runtime/epoch $405.93~\mathrm{s}$. On Fashion-MNIST, spline\_ideal $\mid$ rate $\mid$ 1layer reaches accuracy $0.7873$, F1 $0.7855$, runtime/epoch $764.52~\mathrm{s}$, while STDP $\mid$ rate $\mid$ 1layer gives accuracy $0.1069$, F1 $0.0320$, runtime/epoch $2525.75~\mathrm{s}$ [2508.16216].

The reported ablations state that rate coding outperforms TTFS for linear and device-derived kernels, whereas TTFS is handled well by the ideal spline kernel. The ideal spline kernel provides the best or near-best performance and robustness across encodings, linear SADP is competitive under rate coding but collapses under TTFS, and device-derived spline kernels underperform ideal and linear counterparts, especially under TTFS. The paper’s interpretation is that smoother, bounded kernels improve generalization [2508.16216].

## 5. Supervised SADP and hybrid CNN–SNN architectures

The 2026 supervised extension preserves the agreement-driven hidden-layer mechanism while adding a strictly local output-layer error signal. Inputs are either raw normalized features or frozen CNN embeddings $\mathbf{x}^{(b)}\in[0,1]^{N_{\mathrm{in}}}$, converted to spikes by
$$
S_{\mathrm{in}}^{(b)}(t,i)\sim \mathrm{Bernoulli}(x_i^{(b)}).
$$
LIF dynamics are
$$
\mathbf{V}_l^{(b)}(t)=\lambda \mathbf{V}_l^{(b)}(t-1)+\mathbf{I}_l^{(b)}(t),
$$
$$
\mathbf{s}_l^{(b)}(t)=\mathbb{I}\!\left(\mathbf{V}_l^{(b)}(t)>\boldsymbol{\theta}_l\right),\qquad
\mathbf{V}_l^{(b)}(t)\leftarrow \mathbf{V}_l^{(b)}(t)\bigl(1-\mathbf{s}_l^{(b)}(t)\bigr),
$$
and output decoding uses spike counts
$$
\mathbf{C}^{(b)}=\sum_{t=1}^T \mathbf{s}_{\mathrm{out}}^{(b)}(t),\qquad
\hat y^{(b)}=\arg\max_k C_k^{(b)}.
$$
The architectures are feedforward 1SADP and 2SADP networks, optionally preceded by a frozen CNN encoder comprising three convolutional layers with ReLU, max-pooling, global average pooling, and a dense projection to $N_{\mathrm{in}}=256$ features [2601.08526].

At the output layer, supervision is local:
$$
\mathbf{E}^{(b)}(t)=\mathbf{y}^{(b)}-\mathbf{s}_{\mathrm{out}}^{(b)}(t),
$$
and the mini-batch update is
$$
\Delta W_2=
\frac{1}{B}\sum_{b=1}^B\sum_{t=1}^T
\big(\mathbf{s}_{\mathrm{pre}}^{(b)}(t)\big)^\top \mathbf{E}^{(b)}(t).
$$
Weights are updated with learning rate $\eta_{\mathrm{out}}$, decay $\rho$, and clipping:
$$
W_2\leftarrow \mathrm{clip}\!\left(\rho W_2+\eta_{\mathrm{out}}\Delta W_2,\;[-W_{\max},W_{\max}]\right).
$$
For hidden neurons, agreement is computed against the correct-class output spike train
$$
s_o^{*(b)}(t)=s_{\mathrm{out},\,y^{(b)}}^{(b)}(t).
$$
Observed agreement is
$$
P_{o,j}^{(b)}=\frac{1}{T}\sum_{t=1}^{T}\mathbb{I}\!\left(s_{h,j}^{(b)}(t)=s_o^{*(b)}(t)\right),
$$
chance agreement is
$$
P_{e,j}^{(b)}=
p_{h,j}^{(b)}p_o^{(b)}+
\bigl(1-p_{h,j}^{(b)}\bigr)\bigl(1-p_o^{(b)}\bigr),
$$
and
$$
\kappa_j^{(b)}=
\frac{P_{o,j}^{(b)}-P_{e,j}^{(b)}}{1-P_{e,j}^{(b)}+\varepsilon}.
$$
Input-to-hidden updates are then
$$
\Delta W_1=
\frac{1}{B}\sum_{b=1}^B
\big(\mathbf{x}^{(b)}\big)^\top \boldsymbol{\kappa}^{(b)},
$$
followed by decay and column-wise normalization:
$$
W_1\leftarrow \rho W_1+\eta_{\mathrm{in}}\Delta W_1,\qquad
W_{1,:,j}\leftarrow \frac{W_{1,:,j}}{\|W_{1,:,j}\|_2+\varepsilon}.
$$
The paper emphasizes that this requires no backpropagation, surrogate gradients, or teacher forcing [2601.08526].

Training uses batch size $B=128$, temporal resolutions $T=25$ and $T=100$, learning rates $\eta_{\mathrm{out}}=5\times 10^{-4}$ and $\eta_{\mathrm{in}}=2\times 10^{-4}$, decay $\rho=0.9995$, and 50 epochs in the main benchmarks. In the Poisson-only setting, reported results are: MNIST, 1SADP accuracy $0.8792$, F1 $0.8759$, time/epoch approximately $109$–$110~\mathrm{s}$; Fashion-MNIST, 1SADP accuracy $0.7585$, F1 $0.7422$, time/epoch approximately $81~\mathrm{s}$; CIFAR-10, 1SADP accuracy $0.2359$, F1 $0.1741$, time/epoch approximately $240~\mathrm{s}$. With CNN+Poisson encoding, the reported maxima are: MNIST up to $0.9916$ accuracy with 1SADP at $T=100$; Fashion-MNIST up to $0.8995$ accuracy with 2SADP at $T=100$; CIFAR-10 up to $0.7069$ accuracy with 1SADP at $T=100$. The same framework is also reported on biomedical datasets, reaching up to $0.9985$ on Colon Histopathology, $0.9843$ on Lung Histopathology, and $0.9887$ on Brain Tumor MRI under CNN+Poisson settings [2601.08526].

The paper also reports device-inspired synaptic dynamics derived from iontronic memtransistors. Under those kernels, MNIST accuracy is $0.8238$ with F1 $0.8192$ and time/epoch around $207~\mathrm{s}$, while Fashion-MNIST accuracy is $0.7025$ with F1 $0.6721$ and time/epoch around $207$–$226~\mathrm{s}$. These results are presented as evidence of compatibility with asymmetric, bounded device dynamics, though with some loss relative to idealized updates [2601.08526].

## 6. Limitations, misconceptions, and design implications

Several limitations recur across the SADP literature. In the recurrent-network theory, the analysis relies on weak coupling, $\epsilon=1/N$ scaling, linear response, stationary correlations, and truncation of the Neumann expansion at second-order motifs; strong coupling, nonstationarity, higher-order motifs, multiplicative STDP, triplet rules, or calcium-based models may alter thresholds and weight distributions. The same work also assumes $\int L(s)\,ds\approx 0$ to avoid trivial drift and uses homeostatic inhibitory STDP to keep excitatory firing rates near a target, such as $12~\mathrm{Hz}$, so that timing kernels rather than rate confounds dominate structure formation [2501.09296].

In the delay-learning formulation, the assumptions include fixed homogeneous weights, a single Izhikevich regular-spiking neuron type, integer delays initialized in $[0,40]~\mathrm{ms}$, and a fixed causal window of $10~\mathrm{ms}$. The paper notes that non-separable classes and overtraining can produce homogeneous PGPs that collapse class distinctions, and that an adaptive stopping criterion is needed. Strict PGP thresholds of $90\%$ reduce generalization to unseen classes, whereas relaxed thresholds of $80\%$ improve transfer but increase the risk of merging distinct patterns [2211.08397].

In the explicit $\kappa$-based SADP papers, very sparse spike trains challenge simple linear kernels, highly imbalanced firing rates can bias $\kappa$, and deeper networks may require additional mechanisms for task-aligned feature learning. The supervised variant further depends on a frozen CNN front-end for the strongest results on complex image datasets; strictly spike-native end-to-end training is not addressed there. The 2SADP architecture yields only modest and dataset-dependent gains over 1SADP, and the current experiments focus on feedforward image classification rather than recurrent temporal tasks [2508.16216] [2601.08526].

A second misconception is that “agreement” is always beneficial. The recurrent-assembly analysis demonstrates the opposite for overlapping representations: agreement-based, acausal timing amplifies symmetric correlations produced by shared common input, promotes quadratic growth of both within- and cross-assembly weights, and fuses assemblies beyond a critical overlap. Strictly causal, antisymmetric timing suppresses those symmetric contributions by odd–even cancellation and preserves segregation. This establishes a concrete distinction between agreement as an objective for temporal coherence and agreement as a potential source of representational collapse [2501.09296].

The practical design guidance is therefore conditional rather than universal. For distributed representation without assembly fusion, the recurrent analysis advises strong antisymmetry and near-zero kernel integral, forward LTP only for pre-before-post, LTD for post-before-pre, and avoidance of broad symmetric LTP near $\Delta t=0$. For scalable hardware-ready learning, the explicit SADP papers instead emphasize bounded kernels, clipping to fixed weight ranges, $\epsilon$-floors to prevent synaptic silencing, bit-packing and POPCOUNT for efficient $\kappa$ computation, and smooth spline kernels for sparse codes such as TTFS. For delay plasticity, the design principle is to align causally effective arrivals while preserving locality through synapse-level access only to pre/post spike times and postsynaptic-local averages [2501.09296] [2508.16216] [2211.08397].

Taken together, the cited work indicates that SADP is not a single replacement for STDP but a family of agreement-dependent local rules with sharply different consequences depending on what is made to agree: spike pairs, arrival times, or whole spike trains. In recurrent assemblies, agreement-driven acausal timing can destroy functional specificity; in feedforward temporal coding, delay alignment can consolidate polychronous groups; in window-based synaptic learning, chance-corrected agreement can yield linear-time, hardware-aligned training; and in supervised hybrids, the same agreement principle can be combined with strictly local output errors for fast learning without backpropagation or teacher forcing [2501.09296] [2211.08397] [2508.16216] [2601.08526].

Source: https://www.emergentmind.com/topics/spike-agreement-dependent-plasticity-sadp-d3360c42-e02e-454e-9b7f-3b378efbe752