---
title: Gated Siamese CNN Architectures
url: https://www.emergentmind.com/topics/gated-siamese-cnn
type: topic
---

# Gated Siamese CNN Architectures

Searching arXiv for the cited works and closely related formulations of gated Siamese CNNs.
arxiv_search(query="1607.08378 Gated Siamese Convolutional Neural Network Architecture for Human Re-Identification", max_results=5)
A gated Siamese CNN is a Siamese architecture in which the interaction between paired inputs is modulated by an explicit gate rather than being deferred entirely to a terminal distance or correlation layer. In the literature, the term has been used for several related but non-identical constructions: a learnable Matching Gate that conditions image representations on the counterpart image in human re-identification [1607.08378], a Siamese CNN–ConvGRU model with spatial and temporal attention for video-based person re-identification [1808.01911], a multi-branch Siamese tracker with hard online branch selection [1808.07349], and a Siamese adversarial network for respiratory-gated PET with gate-to-gate consistency learning [2009.06757]. Across these formulations, the common principle is pair-conditioned control of information flow, but the locus of gating differs substantially: mid-level feature channels, recurrent state transitions, branch selection, or acquisition-phase coupling.

## 1. Conceptual scope and major formulations

The Siamese pattern is consistent across these works: two inputs are processed by weight-shared or identical branches, and a similarity, response map, or paired objective is then computed. What varies is how the gate intervenes. In the person re-identification formulation, the gate compares row-aligned mid-level features across the pair and amplifies common local patterns before the final embedding is formed [1607.08378]. In video re-identification, gating is embedded in ConvGRU dynamics and coupled to deterministic spatial attention and temporal pooling so that the network learns which parts and which frames are informative [1808.01911]. In tracking, gating is a hard argmax over branch-wise response-map heuristics, selecting one Siamese branch every \(T\) frames [1808.07349]. In low-dose PET, the Siamese generator processes reference and target respiratory gates jointly, while a pre-trained motion estimator imposes gate-to-gate consistency [2009.06757].

| Work | Siamese structure | Gate mechanism |
|---|---|---|
| Human re-identification [1607.08378] | Two CNN branches for paired pedestrian images | Matching Gate on mid-level row-aligned features |
| Video person re-identification [1808.01911] | Two CNN–ConvGRU–attention streams for paired videos | GRU update/reset gates plus spatial and temporal attention |
| Object tracking [1808.07349] | Multiple Siamese branches over exemplar and search region | Hard online branch selection by response-map score |
| Low-dose gated PET [2009.06757] | Two weight-shared generator branches for reference and target gates | Respiratory gate pairing plus gate-to-gate consistency through frozen registration |

This taxonomy indicates that “gated” is not a single architectural primitive. A plausible implication is that the term is best understood functionally: it denotes selective modulation conditioned on paired evidence, but the mathematical implementation can be residual, recurrent, heuristic, or consistency-based.

## 2. Mid-level matching gates in image-based human re-identification

The clearest canonical use of the phrase appears in “Gated Siamese Convolutional Neural Network Architecture for Human Re-Identification” [1607.08378]. The baseline is a Siamese CNN that resizes person crops to \(128 \times 64\), subtracts the mean training image, and uses per-branch convolutional blocks of Convolution \(\rightarrow\) Batch Normalization \(\rightarrow\) PReLU, followed by three max-pooling layers. Asymmetric convolutions reduce columns while preserving rows, and a final \(16 \times 1\) convolution produces a \(1 \times 1 \times 150\) embedding vector. Euclidean distance in the final 150-D embedding space is used for scoring and ranking.

The proposed Matching Gate (MG) is inserted between layers 4–5, 5–6, and 6–7; on VIPeR, between layers 4–5 and 5–6. Its purpose is to make the representation of each image conditional on the image it is paired with. After resizing, a horizontal row-wise correspondence is assumed. For each row \(r\), the gate summarizes stripe features in both images, computes per-channel similarity, and then boosts the original stripe activations in both branches.

The gate operates in three steps. First, feature summarization along the stripe uses learnable filters \( \mathbf{w} \) and PReLU \(f(\cdot)\):
$$
\mathbf{y}_{r1} = f(\mathbf{w} \ast \mathbf{x}_{r1}), \quad \mathbf{y}_{r2} = f(\mathbf{w} \ast \mathbf{x}_{r2}).
$$
Second, for each channel \(i\), a Gaussian gate is computed:
$$
\mathbf{g}_r^{\,i} = \exp\left(-\frac{(\mathbf{y}_{r1}^{\,i} - \mathbf{y}_{r2}^{\,i})^2}{\mathbf{p}_i^2}\right),
$$
where \(p_i\) is a learnable variance parameter. Third, the scalar gate is repeated horizontally and used for residual boosting:
$$
\mathbf{a}_{r1}^{\,i} = \mathbf{x}_{r1}^{\,i} + \mathbf{x}_{r1}^{\,i} \odot \mathbf{G}_r^{\,i}, \quad
\mathbf{a}_{r2}^{\,i} = \mathbf{x}_{r2}^{\,i} + \mathbf{x}_{r2}^{\,i} \odot \mathbf{G}_r^{\,i}.
$$
After assembling all channels, L2 normalization across channels is applied before the next convolutional block. The design therefore does not merely suppress features; it amplifies local evidence that appears mutually consistent across the pair.

Training is end-to-end from scratch with contrastive loss and margin \(m=1\), with labels \(z=0\) for positive pairs and \(z=1\) for negative pairs:
$$
\mathcal{L} = (1 - z)\, d^2 + z\, \max(0,\, m - d)^2.
$$
Optimization uses RMSProp with decay parameter \(0.95\), initial learning rate \(0.002\), decay by a factor of \(0.9\) after each epoch, batch size \(100\), and early stopping based on validation saturation. The gate parameters are initialized with \(p_i=4\) for all channels.

Empirically, the gate improves the baseline Siamese CNN on multiple benchmarks. On Market-1501 single-query, Rank-1 increases from \(62.32\) to \(65.88\) and mAP from \(36.23\) to \(39.55\); on multi-query, Rank-1 increases from \(72.92\) to \(76.04\) and mAP from \(45.39\) to \(48.45\). On CUHK03 detected, the gated model improves single-query Rank-1 from \(58.1\) to \(61.8\) and mAP from \(48.90\) to \(51.25\), and multi-query Rank-1 from \(63.9\) to \(68.1\) and mAP from \(55.57\) to \(58.84\). On VIPeR, Rank-1 increases from \(36.2\) to \(37.8\) [1607.08378]. The intended effect is especially clear for hard negatives: the gate makes local patterns such as hats or bags more salient when they truly correspond, and less salient when they do not.

A recurrent misconception is that this architecture performs generic cross-attention over all spatial positions. It does not. The mechanism is explicitly row-aligned and channel-wise, and its effectiveness depends on the assumption that horizontal correspondence after resizing is meaningful. The paper identifies this as a limitation under large pose changes, severe misdetections, or occlusions [1607.08378].

## 3. Spatially gated Siamese attention networks for video person re-identification

In “Where-and-When to Look: Deep Siamese Attention Networks for Video-based Person Re-identification” [1808.01911], the gated Siamese idea is extended from static image pairs to video pairs. The task is to determine whether two pedestrian videos from disjoint cameras depict the same identity, and to rank gallery sequences for each probe. The model jointly learns spatiotemporal video representations and their similarity value, rather than treating feature learning and metric learning as separate stages.

Each Siamese stream shares parameters and contains four major components. First, the last convolutional layer of GoogLeNet pre-trained on ImageNet is used as encoder, yielding for frame \(x_t\) a feature cube \( \mathbf{X}_t \in \mathbb{R}^{K \times K \times D} \), instantiated in experiments as \(7 \times 7 \times 1024\). Second, deterministic soft attention predicts a location softmax \(m_t\) over spatial sites conditioned on the previous hidden state:
$$
m_{t,i}=p(\mathbf{L}_t=i|\mathbf{h}_{t-1})=\frac{\exp (W_i^T \mathbf{h}_{t-1})}{\sum_{j=1}^{K\times K} (W_j^T \mathbf{h}_{t-1})}.
$$
The soft-attended input is then
$$
\mathbf{x}_t= m_{t,i} \mathbf{X}_{t,i}.
$$

Third, temporal modeling is performed by spatial gated recurrent units implemented as ConvGRUs. The fully connected GRU equations are standard:
$$
\mathbf{z}_t = \sigma(W_z \mathbf{x}_t + U_z \mathbf{h}_{t-1}), \quad
\mathbf{r}_t = \sigma(W_r \mathbf{x}_t + U_r \mathbf{h}_{t-1}),
$$
$$
\hat{\mathbf{h}}_t = \tanh(W \mathbf{x}_t + U(\mathbf{r}_t \odot \mathbf{h}_{t-1})),
$$
$$
\mathbf{h}_t = (1 - \mathbf{z}_t)\mathbf{h}_{t-1} + \mathbf{z}_t \hat{\mathbf{h}}_t.
$$
To preserve spatial structure, the model replaces matrix multiplications with convolutions:
$$
\mathbf{z}_t^l = \sigma(W_z^l * \mathbf{h}_t^{l-1} + U_z^l * \mathbf{h}_{t-1}^l), \quad
\mathbf{r}_t^l = \sigma(W_r^l * \mathbf{h}_t^{l-1} + U_r^l * \mathbf{h}_{t-1}^l),
$$
$$
\hat{\mathbf{h}}_t^l = \tanh(W^l * \mathbf{h}_t^{l-1} + U^l * (\mathbf{r}_t^l \odot \mathbf{h}_{t-1}^l )),
$$
$$
\mathbf{h}_t^l = (1- \mathbf{z}_t^l)\mathbf{h}_{t-1}^l + \mathbf{z}_t^l \hat{\mathbf{h}}_t^l.
$$
The experimental stack uses \(L=3\) ConvGRU layers with channels \(128, 256, 256\), kernel size \(5 \times 5\), zero padding, and max-pooling between recurrent layers.

Fourth, “when” attention performs temporal pooling:
$$
\bar{\mathbf{h}} = \frac{\alpha_t \mathbf{h}_t^L}{\sum_{t=1}^T \alpha_t \mathbf{h}_t^L},
$$
where \(\alpha_t \ge 0\). This soft selection is intended to reduce bias toward later time steps and to support variable-length sequences by emphasizing informative frames.

Similarity learning is integrated into the Siamese model. Given \(\bar{\mathbf{h}}^a\) and \(\bar{\mathbf{h}}^b\), the similarity score is
$$
s(X^a,X^b) = \frac{1}{1 + e^{-\mathbf{v}^T [\mbox{diag}\left(\bar{\mathbf{h}}^a ( \bar{\mathbf{h}}^b)^{\mathsf{T}}\right)] + c}},
$$
interpreted as the probability that the two sequences belong to the same identity. The training objective combines binary cross-entropy over positive and negative pairs with a doubly-stochastic attention penalty,
$$
\mathcal{L}(\Theta; Z) = \lambda \sum_{i=1}^{K^2} (1- \sum_{t=1}^T m_{t,i})^2
-\left[ \sum_{(a,b)\in Sim} \log s(X^a, X^b)
+ \sum_{(a,b)\in Dis}  \log (1-s(X^a, X^b)) \right],
$$
with \(\lambda=1\) in experiments.

Implementation details are unusually explicit. The model is trained with RMSProp SGD on mini-batches of 10 pairs, gradients clipped to \([-5,5]\), dropout \(0.7\) on hidden activations, and subsequences of length \(T=20\). Data augmentation uses random cropping, mirroring, and random 2D translation within \([-0.05A,0.05A] \times [-0.05B,0.05B]\). Training runs for 1,000 epochs on iLIDS-VID and PRID2011, and 2,000 epochs on MARS. Runtime is approximately \(0.06\) s per pair of frames on a single NVIDIA GTX 980 (12 GB) [1808.01911].

Results on three video re-identification benchmarks show clear dependence on both gating and attention. The full model with GoogLeNet, ConvGRU, and spatial/temporal attention reaches Rank-1/5/10/20 of \(61.2/80.7/90.3/97.3\) on iLIDS-VID, \(74.8/92.6/97.7/98.6\) on PRID2011, and \(69.7/83.4/88.3/93.6\) on MARS. With KISSME on PCA-reduced descriptors, Rank-1 rises to \(61.9\), \(77.0\), and \(73.5\) on the three datasets respectively. When attention is muted by setting \(m_{t,i}=1/K^2\), Rank-1 drops to \(51.8\), \(58.6\), and \(55.2\). Kernel-size ablations show that the \((5 \times 5, 5 \times 5)\) hidden-to-hidden configuration outperforms \((9 \times 9, 1 \times 1)\) and \((1 \times 1, 9 \times 9)\), and the paper concludes that larger hidden kernels improve spatiotemporal modeling [1808.01911].

Here, “gated” refers to recurrent control of state propagation and forgetting at each spatial location. This is distinct from the Matching Gate of image-based re-identification, even though both are Siamese and both make the representation depend on the counterpart input.

## 4. Hard branch gating in Siamese object tracking

“Multi-Branch Siamese Networks with Online Selection for Object Tracking” [1808.07349] adopts a different interpretation. The tracker uses multiple Siamese branches pretrained for different tasks or contexts and chooses one branch online according to the current target characteristics. Each branch applies an identical transformation to an exemplar \(z\) and a search region \(X\), and combines their representations by a cross-correlation layer. For a context-dependent branch \(s_i\),
$$
h_{s_i}(z,X)=\mathrm{corr}\big(f_{s_i}(z),\, f_{s_i}(X)\big),
$$
and for the AlexNet branch,
$$
h_{a}(z,X)=\mathrm{corr}\big(f_{a}(z),\, f_{a}(X)\big).
$$
The inputs have dimensions \(127 \times 127 \times 3\) for \(z\) and \(255 \times 255 \times 3\) for \(X\); the embedding outputs have dimensions \(6 \times 6 \times 256\) and \(22 \times 22 \times 256\), and all response maps are \(17 \times 17\).

The branch ensemble contains \(N_e=N_s+1\) Siamese networks, where the context-dependent group satisfies \(N_s=N_c+1\) with \(N_c=10\). All context-dependent branches and the general branch have the same structure as SiamFC, while the additional branch is AlexNet pretrained on the image classification task with stride modifications so that its response-map dimensions match the others.

The gating mechanism is a hard online choice made every \(T\) frames. Because response magnitudes differ across branches, response weights \(w_i\) are used for normalization. Discriminative power is measured by
$$
R\big(w_i\, h_{B_i}\big)=w_i\Big(P\big(h_{B_i}\big)-M\big(h_{B_i}\big)\Big),
$$
where \(P(\cdot)\) and \(M(\cdot)\) denote the peak and minimum of the response map. The selected branch is
$$
B^{*}=\argmax_{B_i} R\big(w_i\, h_{B_i}\big).
$$
The weights for context-dependent branches are \(1.0\); for the AlexNet branch, grid search from \(8.0\) to \(12.0\) with step \(0.5\) identifies \(10.5\) as best. The optimal branch-selection interval is \(T=7\) frames, with OTB-2013 AUC values \(0.613, 0.616, 0.616, 0.620, 0.616, 0.608\) for \(T=1,3,5,7,10,13\), respectively [1808.07349].

This gating mechanism is neither soft nor learned end-to-end. The paper states that the selection mechanism itself has no learnable parameters, and there is no online learning of branch parameters during tracking. The general SiamFC branch is trained on ILSVRC-2015 video data for 50 epochs with initial learning rate \(0.01\) and decay factor \(\delta=0.869\) after every epoch. The \(N_c\) context-dependent branches are trained after contextual clustering on the low-level feature map from the ImageNet Video dataset, with fine-tuning learning rate \(0.00001\) for 10 epochs. The average testing speed of the full tracker is 17 fps on an Intel i7-3770 3.40 GHz CPU and a Nvidia Titan X GPU [1808.07349].

The full MBST, using general, context-dependent, and AlexNet branches, attains OTB-2013 AUC/precision \(0.620/0.816\), OTB-50 \(0.573/0.773\), and OTB-100 \(0.617/0.811\). Ablation shows that the full multi-branch gated configuration exceeds general-only, context-only, AlexNet-only, and partial combinations. The paper further reports that MBST significantly outperforms SiamFC in the case of deformation, occlusion, and out-of-plane rotations because the contrast between object and background changes and switching to another feature map may give better discriminativity [1808.07349].

A common misunderstanding is to equate all gating with differentiable attention. This example shows otherwise: the gate is a periodic argmax over branch-level heuristics, used for inference-time expert selection rather than for per-feature weighting.

## 5. Respiratory-gated Siamese adversarial networks for low-dose PET

In “Simultaneous Denoising and Motion Estimation for Low-dose Gated PET using a Siamese Adversarial Network with Gate-to-Gate Consistency Learning” [2009.06757], the term “gated” has a dual meaning. It refers first to respiratory gating in PET acquisition, where list-mode data are partitioned into six phase bins per respiratory cycle, and second to a Siamese network that processes a reference gate and a target gate jointly. End-expiration, Gate 4, is used as the reference because it has the least intra-gate motion.

For each gate \(g \in \{1,\dots,6\}\), low-dose and high-dose gated volumes are denoted \(L_g\) and \(H_g\). The Siamese generator \(G\) takes low-dose PET volumes concatenated channel-wise with anatomical prior CT \(\rho\), and two weight-shared branches process a reference and a randomly sampled target gate:
$$
\bar{H}_{ref} = G(L_{ref}, \rho; \theta_G), \quad
\bar{H}_{tgt} = G(L_{tgt}, \rho; \theta_G).
$$
A single adversarial discriminator \(D\) is shared across gates and trained with WGAN-GP. A motion estimation network \(R\), pre-trained on high-dose PET pairs, outputs a distribution over a stationary velocity field \(V\); scaling-and-squaring yields a diffeomorphic transform \(T\), and a differentiable spatial transformer produces the warped volume.

The structure-recovery objective for the Siamese generator combines \(L_1\), SSIM, and adversarial losses:
$$
\mathcal{L}_{SR} = \beta_1 \mathcal{L}_1 + \beta_2 \mathcal{L}_{SSIM} + \beta_3 \mathcal{L}_{adv},
$$
with \(\beta_1=1\), \(\beta_2=1\), and \(\beta_3=0.2\). The adversarial term uses \(\lambda_{gp}=3\). The motion network is first trained on ground-truth HDPET pairs with
$$
\mathcal{L}_R(H_{ref},H_{tgt}) = \frac{1}{K} \sum_{k} \| H_{ref} - (T \circ H_{tgt}) \|_1 + \mathrm{KL}\big[q_{\theta_R}(V\,|\,H_{ref},H_{tgt}) \,\|\, p(V)\big].
$$
After pre-training, \(R\) is frozen and concatenated to the generator, and the gate-to-gate consistency loss becomes
$$
\mathcal{L}_{G2G} = \frac{1}{K} \sum_{k} \| H_{ref} - (\bar{T} \circ \bar{H}_{tgt}) \|_1 + \mathrm{KL}\big[q_{\theta_R}(V\,|\,\bar{H}_{ref},\bar{H}_{tgt}) \,\|\, p(V)\big].
$$
The total objective is
$$
\mathcal{L}_{tot} = \mathcal{L}_{SR} + \mathcal{L}_{G2G}.
$$

The dataset comprises 29 clinical pancreas \(^{18}\)F-FPDTBZ PET/CT exams with respiratory gating using the Anzai system, 22 for training and 7 for testing. HDPET uses 100% of list-mode data; LDPET uses random 1.5%. Volumes reconstructed at \(400 \times 400 \times 109\) are cropped to central \(200 \times 200 \times 109\) and resized to \(128 \times 128 \times 128\). Siamese pairing augments training data by random gate pairs, yielding \(A_6^2 = 6 \times 5 = 30\) ordered pairs per subject [2009.06757].

The reported denoising results, averaged over all six gates on the 7 test studies, are \(28.29\) dB / \(86.4\) for LDPET, \(36.48\) / \(95.9\) for UNet, \(36.67\) / \(96.2\) for 3D cGAN, \(36.92\) / \(96.7\) for SAN-G2G, and \(37.16\) / \(97.0\) for SAN+G2G in PSNR/SSIM\((\times 10^2)\). For motion estimation, the mean vector Euclidean distance (MVED) is \(0.331\) for LDPET, \(0.288\) for SAN-G2G, and \(0.264\) for SAN+G2G, described as approximately 20% average improvement versus the low-dose baseline. The final motion-corrected averaged images show visibly reduced blur and noise relative to uncorrected low-dose gated data [2009.06757].

This formulation differs from the computer-vision uses of “gated Siamese CNN.” The gate here is tied to respiratory phases and inter-gate consistency, while the Siamese component is weight sharing across reference and target gates. The architecture is therefore not a simple transfer of matching-gate or branch-selection ideas into medical imaging.

## 6. Comparative interpretation, misconceptions, and limitations

Across these works, a gated Siamese CNN is not defined by any single gate equation. The term encompasses at least four distinct gating loci. In image re-identification, gating is a learned, differentiable, pair-conditioned modulation of mid-level stripe features [1607.08378]. In video re-identification, gating is the update/reset control intrinsic to ConvGRUs, supplemented by spatial and temporal attention [1808.01911]. In tracking, gating is a hard heuristic selection of one branch among several Siamese experts [1808.07349]. In PET, gating denotes respiratory phase structure, while consistency across paired gates is enforced through a frozen registration network [2009.06757].

Several misconceptions follow from collapsing these variants into a single template. First, “gated” does not necessarily imply recurrent gating; only the video re-identification model uses GRU update and reset gates as the primary mechanism [1808.01911]. Second, gating is not necessarily soft or learned jointly with the backbone; the branch selection mechanism in MBST has no learnable parameters and uses an argmax over a response-map heuristic [1808.07349]. Third, Siamese interaction need not occur only at the output. The Matching Gate paper is explicit that waiting until the final distance can cause the network to miss subtle local cues, so the pair interaction is brought into mid-level layers [1607.08378]. Fourth, in medical imaging, “gated” may refer to the acquisition protocol itself rather than only to an internal gating module [2009.06757].

The limitations are correspondingly heterogeneous. The Matching Gate assumes horizontal correspondence after resizing; large pose changes, severe misdetections, or occlusions can reduce its effectiveness [1607.08378]. The video-based re-identification model exhibits performance drops under cross-dataset transfer, indicating domain shift sensitivity, and can fail when distinct pedestrians share very similar upper-body appearance and gait or when severe occlusion reduces discriminative cues [1808.01911]. MBST identifies a trade-off in gating frequency: selecting every frame can increase the possibility of selecting an inappropriate branch, while selecting too rarely can retain a branch that is no longer discriminative [1808.07349]. The PET framework depends on the quality of the pre-trained motion estimator and requires further validation for other scanners, tracers, extreme low-dose regimes, or different gating strategies [2009.06757].

Taken together, these papers suggest that the enduring value of the gated Siamese CNN idea lies not in a fixed architectural recipe but in a design principle: paired inputs can benefit from explicit, selective interaction before the terminal comparison stage. The particular form of that interaction depends on the structure of the task—row-aligned local evidence for image re-identification, spatiotemporal state propagation for video sequences, expert selection for tracking, or inter-phase consistency for respiratory-gated PET.

Source: https://www.emergentmind.com/topics/gated-siamese-cnn