---
title: 'GSoP-Net2: Advanced Second-Order Pooling'
url: https://www.emergentmind.com/topics/gsop-net2
type: topic
---

# GSoP-Net2: Advanced Second-Order Pooling

Global Second-order Pooling Network 1 (GSoP-Net1) is a deep convolutional network architecture that systematically integrates global second-order pooling into all major intermediate stages of a ResNet-style backbone, augmenting both representational power and non-linear modeling capability compared to first-order pooling approaches. GSoP-Net1 was introduced as the first-order–pooled variant in “Global Second-order Pooling Convolutional Networks” by Gao et al., providing a comprehensive framework for leveraging higher-order statistics—specifically, second-order (covariance) information—throughout a convolutional neural network rather than only at the terminal layer. The architecture is empirically validated on large-scale visual recognition tasks and achieves robust improvements over various established baselines [1811.12006].

## 1. Formal Definition of Global Second-order Pooling

Let $X \in \mathbb{R}^{H \times W \times C}$ denote the output tensor from a convolutional layer, where $x_{ij} \in \mathbb{R}^C$ is the feature vector at spatial location $(i,j)$, with $i=1\ldots H$, $j=1\ldots W$. Denote $N=HW$. The mean feature vector is $\mu = \frac{1}{N} \sum_{i=1}^H \sum_{j=1}^W x_{ij}$. The (centered) channel covariance matrix $C \in \mathbb{R}^{C \times C}$ is computed as:
$$
C = \frac{1}{N} \sum_{i=1}^H \sum_{j=1}^W (x_{ij}-\mu)(x_{ij}-\mu)^T
$$
In practice, $\mu$ may be set to zero as batch normalization (BN) is applied row-wise to $C$ downstream. After obtaining $C$, two subsequent transformations are applied: (1) row-wise BN, $C_{\text{norm}} = \text{BN}_{\text{row}}(C)$; (2) a nonlinear 1$\times$1 row-conv embedding followed by a LeakyReLU activation and logistic sigmoid, producing a channel-wise (or spatial-wise) weighting vector $S \in \mathbb{R}^C$:
\[
S = \sigma(W_2 \cdot \delta(W_1 \cdot \text{BN}_{\text{row}}(C)) + b_2)
\]
where $W_1, b_1, W_2, b_2$ are learned parameters, $\delta$ denotes LeakyReLU with slope 0.1, and $\sigma$ is the element-wise sigmoid. *A plausible implication is that these transformations allow the model to non-linearly re-weigh channels (or spatial locations) as an attention mechanism based on holistic, second-order statistics*.

Optional matrix functions such as the matrix square root $C^{1/2}$ can be incorporated via eigendecomposition or Newton–Schulz iteration, but GSoP-Net1 itself does not utilize these at the final stage, deferring such procedures to the GSoP-Net2 variant.

## 2. GSoP Block Structure: Channel-wise and Spatial-wise Variants

GSoP-Net1 implements both channel-wise and spatial-wise GSoP blocks as modular plug-ins to convolutional backbones.

**Channel-wise GSoP block:**
- Input tensor $X \in \mathbb{R}^{H' \times W' \times C'}$ undergoes a 1$\times$1 convolution reducing channel dimension $C' \rightarrow C$ (typically $C=128$).
- The result $F$ is flattened to an $(N \times C)$ matrix ($N=H'W'$). Covariance $C = \frac{1}{N}F^T F$ is computed.
- Row-wise BN and two-layer row-conv with nonlinearity yield $S \in \mathbb{R}^C$.
- $S$ is expanded and broadcast to rescale the original $X$ across channels:
  $$
  Y_{ij,k} = S_k \cdot X_{ij,k}
  $$
  for all $i$, $j$, $k$.

**Spatial-wise GSoP block:**
- 1$\times$1 convolution reduces $C' \rightarrow C$, producing $F$. Spatial downsampling (e.g., $14 \times 14 \rightarrow 8 \times 8$) yields $\widetilde{F}$.
- Reshape $\widetilde{F}$ to $(M \times C)$ ($M=hw$) and compute spatial covariance $C_{\text{sp}} \in \mathbb{R}^{M \times M}$ as $(1/C) \widetilde{F} \widetilde{F}^T$.
- Nonlinear embedding yields $V \in \mathbb{R}^M$, reshaped and upsampled to $W \in \mathbb{R}^{H' \times W'}$, which is broadcast to rescale $X$ spatially:
  $$
  Y_{ij,k} = \widehat{W}_{ij} X_{ij,k}
  $$

## 3. Placement of GSoP Blocks in ResNet-Style Backbone

GSoP-Net1 adopts a ResNet-50 backbone, inserting channel-wise GSoP blocks after each of the four major residual stages. The precise placements are as follows:

| Stage      | Output Shape     | Bottleneck Structure                         | GSoP Block Insertion     |
|------------|------------------|----------------------------------------------|--------------------------|
| conv2_x    | 56×56×256        | [1×1,64]→[3×3,64]→[1×1,256] × 3             | After final bottleneck   |
| conv3_x    | 28×28×512        | [1×1,128]→[3×3,128]→[1×1,512] × 4           | After stage              |
| conv4_x    | 14×14×1024       | [1×1,256]→[3×3,256]→[1×1,1024] × 6          | After stage              |
| conv5_x    | 7×7×2048         | [1×1,512]→[3×3,512]→[1×1,2048] × 3          | After stage              |

After the last GSoP block, global average pooling reduces the tensor to $1 \times 1 \times 2048$, followed by a fully connected layer for classification.

## 4. Layer-wise Architectural Overview

The forward pass through GSoP-Net1 consists of the following ordered sequence:

1. Input: $224 \times 224 \times 3$
2. Conv1: $7 \times 7$, 64, stride 2 $\rightarrow$ BN $\rightarrow$ ReLU ($112 \times 112 \times 64$)
3. Pool1: $3 \times 3$, max, stride 2 ($56 \times 56 \times 64$)
4. conv2_x: 3 bottlenecks ($56 \times 56 \times 256$)
5. GSoP block ($56 \times 56 \times 256$, channel-wise, $C=128$)
6. conv3_x: 4 bottlenecks ($28 \times 28 \times 512$)
7. GSoP block ($28 \times 28 \times 512$, channel-wise, $C=128$)
8. conv4_x: 6 bottlenecks ($14 \times 14 \times 1024$)
9. GSoP block ($14 \times 14 \times 1024$, channel-wise, $C=128$)
10. conv5_x: 3 bottlenecks ($7 \times 7 \times 2048$)
11. GSoP block ($7 \times 7 \times 2048$, channel-wise, $C=128$)
12. Global average pooling ($1 \times 1 \times 2048$)
13. Fully-connected layer ($2048 \rightarrow 1000$), softmax output

Each channel-wise GSoP block introduces approximately $0.7$ million parameters and $30$ MFLOPs, with an optional spatial-wise branch adding $0.16$ million parameters and $26$ MFLOPs.

## 5. Implementation and Computational Characteristics

Key implementation details of GSoP-Net1 include:
- 1$\times$1 convolutions consistently reduce the channel dimension prior to covariance computation, fixing $C=128$ in all GSoP blocks.
- The channel-wise GSoP operation is efficiently implemented via a matrix multiplication of size $C \times N$ and $N \times C$.
- Row-wise BN is a batch normalization applied along the rows of the $C \times C$ covariance.
- Two 1$\times$1 row-convolutions (with LeakyReLU nonlinearity) effect the embedding/attention computation, acting independently on each row.
- Matrix square-root computation and explicit eigen-decomposition are not used in GSoP-Net1, distinguishing it from the “Net2” variant.
- No large 3D convolutions are introduced. Overhead relative to the ResNet-50 baseline is approximately $10\%$ in parameters and $60\%$ in FLOPs.

## 6. Performance on ImageNet-1K and Comparative Analysis

Empirical evaluation of GSoP-Net1 on ImageNet-1K demonstrates significant improvement relative to both first-order and end-only second-order baselines:

| Model                    | Top-1 Error (%) | Top-5 Error (%) |
|--------------------------|-----------------|-----------------|
| ResNet-50 (baseline)     |      23.85      |       7.13      |
| GSoP-Net1                |   22.32 (↓1.53) |   6.02 (↓1.11)  |
| SE-Net-50                |      23.29      |      6.62       |
| CBAM                     |      22.66      |      6.31       |
| MPN-COV                  |      22.74      |      6.54       |

Ablation studies on scale-reduced ResNet-26 show that single GSoP blocks in early (conv2_x) or late (conv5_x) stages yield Top-1 errors of 18.45% and 18.33%, respectively; applying GSoP blocks throughout all four stages reduces Top-1 error to 17.42% (from a baseline of 19.18%). *This confirms that intermediate second-order pooling delivers cumulative accuracy gains, outperforming first-order (SE, CBAM) and “end-only” (MPN-COV) second-order strategies*.

## 7. Context and Significance

GSoP-Net1 establishes the effectiveness of integrating global second-order pooling throughout the depth of a residual network, as opposed to restricting such pooling to the final layer. The block design leverages covariance-based attention both channel- and spatial-wise without incurring prohibitive computational cost or requiring matrix square-root operations. The architecture demonstrates that holistic second-order statistics, when used as feature recalibration signals at multiple network depths, deliver consistent and non-trivial improvements in large-scale visual recognition tasks [1811.12006]. This approach represents a significant advance in the practical use of higher-order global descriptors in deep convolutional architectures.

Source: https://www.emergentmind.com/topics/gsop-net2