---
title: 'GSoP-Net1: Global Second-Order Pooling in ConvNets'
url: https://www.emergentmind.com/topics/gsop-net1
type: topic
---

# GSoP-Net1: Global Second-Order Pooling in ConvNets

GSoP-Net1 is a convolutional neural network architecture that systematically injects global second-order pooling (GSoP) operations at multiple layers in a deep ConvNet, rather than only at the network’s output. This design, introduced by Gao et al. [1811.12006], addresses the challenge of enhancing non-linear modeling capability in large-scale visual recognition by leveraging higher-order feature representations in both lower and higher layers. The result is a significant improvement over first-order and end-only second-order pooling approaches on tasks such as ImageNet-1K classification.

## 1. Formal Definition of Global Second-Order Pooling

GSoP computes a covariance matrix from the feature map produced by a convolutional layer, providing a holistic second-order representation. Let $X \in \mathbb{R}^{H \times W \times C}$ denote the convolutional output with $x_{ij} \in \mathbb{R}^C$ as the feature at spatial position $(i,j)$, where $i=1\ldots H$, $j=1\ldots W$, and $N=HW$. The mean feature is defined as:
\[
\mu = \frac{1}{N} \sum_{i=1}^H \sum_{j=1}^W x_{ij}.
\]
The covariance matrix $C \in \mathbb{R}^{C \times C}$ is:
\[
C = \frac{1}{N} \sum_{i=1}^H \sum_{j=1}^W (x_{ij} - \mu)(x_{ij} - \mu)^T.
\]
Typically, mean-subtraction is omitted since a subsequent Batch Norm will re-center $C$.

Following this, row-wise Batch Normalization (BN) is applied across each row of $C$, yielding $C_{\text{norm}}$. A two-stage nonlinear transformation forms the channel scaling vector:
\[
Z = W_1 \cdot C_{\text{norm}} + b_1
\]
with $W_1, b_1 \in \mathbb{R}^{M \times C}$, followed by LeakyReLU activation (slope 0.1):
\[
Z' = \text{LeakyReLU}(Z),
\]
and a final linear projection with sigmoid nonlinearity:
\[
S = \sigma(W_2 \cdot Z' + b_2),
\]
where $W_2, b_2 \in \mathbb{R}^{C \times M}$ and $S \in \mathbb{R}^C$. This non-negative scaling vector is used to adaptively weight feature channels.

Matrix square-root normalization of $C$ via eigendecomposition or Newton–Schulz iteration is explored for the GSoP-Net2 variant but is not utilized in GSoP-Net1, which uses standard average pooling in the final stage.

## 2. Structure and Operation of GSoP Blocks

GSoP blocks come in two principal variants: channel-wise and spatial-wise, both designed to extract and apply second-order statistics for tensor scaling.

### Channel-wise GSoP Block

Given input $X \in \mathbb{R}^{H' \times W' \times C'}$, a $1 \times 1$ convolution reduces channels to $C$ (typically $C=128$), producing $F \in \mathbb{R}^{H' \times W' \times C}$. $F$ is reshaped as an $N \times C$ matrix ($N = H'W'$), and $C = (1/N) F^T F$ is computed. After row-wise BN and embedding, the resulting vector $S \in \mathbb{R}^C$ is broadcast and multiplied with $X$ along the channel dimension:
\[
Y_{ij,k} = S_k \cdot X_{ij,k}.
\]
Channels $k > C$ remain unaffected.

### Spatial-wise GSoP Block

For input $X \in \mathbb{R}^{H' \times W' \times C'}$, a $1 \times 1$ convolution again reduces to $C$ channels, resulting in $F \in \mathbb{R}^{H' \times W' \times C}$. $F$ is spatially downsampled (e.g., $14 \times 14 \rightarrow 8 \times 8$) to $\tilde{F} \in \mathbb{R}^{h \times w \times C}$, which is viewed as $M \times C$ ($M = hw$). The spatial covariance $C_\text{sp} \in \mathbb{R}^{M \times M}$ is
\[
C_\text{sp} = \frac{1}{C} \tilde{F} \tilde{F}^T,
\]
normalized row-wise and nonlinearly embedded into $V \in \mathbb{R}^M$, reshaped to $W \in \mathbb{R}^{h \times w}$, upsampled to $\hat{W} \in \mathbb{R}^{H' \times W'}$. Feature scaling is then
\[
Y_{ij,k} = \hat{W}_{ij} X_{ij,k}.
\]

## 3. Block Placement and Network Architecture

GSoP-Net1 is built upon a ResNet-50 backbone, with GSoP channel-wise blocks inserted at the end of each residual stage. Below is the block mapping:

| Stage      | Output Resolution   | Block Placement         |
|------------|--------------------|------------------------|
| conv2_x    | 56 $\times$ 56 $\times$ 256 | After last bottleneck    |
| conv3_x    | 28 $\times$ 28 $\times$ 512 | After the stage         |
| conv4_x    | 14 $\times$ 14 $\times$ 1024 | After the stage         |
| conv5_x    | 7 $\times$ 7 $\times$ 2048  | After the stage         |

The full forward pipeline is as follows:
1. Input: $224 \times 224 \times 3$
2. conv1: $7 \times 7$, 64, stride 2 $\rightarrow$ BN $\rightarrow$ ReLU $\rightarrow 112 \times 112 \times 64$
3. pool1: $3 \times 3$, max, stride 2 $\rightarrow 56 \times 56 \times 64$
4. conv2\_x: three bottleneck blocks $\rightarrow 56 \times 56 \times 256$
5. GSoP block $\rightarrow 56 \times 56 \times 256$
6. conv3\_x: four bottlenecks $\rightarrow 28 \times 28 \times 512$
7. GSoP block $\rightarrow 28 \times 28 \times 512$
8. conv4\_x: six bottlenecks $\rightarrow 14 \times 14 \times 1024$
9. GSoP block $\rightarrow 14 \times 14 \times 1024$
10. conv5\_x: three bottlenecks $\rightarrow 7 \times 7 \times 2048$
11. GSoP block $\rightarrow 7 \times 7 \times 2048$
12. Global average pooling $\rightarrow 1 \times 1 \times 2048$
13. Fully connected (2048 to 1000) $\rightarrow$ softmax

Each GSoP block introduces approximately 0.7 million parameters and 30 million FLOPs (channel-wise) and can optionally include a spatial-wise branch (approximately 0.16 million parameters, 26 million FLOPs).

## 4. Implementation and Computational Considerations

GSoP blocks are implemented efficiently, with uniform $C = 128$ reduction via $1 \times 1$ convolutions prior to pooling. For channel-wise blocks, $F^T F$ is computed via a single matrix multiplication of size $C \times N$ and $N \times C$. Row-wise BN leverages a BatchNorm layer treating the C dimension as channels. The two linear projections (row-convs) are implemented as $1 \times 1$ grouped convolutions, with each output row corresponding to its input. The final rescaling uses a broadcast multiply across tensors.

No large 3D convolutions or eigen-decomposition operations are required in GSoP-Net1, leading to an overhead of approximately 10% in parameters and 60% in FLOPs compared to ResNet-50.

## 5. Experimental Results and Ablation Analysis

On ImageNet-1K with standard $224 \times 224$ training and validation, GSoP-Net1 with blocks in conv2\_x through conv5\_x and final global average pooling achieves:

| Model               | Top-1 Error (%) | Top-5 Error (%) |
|---------------------|----------------|----------------|
| ResNet-50 (baseline)   | 23.85          | 7.13           |
| GSoP-Net1              | 22.32 (↓1.53)  | 6.02 (↓1.11)   |
| SE-Net-50              | 23.29          | 6.62           |
| CBAM                   | 22.66          | 6.31           |
| MPN-COV                | 22.74          | 6.54           |

Ablation studies with ResNet-26 on $1/4$-scale ImageNet indicate that second-order blocks in intermediate layers reduce error incrementally:
- Single channel-wise GSoP at conv2\_x: Top-1 18.45%
- Single at conv5\_x: Top-1 18.33%
- All four stages: Top-1 17.42% (baseline 19.18%)

*This suggests* that incorporating second-order statistics at multiple stages, rather than exclusively at the tail of the network, provides cumulative benefit and outperforms both first-order in-network pooling (SE/CBAM) and classical end-only second-order schemes (MPN-COV).

## 6. Comparison to Related Approaches

GSoP-Net1 is distinct from SE-Net-50 and CBAM, both of which are based on first-order channel or spatial attention mechanisms, and from MPN-COV, which applies second-order pooling only at the network’s output. While SE-Net-50 achieves Top-1/Top-5 errors of 23.29%/6.62% and CBAM obtains 22.66%/6.31%, GSoP-Net1 outperforms these with 22.32%/6.02%. Compared to MPN-COV (Top-1/Top-5 22.74%/6.54%), GSoP-Net1’s strategy of placing GSoP blocks at multiple depths yields superior results [1811.12006].

## 7. Significance and Implications

The methodology of GSoP-Net1 demonstrates that the systematic distribution of global second-order pooling blocks throughout a ConvNet backbone yields substantial gains in non-linear representational capacity and overall recognition performance. By leveraging holistic image statistics beyond first-order pooling and attention mechanisms, and without the need for computationally expensive eigen-decompositions or square-root normalization at every stage, GSoP-Net1 sets a precedent for higher-order representation learning at all layers of deep convolutional architectures. This framework provides a foundation for future advances in hierarchical network design utilizing higher-order pooling mechanisms [1811.12006].

Source: https://www.emergentmind.com/topics/gsop-net1