---
title: 'CoAtNeXt: Hybrid Conv-Transformer for Gastric Histology'
url: https://www.emergentmind.com/topics/coatnext
type: topic
---

# CoAtNeXt: Hybrid Conv-Transformer for Gastric Histology

Searching arXiv for the named paper and its architectural antecedents to ground the article in current literature.
[Tool unavailable in this interface; proceeding using the provided arXiv record and explicitly cited arXiv IDs for related antecedent papers.]
CoAtNeXt is a hybrid convolution–transformer architecture for gastric histopathological image classification that modifies the CoAtNet backbone by replacing early-stage MBConv layers with CBAM-augmented ConvNeXtV2 blocks while preserving deeper ReIPos transformer blocks for long-range contextual modeling. It was introduced for automated analysis of gastric tissue images, motivated by the fact that histopathologic examination remains the diagnostic gold standard yet is performed entirely manually, making evaluations labor-intensive, prone to variability among pathologists, and susceptible to missed findings. In the reported experiments, CoAtNeXt was evaluated on HMU-GC-HE-30K for eight-class classification and GasHisSDB for binary classification, where it achieved 96.47% accuracy and 98.29% accuracy, respectively [2509.09242].

## 1. Clinical and methodological setting

The stated problem setting is gastric tissue classification from histopathological images, with an emphasis on early diagnosis of gastric diseases. The motivating claim is that manual histopathologic assessment is labor-intensive and prone to inter-pathologist variability, while the lack of standard procedures reduces consistency. The model is therefore positioned as an automated, reliable, and efficient method for gastric tissue analysis [2509.09242].

Methodologically, CoAtNeXt belongs to a class of hybrid vision architectures that combine convolutional inductive biases in early stages with transformer-based global context modeling in deeper stages. Its immediate architectural lineage is CoAtNet [2106.04803], while its convolutional replacement units are based on ConvNeXtV2 [2301.00808], and its attention augmentation uses the Convolutional Block Attention Module, CBAM [1807.06521]. This suggests that the design objective was not to discard the CoAtNet hierarchy, but to reallocate representational capacity across local feature extraction and long-range dependency modeling.

A common misconception would be to treat CoAtNeXt as a pure Vision Transformer. The architecture described in the technical overview is explicitly hybrid: convolution-dominant in stages S1 and S2, transformer-based in stages S3 and S4, and stage-coupled by downsampling and channel projection [2509.09242].

## 2. Backbone organization and stage-wise topology

CoAtNeXt retains the five-stage hybrid convolution–transformer layout of CoAtNet, denoted S0–S4. In stages S1 and S2, the original MBConv blocks are replaced by “Improved ConvNeXtV2” blocks, defined as ConvNeXtV2 convolutional units augmented with CBAM. In stages S3 and S4, ReIPos relative-position transformer blocks from CoAtNet are preserved. Between every two stages, a \(1\times1\) linear projection plus stride-2 downsampling halves spatial resolution and doubles channel width [2509.09242].

| Stage | Block type | Output specification |
|---|---|---|
| S0 | Stem: one \(3\times3\) conv, stride 2 | \(112\times112\), \(D=32\) |
| S1 | \(L_1=2\) Improved ConvNeXtV2 blocks | \(56\times56\), \(D_1=64\) |
| S2 | \(L_2=2\) Improved ConvNeXtV2 blocks | \(28\times28\), \(D_2=128\) |
| S3 | \(L_3=4\) ReIPos transformer blocks | \(14\times14\), \(D_3=256\) |
| S4 | \(L_4=2\) ReIPos transformer blocks | \(7\times7\), \(D_4=512\) |
| Head | Global average pooling \(\rightarrow\) FC \(\rightarrow\) softmax or sigmoid | 8-way or binary output |

The stage-scaling rules are stated explicitly as \(D_{s+1}=2\cdot D_s\) for channel doubling and \(H_{s+1}=H_s/2\) for spatial halving. The technical overview specifies an input of a \(224\times224\) RGB image for the stage-wise layout. The head uses softmax for the 8-way setting and sigmoid for the binary setting [2509.09242].

This topology is significant because it preserves the hierarchical multi-scale structure of CoAtNet while changing the early-stage operator family. A plausible implication is that the architecture was designed to strengthen texture-sensitive local encoding without removing the long-range modeling mechanisms that hybrid backbones typically exploit.

## 3. Replacement of MBConv and integration of CBAM

The architectural novelty centers on replacing MBConv with an “Improved ConvNeXtV2” block. The generic MBConv form is given as

$$
y = W_p^\downarrow \bigl(\sigma(\mathrm{SE}(\mathrm{DWConv}_k(\sigma(W_p^\uparrow x))))\bigr) + x,
$$

where \(W_p^\uparrow\) and \(W_p^\downarrow\) are \(1\times1\) expansion and projection, \(\mathrm{DWConv}_k\) is a \(k\times k\) depthwise convolution, and \(\mathrm{SE}(\cdot)\) is squeeze-and-excitation [2509.09242].

The ConvNeXtV2 replacement block is described by the following sequence: a large-kernel depthwise convolution, Layer Normalization per channel, two pointwise (\(1\times1\)) convolutions with a GELU activation in between, Global Response Normalization (GRN) to prevent feature collapse, and residual DropPath. The “Improved ConvNeXtV2” structure is specified as

**Input \(\rightarrow \mathrm{DWConv}_k \rightarrow \mathrm{LN} \rightarrow \mathrm{CBAM} \rightarrow \mathrm{PWConv}_1\) (expansion) \(\rightarrow \mathrm{GELU} \rightarrow \mathrm{GRN} \rightarrow \mathrm{PWConv}_1\) (projection) \(\rightarrow \mathrm{DropPath} \rightarrow +\) Residual** [2509.09242].

For the depthwise convolution, the technical overview gives

$$
(F * K)_{i,j,c} = \sum_{u,v} K_{u,v,c}\,F_{i+u,j+v,c}.
$$

GRN is introduced with channelwise \(L_2\)-norm aggregation and learnable scalars \(\alpha,\beta\). The stated role of GRN is to prevent feature collapse, consistent with the ConvNeXtV2 design lineage [2301.00808].

CBAM is applied sequentially through channel attention and spatial attention:

$$
M_c(F)=\sigma\bigl(\mathrm{MLP}(\mathrm{AvgPool}(F))+\mathrm{MLP}(\mathrm{MaxPool}(F))\bigr),
$$

$$
F' = M_c(F)\odot F,
$$

$$
M_s(F')=\sigma\bigl(f^{7\times7}([\mathrm{AvgPool}(F');\mathrm{MaxPool}(F')])\bigr),
$$

$$
F'' = M_s(F')\odot F'.
$$

Here, \(\sigma\) is the sigmoid, \(f^{7\times7}\) is a \(7\times7\) convolution, and \(\odot\) is element-wise multiplication [1807.06521]. In CoAtNeXt, CBAM is explicitly used to improve local feature extraction through channel and spatial attention mechanisms [2509.09242].

The significance of this substitution lies in the fact that it changes both the convolutional primitive and the attention mechanism used in the early stages. The ablation data indicate that the gain is not attributable to ConvNeXtV2 alone; additional gains appear when CBAM is added on top of the ConvNeXtV2 replacement.

## 4. Data, preprocessing, and optimization protocol

Two public datasets are reported. HMU-GC-HE-30K contains 31,096 patches of size \(224\times224\) spanning eight classes—ADI, DEB, LYM, MUC, MUS, NOR, STR, and TUM—and is described as balanced at approximately 3,887 samples per class. GasHisSDB is reported in its \(80\times80\) subset with 146,651 patches, of which 59,151 are abnormal and 87,500 are normal. HMU-GC-HE-30K is used for eight-class classification and GasHisSDB for binary classification [2509.09242].

Preprocessing is stated as per-channel normalization for both datasets. No advanced augmentations are reported; the technical overview adds that standard practice may include random flips, rotations, and color jitter. This wording is important: it does not document those augmentations as part of the reported experimental protocol [2509.09242].

The implementation stack is TensorFlow 2.14.0 and Keras 2.11.4 on Windows 11 with 128 GB RAM, \(2\times\) NVIDIA RTX 3090 (24 GB), and CUDA 12.7. Optimization uses SGD with momentum (\(m \approx 0.9\)) and no explicit weight decay. The learning rate is selected by grid search, with a typical initial value of 0.01 under a constant schedule. The batch size is 32, training proceeds for 50 epochs per fold, and evaluation is performed with 5-fold cross-validation. The reported losses are Categorical Cross-Entropy for the multiclass task and Binary Cross-Entropy for the binary task [2509.09242].

For reproducibility, the practical notes recommend at least one GPU with at least 24 GB VRAM, report training time of approximately 8–12 hours per 50-epoch run per fold with two GPUs in data-parallel mode, and identify Python 3.9+, TensorFlow 2.x, Keras, and cross-validation scripts as the code stack [2509.09242].

## 5. Quantitative performance and ablation behavior

The evaluation protocol uses confusion-matrix-derived metrics: accuracy, precision, recall, F1-score, and AUC. Results are reported as averages over 5 folds [2509.09242].

| Dataset | CoAtNeXt metrics | Baseline CoAtNet |
|---|---|---|
| HMU-GC-HE-30K | Acc 96.47%, P 96.60%, R 96.47%, F1 96.45%, AUC 99.89% | Acc 93.12%, AUC 99.15% |
| GasHisSDB | Acc 98.29%, P 98.07%, R 98.41%, F1 98.23%, AUC 99.90% | Acc 96.66%, AUC 99.51% |

The paper states that CoAtNeXt outperformed all CNN and ViT models tested and surpassed previous studies in the literature [2509.09242]. Within the scope of the reported experiments, this places the model above both conventional CNN baselines and Vision Transformer baselines, although the full per-model breakdown is not reproduced in the technical overview.

The ablation study isolates the effect of replacing MBConv and varying attention modules. On HMU-GC-HE-30K, the sequence is CoAtNet (MBConv) at 93.12% accuracy, \(+\) ConvNeXtV2 only at 94.50%, \(+\) SE attention at 94.75%, \(+\) ECA attention at 95.64%, and \(+\) CBAM, i.e. CoAtNeXt, at 96.47%. On GasHisSDB, the corresponding sequence is 96.66%, 96.73%, 97.46%, 97.68%, and 98.29% [2509.09242].

These ablations support two bounded conclusions. First, replacing MBConv with ConvNeXtV2 yields a measurable gain on both datasets. Second, among the compared attention augmentations—SE, ECA, and CBAM—the best reported accuracies are obtained with CBAM. A plausible implication is that sequential channel-and-spatial attention is particularly effective in the local feature regime targeted by the early stages of the network.

## 6. Scope, limitations, and plausible extensions

The reported limitations are explicit. The staining modality is restricted to H\&E images, and performance on other stains such as IHC is untested. The eight-class dataset is characterized as relatively new, and more external validation is needed. These constraints mean that the reported results do not establish stain-invariant or broadly cross-institutional performance [2509.09242].

The stated future directions are to incorporate multi-stain or multi-modal data, such as endoscopy plus histology; to extend the model to segmentation or detection of sub-regions such as gland and nuclei; and to explore adaptive learning-rate schedules such as cosine decay and warm restarts together with stronger augmentations including Mixup and CutMix [2509.09242]. These are natural directions because the present study evaluates only classification, uses a constant learning-rate schedule, and does not report advanced augmentation.

In a broader research context, CoAtNeXt exemplifies a design pattern in which a hybrid backbone is preserved while the local operator family is modernized and attention is made more explicit. The paper’s own evidence is confined to gastric histopathology classification, but it suggests that careful substitution of early-stage convolutional blocks can materially affect downstream performance even when the global transformer structure remains intact [2509.09242].

Source: https://www.emergentmind.com/topics/coatnext