Papers
Topics
Authors
Recent
Search
2000 character limit reached

CoAtNeXt: Hybrid Conv-Transformer for Gastric Histology

Updated 10 July 2026
  • CoAtNeXt is a hybrid model that redefines early-stage convolution by replacing MBConv with CBAM-augmented ConvNeXtV2 blocks for improved local feature extraction.
  • It integrates early convolutional inductive biases with deep ReIPos transformer blocks to capture both local textures and long-range contextual information.
  • Evaluated on HMU-GC-HE-30K and GasHisSDB, CoAtNeXt achieved 96.47% and 98.29% accuracy respectively, outperforming traditional CNN and ViT approaches.

Searching arXiv for the named paper and its architectural antecedents to ground the article in current literature. [Tool unavailable in this interface; proceeding using the provided arXiv record and explicitly cited arXiv IDs for related antecedent papers.] CoAtNeXt is a hybrid convolution–transformer architecture for gastric histopathological image classification that modifies the CoAtNet backbone by replacing early-stage MBConv layers with CBAM-augmented ConvNeXtV2 blocks while preserving deeper ReIPos transformer blocks for long-range contextual modeling. It was introduced for automated analysis of gastric tissue images, motivated by the fact that histopathologic examination remains the diagnostic gold standard yet is performed entirely manually, making evaluations labor-intensive, prone to variability among pathologists, and susceptible to missed findings. In the reported experiments, CoAtNeXt was evaluated on HMU-GC-HE-30K for eight-class classification and GasHisSDB for binary classification, where it achieved 96.47% accuracy and 98.29% accuracy, respectively (Yurdakul et al., 11 Sep 2025).

1. Clinical and methodological setting

The stated problem setting is gastric tissue classification from histopathological images, with an emphasis on early diagnosis of gastric diseases. The motivating claim is that manual histopathologic assessment is labor-intensive and prone to inter-pathologist variability, while the lack of standard procedures reduces consistency. The model is therefore positioned as an automated, reliable, and efficient method for gastric tissue analysis (Yurdakul et al., 11 Sep 2025).

Methodologically, CoAtNeXt belongs to a class of hybrid vision architectures that combine convolutional inductive biases in early stages with transformer-based global context modeling in deeper stages. Its immediate architectural lineage is CoAtNet (Dai et al., 2021), while its convolutional replacement units are based on ConvNeXtV2 (Woo et al., 2023), and its attention augmentation uses the Convolutional Block Attention Module, CBAM (Woo et al., 2018). This suggests that the design objective was not to discard the CoAtNet hierarchy, but to reallocate representational capacity across local feature extraction and long-range dependency modeling.

A common misconception would be to treat CoAtNeXt as a pure Vision Transformer. The architecture described in the technical overview is explicitly hybrid: convolution-dominant in stages S1 and S2, transformer-based in stages S3 and S4, and stage-coupled by downsampling and channel projection (Yurdakul et al., 11 Sep 2025).

2. Backbone organization and stage-wise topology

CoAtNeXt retains the five-stage hybrid convolution–transformer layout of CoAtNet, denoted S0–S4. In stages S1 and S2, the original MBConv blocks are replaced by “Improved ConvNeXtV2” blocks, defined as ConvNeXtV2 convolutional units augmented with CBAM. In stages S3 and S4, ReIPos relative-position transformer blocks from CoAtNet are preserved. Between every two stages, a 1×11\times1 linear projection plus stride-2 downsampling halves spatial resolution and doubles channel width (Yurdakul et al., 11 Sep 2025).

Stage Block type Output specification
S0 Stem: one 3×33\times3 conv, stride 2 112×112112\times112, D=32D=32
S1 L1=2L_1=2 Improved ConvNeXtV2 blocks 56×5656\times56, D1=64D_1=64
S2 L2=2L_2=2 Improved ConvNeXtV2 blocks 28×2828\times28, D2=128D_2=128
S3 3×33\times30 ReIPos transformer blocks 3×33\times31, 3×33\times32
S4 3×33\times33 ReIPos transformer blocks 3×33\times34, 3×33\times35
Head Global average pooling 3×33\times36 FC 3×33\times37 softmax or sigmoid 8-way or binary output

The stage-scaling rules are stated explicitly as 3×33\times38 for channel doubling and 3×33\times39 for spatial halving. The technical overview specifies an input of a 112×112112\times1120 RGB image for the stage-wise layout. The head uses softmax for the 8-way setting and sigmoid for the binary setting (Yurdakul et al., 11 Sep 2025).

This topology is significant because it preserves the hierarchical multi-scale structure of CoAtNet while changing the early-stage operator family. A plausible implication is that the architecture was designed to strengthen texture-sensitive local encoding without removing the long-range modeling mechanisms that hybrid backbones typically exploit.

3. Replacement of MBConv and integration of CBAM

The architectural novelty centers on replacing MBConv with an “Improved ConvNeXtV2” block. The generic MBConv form is given as

112×112112\times1121

where 112×112112\times1122 and 112×112112\times1123 are 112×112112\times1124 expansion and projection, 112×112112\times1125 is a 112×112112\times1126 depthwise convolution, and 112×112112\times1127 is squeeze-and-excitation (Yurdakul et al., 11 Sep 2025).

The ConvNeXtV2 replacement block is described by the following sequence: a large-kernel depthwise convolution, Layer Normalization per channel, two pointwise (112×112112\times1128) convolutions with a GELU activation in between, Global Response Normalization (GRN) to prevent feature collapse, and residual DropPath. The “Improved ConvNeXtV2” structure is specified as

Input 112×112112\times1129 (expansion) D=32D=320 (projection) D=32D=321 Residual (Yurdakul et al., 11 Sep 2025).

For the depthwise convolution, the technical overview gives

D=32D=322

GRN is introduced with channelwise D=32D=323-norm aggregation and learnable scalars D=32D=324. The stated role of GRN is to prevent feature collapse, consistent with the ConvNeXtV2 design lineage (Woo et al., 2023).

CBAM is applied sequentially through channel attention and spatial attention:

D=32D=325

D=32D=326

D=32D=327

D=32D=328

Here, D=32D=329 is the sigmoid, L1=2L_1=20 is a L1=2L_1=21 convolution, and L1=2L_1=22 is element-wise multiplication (Woo et al., 2018). In CoAtNeXt, CBAM is explicitly used to improve local feature extraction through channel and spatial attention mechanisms (Yurdakul et al., 11 Sep 2025).

The significance of this substitution lies in the fact that it changes both the convolutional primitive and the attention mechanism used in the early stages. The ablation data indicate that the gain is not attributable to ConvNeXtV2 alone; additional gains appear when CBAM is added on top of the ConvNeXtV2 replacement.

4. Data, preprocessing, and optimization protocol

Two public datasets are reported. HMU-GC-HE-30K contains 31,096 patches of size L1=2L_1=23 spanning eight classes—ADI, DEB, LYM, MUC, MUS, NOR, STR, and TUM—and is described as balanced at approximately 3,887 samples per class. GasHisSDB is reported in its L1=2L_1=24 subset with 146,651 patches, of which 59,151 are abnormal and 87,500 are normal. HMU-GC-HE-30K is used for eight-class classification and GasHisSDB for binary classification (Yurdakul et al., 11 Sep 2025).

Preprocessing is stated as per-channel normalization for both datasets. No advanced augmentations are reported; the technical overview adds that standard practice may include random flips, rotations, and color jitter. This wording is important: it does not document those augmentations as part of the reported experimental protocol (Yurdakul et al., 11 Sep 2025).

The implementation stack is TensorFlow 2.14.0 and Keras 2.11.4 on Windows 11 with 128 GB RAM, L1=2L_1=25 NVIDIA RTX 3090 (24 GB), and CUDA 12.7. Optimization uses SGD with momentum (L1=2L_1=26) and no explicit weight decay. The learning rate is selected by grid search, with a typical initial value of 0.01 under a constant schedule. The batch size is 32, training proceeds for 50 epochs per fold, and evaluation is performed with 5-fold cross-validation. The reported losses are Categorical Cross-Entropy for the multiclass task and Binary Cross-Entropy for the binary task (Yurdakul et al., 11 Sep 2025).

For reproducibility, the practical notes recommend at least one GPU with at least 24 GB VRAM, report training time of approximately 8–12 hours per 50-epoch run per fold with two GPUs in data-parallel mode, and identify Python 3.9+, TensorFlow 2.x, Keras, and cross-validation scripts as the code stack (Yurdakul et al., 11 Sep 2025).

5. Quantitative performance and ablation behavior

The evaluation protocol uses confusion-matrix-derived metrics: accuracy, precision, recall, F1-score, and AUC. Results are reported as averages over 5 folds (Yurdakul et al., 11 Sep 2025).

Dataset CoAtNeXt metrics Baseline CoAtNet
HMU-GC-HE-30K Acc 96.47%, P 96.60%, R 96.47%, F1 96.45%, AUC 99.89% Acc 93.12%, AUC 99.15%
GasHisSDB Acc 98.29%, P 98.07%, R 98.41%, F1 98.23%, AUC 99.90% Acc 96.66%, AUC 99.51%

The paper states that CoAtNeXt outperformed all CNN and ViT models tested and surpassed previous studies in the literature (Yurdakul et al., 11 Sep 2025). Within the scope of the reported experiments, this places the model above both conventional CNN baselines and Vision Transformer baselines, although the full per-model breakdown is not reproduced in the technical overview.

The ablation study isolates the effect of replacing MBConv and varying attention modules. On HMU-GC-HE-30K, the sequence is CoAtNet (MBConv) at 93.12% accuracy, L1=2L_1=27 ConvNeXtV2 only at 94.50%, L1=2L_1=28 SE attention at 94.75%, L1=2L_1=29 ECA attention at 95.64%, and 56×5656\times560 CBAM, i.e. CoAtNeXt, at 96.47%. On GasHisSDB, the corresponding sequence is 96.66%, 96.73%, 97.46%, 97.68%, and 98.29% (Yurdakul et al., 11 Sep 2025).

These ablations support two bounded conclusions. First, replacing MBConv with ConvNeXtV2 yields a measurable gain on both datasets. Second, among the compared attention augmentations—SE, ECA, and CBAM—the best reported accuracies are obtained with CBAM. A plausible implication is that sequential channel-and-spatial attention is particularly effective in the local feature regime targeted by the early stages of the network.

6. Scope, limitations, and plausible extensions

The reported limitations are explicit. The staining modality is restricted to H&E images, and performance on other stains such as IHC is untested. The eight-class dataset is characterized as relatively new, and more external validation is needed. These constraints mean that the reported results do not establish stain-invariant or broadly cross-institutional performance (Yurdakul et al., 11 Sep 2025).

The stated future directions are to incorporate multi-stain or multi-modal data, such as endoscopy plus histology; to extend the model to segmentation or detection of sub-regions such as gland and nuclei; and to explore adaptive learning-rate schedules such as cosine decay and warm restarts together with stronger augmentations including Mixup and CutMix (Yurdakul et al., 11 Sep 2025). These are natural directions because the present study evaluates only classification, uses a constant learning-rate schedule, and does not report advanced augmentation.

In a broader research context, CoAtNeXt exemplifies a design pattern in which a hybrid backbone is preserved while the local operator family is modernized and attention is made more explicit. The paper’s own evidence is confined to gastric histopathology classification, but it suggests that careful substitution of early-stage convolutional blocks can materially affect downstream performance even when the global transformer structure remains intact (Yurdakul et al., 11 Sep 2025).

Topic to Video (Beta)

No one has generated a video about this topic yet.

Whiteboard

No one has generated a whiteboard explanation for this topic yet.

Follow Topic

Get notified by email when new papers are published related to CoAtNeXt.