- The paper introduces SWNet, a bimodal RGB–NIR encoder–decoder with gated fusion, CBAM attention, deep supervision, and edge refinement for detecting weeds that visually blend with banana plants.
- The fused model achieves a weighted F-measure of 0.8767, MAE of 0.0070, and mean E-measure of 0.9860, outperforming ten camouflaged-object detection baselines on the 272-image Weeds-Banana dataset.
- The results show that cross-spectral information drives most of the gains, while limited crop diversity, 42.32M parameters, sensor requirements, and untested edge deployment constrain real-world generalization.
Motivation and problem setting
SWNet addresses camouflaged weed detection in dense agricultural canopies, where invasive weeds exhibit homochromatic blending with the primary crop (banana) — matching leaf shape, texture, and coloration so closely that visible-spectrum (RGB) methods become unreliable. The paper frames this as an instance of Camouflaged Object Detection (COD) and argues that the physiological differences in chlorophyll reflectance and cellular structure captured in the Near-Infrared (NIR) spectrum provide the discriminating signal that RGB alone cannot supply. The work is evaluated on the Weeds-Banana dataset, a multispectral benchmark of 272 high-resolution (1024 × 1024) RGB/NIR/mask image pairs, and benchmarks against ten COD architectures under standardized training configurations.
Architecture
SWNet follows a bimodal encoder–decoder design with four main components:
- Backbone: PVTv2-B2, pretrained on ImageNet, extracting four-stage multi-scale features with strides of 4–32. The transformer self-attention captures long-range dependencies relevant when the target mimics its immediate surroundings.
- Bimodal Gated Fusion Module: modality-specific gates computed via global average pooling and 1 × 1 convolutions dynamically weight RGB and NIR contributions; the gated features are concatenated, projected, and passed through a CBAM (channel gate with average/max pooling and a shared MLP, followed by a 7 × 7 spatial gate) to suppress sensor-specific noise.
- Decoder with deep supervision: bilinear upsampling, skip-connection aggregation, and convolutional smoothing; four auxiliary segmentation heads produce intermediate masks that are averaged at inference.
- Edge-Aware Refinement: an edge head predicts contours from the final decoder stage, and the final output is computed as Ofinal=Mask×(1+σ(Edge)), sharpening boundary transitions.
Training uses AdamW (LR 1×10−4, weight decay 1×10−4), cosine annealing, 416 × 416 inputs, 200 epochs, batch size 10, on a single RTX 4090. The loss is the F3Net-style structure loss (weighted BCE + weighted IoU) summed over four stages, plus a BCE edge loss against ground-truth edges derived by differencing local max and min pooling of the mask.
Benchmark design
The paper evaluates ten COD methods — SINet-v2, BGNet, C²F-Net, OCENet, EAMNet, DGNet, HitNet, ARNet, CHNet, and ARNet-v2 — spanning Res2Net-50, ResNet-50, EfficientNet, SMT-Tiny, and PVTv2 backbones (8.30M to 77.80M parameters; SWNet itself is 42.32M). Training configurations (optimizer, learning rate, batch size, scheduler, loss) are reported per method to standardize the comparison. Metrics include Sα, weighted F-measure (Fβw), mean absolute error (M), and adaptive/mean/max variants of E-measure and F-measure.
Quantitative results
The central result is that the fused Vis+NIR configuration of SWNet outperforms all ten baselines on most metrics, while its single-modality variants are merely competitive:
| Configuration |
Sα |
Fβw |
M |
Eϕmean |
1×10−40 |
| ARNet (Vis, prior best 1×10−41) |
0.8800 |
0.8131 |
0.0091 |
0.9604 |
0.8492 |
| ARNet-v2 (Vis) |
0.9027 |
0.8229 |
0.0086 |
0.9667 |
0.8425 |
| HitNet (Vis) |
0.8773 |
0.8090 |
0.0088 |
0.9652 |
0.7970 |
| SWNet (Vis only) |
0.7971 |
0.7227 |
0.0122 |
0.9338 |
0.7175 |
| SWNet (NIR only) |
0.8413 |
0.7624 |
0.0108 |
0.9325 |
0.7460 |
| SWNet (Vis+NIR) |
0.8966 |
0.8767 |
0.0070 |
0.9860 |
0.8590 |
Three observations follow directly. First, the cross-spectral fusion yields a gain of roughly +6.4 points in weighted F-measure over the best visible-light baseline (0.8767 vs. 0.8131) and a record-low MAE of 0.0070, which the authors attribute to the Bimodal Gated Fusion Module exploiting NIR chlorophyll-reflectance contrast. Second, the single-modality SWNet results are notably weaker than ARNet-v2's Vis-only performance on 1×10−42 (0.7971 vs. 0.9027), so the fused model's superiority rests almost entirely on the multimodal input rather than on the backbone or decoder design — a point the paper does not explicitly confront. Third, ARNet-v2 retains the best Vis-only 1×10−43 (0.9027) and 1×10−44 (0.8229) among baselines, and SWNet's fused 1×10−45 (0.8966) does not exceed it; the claim that SWNet "exceeds the results of ARNet-v2 across multiple structural and pixel-wise metrics" holds for most metrics but not for 1×10−46.
Qualitatively, the fused model produces masks with fewer false positives and negatives in high-clutter regions where BGNet, HitNet, ARNet-v2, and CHNet show over- or under-segmentation, with the Edge-Aware Refinement module credited for sharp boundary transitions.
Ablation study
Ablations over the two enhancement modules confirm their complementarity:
| Configuration |
1×10−47 |
1×10−48 |
1×10−49 |
| Edge only |
0.8797 |
0.8332 |
0.0071 |
| CBAM only |
0.8714 |
0.8162 |
0.0077 |
| Edge + CBAM |
0.8966 |
0.8767 |
0.0070 |
The combined configuration improves 1×10−40 by roughly 4.4 points over Edge-only and 6.1 points over CBAM-only, supporting the claim that spatial-channel attention and explicit boundary refinement are jointly necessary. The ablation is limited to these two modules; the contribution of the gated fusion mechanism itself, the deep supervision, or the choice of PVTv2 over a CNN backbone is not isolated experimentally.
Limitations and open questions
The paper concedes several constraints. The dataset is small (272 image pairs) and crop-specific (banana), so generalization across crops, growth stages, and illumination regimes is untested. At 42.32M parameters, SWNet is heavier than compact baselines such as DGNet (8.30M), and the authors acknowledge that real-time inference on edge platforms for autonomous agricultural robotics remains unaddressed. The NIR advantage depends on the availability of registered multispectral sensors, which adds hardware cost not analyzed in the paper. Finally, the mechanism-level attribution of gains to the gated fusion module is supported by the Vis-vs-fused comparison but not by a direct ablation against a simpler concatenation-based fusion baseline.
Conclusion
SWNet demonstrates that bimodal RGB–NIR fusion with gated attention and edge-aware refinement substantially improves camouflaged weed segmentation on the Weeds-Banana benchmark, achieving the best reported 1×10−41 (0.8767), MAE (0.0070), and E-measure (0.9860) against ten COD baselines. The evidence indicates that cross-spectral information, rather than architectural novelty alone, drives most of the improvement. The open questions left by the paper concern cross-crop generalization, computational efficiency for field deployment, and component-level validation of the fusion design.