Papers
Topics
Authors
Recent
Search
2000 character limit reached

Wavelet-guided Misalignment-aware Network (WMNet)

Updated 7 July 2026
  • The paper presents a novel WMNet framework that integrates discrete wavelet transforms into a dual-stream YOLOv11 detector to handle cross-modal misalignment.
  • It leverages key components like WU-Net for RGB restoration, SAWF for self-adaptive fusion, and a wavelet-based enhancer to target spatial offsets and resolution discrepancies.
  • Empirical results on DVTOD, DroneVehicle, and M³FD benchmarks demonstrate state-of-the-art mAP performance, underlining its robustness in challenging environments.

Wavelet-guided Misalignment-aware Network (WMNet) is a visible-infrared object detection framework introduced for cross-modal settings in which visible and infrared image pairs are not perfectly registered because of sensor resolution, viewpoint, focal length, capture timing, and modality-specific appearance differences. In this formulation, misalignment is not treated as a preprocessing nuisance but as a central fusion problem. WMNet addresses this problem through wavelet-domain multi-frequency analysis and modality-aware fusion within a dual-stream YOLOv11 detector, with the infrared image used as the reference modality. Its design explicitly targets spatial offset, resolution discrepancy, and modality deficiency, and it was evaluated on DVTOD, DroneVehicle, and M3^3FD, where it achieved state-of-the-art results on misaligned cross-modal detection tasks (Zhang et al., 27 Jul 2025).

1. Misalignment in visible-infrared detection

Visible-infrared object detection seeks to exploit the complementary information of RGB and infrared imagery to improve robustness. In real drone and remote-sensing deployments, however, the two modalities frequently differ in sensor resolution, spatial position, focal characteristics, capture timing, and appearance. Corresponding objects may therefore be shifted, differently scaled, or present in only one modality. Under these conditions, naïve feature fusion can blur boundaries, merge nearby targets, and inject false cues into the detector.

WMNet is organized around the claim that different forms of misalignment manifest differently across frequency bands. Low-frequency components encode coarse structure and global layout, and are relatively stable under scale mismatch. High-frequency components encode edges, boundaries, and fine detail, which are useful for separating nearby objects but are also more sensitive to local displacement and noise. The framework therefore inserts discrete wavelet transform operations directly into the network rather than relying on fusion only in the raw spatial domain.

A common misconception in this area is that effective cross-modal fusion requires approximate pixel-level correspondence before the detector can operate well. WMNet instead adopts asymmetric, reference-guided interaction, using infrared as an anchor and allowing RGB information to enter selectively. This design is intended to avoid the failure mode in which shifted RGB evidence directly corrupts the more stable reference stream.

2. System architecture and design logic

WMNet is built on a dual-stream YOLOv11 backbone and consists of three main components: WU-Net, SAWF, and a wavelet-based correlated feature enhancer. The processing order is explicit. First, WU-Net enhances the RGB stream under infrared guidance. Second, SAWF performs misalignment-aware fusion at multiple backbone stages. Third, correlated feature enhancement reconstructs fused features in the wavelet domain to sharpen object boundaries and suppress ambiguous background interference.

Component Function
WU-Net Wavelet U-Net-style enhancement of RGB guided by infrared
SAWF Self-Adaptive Wavelet Fusion for misalignment-aware multi-scale fusion
Wavelet-based correlated feature enhancer Wavelet-domain reconstruction for contour sharpening and background suppression

This modular division reflects the three error sources emphasized by the method. WU-Net primarily addresses modality-specific degradation such as low light, noise, and texture inconsistency. SAWF handles cross-modal interaction under spatial and resolution mismatch. The correlated feature enhancer targets boundary ambiguity and false merging after fusion. A plausible implication is that the architecture is designed to distribute robustness across restoration, interaction, and reconstruction rather than relying on a single alignment mechanism.

3. WU-Net and frequency-aware RGB restoration

WU-Net is a wavelet U-Net-style enhancement module that improves the visible branch before cross-modal fusion. For an input image I∈RH×W×C\mathbf{I} \in \mathbb{R}^{H \times W \times C}, the module repeatedly applies the discrete wavelet transform, recursively propagating the low-frequency stream while retaining the three remaining subbands as high-frequency information:

F0,LL=I,Fk,LL,Fk,H=Split(WT(Fk−1,LL)).\mathbf{F}_{0,LL} = \mathbf{I}, \quad \mathbf{F}_{k,LL}, \mathbf{F}_{k,H} = \mathrm{Split}\left(\mathrm{WT}\left(\mathbf{F}_{k-1,LL}\right)\right).

At the deepest level, WU-Net performs frequency-aware attention on the decomposed RGB and infrared features. Low-frequency and high-frequency features from both modalities are concatenated and used to construct an attention weight explicitly modulated by coarse and fine relations:

WA=Softmax(qLLkLL⊤d)⊙[1+Softmax(qHkH⊤d)].\mathbf{W}_A = \mathrm{Softmax}\left(\frac{\mathbf{q}_{LL}\mathbf{k}_{LL}^\top}{\sqrt{d}}\right) \odot \left[ 1+\mathrm{Softmax}\left(\frac{\mathbf{q}_H\mathbf{k}_H^\top}{\sqrt{d}}\right) \right].

In this mechanism, low-frequency attention captures coarse semantic alignment, while the high-frequency term adjusts it with local detail sensitivity. The decoder then reconstructs the enhanced RGB image progressively via inverse wavelet transform, and the final output is formed with a residual scaled-convolution term.

The practical effect reported for WU-Net is that the RGB branch becomes more robust to illumination changes, noise, and weak textures, while infrared injects reliable structural cues. The ablation analysis attributes major gains to this module alone, indicating that pre-fusion enhancement materially improves downstream detection.

4. SAWF, asymmetric fusion, and correlated feature enhancement

The central fusion mechanism in WMNet is SAWF, or Self-Adaptive Wavelet Fusion. SAWF is inserted into the multi-scale feature fusion stage and is designed to treat fusion and misalignment jointly rather than sequentially.

The first step is asymmetric, reference-guided interaction. RGB features are wavelet-decomposed and modulated by a learnable channel-wise weight W1\mathbf{W}_1, initialized to $0.1$:

XrgbLL,XrgbH=Split(W1⊙Conv3×3(WT(Xrgb))).\mathbf{X}_{rgb_{LL}}, \mathbf{X}_{rgb_H} = \mathrm{Split}\left( \mathbf{W}_1 \odot \mathrm{Conv}_{3\times3}\big(\mathrm{WT}(\mathbf{X}_{rgb})\big) \right).

This functions as a modality-decoupling mechanism: potentially unstable RGB evidence is attenuated during the most sensitive part of the fusion process, while useful complementary content remains available.

The core interaction engine is the Cross-Modality Fusion Mamba (CFM) module, which uses a Mamba2 state-space block rather than a transformer. Its state-space dynamics are written as

h(t)=Ath(t−1)+Btxt,y(t)=Cth(t)+Dtxt.h(t) = \mathbf{A}_t h(t-1) + \mathbf{B}_t x_t, \qquad y(t) = \mathbf{C}_t h(t) + \mathbf{D}_t x_t.

To operate on 2D features, feature maps are pooled to 16×1616 \times 16, flattened into sequences, concatenated, and then processed by Mamba:

Xf=mamba(Concat(Xir,XrgbLL)),Xir′,Xrgb′=Split(Xf).\mathbf{X}_f = \mathrm{mamba}\big(\mathrm{Concat}(\mathbf{X}_{ir}, \mathbf{X}_{rgb_{LL}})\big), \qquad \mathbf{X}_{ir}', \mathbf{X}_{rgb}' = \mathrm{Split}(\mathbf{X}_f).

This reference-guided asymmetry is intended to be more robust than fusion methods that assume dense spatial correspondence.

SAWF then applies an outer correlated feature enhancement path in the wavelet domain. Updated RGB low-frequency features and retained RGB frequency components are reconstructed with inverse wavelet transform:

I∈RH×W×C\mathbf{I} \in \mathbb{R}^{H \times W \times C}0

A weighted skip connection preserves modality-specific detail:

I∈RH×W×C\mathbf{I} \in \mathbb{R}^{H \times W \times C}1

In qualitative examples, this path is reported to be particularly effective for pedestrians and small vehicles, where misalignment often causes merged detections or foreground-background confusion (Zhang et al., 27 Jul 2025).

5. Detection pipeline, training protocol, and datasets

After SAWF, each modality stream is refined by shared lightweight layers,

I∈RH×W×C\mathbf{I} \in \mathbb{R}^{H \times W \times C}2

and then fused by a learnable Scaled Add strategy:

I∈RH×W×C\mathbf{I} \in \mathbb{R}^{H \times W \times C}3

The modality-specific outputs are added,

I∈RH×W×C\mathbf{I} \in \mathbb{R}^{H \times W \times C}4

and the latter three fused pyramid features I∈RH×W×C\mathbf{I} \in \mathbb{R}^{H \times W \times C}5 are sent to the YOLO detection head for multi-scale prediction. The detector uses a standard multi-scale YOLO-style head and is trained with the usual detection losses under the YOLO framework.

The evaluation protocol spans three public datasets representing different alignment regimes:

Dataset Regime Resolution
DVTOD Drone-based, heavily misaligned, large RGB-IR offsets RGB I∈RH×W×C\mathbf{I} \in \mathbb{R}^{H \times W \times C}6, IR I∈RH×W×C\mathbf{I} \in \mathbb{R}^{H \times W \times C}7
DroneVehicle Smaller misalignment and modality inconsistency I∈RH×W×C\mathbf{I} \in \mathbb{R}^{H \times W \times C}8
MI∈RH×W×C\mathbf{I} \in \mathbb{R}^{H \times W \times C}9FD More generally aligned, multi-scene F0,LL=I,Fk,LL,Fk,H=Split(WT(Fk−1,LL)).\mathbf{F}_{0,LL} = \mathbf{I}, \quad \mathbf{F}_{k,LL}, \mathbf{F}_{k,H} = \mathrm{Split}\left(\mathrm{WT}\left(\mathbf{F}_{k-1,LL}\right)\right).0

The reported metrics are F0,LL=I,Fk,LL,Fk,H=Split(WT(Fk−1,LL)).\mathbf{F}_{0,LL} = \mathbf{I}, \quad \mathbf{F}_{k,LL}, \mathbf{F}_{k,H} = \mathrm{Split}\left(\mathrm{WT}\left(\mathbf{F}_{k-1,LL}\right)\right).1 and COCO-style F0,LL=I,Fk,LL,Fk,H=Split(WT(Fk−1,LL)).\mathbf{F}_{0,LL} = \mathbf{I}, \quad \mathbf{F}_{k,LL}, \mathbf{F}_{k,H} = \mathrm{Split}\left(\mathrm{WT}\left(\mathbf{F}_{k-1,LL}\right)\right).2 averaged over IoU thresholds from F0,LL=I,Fk,LL,Fk,H=Split(WT(Fk−1,LL)).\mathbf{F}_{0,LL} = \mathbf{I}, \quad \mathbf{F}_{k,LL}, \mathbf{F}_{k,H} = \mathrm{Split}\left(\mathrm{WT}\left(\mathbf{F}_{k-1,LL}\right)\right).3 to F0,LL=I,Fk,LL,Fk,H=Split(WT(Fk−1,LL)).\mathbf{F}_{0,LL} = \mathbf{I}, \quad \mathbf{F}_{k,LL}, \mathbf{F}_{k,H} = \mathrm{Split}\left(\mathrm{WT}\left(\mathbf{F}_{k-1,LL}\right)\right).4 in steps of F0,LL=I,Fk,LL,Fk,H=Split(WT(Fk−1,LL)).\mathbf{F}_{0,LL} = \mathbf{I}, \quad \mathbf{F}_{k,LL}, \mathbf{F}_{k,H} = \mathrm{Split}\left(\mathrm{WT}\left(\mathbf{F}_{k-1,LL}\right)\right).5. The implementation uses a modified CFT-style dual-stream design based on YOLOv11, SGD, 300 epochs for DVTOD and MF0,LL=I,Fk,LL,Fk,H=Split(WT(Fk−1,LL)).\mathbf{F}_{0,LL} = \mathbf{I}, \quad \mathbf{F}_{k,LL}, \mathbf{F}_{k,H} = \mathrm{Split}\left(\mathrm{WT}\left(\mathbf{F}_{k-1,LL}\right)\right).6FD, 150 epochs for DroneVehicle, batch size 4, input size F0,LL=I,Fk,LL,Fk,H=Split(WT(Fk−1,LL)).\mathbf{F}_{0,LL} = \mathbf{I}, \quad \mathbf{F}_{k,LL}, \mathbf{F}_{k,H} = \mathrm{Split}\left(\mathrm{WT}\left(\mathbf{F}_{k-1,LL}\right)\right).7, NVIDIA GeForce RTX 4080S hardware, and no pretrained weights.

6. Empirical performance and ablation findings

WMNet achieves the best reported results on all three datasets:

Dataset F0,LL=I,Fk,LL,Fk,H=Split(WT(Fk−1,LL)).\mathbf{F}_{0,LL} = \mathbf{I}, \quad \mathbf{F}_{k,LL}, \mathbf{F}_{k,H} = \mathrm{Split}\left(\mathrm{WT}\left(\mathbf{F}_{k-1,LL}\right)\right).8 F0,LL=I,Fk,LL,Fk,H=Split(WT(Fk−1,LL)).\mathbf{F}_{0,LL} = \mathbf{I}, \quad \mathbf{F}_{k,LL}, \mathbf{F}_{k,H} = \mathrm{Split}\left(\mathrm{WT}\left(\mathbf{F}_{k-1,LL}\right)\right).9
DVTOD 86.3 50.1
DroneVehicle 85.8 63.1
MWA=Softmax(qLLkLL⊤d)⊙[1+Softmax(qHkH⊤d)].\mathbf{W}_A = \mathrm{Softmax}\left(\frac{\mathbf{q}_{LL}\mathbf{k}_{LL}^\top}{\sqrt{d}}\right) \odot \left[ 1+\mathrm{Softmax}\left(\frac{\mathbf{q}_H\mathbf{k}_H^\top}{\sqrt{d}}\right) \right].0FD 88.3 58.5

On DVTOD, described as the hardest benchmark, WMNet outperforms CFT, CMA-Det, and ICAFusion. The lightweight variant WMNet-Lite is also reported to perform strongly with only 17.58M parameters. On DroneVehicle, WMNet surpasses SuperYOLO, GHOST, ICA-Fusion, GM-DETR, and CMA-Det while remaining relatively compact at 17.58M parameters. On MWA=Softmax(qLLkLL⊤d)⊙[1+Softmax(qHkH⊤d)].\mathbf{W}_A = \mathrm{Softmax}\left(\frac{\mathbf{q}_{LL}\mathbf{k}_{LL}^\top}{\sqrt{d}}\right) \odot \left[ 1+\mathrm{Softmax}\left(\frac{\mathbf{q}_H\mathbf{k}_H^\top}{\sqrt{d}}\right) \right].1FD, which is more aligned, WMNet still leads, and this is used to argue that the method is not over-specialized to severe misalignment.

The ablation study highlights four points. First, WU-Net alone contributes major gains by improving the RGB modality before fusion. Second, SAWF alone is not intended to replace the full fusion backbone. Third, CFM is parameter-efficient, and replacing a transformer fusion block with CFM significantly reduces parameters. Fourth, the best DVTOD result arises from using WU-Net, SAWF, and CFM together, reaching 86.3 WA=Softmax(qLLkLL⊤d)⊙[1+Softmax(qHkH⊤d)].\mathbf{W}_A = \mathrm{Softmax}\left(\frac{\mathbf{q}_{LL}\mathbf{k}_{LL}^\top}{\sqrt{d}}\right) \odot \left[ 1+\mathrm{Softmax}\left(\frac{\mathbf{q}_H\mathbf{k}_H^\top}{\sqrt{d}}\right) \right].2 and 50.1 WA=Softmax(qLLkLL⊤d)⊙[1+Softmax(qHkH⊤d)].\mathbf{W}_A = \mathrm{Softmax}\left(\frac{\mathbf{q}_{LL}\mathbf{k}_{LL}^\top}{\sqrt{d}}\right) \odot \left[ 1+\mathrm{Softmax}\left(\frac{\mathbf{q}_H\mathbf{k}_H^\top}{\sqrt{d}}\right) \right].3. Weight analysis for WA=Softmax(qLLkLL⊤d)⊙[1+Softmax(qHkH⊤d)].\mathbf{W}_A = \mathrm{Softmax}\left(\frac{\mathbf{q}_{LL}\mathbf{k}_{LL}^\top}{\sqrt{d}}\right) \odot \left[ 1+\mathrm{Softmax}\left(\frac{\mathbf{q}_H\mathbf{k}_H^\top}{\sqrt{d}}\right) \right].4 and WA=Softmax(qLLkLL⊤d)⊙[1+Softmax(qHkH⊤d)].\mathbf{W}_A = \mathrm{Softmax}\left(\frac{\mathbf{q}_{LL}\mathbf{k}_{LL}^\top}{\sqrt{d}}\right) \odot \left[ 1+\mathrm{Softmax}\left(\frac{\mathbf{q}_H\mathbf{k}_H^\top}{\sqrt{d}}\right) \right].5 further shows that reducing RGB dominance in the inner fusion path is beneficial under severe misalignment; the best DVTOD setting is reported as WA=Softmax(qLLkLL⊤d)⊙[1+Softmax(qHkH⊤d)].\mathbf{W}_A = \mathrm{Softmax}\left(\frac{\mathbf{q}_{LL}\mathbf{k}_{LL}^\top}{\sqrt{d}}\right) \odot \left[ 1+\mathrm{Softmax}\left(\frac{\mathbf{q}_H\mathbf{k}_H^\top}{\sqrt{d}}\right) \right].6 when WA=Softmax(qLLkLL⊤d)⊙[1+Softmax(qHkH⊤d)].\mathbf{W}_A = \mathrm{Softmax}\left(\frac{\mathbf{q}_{LL}\mathbf{k}_{LL}^\top}{\sqrt{d}}\right) \odot \left[ 1+\mathrm{Softmax}\left(\frac{\mathbf{q}_H\mathbf{k}_H^\top}{\sqrt{d}}\right) \right].7 (Zhang et al., 27 Jul 2025).

7. Practical significance, limitations, and acronym ambiguity

WMNet is presented as a unified, lightweight, and generalizable solution for visible-infrared detection under real-world misalignment. Its practical significance lies in integrating alignment-awareness directly into the detector rather than assuming perfect registration or relying on expensive alignment preprocessing. The wavelet-domain design provides a structured way to separate coarse layout, boundary detail, and noise, which is particularly relevant in drone and remote-sensing scenarios.

The limitations stated or implied in the source are comparatively narrow but important. The framework is specialized to visible-infrared detection, relies on carefully designed modality guidance and wavelet operations, and leaves room for extension to more modalities and improved edge deployment. This suggests that the current formulation is best understood as a targeted solution for RGB-IR fusion rather than a modality-agnostic architecture.

The acronym “WMNet” is also used by an unrelated HDR video reconstruction model, titled “Wavelet-Domain Masked Image Modeling for Color-Consistent HDR Video Reconstruction,” which combines W-MIM, T-MoE, and DMM for LDR-to-HDR video processing (Zhang et al., 7 Feb 2026). This suggests that the acronym alone is not a reliable identifier across subfields. In the context of visible-infrared object detection, however, WMNet specifically denotes the Wavelet-guided Misalignment-aware Network and its central methodological lesson: misalignment-aware fusion benefits from modeling frequency structure, modality asymmetry, and selective interaction jointly.

Topic to Video (Beta)

No one has generated a video about this topic yet.

Whiteboard

No one has generated a whiteboard explanation for this topic yet.

Follow Topic

Get notified by email when new papers are published related to Wavelet-guided Misalignment-aware Network (WMNet).