Papers
Topics
Authors
Recent
Search
2000 character limit reached

AAW-YOLO: Wavelet & Attention for Artery Segmentation

Updated 3 July 2026
  • The paper presents a novel deep learning architecture that integrates wavelet transforms and attention modules to achieve high-precision cerebral artery segmentation from TCCD ultrasound images.
  • It employs a YOLO-based detection framework enhanced with wavelet-based convolution blocks for multi-scale feature extraction and a linear attention module to refine small, low-contrast vessel features.
  • Model evaluation on a prospectively collected TCCD dataset shows outstanding performance with a Dice of 0.901 and mAP of 0.953, and an inference speed of ~70 FPS, making it suitable for clinical deployment.

Attention-Augmented Wavelet YOLO (AAW-YOLO) is a real-time deep learning system for automated segmentation and detection of cerebral arteries within the Circle of Willis (CoW) from transcranial color-coded Doppler (TCCD) ultrasound. Designed specifically for cerebrovascular analysis, AAW-YOLO integrates wavelet-domain convolutional operations and attention mechanisms into a lightweight YOLO-based detection architecture, achieving high-precision instance segmentation with rapid inference suitable for clinical deployment. It is the first reported approach for AI-driven CoW segmentation on TCCD, addressing challenges in operator-dependent landmark identification and angle correction inherent to ultrasound-based neuroimaging (Zhang et al., 19 Aug 2025).

1. Network Architecture and Wavelet Integration

The AAW-YOLO architecture builds on a YOLO-11 backbone composed of convolutional (Conv) layers and C3K2 bottleneck blocks. To increase the effective receptive field efficiently, AAW-YOLO replaces selected bottlenecks with a wavelet-aided block (WTC2f). This block decomposes input feature maps X∈RC×H×W\mathbf{X} \in \mathbb{R}^{C \times H \times W} using a one-level 2D discrete wavelet transform (DWT) into four subbands:

(XLL,XLH,XHL,XHH)=DWT(X)(\mathbf{X}_{LL},\mathbf{X}_{LH},\mathbf{X}_{HL},\mathbf{X}_{HH}) = \mathrm{DWT}(\mathbf{X})

with:

  • XLL:=(X∗ϕ2D)↓2\mathbf{X}_{LL}:=(\mathbf{X} * \phi_{2D})\downarrow_2
  • XLH:=(X∗ϕh⊗ψv)↓2\mathbf{X}_{LH}:=(\mathbf{X} * \phi_h \otimes \psi_v)\downarrow_2
  • XHL:=(X∗ψh⊗ϕv)↓2\mathbf{X}_{HL}:=(\mathbf{X} * \psi_h \otimes \phi_v)\downarrow_2
  • XHH:=(X∗ψ2D)↓2\mathbf{X}_{HH}:=(\mathbf{X} * \psi_{2D})\downarrow_2

Here, ϕ\phi and ψ\psi represent the low/high-pass Haar wavelet kernels and ↓2\downarrow_2 denotes 2× downsampling. Each subband undergoes channel-wise convolution followed by feature reconstruction via the inverse DWT (IDWT\mathrm{IDWT}). This mechanism introduces multi-scale context, facilitates small vessel recognition, and delivers minimal increase in computational cost.

2. Attention-Augmented Bottleneck

To improve refinement of fine vascular features, especially those associated with small and low-contrast (contralateral) vessels, AAW-YOLO incorporates a linear attention module within a C2F-style bottleneck. The input tensor (XLL,XLH,XHL,XHH)=DWT(X)(\mathbf{X}_{LL},\mathbf{X}_{LH},\mathbf{X}_{HL},\mathbf{X}_{HH}) = \mathrm{DWT}(\mathbf{X})0 is split by channel. For the second half, query, key, and value projections are computed:

(XLL,XLH,XHL,XHH)=DWT(X)(\mathbf{X}_{LL},\mathbf{X}_{LH},\mathbf{X}_{HL},\mathbf{X}_{HH}) = \mathrm{DWT}(\mathbf{X})1

where (XLL,XLH,XHL,XHH)=DWT(X)(\mathbf{X}_{LL},\mathbf{X}_{LH},\mathbf{X}_{HL},\mathbf{X}_{HH}) = \mathrm{DWT}(\mathbf{X})2 flattens spatial dimensions to a sequence of length (XLL,XLH,XHL,XHH)=DWT(X)(\mathbf{X}_{LL},\mathbf{X}_{LH},\mathbf{X}_{HL},\mathbf{X}_{HH}) = \mathrm{DWT}(\mathbf{X})3. Linear attention is performed:

(XLL,XLH,XHL,XHH)=DWT(X)(\mathbf{X}_{LL},\mathbf{X}_{LH},\mathbf{X}_{HL},\mathbf{X}_{HH}) = \mathrm{DWT}(\mathbf{X})4

with (XLL,XLH,XHL,XHH)=DWT(X)(\mathbf{X}_{LL},\mathbf{X}_{LH},\mathbf{X}_{HL},\mathbf{X}_{HH}) = \mathrm{DWT}(\mathbf{X})5 as softmax or other normalization. The resultant attended features are reshaped, processed by a lightweight bottleneck, and concatenated with the unchanged first half of the channels, yielding more discriminative features for segmentation and detection.

3. Multi-Task Training Objective

AAW-YOLO's training objective jointly optimizes detection and instance segmentation by combining four loss components:

(XLL,XLH,XHL,XHH)=DWT(X)(\mathbf{X}_{LL},\mathbf{X}_{LH},\mathbf{X}_{HL},\mathbf{X}_{HH}) = \mathrm{DWT}(\mathbf{X})6

  • Bounding box regression uses mean-squared error on object center, width, and height:

(XLL,XLH,XHL,XHH)=DWT(X)(\mathbf{X}_{LL},\mathbf{X}_{LH},\mathbf{X}_{HL},\mathbf{X}_{HH}) = \mathrm{DWT}(\mathbf{X})7

  • Objectness is trained via binary cross-entropy:

(XLL,XLH,XHL,XHH)=DWT(X)(\mathbf{X}_{LL},\mathbf{X}_{LH},\mathbf{X}_{HL},\mathbf{X}_{HH}) = \mathrm{DWT}(\mathbf{X})8

  • Classification employs multi-class cross-entropy:

(XLL,XLH,XHL,XHH)=DWT(X)(\mathbf{X}_{LL},\mathbf{X}_{LH},\mathbf{X}_{HL},\mathbf{X}_{HH}) = \mathrm{DWT}(\mathbf{X})9

  • Segmentation is guided by Dice loss:

XLL:=(X∗ϕ2D)↓2\mathbf{X}_{LL}:=(\mathbf{X} * \phi_{2D})\downarrow_20

This design enables simultaneous instance detection and fine-grained vascular mask prediction.

4. Dataset Composition and Preprocessing Protocols

AAW-YOLO was trained and evaluated on a prospectively collected TCCD dataset comprising 29 videos (left and right insonation) from 15 subjects. Manual annotation yielded 738 frames and 3,419 labeled artery instances, with vessel classes including ipsilateral/contralateral ACA_A1, ACA_A2, MCA_M1, PCA_P1, and PCA_P2. Annotations adhered to established cerebrovascular segmentation criteria. Image-level preprocessing consisted of normalization; data augmentation involved random flip, minor rotation, and intensity/contrast jittering to promote robustness to scanning angle and appearance variation. No contrast agent was administered, and data originated from a single ultrasound platform, with all images manually segmented.

5. Performance Metrics and Experimental Outcomes

Model evaluation leveraged multiple metrics:

  • Dice coefficient:

XLL:=(X∗ϕ2D)↓2\mathbf{X}_{LL}:=(\mathbf{X} * \phi_{2D})\downarrow_21

XLL:=(X∗ϕ2D)↓2\mathbf{X}_{LL}:=(\mathbf{X} * \phi_{2D})\downarrow_22

  • Pixel-level Precision and Recall:

XLL:=(X∗ϕ2D)↓2\mathbf{X}_{LL}:=(\mathbf{X} * \phi_{2D})\downarrow_23

XLL:=(X∗ϕ2D)↓2\mathbf{X}_{LL}:=(\mathbf{X} * \phi_{2D})\downarrow_24

On the held-out test set, AAW-YOLO achieved:

  • Dice: XLL:=(X∗ϕ2D)↓2\mathbf{X}_{LL}:=(\mathbf{X} * \phi_{2D})\downarrow_25
  • IoU: XLL:=(X∗ϕ2D)↓2\mathbf{X}_{LL}:=(\mathbf{X} * \phi_{2D})\downarrow_26
  • Precision: XLL:=(X∗ϕ2D)↓2\mathbf{X}_{LL}:=(\mathbf{X} * \phi_{2D})\downarrow_27
  • Recall: XLL:=(X∗ϕ2D)↓2\mathbf{X}_{LL}:=(\mathbf{X} * \phi_{2D})\downarrow_28
  • mAP: XLL:=(X∗ϕ2D)↓2\mathbf{X}_{LL}:=(\mathbf{X} * \phi_{2D})\downarrow_29

Inference time per XLH:=(X∗ϕh⊗ψv)↓2\mathbf{X}_{LH}:=(\mathbf{X} * \phi_h \otimes \psi_v)\downarrow_20 frame was XLH:=(X∗ϕh⊗ψv)↓2\mathbf{X}_{LH}:=(\mathbf{X} * \phi_h \otimes \psi_v)\downarrow_21 ms on RTX 4070 GPUs (≈XLH:=(X∗ϕh⊗ψv)↓2\mathbf{X}_{LH}:=(\mathbf{X} * \phi_h \otimes \psi_v)\downarrow_22 FPS), surpassing clinical real-time requirements (20–50 FPS).

6. Ablation Study and Comparative Evaluation

Ablation experiments compared variants: | Model | Dice | mAP | |-----------------------------|-------|-------| | YOLO-11 Baseline | 0.860 | 0.922 | | Baseline + WTConv | 0.870 | 0.930 | | AA-YOLO (attention only) | 0.880 | 0.941 | | AA-YOLO + WTConv | 0.895 | 0.949 | | AAW-YOLO (attention + wavelet) | 0.901 | 0.953 |

Integrating WTConv improved Dice and mAP over the baseline. The addition of attention (AA-YOLO) further boosted both metrics. Full AAW-YOLO yielded the highest results. Subgroup analysis demonstrated the difficulty of segmenting contralateral arteries, but AAW-YOLO reduced the ipsilateral/contralateral Dice gap to XLH:=(X∗ϕh⊗ψv)↓2\mathbf{X}_{LH}:=(\mathbf{X} * \phi_h \otimes \psi_v)\downarrow_23 compared to baseline (XLH:=(X∗ϕh⊗ψv)↓2\mathbf{X}_{LH}:=(\mathbf{X} * \phi_h \otimes \psi_v)\downarrow_24), consistent with improvements for small, low-contrast vessels.

7. Strengths, Limitations, and Future Prospects

AAW-YOLO demonstrates several strengths:

  • First real-time deep learning solution for CoW segmentation in TCCD combining detection and mask prediction.
  • Network parameters: XLH:=(X∗ϕh⊗ψv)↓2\mathbf{X}_{LH}:=(\mathbf{X} * \phi_h \otimes \psi_v)\downarrow_25M, computational cost: XLH:=(X∗ϕh⊗ψv)↓2\mathbf{X}_{LH}:=(\mathbf{X} * \phi_h \otimes \psi_v)\downarrow_26 GFLOPs, and operational throughput suitable for clinical use.
  • Enhanced detection of small vessels via wavelet and attention integration.

Limitations:

  • Analysis is single-frame; does not leverage temporal (video) consistency.
  • Unilateral analysis precludes bilateral anatomical referencing.
  • No contrast enhancement for fine vessel visibility.
  • Entire dataset from a single clinical site/platform.

Future research directions include integration of video-based temporal modeling (e.g. Space-Time Memory networks), bilateral vessel modeling, exploration of contrast-enhanced TCCD, and multi-center, multi-platform validation.

Reference: "A Novel Attention-Augmented Wavelet YOLO System for Real-time Brain Vessel Segmentation on Transcranial Color-coded Doppler" (Zhang et al., 19 Aug 2025)

Definition Search Book Streamline Icon: https://streamlinehq.com
References (1)

Topic to Video (Beta)

No one has generated a video about this topic yet.

Whiteboard

No one has generated a whiteboard explanation for this topic yet.

Follow Topic

Get notified by email when new papers are published related to Attention-Augmented Wavelet YOLO (AAW-YOLO).