---
title: 'AAW-YOLO: Wavelet & Attention for Artery Segmentation'
url: https://www.emergentmind.com/topics/attention-augmented-wavelet-yolo-aaw-yolo
type: topic
---

# AAW-YOLO: Wavelet & Attention for Artery Segmentation

Attention-Augmented Wavelet YOLO (AAW-YOLO) is a real-time deep learning system for automated segmentation and detection of cerebral arteries within the Circle of Willis (CoW) from transcranial color-coded Doppler (TCCD) ultrasound. Designed specifically for cerebrovascular analysis, AAW-YOLO integrates wavelet-domain convolutional operations and attention mechanisms into a lightweight YOLO-based detection architecture, achieving high-precision instance segmentation with rapid inference suitable for clinical deployment. It is the first reported approach for AI-driven CoW segmentation on TCCD, addressing challenges in operator-dependent landmark identification and angle correction inherent to ultrasound-based neuroimaging [2508.13875].

## 1. Network Architecture and Wavelet Integration

The AAW-YOLO architecture builds on a YOLO-11 backbone composed of convolutional (Conv) layers and C3K2 bottleneck blocks. To increase the effective receptive field efficiently, AAW-YOLO replaces selected bottlenecks with a wavelet-aided block (WTC2f). This block decomposes input feature maps $\mathbf{X} \in \mathbb{R}^{C \times H \times W}$ using a one-level 2D discrete wavelet transform (DWT) into four subbands:
$$(\mathbf{X}_{LL},\mathbf{X}_{LH},\mathbf{X}_{HL},\mathbf{X}_{HH}) = \mathrm{DWT}(\mathbf{X})$$
with:
- $\mathbf{X}_{LL}:=(\mathbf{X} * \phi_{2D})\downarrow_2$
- $\mathbf{X}_{LH}:=(\mathbf{X} * \phi_h \otimes \psi_v)\downarrow_2$
- $\mathbf{X}_{HL}:=(\mathbf{X} * \psi_h \otimes \phi_v)\downarrow_2$
- $\mathbf{X}_{HH}:=(\mathbf{X} * \psi_{2D})\downarrow_2$

Here, $\phi$ and $\psi$ represent the low/high-pass Haar wavelet kernels and $\downarrow_2$ denotes 2× downsampling. Each subband undergoes channel-wise convolution followed by feature reconstruction via the inverse DWT ($\mathrm{IDWT}$). This mechanism introduces multi-scale context, facilitates small vessel recognition, and delivers minimal increase in computational cost.

## 2. Attention-Augmented Bottleneck

To improve refinement of fine vascular features, especially those associated with small and low-contrast (contralateral) vessels, AAW-YOLO incorporates a linear attention module within a C2F-style bottleneck. The input tensor $\mathbf{F} \in \mathbb{R}^{C \times H \times W}$ is split by channel. For the second half, query, key, and value projections are computed:
$$
\mathbf{Q}=W_Q\,\mathcal{R}(\mathbf{F}_2), \quad
\mathbf{K}=W_K\,\mathcal{R}(\mathbf{F}_2), \quad
\mathbf{V}=W_V\,\mathcal{R}(\mathbf{F}_2)
$$
where $\mathcal{R}$ flattens spatial dimensions to a sequence of length $H \times W$. Linear attention is performed:
$$
\mathbf{A} = \sigma(\mathbf{Q}\mathbf{K}^\top)\mathbf{V}
$$
with $\sigma$ as softmax or other normalization. The resultant attended features are reshaped, processed by a lightweight bottleneck, and concatenated with the unchanged first half of the channels, yielding more discriminative features for segmentation and detection.

## 3. Multi-Task Training Objective

AAW-YOLO's training objective jointly optimizes detection and instance segmentation by combining four loss components:
$$
\mathcal{L} = \mathcal{L}_\mathrm{box} + \mathcal{L}_\mathrm{obj} + \mathcal{L}_\mathrm{cls} + \lambda_\mathrm{dice} \mathcal{L}_\mathrm{dice}
$$
- Bounding box regression uses mean-squared error on object center, width, and height:
  $$
  \mathcal{L}_{\mathrm{box}} = \sum_{i\in\Omega}\|\mathbf{b}_i - \mathbf{b}_i^*\|^2
  $$
- Objectness is trained via binary cross-entropy:
  $$
  \mathcal{L}_\mathrm{obj} = -\sum_{i}[o_i \log p_i + (1-o_i)\log(1-p_i)]
  $$
- Classification employs multi-class cross-entropy:
  $$
  \mathcal{L}_\mathrm{cls} = -\sum_{i,c} t_{i,c}\log s_{i,c}
  $$
- Segmentation is guided by Dice loss:
  $$
  \mathcal{L}_\mathrm{dice} = 1 - \frac{2\sum_{p}\hat{M}_p M_p + \epsilon}{\sum_{p}\hat{M}_p + \sum_{p}M_p + \epsilon}
  $$

This design enables simultaneous instance detection and fine-grained vascular mask prediction.

## 4. Dataset Composition and Preprocessing Protocols

AAW-YOLO was trained and evaluated on a prospectively collected TCCD dataset comprising 29 videos (left and right insonation) from 15 subjects. Manual annotation yielded 738 frames and 3,419 labeled artery instances, with vessel classes including ipsilateral/contralateral ACA_A1, ACA_A2, MCA_M1, PCA_P1, and PCA_P2. Annotations adhered to established cerebrovascular segmentation criteria. Image-level preprocessing consisted of normalization; data augmentation involved random flip, minor rotation, and intensity/contrast jittering to promote robustness to scanning angle and appearance variation. No contrast agent was administered, and data originated from a single ultrasound platform, with all images manually segmented.

## 5. Performance Metrics and Experimental Outcomes

Model evaluation leveraged multiple metrics:
- Dice coefficient:
  $$
  \mathrm{Dice} = \frac{2\,\mathrm{TP}}{2\,\mathrm{TP} + \mathrm{FP} + \mathrm{FN}}
  $$
- Intersection over Union (IoU):
  $$
  \mathrm{IoU} = \frac{\mathrm{TP}}{\mathrm{TP} + \mathrm{FP} + \mathrm{FN}}
  $$
- Pixel-level Precision and Recall:
  $$
  \mathrm{Precision} = \frac{\mathrm{TP}}{\mathrm{TP} + \mathrm{FP}}, \quad
  \mathrm{Recall} = \frac{\mathrm{TP}}{\mathrm{TP} + \mathrm{FN}}
  $$
- Mean average precision (mAP) over predicted classes:
  $$
  \mathrm{mAP} = \frac{1}{C}\sum_{i=1}^C \mathrm{AP}_i
  $$

On the held-out test set, AAW-YOLO achieved:
- Dice: $0.901$
- IoU: $0.823$
- Precision: $0.882$
- Recall: $0.926$
- mAP: $0.953$

Inference time per $640\times640$ frame was $14.199$ ms on RTX 4070 GPUs (≈$70.4$ FPS), surpassing clinical real-time requirements (20–50 FPS).

## 6. Ablation Study and Comparative Evaluation

Ablation experiments compared variants:
| Model                       | Dice  | mAP   |
|-----------------------------|-------|-------|
| YOLO-11 Baseline            | 0.860 | 0.922 |
| Baseline + WTConv           | 0.870 | 0.930 |
| AA-YOLO (attention only)    | 0.880 | 0.941 |
| AA-YOLO + WTConv            | 0.895 | 0.949 |
| AAW-YOLO (attention + wavelet) | 0.901 | 0.953 |

Integrating WTConv improved Dice and mAP over the baseline. The addition of attention (AA-YOLO) further boosted both metrics. Full AAW-YOLO yielded the highest results. Subgroup analysis demonstrated the difficulty of segmenting contralateral arteries, but AAW-YOLO reduced the ipsilateral/contralateral Dice gap to $0.026$ compared to baseline ($0.048$), consistent with improvements for small, low-contrast vessels.

## 7. Strengths, Limitations, and Future Prospects

AAW-YOLO demonstrates several strengths:
- First real-time deep learning solution for CoW segmentation in TCCD combining detection and mask prediction.
- Network parameters: $2.76$M, computational cost: $10.4$ GFLOPs, and operational throughput suitable for clinical use.
- Enhanced detection of small vessels via wavelet and attention integration.

Limitations:
- Analysis is single-frame; does not leverage temporal (video) consistency.
- Unilateral analysis precludes bilateral anatomical referencing.
- No contrast enhancement for fine vessel visibility.
- Entire dataset from a single clinical site/platform.

Future research directions include integration of video-based temporal modeling (e.g. Space-Time Memory networks), bilateral vessel modeling, exploration of contrast-enhanced TCCD, and multi-center, multi-platform validation.

*Reference: "A Novel Attention-Augmented Wavelet YOLO System for Real-time Brain Vessel Segmentation on Transcranial Color-coded Doppler" [2508.13875]*

Source: https://www.emergentmind.com/topics/attention-augmented-wavelet-yolo-aaw-yolo