---
title: 'AGFA-Net: 3D Coronary Segmentation'
url: https://www.emergentmind.com/topics/agfa-net
type: topic
---

# AGFA-Net: 3D Coronary Segmentation

AGFA-Net is a 3D deep neural network architecture designed for coronary artery segmentation in computed tomography angiography (CCTA) volumes, particularly addressing the challenges of low image contrast, complex vessel anatomy, and varying vessel sizes. The network is characterized by an encoder-decoder topology augmented with attention and feature aggregation modules to enhance the extraction and fusion of contextually salient features. AGFA-Net achieved an average Dice coefficient of 86.74% and a Hausdorff distance of 0.23 mm on a dataset of 1,000 CCTA scans using 5-fold cross-validation, surpassing conventional benchmarks such as U-Net3D and V-Net [2406.08724].

## 1. Architectural Pipeline

AGFA-Net follows a U-shaped 3D encoder-decoder paradigm with three principal modules: Feature Refinement Module (FRM), Scale-Adaptive Feature Augmentation (SAFA), and Hierarchical Feature Integration Module (HFIM). Each module is strategically placed to optimize context-aware feature extraction and multi-scale information fusion.

- **Input:** 3D CCTA subvolume of size 128×160×160 voxels.
- **Encoder:** Four convolutional encoding blocks (E1–E4), each followed by a FRM and downsampling.
- **Bottleneck:** Fifth encoding block (E5) with appended SAFA for multi-resolution attention processing.
- **Decoder:** Four upsampling decoding blocks (D4–D1) that utilize skip connections, integrate FRM, and fuse features using HFIM.
- **Output:** Final $1\times1\times1$ convolution and sigmoid produce a per-voxel probability for the “coronary artery” class.

High-level pseudocode illustrates the stepwise progression through the network:
```python
def AGFA_Net(X):
    # Encoder
    for i in range(1, 5):
        Xi = ConvBlock_i(X_{i-1})   # E1–E4
        Xi = FRM(Xi)
        X_{i} = Downsample(Xi)
    X5 = ConvBlock_5(X4)
    X5 = SAFA(X5)
    # Decoder
    for i in range(4, 0, -1):
        U_i = Upsample(X_{i+1})
        C_i = Concat(U_i, X_i)
        C_i = FRM(C_i)
        X_i_prime = HFIM(C_i, X_{i+1}_prime)
    Y = Conv1x1x1(X_1_prime)
    return Sigmoid(Y)
```

## 2. Attention and Feature Enhancement Modules

### Feature Refinement Module (FRM)

FRM operates after each encoder and decoder block. It combines channel and spatial attention:

- **Channel attention:** On input $Y\in\mathbb{R}^{C\times D\times H\times W}$, global average and max pooling are applied, followed by a shared MLP projection to modulate each channel:
  $$
  A_c(Y) = \sigma(\text{MLP}(Y_\text{avg}) + \text{MLP}(Y_\text{max})) \in [0,1]^C
  $$
- **Spatial attention:** Channel pooling yields two $1\times D\times H\times W$ maps concatenated and passed through a $7^3$ convolution:
  $$
  A_s(Y) = \sigma(\text{Conv}_{7^3}([Y_\text{avg}^s; Y_\text{max}^s]))
  $$
- **Refinement:** Outputs are sequentially modulated:
  $$
  Y' = A_c(Y)\odot Y,\quad
  Y'' = A_s(Y')\odot Y'
  $$

### Scale-Adaptive Feature Augmentation (SAFA)

SAFA at the bottleneck splits feature channels into four groups, each processed by a dilated $3^3$ convolution with different rates ($d_n=1,2,3,4$) and channel gating. The outputs are concatenated:
$$
\widetilde{Y} = \text{Concat}(\widetilde{Y}^{(1)},\widetilde{Y}^{(2)},\widetilde{Y}^{(3)},\widetilde{Y}^{(4)})
$$
A Transformer-style self-attention mechanism constructs $Q,K,V$ using $W_Q,W_K,W_V$ (shape $3\times1\times1$, etc.), computes attention maps after reshaping, and fuses via:
$$
S = \text{softmax}(Q^T K), \quad \text{Attention}(Y) = S V
$$
The final bottleneck feature is
$$
Y_\text{SAFA} = \widetilde{Y} + \text{Attention}(\widetilde{Y})
$$

## 3. Hierarchical Feature Integration and Decoder Fusion

The Hierarchical Feature Integration Module (HFIM) fuses adjacent scales in the decoder, operating on the concatenation of current scale features $Y_i$ and the upsampled deeper decoder features $U_{i+1}$:
$$
Z_i = \text{Concat}(Y_i, U_{i+1}), \quad M_i = \sigma(Z_i), \quad Y_i' = \text{Conv}_{3^3}(M_i \odot Y_i)
$$
This iterative fusion enhances multi-scale context propagation toward the output.

## 4. Loss Function Formulation

The composite segmentation loss is a convex combination of weighted cross-entropy (WCE) and soft Dice loss, with $\alpha=0.6$:
$$
\mathcal{L} = \alpha \mathcal{L}_{\text{WCE}} + (1-\alpha)\mathcal{L}_{\text{Dice}}
$$
- **Weighted cross-entropy:**
  $$
  \mathcal{L}_{\mathrm{WCE}} = -\frac{1}{N}\sum_{n=1}^N [w\,g_n \log(p_n) + (1-g_n)\log(1-p_n)]
  $$
- **Soft Dice:**
  $$
  \mathcal{L}_{\text{Dice}} = 1-\frac{2\sum_{n=1}^N g_n p_n + \varepsilon}{\sum_{n=1}^N g_n + \sum_{n=1}^N p_n + \varepsilon},\qquad \varepsilon=1.0
  $$

## 5. Data Preparation and Training Protocol

- **Data:** 1,000 CCTA volumes; pixel spacing 0.28–0.41 mm, slice thickness 0.5–1.0 mm, slices per volume 199–275, patient ages 46–78.
- **Preprocessing:** Intensity normalization, random subvolume crops (128×160×160), rotation ($\pm20^{\circ}$), horizontal flips.
- **Cross-validation:** 5-fold protocol with 800/200 splits per fold; within-training held-out 15% validation subset.
- **Optimization:** Adam optimizer (lr$=3\times10^{-3}$, weight decay $=1\times10^{-6}$); CosineAnnealingWarmRestarts, batch size 16 across four Tesla V100 GPUs, 500 epochs per fold.

## 6. Quantitative Performance and Comparative Results

Evaluation utilized Dice coefficient, recall, precision, and Hausdorff distance:
$$
\text{Dice} = \frac{2\,\text{TP}}{2\,\text{TP} + \text{FP} + \text{FN}}
$$
$$
\text{Recall} = \frac{\text{TP}}{\text{TP} + \text{FN}},\qquad \text{Precision} = \frac{\text{TP}}{\text{TP} + \text{FP}}
$$
$$
H(A,B) = \max\{\sup_{a\in A}\inf_{b\in B} d(a,b),\; \sup_{b\in B}\inf_{a\in A} d(b,a)\}
$$

| Network   | Dice    | Recall   | Precision | Hausdorff (mm) |
|-----------|---------|----------|-----------|----------------|
| U-Net3D   | 81.06%  | 96.67%   | 83.48%    | —              |
| V-Net     | 83.43%  | 97.14%   | 85.27%    | —              |
| AGFA-Net  | 86.74%  | 98.79%   | 90.53%    | 0.23           |

AGFA-Net demonstrated improved Dice (by 3.31–5.68 points), recall, precision, and boundary accuracy over established 3D segmentation approaches.

## 7. Ablation Study and Module Contributions

Module contribution was quantified by testing nine architectural variants:

| Model      | FRM | SAFA | HFIM | Dice    | Recall  | Precision |
|------------|-----|------|------|---------|---------|-----------|
| Baseline   |  –  |  –   |  –   | 79.91%  | 96.58%  | 83.13%    |
| Net1       |  ✓  |  –   |  –   | 83.31%  | 96.85%  | 85.22%    |
| Net4       |  –  | ✓    | –    | 84.27%  | 97.48%  | 87.79%    |
| Net6       |  –  |  –   | ✓    | 84.14%  | 97.68%  | 86.44%    |
| Net7 (FRM+SAFA) | ✓  | ✓    | –    | 85.23%  | 98.54%  | 89.27%    |
| Net8 (FRM+HFIM) | ✓  | –    | ✓    | 85.09%  | 98.28%  | 88.89%    |
| Net9 (SAFA+HFIM)| –  | ✓    | ✓    | 85.41%  | 98.45%  | 89.36%    |
| AGFA-Net   | ✓   | ✓    | ✓    | 86.74%  | 98.79%  | 90.53%    |

- FRM alone: +3.40% Dice
- SAFA alone: +4.36% Dice
- HFIM alone: +4.23% Dice
- Any two modules combined: +5–5.5% Dice
- All three modules (AGFA-Net): +6.83% Dice over baseline

The coordinated deployment of FRM, SAFA, and HFIM provided the most consistent and substantial improvements across sensitivity, precision, and boundary metrics [2406.08724].

Source: https://www.emergentmind.com/topics/agfa-net