Learnable Edge Kernels in Medical Registration
- Learnable edge kernels are convolutional filters initialized with a Laplacian edge-detection prior and refined through gradient descent to capture essential anatomical features.
- They are integrated into both rigid and non-rigid registration architectures to enhance alignment by focusing on structural rather than intensity variations.
- Empirical evaluations demonstrate state-of-the-art performance, with improvements in metrics like Dice similarity and robustness across multi-modal medical datasets.
Learnable edge kernels are a class of convolutional kernels used within neural networks for medical image registration, designed to address the challenges of aligning images from different modalities, time points, or subjects. Unlike conventional convolutional layers that initialize and optimize filters without explicit structural bias, learnable edge kernels are seeded with a predefined edge detection prior—typically the Laplacian—and are then perturbed and adapted via gradient-based training. This approach enables the kernel to extract optimal, data-driven edge features, improving robustness to intensity and contrast variations and enhancing the detection of anatomical boundaries. Learnable edge kernels have been integrated into both rigid and non-rigid registration architectures, demonstrating state-of-the-art performance on multi-modal and anatomical alignment tasks (Siyal et al., 1 Dec 2025).
1. Definition and Initialization of Learnable Edge Kernels
The learnable edge kernel begins with a classical edge-detection operator as its initialization. In the standard implementation, a (2D) or (3D) discrete Laplacian kernel is used:
(2D)
For 3D, the mask is defined such that the center voxel receives , and each of its six face-neighbors , all others zero:
To introduce controlled variation, each weight is multiplicatively perturbed by i.i.d. Gaussian noise: with perturbation scale . This generates the initial kernel as .
2. Integration into Registration Architectures
Learnable edge kernels are incorporated as a dedicated convolutional branch within both rigid and non-rigid medical image registration networks. Each input volume (moving or fixed) is first processed by this edge-kernel layer. In variant architectures, this may be a fixed Laplacian (variant 1) or a learnable edge kernel (variants 2–4). The outputs from the edge-kernel branch are concatenated, and then the top 16 channels—ranked by mean activation—are forwarded to the subsequent encoder blocks of the main registration backbone, which often follows a U-Net or residual architecture.
Convolution is standard: and analogously in 3D.
3. Learning and Optimization
The edge-kernel branch is implemented as a standard convolutional layer, ensuring its weights are adapted during network training through stochastic gradient descent. Gradients from the global registration loss 0 propagate to the kernel weights: 1 where 2 is the learning rate, typically handled by the Adam optimizer within PyTorch.
The key advantage is that training drives the kernel to extract edge features optimal for the registration objective, potentially capturing both generic and anatomy-specific boundaries critical in medical imaging.
4. Architectural Variants
Eight total architecture variants were constructed to systematically assess the role and location of learnable edge kernels, split across rigid and non-rigid registration:
| Variant | Rigid (Reg-LEdge-Model) | Non-rigid (Reg-LEdge-U-Model) |
|---|---|---|
| 1 | Fixed Laplacian; concat, residual encoder, upsample | Residual block, edge module (every encoder) |
| 2 | Learnable edge module replaces Laplacian | Edge + dense-fusion + pooling; rest residual |
| 3 | Edge module in 1st downsample block only | Residual→dense-fusion+pooling (no edge deeper) |
| 4 | Edge module in every downsample block | Edge module + inception + residual + pooling (every encoder) |
Variant 4 for both rigid and non-rigid registration places the learnable edge module pervasively throughout the encoder, which empirically yields the best performance.
5. Loss Functions and Training Objectives
Rigid registration networks regress 12 affine parameters 3 by optimizing an unsupervised objective based on image similarity after warping: 4 where 5 can be negative local mutual information (LMI).
Non-rigid registration is trained via: 6 with 7 (bending-energy) as a diffusion-style regularizer. The (local) mutual information term is: 8
6. Empirical Evaluation
The methods were tested on 60 intra-patient 9 MRI pairs (split 30/10/20 for training, validation, test) from the Medical University Innsbruck, as well as IXI (atlas→patient) and OASIS (inter-patient) public datasets. Rigid experiments included both skull-intact and skull-removed images; non-rigid always used skull-stripped data. Baseline methods included ANTs/SyN, LDDMM, VoxelMorph (affine/deform), ViT-V-Net, TransMorph, and CycleMorph. Key metrics were Dice similarity (WM–GM, brain mask) and the proportion of voxels with non-positive Jacobian determinant 0.
Selected results:
| Method | Dice WM–GM (1sd) | Dice Brain Mask (2sd) | 3 | Scenario |
|---|---|---|---|---|
| ANTs | 0.742 ± 0.122 | 0.924 ± 0.136 | – | Rigid, skull intact |
| VoxelMorph-Aff | 0.732 ± 0.133 | 0.834 ± 0.121 | – | Rigid, skull intact |
| Reg-LEdge-Var-4 | 0.784 ± 0.125 | 0.917 ± 0.101 | – | Rigid, skull intact |
| Reg-LE-4 | 0.757 ± 0.116 | 0.912 ± 0.161 | – | Rigid, skull removed |
| SyN | 0.766 ± 0.136 | – | 4 | Non-rigid |
| LEdge-U-Var-4 | 0.798 ± 0.116 | – | 5 | Non-rigid |
On IXI and OASIS data, learnable edge kernel models matched or outperformed recent transformer-based and CNN-based registration methods (TransMorph, VoxelMorph, nn-Former), while maintaining diffeomorphic guarantees.
7. Insights and Future Directions
Seeding the network with explicit edge-detection priors and permitting data-driven adaptation biases early feature extraction toward anatomically meaningful boundaries. This stabilizes multi-modal similarity (as edges vary less than intensity profiles across modalities such as T1/T2), improves robustness to noise and intensity/contrast perturbations, and guides both global (rigid) and local (non-rigid) alignment toward structural congruence rather than intensity correspondence alone. Pervasive deployment of edge modules (variant 4) yields the most significant performance gains, confirming the importance of edge-aware feature learning.
Potential extensions include multi-scale edge kernels, orientation-specific learned filters (such as Sobel-like banks), and adaptation to transformer-based backbones for registration (Siyal et al., 1 Dec 2025).