Papers
Topics
Authors
Recent
Search
2000 character limit reached

Efficient Restormer: Optimized Image Restoration

Updated 27 April 2026
  • The paper demonstrates that Efficient Restormers reduce computational cost by over 200x in multiply-add operations and 10,000x in memory usage via channel-axis attention and depthwise convolutions.
  • The models utilize a U-shaped encoder–decoder architecture with MDTA and GDFN modules to efficiently extract features for high-resolution image restoration and 7T MRI synthesis.
  • Effective training paradigms, including progressive learning and data augmentation, yield superior metrics (e.g., 64% NMSE reduction) compared to traditional CNNs and transformer baselines.

Efficient Restormer models are a class of transformer-based architectures designed for image restoration and related regression tasks, achieving state-of-the-art performance while substantially reducing computational complexity and parameter count. These models leverage modified self-attention mechanisms—including channel-wise attention, depthwise convolutions for local context, and gated feed-forward networks—to process high-resolution images efficiently. Notably, in medical imaging, they enable accurate synthesis of high-quality 7T MRI T1 maps from routine clinical 1.5T or 3T scans, as demonstrated by the 7T-Restormer (Eidex et al., 11 Jul 2025). Efficient Restormers build on the foundational Restormer design (Zamir et al., 2021), extending and adapting its principles to diverse application domains.

1. Architectural Principles and Core Modules

Efficient Restormers employ a U-shaped encoder–decoder architecture with multi-scale processing. The network input is typically a single-channel (e.g., MRI slice) or multi-channel (natural image) tensor x∈RH×Wx \in \mathbb{R}^{H \times W}, processed at multiple spatial scales via strided downsampling (such as pixel-unshuffle) in the encoder, and corresponding upsampling in the decoder.

The architectural hallmarks are as follows:

  • Multi-Dconv Head Transposed Attention (MDTA): Instead of conventional self-attention, MDTA applies layer normalization, followed by a sequence of 1×11\times1 pointwise convolutions and 3×33\times3 depthwise convolutions to produce the query, key, and value matrices. Attention is computed along the channel dimension CC rather than the spatial dimension N=Hâ‹…WN=H \cdot W—drastically reducing both computational cost and memory. The output is:

A=softmax(QKTdk)A = \mathrm{softmax}\left( \frac{QK^{T}}{\sqrt{d_k}} \right)

with channel grouping for multi-headed attention.

  • Gated-Dconv Feed-Forward Network (GDFN): The feed-forward sublayer features two depthwise convolution branches after normalization: one branch applies GELU and the two outputs are fused via elementwise multiplication (gating). A final 1×11\times1 convolution projects the result, followed by a residual connection.
  • Skip Connections: At each spatial scale, features from the encoder are concatenated into the corresponding decoder stage, preserving both local and global details.
  • Downsampling/Upsampling: Pixel-unshuffle (down) and pixel-shuffle (up) operations alter spatial resolution and channel count efficiently. For example, pixel-unshuffle halves each spatial dimension while quadrupling channels.

2. Computational Complexity and Parameter Efficiency

The efficiency of Restormer-based architectures is rooted in several design choices that minimize both the theoretical and practical cost:

  • Channel-Axis Attention: Conventional Vision Transformer attention scales as O(N2C)O(N^2C); MDTA reduces this to O(NC2)O(NC^2). For image sizes H=W=256H=W=256 and 1×11\times10, this results in over 200x reduction in multiply-add operations and a 10,000x reduction in attention-map memory compared to standard attention (Zamir et al., 2021).
  • Depthwise Separable Convolutions: All main blocks—MDTA and GDFN—use depthwise convolutions to reduce parameters and emphasize local context. The combination of 1×11\times11 and 1×11\times12 convolutions allows both channel mixing and spatial filtering.
  • Model Size and Inference Speed: In clinical MRI synthesis, the 7T-Restormer achieves its performance with only 10.5 million parameters and an inference time of 0.27 seconds per 1×11\times13 slice, compared to 56.7M/2.7s for ResShift and 70.4M for ResViT (Eidex et al., 11 Jul 2025).
Model Params (M) Inference Time (s, 256×384) Relative Params
7T-Restormer 10.5 0.27 1×
ResShift 56.7 ≈2.7 5.4×
ResViT 70.4 — 6.7×

A plausible implication is that such linear complexity enables training and inference on full-resolution images without aggressive downsampling or patching.

3. Training Paradigms and Optimization Schemes

Restormer variants are typically trained end-to-end using 1×11\times14 loss (voxel-wise for MRI, pixel-wise for images):

1×11\times15

Other salient aspects:

  • AdamW optimizer with a learning rate 1×11\times16, 1×11\times17, 1×11\times18, and weight decay 1×11\times19.
  • Batch size and number of epochs adapted to GPU resources; e.g., 8 and 50, respectively, on a single NVIDIA A6000 ADA (Eidex et al., 11 Jul 2025).
  • Data Augmentation: Random horizontal flips and center cropping are standard.
  • Progressive Learning: Patch size is increased and batch size reduced in stages for efficiency (Zamir et al., 2021).
  • Dataset Stratification: Comprehensive validation of MRI generation is performed on a mixed cohort (105 train, 19 val, 17 test patients; 32,128 total slices), with better generalization when training on combined 1.5T + 3T data (Eidex et al., 11 Jul 2025).

4. Quantitative Performance and Benchmarking

Efficient Restormers deliver substantial empirical benefits across domains, achieving superior image quality at a fraction of the cost:

  • MRI Synthesis (7T-Restormer): On the test set, for 1.5T input,
    • NMSE: 3×33\times30
    • PSNR: 3×33\times31 dB
    • SSIM: 3×33\times32
    • For 3T input, PSNR: 3×33\times33 dB, SSIM: 3×33\times34 (Eidex et al., 11 Jul 2025).
  • Relative Improvement over Baselines:
    • 3×33\times35 NMSE vs. ResShift (0.019 vs. 0.052 at 1.5T),
    • 3×33\times36 NMSE vs. ResViT (0.019 vs. 0.032 at 1.5T).
  • Natural Image Restoration: For deraining, deblurring, and denoising, Restormer achieves PSNR and SSIM improvements over CNNs and windowed transformer baselines, e.g., 33.96 dB for deraining vs. 32.91 dB for SPAIR, and >40 dB for real denoising on SIDD and DND (Zamir et al., 2021).
  • Ablation Studies: Removing MDTA or GDFN blocks, or substituting standard self-attention, consistently reduces accuracy, confirming the critical roles of these efficient designs.

5. Comparative and Ablative Analyses

Efficient Restormer architectures have been favorably compared to both convolutional baselines and competing transformers:

  • Vs. Windowed/Spatial Transformers: Dual-former, Uformer, and related hybrids localize all expensive self-attention operations to the coarse bottleneck, relegating higher resolution stages to pure convolution. Dual-former achieves comparable or higher PSNR/SSIM at 3×33\times37–3×33\times38 lower GFLOPs than MAXIM, Uformer, or Restormer (Chen et al., 2022).
  • Vs. Convolutional Backbones: Systems adopting transformer-like convolution blocks (e.g., RSFormer) further replace channel attention with large-kernel convolutional attention at intra-stage blocks, yielding higher PSNR and faster inference than pure Restormer while slightly reducing parameter count (Gao et al., 2023).

Key architecture variants implementing the efficient Restormer paradigm include:

Method Efficiency Strategy PSNR/SSIM Gain GFLOPs Reduction Reference
Restormer Channel-transposed attention everywhere +1.0 dB, +0.32dB 1× (Zamir et al., 2021)
Dual-former Single low-res hybrid Transformer block +1.91 dB 4–20× (Chen et al., 2022)
RSFormer TCB + global-local sampling +1.02 dB 15.6% faster (Gao et al., 2023)
7T-Restormer MDTA/GDFN for MRI synthesis –64% NMSE 5–7× parameter savings (Eidex et al., 11 Jul 2025)

6. Limitations and Future Directions

Efficient Restormers, while minimizing computation and parameter count, exhibit some limitations:

  • 2D-Only Processing: Current implementations (including 7T-Restormer) operate on 2D slices; extension to 3D volumetric transformers could enhance spatial coherence but demands larger datasets (Eidex et al., 11 Jul 2025).
  • Single-Scale Attention: Most variants apply self-attention at only one scale, potentially missing multi-scale long-range dependencies.
  • Domain Generalization: MRI applications to date have focused on single-vendor, single-diagnosis cohorts; broader validation on multi-site, multi-pathology data is required.
  • Potential Directions: Future work includes foundation model pretraining on large medical corpora, extending architectures to different imaging modalities, and further combining global-local or cross-dimensional attention mechanisms.

7. Significance and Impact

The efficient Restormer family of architectures provides a compelling solution for high-resolution image restoration, MRI synthesis, and related regression tasks where both computational resources and data volume are at a premium. The essential technical innovation is the reformulation of self-attention as a channel-axis operation, coupled with depthwise convolutions for local context and gating mechanisms for nonlinearity and selective feature propagation. These models have set new baselines in medical imaging (notably, 7T MRI synthesis)—delivering significant NMSE reductions and parameter savings relative to convolutional GANs, diffusion models, and classical ViT-based frameworks—while remaining broadly transferable to general image restoration domains (Eidex et al., 11 Jul 2025, Zamir et al., 2021, Chen et al., 2022, Gao et al., 2023).

A plausible implication is that the underlying efficiency principles of Restormer—channelized attention, depth-separable convolutions, and multi-scale encoder–decoder layout—will continue to inform transformer designs in computationally constrained imaging environments and domains with limited annotated data.

Topic to Video (Beta)

No one has generated a video about this topic yet.

Whiteboard

No one has generated a whiteboard explanation for this topic yet.

Follow Topic

Get notified by email when new papers are published related to Efficient Restormer.