Papers
Topics
Authors
Recent
Search
2000 character limit reached

EfficientViT-L1: Efficient Vision Transformer

Updated 14 January 2026
  • EfficientViT-L1 is a hierarchical vision transformer backbone that uses multi-scale linear attention to balance global context and local detail for dense prediction tasks.
  • It combines a convolutional stem with progressive transformer stages featuring structured downsampling, ensuring both computational efficiency and high throughput.
  • The architecture delivers significant latency reductions and performance gains in semantic segmentation and super-resolution, ideal for resource-constrained deployments.

EfficientViT-L1 is a hierarchical high-resolution vision transformer backbone optimized for dense prediction tasks under stringent computational budgets. It introduces multi-scale linear attention to achieve the global receptive field and multi-scale processing—key for semantic segmentation, super-resolution, and similar vision workloads—while utilizing only lightweight, hardware-efficient operations. The EfficientViT-L1 configuration emphasizes structured downsampling, localized convolutional aggregation, and global linear kernel-based attention to provide high throughput and competitive model accuracy in real-world deployment scenarios (Cai et al., 2022).

1. Architectural Design

EfficientViT-L1 employs a five-part architecture: an initial convolutional “stem” followed by a four-stage transformer backbone. Stages 3 and 4 integrate the model’s defining multi-scale linear attention (MSLA) blocks, while earlier stages focus on feature extraction via convolutional and feed-forward projections. Downsampling occurs at each stage boundary, resulting in progressive reduction of resolution and increased abstraction. The input is assumed to be 224×224224 \times 224 pixels for benchmarking purposes. The overall architectural flow is as follows:

  1. Stem: Two 3×3 convolutions with stride 2 successively lower resolution to 56×5656 \times 56.
  2. Stage 1: Processes at 56×5656 \times 56; primarily feed-forward, no attention.
  3. Stage 2: Downsamples to 28×2828 \times 28; again, no attention.
  4. Stage 3: Operates at 14×1414 \times 14; MSLA introduced.
  5. Stage 4: Operates at 7×77 \times 7; MSLA present.

Both MSLA stages use a combination of ordinary (1×11\times1 identity) and 5×55\times5 depth-wise separable (DWS) convolutional branches, aggregating both global and local context (Cai et al., 2022).

2. Per-Stage Configuration

EfficientViT-L1’s stage-wise block depth, channel width, and expansion ratios can be summarized as follows (exact values are representative estimates based on GitHub configuration, as the original paper does not explicitly list all hyperparameters for L1):

Stage Output Resolution # Blocks (BsB_s) Hidden Channels (CsC_s) FFN Expansion (56×5656 \times 560) MS-Linear-Attn
1 56×5656 \times 561 2 64 4 None
2 56×5656 \times 562 2 128 4 None
3 56×5656 \times 563 6 256 4 56×5656 \times 564 (two-branch, heads 56×5656 \times 565)
4 56×5656 \times 566 2 512 4 56×5656 \times 567 (same as Stage 3)

A plausible implication is that attention heads per MSLA stage are in the range of 56×5656 \times 568, giving a per-head dimension 56×5656 \times 569 for each attention split (Cai et al., 2022).

3. Multi-Scale Linear Attention Mechanism

The core contribution of EfficientViT-L1 is the MSLA mechanism, first applied in Stages 3 and 4. This reformulates the standard softmax self-attention with a linear ReLU kernel 56×5656 \times 560, reducing memory and compute complexity. Attention is computed as follows:

  1. Input projections: Keys 56×5656 \times 561, queries 56×5656 \times 562, and values 56×5656 \times 563 generated from input tokens (56×5656 \times 564), each of size 56×5656 \times 565.
  2. Multi-scale aggregation: 56×5656 \times 566, 56×5656 \times 567, 56×5656 \times 568 each are processed by (a) identity and (b) 56×5656 \times 569 DWS conv, creating a set 28×2828 \times 280 for 28×2828 \times 281.
  3. Kernel-feature computation: For each scale, compute summary statistics

28×2828 \times 282

Aggregate across all scales via learned weights 28×2828 \times 283, then output for each token 28×2828 \times 284:

28×2828 \times 285

where

28×2828 \times 286

This architectural innovation enables both global and local contextualization without expensive full-rank softmax attention, providing the model with requisite expressivity for dense tasks while maintaining efficiency (Cai et al., 2022).

4. Computational Efficiency and Hyperparameters

EfficientViT-L1 is explicitly designed for high-throughput inference and deployment on resource-constrained hardware. Reported statistics:

  • Parameter count: Approximately 53 million
  • Multiply-Accumulates (MACs): 5.3 G for 28×2828 \times 287 inputs
  • Edge GPU (Jetson AGX Orin, TensorRT fp16, batch size 1): 2.6 ms per image
  • Cloud GPU (A100, TensorRT fp16): 6,207 images/sec (≈0.16 ms/img)
  • Mobile CPU (Snapdragon 8 Gen 1, TFLite fp32): B-series at comparable size runs in 30–50 ms (not reported directly for L1)

A plausible implication is that EfficientViT-L1’s latency and compute characteristics enable near real-time performance in edge and embedded vision applications, with throughput far exceeding contemporary softmax-attention-based and large-kernel convolution backbones at similar accuracy (Cai et al., 2022).

5. Functional Workflow and Pseudocode

The EfficientViTBlock, as applied in MSLA stages, is captured by the following workflow:

14×1414 \times 147

In this workflow, linear attention and feed-forward processing are interleaved, separated by local DWS convolution to increase inductive bias toward spatial priors while retaining the efficiency of linear attention (Cai et al., 2022).

6. Application Domains and Comparative Performance

EfficientViT-L1 and its family are targeted primarily at high-resolution dense prediction tasks such as semantic segmentation (Cityscapes), image super-resolution, and general vision backbones for downstream transfer. Notable observations from benchmarking:

  • Up to 28×2828 \times 288 GPU latency reduction over SegFormer, 28×2828 \times 289 over SegNeXt, with no loss in segmentation accuracy (Cityscapes).
  • For super-resolution, 14×1414 \times 140 speedup over Restormer with a measured PSNR gain (14×1414 \times 141 dB).
  • In zero-shot instance segmentation (COCO, Segment Anything), 14×1414 \times 142 throughput improvement on A100 GPU.
  • Across tasks and hardware, EfficientViT-L1 and its configuration deliver substantial gains in throughput versus softmax-attention and large-kernel CNN backbones at comparable or superior accuracy levels (Cai et al., 2022).

7. Model Scope and Configurability

EfficientViT-L1 represents the smallest "L-series" configuration among EfficientViT variants described in (Cai et al., 2022). Deeper and wider configurations, such as L2 and L3, increase both block counts and hidden channel dimensions for heightened accuracy; for example, going from 14×1414 \times 143 to 14×1414 \times 144 blocks per stage, and 14×1414 \times 145 to 14×1414 \times 146 hidden channels. All EfficientViT-L models retain the same core attention and stage design, emphasizing modular configurability. Exact per-model hyperparameters are maintained in the official GitHub repository referenced by the authors (Cai et al., 2022). This design principle facilitates model selection across deployment scenarios and computational environments.


EfficientViT-L1 exemplifies the trend toward vision transformers optimized for memory and computational efficiency, balancing global receptive field, local spatial aggregation, and practical deployability for state-of-the-art dense prediction tasks. All technical details are traceable to (Cai et al., 2022).

Definition Search Book Streamline Icon: https://streamlinehq.com
References (1)

Topic to Video (Beta)

No one has generated a video about this topic yet.

Whiteboard

No one has generated a whiteboard explanation for this topic yet.

Follow Topic

Get notified by email when new papers are published related to EfficientViT-L1 Backbone.