TuNAS: Tunable NAS for Latency-Aware CNNs
- TuNAS is a tunable NAS algorithm that uses an RL-based controller and weight-sharing techniques to efficiently navigate enormous CNN search spaces with latency constraints.
- It employs a one-shot supernetwork with aggressive parameter sharing methods such as operation collapsing and channel masking to reduce computational redundancy.
- Empirical evaluations on ImageNet and MS COCO show TuNAS achieves 1–2.5% higher accuracy with lower compute cost than random search in large-scale settings.
TuNAS (“Tunable NAS”) is a scalable, weight-sharing neural architecture search (NAS) algorithm designed to identify high-quality, resource-aware convolutional neural network (CNN) architectures within extremely large search spaces while adhering to strict latency constraints. TuNAS employs a reinforcement learning (RL)–based controller paired with a supernetwork, leveraging aggressive parameter sharing and latency-aware reward mechanisms to efficiently explore vast architectural regimes. Comprehensive experimental analyses on ImageNet and COCO benchmarks demonstrate that TuNAS yields architectures outperforming those found by random search, particularly as the search space size escalates (Bender et al., 2020).
1. Algorithmic Foundations
TuNAS maintains a supernetwork (one-shot model) parameterized by shared weights and an RL-based controller that defines a multinomial policy over architectural choices. Each step in the search alternates between two core operations:
- Policy Update: The controller is updated using a policy-gradient approach (REINFORCE) based on a scalar reward function evaluated over validation performance and latency compliance.
- Weight Sharing: The shared weights are optimized (SGD) using the loss from a single subnetwork, sampled according to .
The algorithm can be summarized as follows:
- Initialize shared weights and controller (uniform distribution).
- Repeat for :
- Sample architecture .
- Compute training loss .
- Evaluate validation quality 0.
- Calculate latency 1 and reward 2, for 3.
- Update 4 via REINFORCE: 5, where 6 is a moving baseline.
- Update 7 by gradient descent on 8.
- Select the final architecture 9 by maximizing 0 over each decision.
This framework transforms the constrained optimization (maximizing 1 subject to 2) into an unconstrained RL objective using a reward function that enables direct and stable targeting of the latency constraint.
2. Search Space Construction and Architectural Choices
TuNAS operates over three primary mobile-CNN search spaces, each defined as sequences of inverted-bottleneck blocks in the MobileNetV2/V3 style. Each block allows categorical variation along these axes:
| Search Space | Expansion Ratios | Depthwise Kernel Sizes | Output Channels | Squeeze-and-Excite | Activation |
|---|---|---|---|---|---|
| ProxylessNAS | 3, 6 | 3, 5, 7 | base MobileNetV2/V3 | No | ReLU |
| ProxylessNAS-Enlarged | 3, 6 | 3, 5, 7 | Multiplier ×3 | No | ReLU |
| MobileNetV3-Like | 1–6 | 3, 5, 7 | Same as above | Yes/No | ReLU/Swish |
No novel operators beyond the MobileNetV2/V3 family are introduced. Rather, TuNAS's contribution is the ability to explore these vast combinations with significant computational efficiency. Cardinalities of the evaluated search spaces are approximately 4, 5, and 6 respectively.
3. Weight-Sharing Strategies
TuNAS introduces aggressive parameter sharing to make supernetwork training tractable:
- Operation Collapsing: The same 1×1 expand convolutional weights are used across differing depthwise kernel sizes, reducing redundancy.
- Channel Masking: Instead of allocating unique weight tensors for each output channel configuration, convolutional layers are implemented up to a maximum width 7. Narrower configurations are realized by masking off excess channels, making the supernetwork's parameter count linear rather than exponential in categorical options.
Sampling proceeds by executing and updating weights along a single path 8 per search iteration. Only the active path's weights receive gradient signal, promoting greater parameter reuse and diminishing interference between rare operations.
4. Hyperparameter Specification and Quality Improvement Techniques
Key hyperparameter values are as follows:
- Shared-weight training: RMSProp (momentum=0.9, decay=0.9, 9), max learning rate 2.64 (warm-up 2.5%), cosine decay, batch size 4096 (utilizing 128×32 TPU cores), and L2 regularization 0.
- Controller: Adam (learning rate 1 base, 2, 3, 4). Controller learning rate zero for the first 25% of steps (RL disabled for warmup), then either constant or exponentially increased to 5.
- Search duration: 6–7 hours on TPU-v3-32, corresponding to 8 epochs.
To safeguard training signal propagation to rare operations and filters, warmup schedules are implemented:
- Filter Warmup: During the first 25%, all 9 filters are stochastically enabled (linear anneal 0), ensuring rare channel slices receive updates.
- Operation Warmup: Similarly, all operations in each block are enabled with linearly annealed probability, averaging their outputs during initial training steps.
- Absolute-Value Reward: Direct reward 1 eliminates the need for extensive tuning of cost exponents 2, with the same value performing robustly across spaces and hardware.
These innovations reduce manual hyperparameter tuning and stabilize optimization, particularly in early training epochs.
5. Empirical Evaluation and Comparative Analysis
TuNAS was benchmarked on ImageNet-1K classification and MS COCO object detection. Architectures are searched using 90 epochs of weight-sharing training; final selected architectures are retrained from scratch for up to 360 epochs.
ImageNet-1K Classification
| Search Space | Random Search (N=50) | TuNAS (90 epochs) | TuNAS (360 epochs) | Reference Model (training setup) |
|---|---|---|---|---|
| ProxylessNAS | 3 | 4 | – | 5 |
| ProxylessNAS-Enl. | 6 | 7 | – | 8 |
| MobileNetV3-Like | 9 | 0 | 1 | 2 (at 58 ms) |
TuNAS yields a 3–4 percentage point absolute improvement in top-1 accuracy compared to random search, with gains magnifying for larger search spaces (up to 5 at 6 models). Notably, random search incurs 7–8 greater compute cost per result, as each trial involves full stand-alone training.
MS COCO Detection
| Backbone | mAP (%) | Latency (ms) |
|---|---|---|
| MobileNetV2+SSDLite | 20.7 | 126 |
| MnasNet+SSDLite | 21.3 | 129 |
| ProxylessNAS+SSDLite | 21.8 | 140 |
| MobileNetV3+SSDLite | 22.0 | 106 |
| TuNAS+SSDLite | 22.5 | 106 |
TuNAS+SSDLite achieves the highest mAP (22.5) at a 106 ms latency target.
6. Key Insights and Practical Recommendations
TuNAS provides evidence that weight-sharing NAS—with the proposed architectural parameter sharing, RL-based controller, absolute-value latency reward, and warmup schedules—is especially advantageous in large-scale search spaces where computational resources are constrained.
- On search spaces up to 9 models, TuNAS finds CNNs that are 0–1 percentage points absolutely more accurate than those found by random search under identical hardware and latency constraints.
- Hyperparameter robustness is improved: the same settings (notably 2) are directly transferable between search spaces, mitigating the need for costly per-task grid search.
- For practical use: set the target latency 3, reuse the published hyperparameters for 4 and 5, and execute five independent search runs to estimate variance, selecting the final architecture by argmax-probability under 6 and retraining from scratch.
A plausible implication is that as search space complexity and resource constraints become more salient, the advantages of weight-sharing approaches—when properly tuned—are magnified, reinforcing the rationale for their deployment in industrial NAS applications (Bender et al., 2020).