SAFE-KD: Risk-Controlled Early Exit for Vision Models
- The framework SAFE-KD is a universal early-exit system that combines hierarchical knowledge distillation with conformal risk control to guarantee statistically bounded selective risk.
- It attaches intermediate classifier exits to any vision backbone (CNN or ViT) and employs decoupled knowledge distillation alongside consistency regularization for calibrated risk control.
- Empirical results show up to a 45% reduction in expected inference depth while maintaining or surpassing full-inference accuracy, ensuring efficiency and robustness.
SAFE-KD offers a universal, risk-controlled early-exit framework for modern vision backbones, combining hierarchical knowledge distillation with conformal risk control (CRC) to achieve statistically guaranteed bounds on selective misclassification risk for early-exit architectures. It enables substantial reductions in inference cost via early stopping for "easy" samples, while maintaining user-specified upper bounds on misclassification risk at each exit, calibrated on finite data. SAFE-KD is model-agnostic and deploys on a variety of convolutional (CNN) and transformer-based (ViT) image models (Khazem, 3 Feb 2026).
1. Architecture and Components
SAFE-KD is structured as a lightweight "wrapper" atop any standard vision backbone, supporting both CNNs and Vision Transformers. Its architecture comprises:
- Base Backbone: Any pretrained or trainable vision model (e.g., ResNet, ConvNeXt, ViT, Swin).
- Intermediate Exit Heads: At select depths, SAFE-KD attaches classifiers producing logits (), class probabilities , and a confidence score (typically, Maximum Softmax Probability ). For CNNs, exits use global average pooling and an optional MLP before a fully connected (FC) layer; for ViTs, exits use CLS or mean token pooling, optional LayerNorm, then FC.
- Teacher Network: An Exponential Moving Average (EMA) of the full model serves as the teacher for knowledge transfer.
This configuration allows SAFE-KD to operate agnostically across architectures, minimally increasing inference overhead.
2. Decoupled Knowledge Distillation and Consistency
Training leverages hierarchical Decoupled Knowledge Distillation (DKD), coupled with deep-to-shallow consistency regularization:
- DKD Loss: For each exit , knowledge is distilled from the teacher logits to using a split-KL loss:
where 0 and 1 are the (teacher, student) probabilities on the target class 2 (ground-truth label), and 3 and 4 are normalized distributions over all non-target classes.
- Consistency Regularization: To align intermediate exits with the final head, SAFE-KD adds
5
with weighting 6, regularizing posterior agreement between each intermediate exit and the ultimate exit (7).
- Total Loss: For weights 8 summing to 9, full training minimizes:
0
where 1 is a scaling factor for DKD.
This hierarchical approach increases calibration, depth-to-exit consistency, and maintains high accuracy at all exits.
3. Conformal Risk Control for Early-Exit Thresholds
At inference, SAFE-KD employs Conformal Risk Control (CRC) to set data-driven confidence thresholds at each exit, guaranteeing a user-specified selective risk:
- Nonconformity Score: 2 at exit 3.
- Acceptance Set: 4.
- Selective Misclassification Risk: 5.
- Threshold Calibration: Using a held-out calibration set 6, thresholds 7 are chosen so the conformal upper bound:
8
does not exceed the desired risk level 9. Here 0.
CRC, under the exchangeability assumption, ensures
1
for each exit, providing finite-sample statistical guarantees.
4. Safe Inference Policy and Practical Deployment
At test time, early exit is governed by the following procedure:
- Proceed through exits 2, checking at each if 3.
- The first such 4 is used for prediction. If none, inference proceeds to the final exit 5.
- For every exit 6, the empirical misclassification risk among samples exiting there is guaranteed not to exceed 7 (up to sampling correction).
This allows the system designer to select 8 according to operational requirements, trading off computational savings for tightly controlled selective risk. All calibration is based on a held-out set.
5. Empirical Evaluation and Results
SAFE-KD has been empirically validated across six architectures (ResNet-50, MobileNetV3-S, EfficientNet-B0, ConvNeXt-T, ViT-S, Swin-T) and multiple image datasets (CIFAR-10/100, STL-10, Pets, Flowers102, Aircraft), delivering:
- Compute-Accuracy Trade-offs: At 9 risk, SAFE-KD achieves 0--1 lower expected depth while matching or surpassing full-inference accuracy. Baseline methods (fixed MSP or entropy thresholds) violate the risk constraint, with observed risks 2--3.
- Calibration: SAFE-KD reduces negative log-likelihood (NLL) and expected calibration error (ECE) at all exits.
- Risk Guarantee: Across sweeps in 4, the observed per-exit risk tracks the theoretical bound 5, confirming 6 tightness.
- Robustness: On CIFAR-10-C corrupted data (severity 3), SAFE-KD attains lower mean corruption error (mCE) at both shallowest and deepest exits compared to comparable multi-exit and DKD-based models, e.g., mCE at exit 1: SAFE-KD 7, DKD 8, MultiExit 9; at final exit: SAFE-KD 0.
- Ablation Findings: Removing DKD degrades accuracy by 1 and forces safer (more conservative) thresholds, increasing average depth. Removing consistency (2) triggers higher exit-variance, though risk guarantees persist.
- Example Table (for CIFAR-100, ResNet-50, 3):
| Method | Accuracy | Exp. Depth | Observed Risk |
|---|---|---|---|
| Fixed MSP | 81.5% | 0.72 | 6.8% (Unsafe) |
| Entropy gate | 80.9% | 0.65 | 7.5% (Unsafe) |
| SAFE-KD (CRC) | 82.3% | 0.59 | 4.8% (Safe) |
SAFE-KD consistently defines the empirical Pareto frontier for the target risk constraint across tasks.
6. Calibration, Robustness, and Risk Guarantees
SAFE-KD's deployment of CRC uniquely enables it to deliver finite-sample, statistically-tight risk control not attainable with heuristic thresholds. Reliability diagrams confirm alignment of observed risk to the target 4 across exit depths. Selective risk curves for a sweep of 5 show empirical 6 at or just under 7.
For corrupted or hard samples, the framework naturally "defers" to deeper exits, preserving guaranteed selective risk at cost of additional computation. This property, together with out-of-the-box calibration from DKD and consistency, distinguishes SAFE-KD from prior early-exit and distillation methods without such formal risk control.
7. Summary and Broader Impact
SAFE-KD constitutes a general-purpose, modular extension to vision models requiring minimal architectural modification and no retraining of the backbone. Its integration of CRC, DKD, and deep-to-shallow consistency provides user-tunable, quantifiable risk guarantees for early exiting, fine-grained calibration, and enhanced robustness under dataset shift or corruption. The framework supports frequent regression testing and online adaptation to evolving operational requirements, making it especially suitable for resource-constrained or safety-critical deployments (Khazem, 3 Feb 2026).