TCUQ: Single-Pass Uncertainty for TinyML
- The paper introduces TCUQ, a single-pass, label-free uncertainty monitor that transforms short-horizon temporal signals into calibrated risk scores for streaming TinyML applications.
- It employs a lightweight streaming conformal layer with O(1) per step updates to enforce an accept/abstain decision rule under fluctuating distribution shifts.
- Empirical evaluations demonstrate that TCUQ reduces latency and memory usage compared to ensemble methods while maintaining robust detection and calibration performance on microcontrollers.
TCUQ, introduced as “Single-Pass Uncertainty Quantification from Temporal Consistency with Streaming Conformal Calibration for TinyML,” is a single pass, label free uncertainty monitor for streaming TinyML that converts short horizon temporal consistency captured via lightweight signals on posteriors and features into a calibrated risk score with an ring buffer and per step updates. It is designed for battery-powered microcontrollers operating on continuous input streams, where in-distribution, corrupted-in-distribution, and out-of-distribution samples may interleave, and where uncertainty quantification must be obtained without online labels, extra forward passes, or large memory overheads. The method couples temporal-consistency features with a streaming conformal layer to produce a budgeted accept/abstain rule, yielding calibrated behavior on kilobyte-scale devices (Lamaakal et al., 18 Aug 2025).
1. Problem formulation and operating regime
TCUQ is motivated by the fact that TinyML systems must satisfy tight memory, compute, and energy budgets while remaining robust to distribution shift in continuous streams. The target setting includes three stream components: in-distribution (ID) samples, corrupted-in-distribution (CID) variants such as noise, blur, and fog, and out-of-distribution (OOD) inputs. In this setting, deep networks that may appear calibrated on static test sets can become overconfident under CID and OOD conditions, which creates a need for on-device uncertainty quantification that detects accuracy drops in real time, requires no extra forward passes or labels at inference, fits in of flash and RAM, and updates in time per input (Lamaakal et al., 18 Aug 2025).
The method is explicitly positioned against several established approaches. Deep ensembles, MC Dropout, and early-exit heads either multiply inference cost by passes or add heavyweight heads. Post-hoc calibration is described as failing under dynamic shifts. Standard conformal predictors provide distribution-free risk control, but standard implementations store all scores or require labels. TCUQ addresses this combination of constraints by extracting a single uncertainty signal from short-horizon temporal consistency of posteriors and features, calibrating it on the fly via a memory-light streaming conformal layer, and enforcing a user-specified accept/abstain budget.
A plausible implication is that TCUQ is not primarily an alternative classifier architecture; rather, it is an uncertainty monitor layered on top of a trained backbone. This distinction matters because its computational and memory costs are additive to a single existing forward pass rather than requiring architectural replication or multi-pass inference.
2. Temporal-consistency signals and uncertainty scoring
TCUQ assumes a trained backbone with final feature vector and class probabilities . The hard decision and two instantaneous confidence statistics are defined as
where are the sorted probabilities (Lamaakal et al., 18 Aug 2025).
A ring buffer of window size 0 stores recent posterior-feature pairs 1. From this buffer, TCUQ computes four lightweight temporal signals. Let 2 be a fixed lag set, for example 3, with weights 4 and 5.
The first signal is a multi-lag predictive divergence based on Jensen–Shannon divergence: 6 with
7
The second signal is feature stability, computed through cosine similarity,
8
and used via the instability term 9.
The third signal is decision persistence,
0
and is used through the inconsistency term 1.
The fourth signal is an instantaneous confidence proxy blending inverse maximum probability and inverse margin: 2
These are stacked into
3
TCUQ then applies a tiny logistic combiner fitted offline: 4 where 5 are learned once on a small labeled development set mixing ID, CID, and OOD examples. At inference, this reduces to one inner product and one sigmoid.
This design makes the uncertainty estimate explicitly temporal rather than purely instantaneous. The paper’s formulation suggests that TCUQ exploits local temporal smoothness or consistency as an operational prior: when posteriors, features, and decisions fluctuate over short lags, risk increases even if a single-step confidence score appears high.
3. Streaming conformal calibration and abstention control
The stepwise uncertainty estimate is transformed into a scalar nonconformity score
6
A streaming conformal layer maintains an online estimate 7 of the 8-quantile of 9 using a memory-constant quantile tracker (Lamaakal et al., 18 Aug 2025).
One concrete update rule given is the Robbins–Monro update: 0 Only the scalar 1 and one conditional are required per step. During warm-up, a tiny buffer of size 2 is used to initialize 3 to the empirical quantile, after which the streaming update is used.
The decision rule is budgeted accept/abstain. If 4 and the abstention controller still has budget, the system abstains; otherwise it accepts and returns 5. The controller can enforce a long-run abstention fraction near a user budget 6 by allowing at most 7 abstentions in 8 steps.
The paper states a formal calibration guarantee under the standard conformal exchangeability assumption on 9: 0 For the streaming setting, it further states via Hoeffding’s inequality on exceedance indicators 1 that for any 2, with probability at least 3,
4
Thus, as 5, the empirical reject rate converges to 6 with high-confidence error 7.
A common misconception would be to treat TCUQ as a fully label-free conformal predictor in the classical offline sense. The paper is more specific: the monitor is label free online, while the logistic combiner is fitted offline on a small labeled development set. The conformal component then calibrates the resulting score during deployment without online labels or score storage beyond constant memory.
4. Data structures, computational profile, and embedded implementation
The central data structure is a ring buffer of size 8 that stores the last 9 posterior vectors of length 0 and feature vectors of length 1. This yields memory complexity
2
The paper notes that in practice 3 can be reduced to 4 by an offline-learned 5 projection or global pooling (Lamaakal et al., 18 Aug 2025).
Per input 6, the method performs one forward pass through the backbone, one buffer overwrite at pointer 7, computation of 8, 9, 0, and 1 by retrieving 2 past entries, one vector dot product to obtain 3, one scalar blend to obtain 4, one streaming-quantile update, and one budget-controller check. Since 5 is constant, for example 6, all additional monitoring operations are 7 per step aside from the fixed-cost forward pass.
The implementation targets microcontrollers directly. TCUQ was compiled on two representative MCUs: STM32F767ZI, described as “Big-MCU” with 512 KB Flash and a 216 MHz Cortex-M7, and STM32L432KC, described as “Small-MCU” with 256 KB Flash and an 80 MHz Cortex-M4. On microcontrollers, TCUQ is reported to fit comfortably on kilobyte scale devices, adding less than 4 KB RAM for 8 and less than 1 KB Flash for the buffer and quantile code.
This computational profile is significant because it separates the method from heavier uncertainty-estimation schemes. The paper’s framing suggests that TCUQ’s principal engineering contribution lies not only in detection quality, but in the fact that the monitor remains viable under deployment constraints where alternative methods do not fit.
5. Empirical behavior on microcontrollers and streaming benchmarks
The experimental evaluation uses SpeechCommands with DSCNN, MNIST, CIFAR-10 with ResNet-8, and TinyImageNet with MobileNetV2. Baselines include BASE confidence threshold, MC Dropout with 9, Deep ensembles with 0, Early-exit ensemble with a 3-head ensemble, G-ODIN, and HYDRA (Lamaakal et al., 18 Aug 2025).
For footprint and latency, the reported summary is that on microcontrollers TCUQ reduces footprint and latency versus early exit and deep ensembles, “typically about 50 to 60% smaller and about 30 to 45% faster,” while methods of similar accuracy often run out of memory. More specific device-level results are also reported. On the Big-MCU, CIFAR-10 TCUQ is 38% faster than DEEP, 29% faster than EE-ens, and 52% smaller in flash. On the Small-MCU, EE-ens and DEEP do not fit CIFAR-10, and on SpeechCommands TCUQ is 44% faster than DEEP and uses 22% less flash.
For accuracy-drop detection under CID streams, the evaluation streams clean ID data followed by corrupted samples from MNIST-C, CIFAR-10-C, TinyImageNet-C, and SpeechCmd-C at all severities. A drop event is labeled when a moving-window accuracy falls 1 below the ID mean, and the nonconformity score 2 is evaluated by AUPRC. The reported averages include MNIST-C with BASE 3, EE-ens 4, DEEP 5, and TCUQ 6; SpeechCmd-C with BASE 7, EE-ens 8, DEEP 9, and TCUQ 0; CIFAR-10-C at severity 5 with BASE 1, EE-ens 2, DEEP 3, and TCUQ 4; and TinyImageNet-C at severity 5 with TCUQ reaching up to 5 AUPRC.
For failure detection, the paper reports AUROC results for both ID-correct versus ID-incorrect classification and ID-correct versus OOD discrimination. In the ID-correct versus ID-incorrect setting, TCUQ reaches 6 on MNIST, 7 on SpeechCommands, and 8 on CIFAR-10, where it ties MC Dropout. In the ID-correct versus OOD setting, it reports 9 on MNIST, 0 on SpeechCommands, where it ties DEEP, and 1 on CIFAR-10.
For in-distribution calibration, the reported proper-score metrics include MNIST with TCUQ NLL 2 versus BASE 3, Brier 4 versus 5, and 6 versus 7; SpeechCommands with NLL 8 versus 9 and ECE 00 versus 01; and CIFAR-10 for a capacity-matched “TCUQ+” with NLL 02 versus DEEP 03 and Brier 04.
6. Trade-offs, limitations, and extensions
The method exposes several explicit design trade-offs. Increasing the window size 05 provides more stable signals but increases memory as 06. Enlarging the lag set 07 captures longer-term consistency but introduces slight extra delay and 08 operations. The blend parameter 09 controls the relative contribution of temporal cue 10 and instantaneous confidence (Lamaakal et al., 18 Aug 2025).
The paper also states several limitations. Hyperparameter sensitivity to 11 remains present, although it states that small sweeps on a development set suffice. Very rapid distribution jumps can transiently mis-calibrate the method until the quantile adapts. On very large backbones, ultra-small error rates may still favor deep-ensemble calibration.
Proposed extensions include adaptive windowing or learnable lag schedules to shrink state further, integer-only implementations of temporal kernels such as cosine similarity and JSD for sub-12s inference on tiny cores, hybridization with lightweight OOD scores such as Mahalanobis or energy when extra headroom permits, extension to streaming Transformers using the [CLS] token for feature stability and token-wise divergences, and long-horizon field studies under real-world drift.
Taken together, these points situate TCUQ as a resource-constrained uncertainty monitor rather than a general-purpose replacement for all calibration strategies. The paper’s results suggest that temporal consistency, when coupled with streaming conformal calibration, can provide a practical and resource efficient foundation for on-device monitoring in TinyML, particularly in regimes where multi-pass or ensemble-based methods are too costly to deploy.