Papers
Topics
Authors
Recent
Search
2000 character limit reached

TCUQ: Single-Pass Uncertainty for TinyML

Updated 8 July 2026
  • The paper introduces TCUQ, a single-pass, label-free uncertainty monitor that transforms short-horizon temporal signals into calibrated risk scores for streaming TinyML applications.
  • It employs a lightweight streaming conformal layer with O(1) per step updates to enforce an accept/abstain decision rule under fluctuating distribution shifts.
  • Empirical evaluations demonstrate that TCUQ reduces latency and memory usage compared to ensemble methods while maintaining robust detection and calibration performance on microcontrollers.

TCUQ, introduced as “Single-Pass Uncertainty Quantification from Temporal Consistency with Streaming Conformal Calibration for TinyML,” is a single pass, label free uncertainty monitor for streaming TinyML that converts short horizon temporal consistency captured via lightweight signals on posteriors and features into a calibrated risk score with an O(W)O(W) ring buffer and O(1)O(1) per step updates. It is designed for battery-powered microcontrollers operating on continuous input streams, where in-distribution, corrupted-in-distribution, and out-of-distribution samples may interleave, and where uncertainty quantification must be obtained without online labels, extra forward passes, or large memory overheads. The method couples temporal-consistency features with a streaming conformal layer to produce a budgeted accept/abstain rule, yielding calibrated behavior on kilobyte-scale devices (Lamaakal et al., 18 Aug 2025).

1. Problem formulation and operating regime

TCUQ is motivated by the fact that TinyML systems must satisfy tight memory, compute, and energy budgets while remaining robust to distribution shift in continuous streams. The target setting includes three stream components: in-distribution (ID) samples, corrupted-in-distribution (CID) variants such as noise, blur, and fog, and out-of-distribution (OOD) inputs. In this setting, deep networks that may appear calibrated on static test sets can become overconfident under CID and OOD conditions, which creates a need for on-device uncertainty quantification that detects accuracy drops in real time, requires no extra forward passes or labels at inference, fits in O(kilobytes)O(\text{kilobytes}) of flash and RAM, and updates in O(1)O(1) time per input (Lamaakal et al., 18 Aug 2025).

The method is explicitly positioned against several established approaches. Deep ensembles, MC Dropout, and early-exit heads either multiply inference cost by KK passes or add heavyweight heads. Post-hoc calibration is described as failing under dynamic shifts. Standard conformal predictors provide distribution-free risk control, but standard implementations store all scores or require labels. TCUQ addresses this combination of constraints by extracting a single uncertainty signal from short-horizon temporal consistency of posteriors and features, calibrating it on the fly via a memory-light streaming conformal layer, and enforcing a user-specified accept/abstain budget.

A plausible implication is that TCUQ is not primarily an alternative classifier architecture; rather, it is an uncertainty monitor layered on top of a trained backbone. This distinction matters because its computational and memory costs are additive to a single existing forward pass rather than requiring architectural replication or multi-pass inference.

2. Temporal-consistency signals and uncertainty scoring

TCUQ assumes a trained backbone pϕ(y∣x)p_\phi(y \mid x) with final feature vector ft=f(xt)∈Rdf_t = f(x_t) \in \mathbb{R}^d and class probabilities pt=pϕ(⋅∣xt)∈ΔL−1p_t = p_\phi(\cdot \mid x_t) \in \Delta^{L-1}. The hard decision and two instantaneous confidence statistics are defined as

y^t=arg⁡max⁡ℓ pt[ℓ],Ct=max⁡ℓ pt[ℓ],Δt=pt(1)−pt(2),\hat y_t = \arg\max_\ell\, p_t[\ell], \qquad C_t = \max_\ell\, p_t[\ell], \qquad \Delta_t = p_t^{(1)} - p_t^{(2)},

where pt(1)≥pt(2)≥⋯p_t^{(1)} \ge p_t^{(2)} \ge \cdots are the sorted probabilities (Lamaakal et al., 18 Aug 2025).

A ring buffer of window size O(1)O(1)0 stores recent posterior-feature pairs O(1)O(1)1. From this buffer, TCUQ computes four lightweight temporal signals. Let O(1)O(1)2 be a fixed lag set, for example O(1)O(1)3, with weights O(1)O(1)4 and O(1)O(1)5.

The first signal is a multi-lag predictive divergence based on Jensen–Shannon divergence: O(1)O(1)6 with

O(1)O(1)7

The second signal is feature stability, computed through cosine similarity,

O(1)O(1)8

and used via the instability term O(1)O(1)9.

The third signal is decision persistence,

O(kilobytes)O(\text{kilobytes})0

and is used through the inconsistency term O(kilobytes)O(\text{kilobytes})1.

The fourth signal is an instantaneous confidence proxy blending inverse maximum probability and inverse margin: O(kilobytes)O(\text{kilobytes})2

These are stacked into

O(kilobytes)O(\text{kilobytes})3

TCUQ then applies a tiny logistic combiner fitted offline: O(kilobytes)O(\text{kilobytes})4 where O(kilobytes)O(\text{kilobytes})5 are learned once on a small labeled development set mixing ID, CID, and OOD examples. At inference, this reduces to one inner product and one sigmoid.

This design makes the uncertainty estimate explicitly temporal rather than purely instantaneous. The paper’s formulation suggests that TCUQ exploits local temporal smoothness or consistency as an operational prior: when posteriors, features, and decisions fluctuate over short lags, risk increases even if a single-step confidence score appears high.

3. Streaming conformal calibration and abstention control

The stepwise uncertainty estimate is transformed into a scalar nonconformity score

O(kilobytes)O(\text{kilobytes})6

A streaming conformal layer maintains an online estimate O(kilobytes)O(\text{kilobytes})7 of the O(kilobytes)O(\text{kilobytes})8-quantile of O(kilobytes)O(\text{kilobytes})9 using a memory-constant quantile tracker (Lamaakal et al., 18 Aug 2025).

One concrete update rule given is the Robbins–Monro update: O(1)O(1)0 Only the scalar O(1)O(1)1 and one conditional are required per step. During warm-up, a tiny buffer of size O(1)O(1)2 is used to initialize O(1)O(1)3 to the empirical quantile, after which the streaming update is used.

The decision rule is budgeted accept/abstain. If O(1)O(1)4 and the abstention controller still has budget, the system abstains; otherwise it accepts and returns O(1)O(1)5. The controller can enforce a long-run abstention fraction near a user budget O(1)O(1)6 by allowing at most O(1)O(1)7 abstentions in O(1)O(1)8 steps.

The paper states a formal calibration guarantee under the standard conformal exchangeability assumption on O(1)O(1)9: KK0 For the streaming setting, it further states via Hoeffding’s inequality on exceedance indicators KK1 that for any KK2, with probability at least KK3,

KK4

Thus, as KK5, the empirical reject rate converges to KK6 with high-confidence error KK7.

A common misconception would be to treat TCUQ as a fully label-free conformal predictor in the classical offline sense. The paper is more specific: the monitor is label free online, while the logistic combiner is fitted offline on a small labeled development set. The conformal component then calibrates the resulting score during deployment without online labels or score storage beyond constant memory.

4. Data structures, computational profile, and embedded implementation

The central data structure is a ring buffer of size KK8 that stores the last KK9 posterior vectors of length pϕ(y∣x)p_\phi(y \mid x)0 and feature vectors of length pϕ(y∣x)p_\phi(y \mid x)1. This yields memory complexity

pϕ(y∣x)p_\phi(y \mid x)2

The paper notes that in practice pϕ(y∣x)p_\phi(y \mid x)3 can be reduced to pϕ(y∣x)p_\phi(y \mid x)4 by an offline-learned pϕ(y∣x)p_\phi(y \mid x)5 projection or global pooling (Lamaakal et al., 18 Aug 2025).

Per input pϕ(y∣x)p_\phi(y \mid x)6, the method performs one forward pass through the backbone, one buffer overwrite at pointer pϕ(y∣x)p_\phi(y \mid x)7, computation of pϕ(y∣x)p_\phi(y \mid x)8, pϕ(y∣x)p_\phi(y \mid x)9, ft=f(xt)∈Rdf_t = f(x_t) \in \mathbb{R}^d0, and ft=f(xt)∈Rdf_t = f(x_t) \in \mathbb{R}^d1 by retrieving ft=f(xt)∈Rdf_t = f(x_t) \in \mathbb{R}^d2 past entries, one vector dot product to obtain ft=f(xt)∈Rdf_t = f(x_t) \in \mathbb{R}^d3, one scalar blend to obtain ft=f(xt)∈Rdf_t = f(x_t) \in \mathbb{R}^d4, one streaming-quantile update, and one budget-controller check. Since ft=f(xt)∈Rdf_t = f(x_t) \in \mathbb{R}^d5 is constant, for example ft=f(xt)∈Rdf_t = f(x_t) \in \mathbb{R}^d6, all additional monitoring operations are ft=f(xt)∈Rdf_t = f(x_t) \in \mathbb{R}^d7 per step aside from the fixed-cost forward pass.

The implementation targets microcontrollers directly. TCUQ was compiled on two representative MCUs: STM32F767ZI, described as “Big-MCU” with 512 KB Flash and a 216 MHz Cortex-M7, and STM32L432KC, described as “Small-MCU” with 256 KB Flash and an 80 MHz Cortex-M4. On microcontrollers, TCUQ is reported to fit comfortably on kilobyte scale devices, adding less than 4 KB RAM for ft=f(xt)∈Rdf_t = f(x_t) \in \mathbb{R}^d8 and less than 1 KB Flash for the buffer and quantile code.

This computational profile is significant because it separates the method from heavier uncertainty-estimation schemes. The paper’s framing suggests that TCUQ’s principal engineering contribution lies not only in detection quality, but in the fact that the monitor remains viable under deployment constraints where alternative methods do not fit.

5. Empirical behavior on microcontrollers and streaming benchmarks

The experimental evaluation uses SpeechCommands with DSCNN, MNIST, CIFAR-10 with ResNet-8, and TinyImageNet with MobileNetV2. Baselines include BASE confidence threshold, MC Dropout with ft=f(xt)∈Rdf_t = f(x_t) \in \mathbb{R}^d9, Deep ensembles with pt=pϕ(⋅∣xt)∈ΔL−1p_t = p_\phi(\cdot \mid x_t) \in \Delta^{L-1}0, Early-exit ensemble with a 3-head ensemble, G-ODIN, and HYDRA (Lamaakal et al., 18 Aug 2025).

For footprint and latency, the reported summary is that on microcontrollers TCUQ reduces footprint and latency versus early exit and deep ensembles, “typically about 50 to 60% smaller and about 30 to 45% faster,” while methods of similar accuracy often run out of memory. More specific device-level results are also reported. On the Big-MCU, CIFAR-10 TCUQ is 38% faster than DEEP, 29% faster than EE-ens, and 52% smaller in flash. On the Small-MCU, EE-ens and DEEP do not fit CIFAR-10, and on SpeechCommands TCUQ is 44% faster than DEEP and uses 22% less flash.

For accuracy-drop detection under CID streams, the evaluation streams clean ID data followed by corrupted samples from MNIST-C, CIFAR-10-C, TinyImageNet-C, and SpeechCmd-C at all severities. A drop event is labeled when a moving-window accuracy falls pt=pϕ(⋅∣xt)∈ΔL−1p_t = p_\phi(\cdot \mid x_t) \in \Delta^{L-1}1 below the ID mean, and the nonconformity score pt=pϕ(⋅∣xt)∈ΔL−1p_t = p_\phi(\cdot \mid x_t) \in \Delta^{L-1}2 is evaluated by AUPRC. The reported averages include MNIST-C with BASE pt=pϕ(⋅∣xt)∈ΔL−1p_t = p_\phi(\cdot \mid x_t) \in \Delta^{L-1}3, EE-ens pt=pϕ(⋅∣xt)∈ΔL−1p_t = p_\phi(\cdot \mid x_t) \in \Delta^{L-1}4, DEEP pt=pϕ(⋅∣xt)∈ΔL−1p_t = p_\phi(\cdot \mid x_t) \in \Delta^{L-1}5, and TCUQ pt=pϕ(⋅∣xt)∈ΔL−1p_t = p_\phi(\cdot \mid x_t) \in \Delta^{L-1}6; SpeechCmd-C with BASE pt=pϕ(⋅∣xt)∈ΔL−1p_t = p_\phi(\cdot \mid x_t) \in \Delta^{L-1}7, EE-ens pt=pϕ(⋅∣xt)∈ΔL−1p_t = p_\phi(\cdot \mid x_t) \in \Delta^{L-1}8, DEEP pt=pϕ(⋅∣xt)∈ΔL−1p_t = p_\phi(\cdot \mid x_t) \in \Delta^{L-1}9, and TCUQ y^t=arg⁡max⁡ℓ pt[ℓ],Ct=max⁡ℓ pt[ℓ],Δt=pt(1)−pt(2),\hat y_t = \arg\max_\ell\, p_t[\ell], \qquad C_t = \max_\ell\, p_t[\ell], \qquad \Delta_t = p_t^{(1)} - p_t^{(2)},0; CIFAR-10-C at severity 5 with BASE y^t=arg⁡max⁡ℓ pt[ℓ],Ct=max⁡ℓ pt[ℓ],Δt=pt(1)−pt(2),\hat y_t = \arg\max_\ell\, p_t[\ell], \qquad C_t = \max_\ell\, p_t[\ell], \qquad \Delta_t = p_t^{(1)} - p_t^{(2)},1, EE-ens y^t=arg⁡max⁡ℓ pt[ℓ],Ct=max⁡ℓ pt[ℓ],Δt=pt(1)−pt(2),\hat y_t = \arg\max_\ell\, p_t[\ell], \qquad C_t = \max_\ell\, p_t[\ell], \qquad \Delta_t = p_t^{(1)} - p_t^{(2)},2, DEEP y^t=arg⁡max⁡ℓ pt[ℓ],Ct=max⁡ℓ pt[ℓ],Δt=pt(1)−pt(2),\hat y_t = \arg\max_\ell\, p_t[\ell], \qquad C_t = \max_\ell\, p_t[\ell], \qquad \Delta_t = p_t^{(1)} - p_t^{(2)},3, and TCUQ y^t=arg⁡max⁡ℓ pt[ℓ],Ct=max⁡ℓ pt[ℓ],Δt=pt(1)−pt(2),\hat y_t = \arg\max_\ell\, p_t[\ell], \qquad C_t = \max_\ell\, p_t[\ell], \qquad \Delta_t = p_t^{(1)} - p_t^{(2)},4; and TinyImageNet-C at severity 5 with TCUQ reaching up to y^t=arg⁡max⁡ℓ pt[ℓ],Ct=max⁡ℓ pt[ℓ],Δt=pt(1)−pt(2),\hat y_t = \arg\max_\ell\, p_t[\ell], \qquad C_t = \max_\ell\, p_t[\ell], \qquad \Delta_t = p_t^{(1)} - p_t^{(2)},5 AUPRC.

For failure detection, the paper reports AUROC results for both ID-correct versus ID-incorrect classification and ID-correct versus OOD discrimination. In the ID-correct versus ID-incorrect setting, TCUQ reaches y^t=arg⁡max⁡ℓ pt[ℓ],Ct=max⁡ℓ pt[ℓ],Δt=pt(1)−pt(2),\hat y_t = \arg\max_\ell\, p_t[\ell], \qquad C_t = \max_\ell\, p_t[\ell], \qquad \Delta_t = p_t^{(1)} - p_t^{(2)},6 on MNIST, y^t=arg⁡max⁡ℓ pt[ℓ],Ct=max⁡ℓ pt[ℓ],Δt=pt(1)−pt(2),\hat y_t = \arg\max_\ell\, p_t[\ell], \qquad C_t = \max_\ell\, p_t[\ell], \qquad \Delta_t = p_t^{(1)} - p_t^{(2)},7 on SpeechCommands, and y^t=arg⁡max⁡ℓ pt[ℓ],Ct=max⁡ℓ pt[ℓ],Δt=pt(1)−pt(2),\hat y_t = \arg\max_\ell\, p_t[\ell], \qquad C_t = \max_\ell\, p_t[\ell], \qquad \Delta_t = p_t^{(1)} - p_t^{(2)},8 on CIFAR-10, where it ties MC Dropout. In the ID-correct versus OOD setting, it reports y^t=arg⁡max⁡ℓ pt[ℓ],Ct=max⁡ℓ pt[ℓ],Δt=pt(1)−pt(2),\hat y_t = \arg\max_\ell\, p_t[\ell], \qquad C_t = \max_\ell\, p_t[\ell], \qquad \Delta_t = p_t^{(1)} - p_t^{(2)},9 on MNIST, pt(1)≥pt(2)≥⋯p_t^{(1)} \ge p_t^{(2)} \ge \cdots0 on SpeechCommands, where it ties DEEP, and pt(1)≥pt(2)≥⋯p_t^{(1)} \ge p_t^{(2)} \ge \cdots1 on CIFAR-10.

For in-distribution calibration, the reported proper-score metrics include MNIST with TCUQ NLL pt(1)≥pt(2)≥⋯p_t^{(1)} \ge p_t^{(2)} \ge \cdots2 versus BASE pt(1)≥pt(2)≥⋯p_t^{(1)} \ge p_t^{(2)} \ge \cdots3, Brier pt(1)≥pt(2)≥⋯p_t^{(1)} \ge p_t^{(2)} \ge \cdots4 versus pt(1)≥pt(2)≥⋯p_t^{(1)} \ge p_t^{(2)} \ge \cdots5, and pt(1)≥pt(2)≥⋯p_t^{(1)} \ge p_t^{(2)} \ge \cdots6 versus pt(1)≥pt(2)≥⋯p_t^{(1)} \ge p_t^{(2)} \ge \cdots7; SpeechCommands with NLL pt(1)≥pt(2)≥⋯p_t^{(1)} \ge p_t^{(2)} \ge \cdots8 versus pt(1)≥pt(2)≥⋯p_t^{(1)} \ge p_t^{(2)} \ge \cdots9 and ECE O(1)O(1)00 versus O(1)O(1)01; and CIFAR-10 for a capacity-matched “TCUQ+” with NLL O(1)O(1)02 versus DEEP O(1)O(1)03 and Brier O(1)O(1)04.

6. Trade-offs, limitations, and extensions

The method exposes several explicit design trade-offs. Increasing the window size O(1)O(1)05 provides more stable signals but increases memory as O(1)O(1)06. Enlarging the lag set O(1)O(1)07 captures longer-term consistency but introduces slight extra delay and O(1)O(1)08 operations. The blend parameter O(1)O(1)09 controls the relative contribution of temporal cue O(1)O(1)10 and instantaneous confidence (Lamaakal et al., 18 Aug 2025).

The paper also states several limitations. Hyperparameter sensitivity to O(1)O(1)11 remains present, although it states that small sweeps on a development set suffice. Very rapid distribution jumps can transiently mis-calibrate the method until the quantile adapts. On very large backbones, ultra-small error rates may still favor deep-ensemble calibration.

Proposed extensions include adaptive windowing or learnable lag schedules to shrink state further, integer-only implementations of temporal kernels such as cosine similarity and JSD for sub-O(1)O(1)12s inference on tiny cores, hybridization with lightweight OOD scores such as Mahalanobis or energy when extra headroom permits, extension to streaming Transformers using the [CLS] token for feature stability and token-wise divergences, and long-horizon field studies under real-world drift.

Taken together, these points situate TCUQ as a resource-constrained uncertainty monitor rather than a general-purpose replacement for all calibration strategies. The paper’s results suggest that temporal consistency, when coupled with streaming conformal calibration, can provide a practical and resource efficient foundation for on-device monitoring in TinyML, particularly in regimes where multi-pass or ensemble-based methods are too costly to deploy.

Definition Search Book Streamline Icon: https://streamlinehq.com
References (1)

Topic to Video (Beta)

No one has generated a video about this topic yet.

Whiteboard

No one has generated a whiteboard explanation for this topic yet.

Follow Topic

Get notified by email when new papers are published related to TCUQ.