Papers
Topics
Authors
Recent
Search
2000 character limit reached

Streaming Unlearning Techniques

Updated 19 July 2026
  • Streaming unlearning is a process that removes data influence in real time by interleaving learning, deletion, and prediction on evolving data streams.
  • It addresses challenges such as non-stationary distributions, finite-memory constraints, and counterfactual state alignment in online and federated environments.
  • Recent approaches—from online L-BFGS and SAFE to hardware-accelerated methods—emphasize efficient, certified deletion while sustaining continual model performance.

Searching arXiv for papers on streaming unlearning and closely related online/federated/continual unlearning. Streaming unlearning denotes machine unlearning in settings where data and deletion requests arrive over time, rather than as a single post hoc batch operation. In this regime, learning, prediction, and forgetting are interleaved on a live event stream, and the target of unlearning is typically defined relative to a counterfactual model or optimizer state that would have arisen had the deleted data never been processed. Recent work frames the problem through online regret, distribution shift, deletion capacity, finite-memory optimizer dynamics, and persistent erasure under continued training, thereby moving beyond the static, i.i.d. assumptions that dominate earlier batch unlearning formulations (Stewart, 13 Aug 2025, Shen et al., 21 Jul 2025).

1. From batch deletion to event-driven forgetting

Most prior machine unlearning formulations assume a fixed dataset SS, a later deletion request U⊂SU \subset S, and a one-shot post hoc unlearning procedure that does not continue learning afterward. The stream-native view departs from this by treating production learning systems as live streams in which inserts, updates, deletes, and predictions occur continuously. In "Mo' Memory, Mo' Problems: Stream-Native Machine Unlearning" (Stewart, 13 Aug 2025), the event stream is formalized as

{Et}t=1N,Et∈{insert(xt,yt), delete(ut)},\{E_t\}_{t=1}^N,\qquad E_t \in \{\text{insert}(x_t,y_t),\ \text{delete}(u_t)\},

with deletes interleaved arbitrarily with future inserts. This setting is motivated by recommender systems, ad ranking, and fraud detection, where the underlying distribution is non-stationary and methods that store gradients or Hessians for all training points become infeasible over long horizons (Stewart, 13 Aug 2025).

A complementary formulation appears in "Machine Unlearning for Streaming Forgetting" (Shen et al., 21 Jul 2025), which models a sequence of forgetting requests

F1,F2,…,FT,F_1, F_2, \dots, F_T,

over an evolving retained dataset

Dt=Dt−1∖Ft.D_t = D_{t-1} \setminus F_t.

There, streaming unlearning is treated as a sequence of deletion-induced distribution shifts, and the objective after each request is to track the retrained model

wt∗=arg⁡min⁡wL(Dt,w)w_t^* = \arg\min_w \mathcal{L}(D_t, w)

without re-accessing the full original training set (Shen et al., 21 Jul 2025).

These formulations differ in emphasis but share a common core: streaming unlearning is not merely repeated batch deletion. It is an online process in which the model must remain usable while deletion requests are honored under memory, computation, and data-retention constraints. A plausible implication is that any unlearning method whose computational or storage cost scales with the full training history is structurally mismatched to the intended deployment regime.

2. Formal models: regret, indistinguishability, and counterfactual state

One formal line of work defines streaming unlearning through online learning. In (Stewart, 13 Aug 2025), the central object is a Memory Pair (A,Aˉ)(A,\bar A), a coupled learner–unlearner acting on the event stream. If EtE_t is an insert event, the learner updates shared state θt\theta_t by

θt=A(θt−1,Et),\theta_t = A(\theta_{t-1}, E_t),

whereas on a delete event the unlearner produces

U⊂SU \subset S0

The gold standard is the ideal replay model U⊂SU \subset S1, defined by replaying only the insert events that have not been deleted, in correct order, from scratch (Stewart, 13 Aug 2025). A U⊂SU \subset S2-memory pair is then required to satisfy online accuracy, online unlearning via zCDP, and U⊂SU \subset S3-deletion capacity. This construction translates the batch requirement of being indistinguishable from retraining into a continual streaming condition (Stewart, 13 Aug 2025).

A second line of work formulates the problem in online convex optimization. "Online Learning and Unlearning" (Hu et al., 13 May 2025) defines an online learner–unlearner (OLU) that must process deletion requests interleaved with loss observations. Its online unlearning guarantee is stated using U⊂SU \subset S4-Rènyi divergence over output trajectories: after a deletion request for U⊂SU \subset S5 is processed at time U⊂SU \subset S6, all future outputs on the interval U⊂SU \subset S7 must be statistically indistinguishable from those produced by retraining on the sequence where the deleted losses are replaced by a skip symbol U⊂SU \subset S8 (Hu et al., 13 May 2025). The corresponding regret notion is interval-wise dynamic regret against the best comparator after the first U⊂SU \subset S9 deletions.

A third formulation treats the target of unlearning as an optimizer state, not just a parameter vector. "Form and Function: Machine Unlearning as a Problem of Misaligned States" (Stewart, 17 May 2026) defines an actual event history

{Et}t=1N,Et∈{insert(xt,yt), delete(ut)},\{E_t\}_{t=1}^N,\qquad E_t \in \{\text{insert}(x_t,y_t),\ \text{delete}(u_t)\},0

and a deletion-edited counterfactual history

{Et}t=1N,Et∈{insert(xt,yt), delete(ut)},\{E_t\}_{t=1}^N,\qquad E_t \in \{\text{insert}(x_t,y_t),\ \text{delete}(u_t)\},1

where each event has the form

{Et}t=1N,Et∈{insert(xt,yt), delete(ut)},\{E_t\}_{t=1}^N,\qquad E_t \in \{\text{insert}(x_t,y_t),\ \text{delete}(u_t)\},2

For online L-BFGS, the optimizer state is

{Et}t=1N,Et∈{insert(xt,yt), delete(ut)},\{E_t\}_{t=1}^N,\qquad E_t \in \{\text{insert}(x_t,y_t),\ \text{delete}(u_t)\},3

with {Et}t=1N,Et∈{insert(xt,yt), delete(ut)},\{E_t\}_{t=1}^N,\qquad E_t \in \{\text{insert}(x_t,y_t),\ \text{delete}(u_t)\},4 the current parameters and {Et}t=1N,Et∈{insert(xt,yt), delete(ut)},\{E_t\}_{t=1}^N,\qquad E_t \in \{\text{insert}(x_t,y_t),\ \text{delete}(u_t)\},5 the curvature memory. The counterfactual target is the full state {Et}t=1N,Et∈{insert(xt,yt), delete(ut)},\{E_t\}_{t=1}^N,\qquad E_t \in \{\text{insert}(x_t,y_t),\ \text{delete}(u_t)\},6 produced by running the same optimizer on the deletion-edited stream. The state-alignment error is

{Et}t=1N,Et∈{insert(xt,yt), delete(ut)},\{E_t\}_{t=1}^N,\qquad E_t \in \{\text{insert}(x_t,y_t),\ \text{delete}(u_t)\},7

where {Et}t=1N,Et∈{insert(xt,yt), delete(ut)},\{E_t\}_{t=1}^N,\qquad E_t \in \{\text{insert}(x_t,y_t),\ \text{delete}(u_t)\},8 is the post-unlearning state (Stewart, 17 May 2026). This explicitly rejects the assumption that streaming unlearning can be evaluated solely in parameter space.

Related continual-learning work does not usually use the term unlearning, but it sharpens the distinction between situations where forgetting is desirable and those where it is harmful. "A Practical Guide to Streaming Continual Learning" (Cossu et al., 2 Mar 2026) distinguishes real drift, in which {Et}t=1N,Et∈{insert(xt,yt), delete(ut)},\{E_t\}_{t=1}^N,\qquad E_t \in \{\text{insert}(x_t,y_t),\ \text{delete}(u_t)\},9 changes and some past knowledge must be updated or suppressed, from virtual drift, in which F1,F2,…,FT,F_1, F_2, \dots, F_T,0 changes while previous decision rules remain valid. The paper states that forgetting should be selective and context-aware rather than indiscriminate (Cossu et al., 2 Mar 2026). This suggests that streaming unlearning is naturally aligned with real-drift regions, privacy deletion requests, and other settings in which previously useful information has become invalid or impermissible.

3. Algorithmic families and update mechanisms

Recent algorithms for streaming unlearning differ chiefly in what object they update: parameters, optimizer state, sufficient statistics, or hardware-resident layer weights.

In (Stewart, 13 Aug 2025), the proposed stream-native Memory Pair uses online L-BFGS as both learner and unlearner. On an insert event for F1,F2,…,FT,F_1, F_2, \dots, F_T,1, the learner computes the gradient

F1,F2,…,FT,F_1, F_2, \dots, F_T,2

obtains an L-BFGS direction

F1,F2,…,FT,F_1, F_2, \dots, F_T,3

updates parameters

F1,F2,…,FT,F_1, F_2, \dots, F_T,4

and appends the curvature pair

F1,F2,…,FT,F_1, F_2, \dots, F_T,5

to bounded memory. On a delete event for F1,F2,…,FT,F_1, F_2, \dots, F_T,6, it computes

F1,F2,…,FT,F_1, F_2, \dots, F_T,7

removes influence by

F1,F2,…,FT,F_1, F_2, \dots, F_T,8

and adds calibrated Gaussian noise

F1,F2,…,FT,F_1, F_2, \dots, F_T,9

to obtain zCDP-based unlearning guarantees (Stewart, 13 Aug 2025). The key engineering substitution is replacing explicit Hessian inversion with limited-memory quasi-Newton structure.

SAFE, introduced in (Shen et al., 21 Jul 2025), takes a different route. It decomposes the per-step risk into a retention term and a forgetting term and avoids revisiting the full original dataset by storing a precomputed gradient of the original risk at Dt=Dt−1∖Ft.D_t = D_{t-1} \setminus F_t.0 and maintaining lightweight distributional statistics. The retention gradient is updated recursively as

Dt=Dt−1∖Ft.D_t = D_{t-1} \setminus F_t.1

while the forgetting term is approximated by reweighting the original predictor using estimates of label shift and class-conditional shift (Shen et al., 21 Jul 2025). SAFE uses class counts Dt=Dt−1∖Ft.D_t = D_{t-1} \setminus F_t.2, random-projection latent Gaussians, and incremental updates of means and covariances to approximate the post-deletion conditional distribution. It then performs a single normalized gradient step from Dt=Dt−1∖Ft.D_t = D_{t-1} \setminus F_t.3: Dt=Dt−1∖Ft.D_t = D_{t-1} \setminus F_t.4 where Dt=Dt−1∖Ft.D_t = D_{t-1} \setminus F_t.5 combines retention and forgetting gradients (Shen et al., 21 Jul 2025).

The OLU framework of (Hu et al., 13 May 2025) presents two OGD-based algorithms. Passive OLU leaves the online update rule unchanged except at deletion times, where it injects Gaussian noise calibrated to the residual influence of the deleted point under contractive dynamics. Active OLU instead invokes an offline unlearning procedure derived from Descent-to-Delete at deletion times, first moving toward the ERM on all past points and then toward the ERM excluding deleted points, followed by Gaussian noise injection (Hu et al., 13 May 2025). The passive method has essentially no additional optimization cost beyond OGD, while the active method pays extra gradient-descent passes over past data.

Unlearning can also be used to simulate sliding-window adaptation under concept drift. "Unlearning-based sliding window for continual learning under concept drift" (Wozniak et al., 15 Mar 2026) proposes UIL, which after the window fills replaces retraining on the current window by the update

Dt=Dt−1∖Ft.D_t = D_{t-1} \setminus F_t.6

In this formulation, unlearning removes the oldest chunk leaving the window, and incremental training incorporates the newest chunk (Wozniak et al., 15 Mar 2026). The paper explicitly positions this as a computationally cheaper approximation to standard sliding-window retraining.

At the edge-hardware level, FiCABU realizes streaming-style unlearning within a GEMM-centric accelerator pipeline. It computes Fisher diagonals on the forget set, applies SSD-style parameter dampening, begins at back-end layers, and stops when forget accuracy reaches a target Dt=Dt−1∖Ft.D_t = D_{t-1} \setminus F_t.7 (Cho et al., 6 Nov 2025). Although the paper focuses on retraining-free edge unlearning rather than formal online regret, it explicitly treats unlearning as a lightweight incremental editing phase that runs through the same streaming pipeline as inference (Cho et al., 6 Nov 2025).

4. Guarantees: regret, dynamic tracking, and deletion capacity

Streaming unlearning theory has largely adopted regret-style guarantees because the target model changes over time.

In (Stewart, 13 Aug 2025), the online learner is evaluated using cumulative regret

Dt=Dt−1∖Ft.D_t = D_{t-1} \setminus F_t.8

with vanishing average regret as the online criterion. Under Dt=Dt−1∖Ft.D_t = D_{t-1} \setminus F_t.9-strong convexity, wt∗=arg⁡min⁡wL(Dt,w)w_t^* = \arg\min_w \mathcal{L}(D_t, w)0-Lipschitz losses, and uniformly bounded L-BFGS preconditioners wt∗=arg⁡min⁡wL(Dt,w)w_t^* = \arg\min_w \mathcal{L}(D_t, w)1, the paper proves logarithmic static regret

wt∗=arg⁡min⁡wL(Dt,w)w_t^* = \arg\min_w \mathcal{L}(D_t, w)2

and dynamic regret

wt∗=arg⁡min⁡wL(Dt,w)w_t^* = \arg\min_w \mathcal{L}(D_t, w)3

where

wt∗=arg⁡min⁡wL(Dt,w)w_t^* = \arg\min_w \mathcal{L}(D_t, w)4

is the comparator path length (Stewart, 13 Aug 2025). The paper identifies the logarithmic wt∗=arg⁡min⁡wL(Dt,w)w_t^* = \arg\min_w \mathcal{L}(D_t, w)5 result as the first such regret bound for a certified online unlearning algorithm (Stewart, 13 Aug 2025).

SAFE instead establishes a non-convex dynamic-regret-style guarantee. Defining

wt∗=arg⁡min⁡wL(Dt,w)w_t^* = \arg\min_w \mathcal{L}(D_t, w)6

it proves cumulative streaming unlearning regret

wt∗=arg⁡min⁡wL(Dt,w)w_t^* = \arg\min_w \mathcal{L}(D_t, w)7

without assuming convex losses (Shen et al., 21 Jul 2025). This directly tracks the movement of the retrained comparator sequence rather than a fixed optimal model.

OLU algorithms in (Hu et al., 13 May 2025) show that online unlearning can preserve regret rates comparable to standard OGD. For strongly convex losses, passive OLU achieves regret with the same wt∗=arg⁡min⁡wL(Dt,w)w_t^* = \arg\min_w \mathcal{L}(D_t, w)8 dependence as standard OGD up to deletion-dependent overheads, while active OLU attains

wt∗=arg⁡min⁡wL(Dt,w)w_t^* = \arg\min_w \mathcal{L}(D_t, w)9

under strong convexity, smoothness, and an alignment assumption on interval-wise minimizers (Hu et al., 13 May 2025). The significance is not a universal best bound but the preservation of online-learning competitiveness under certified deletion.

Deletion capacity is the other recurrent theoretical quantity. In (Stewart, 13 Aug 2025), sample complexity (A,Aˉ)(A,\bar A)0 is defined as the smallest horizon (A,Aˉ)(A,\bar A)1 such that average regret remains below (A,Aˉ)(A,\bar A)2 after at most (A,Aˉ)(A,\bar A)3 deletions. The corresponding (A,Aˉ)(A,\bar A)4-Deletion Capacity theorem states that average regret remains (A,Aˉ)(A,\bar A)5 after (A,Aˉ)(A,\bar A)6 events if

(A,Aˉ)(A,\bar A)7

with worst-case dependence (A,Aˉ)(A,\bar A)8 when (A,Aˉ)(A,\bar A)9 (Stewart, 13 Aug 2025). This quantity measures how many streaming deletions can be absorbed before retraining is required.

Older convex unlearning theory offers a related but batch-defined notion of deletion capacity. "Remember What You Want to Forget: Algorithms for Machine Unlearning" (Sekhari et al., 2021) defines EtE_t0 as the maximum number of deleted samples for which population risk remains small under worst-case deletions, and proves capacity on the order of

EtE_t1

for convex Hessian-Lipschitz losses, separating unlearning-specific methods from differentially private learning-based baselines (Sekhari et al., 2021). Although this paper is not stream-native, it established deletion capacity as a central analytical primitive later adapted to streaming settings.

5. Memory, optimizer state, and systems considerations

A defining concern in streaming unlearning is that state grows with time unless the algorithm is explicitly memory-bounded. This is one of the strongest motivations for online L-BFGS and related finite-state methods.

In (Stewart, 13 Aug 2025), limited-memory L-BFGS yields both time and memory complexity

EtE_t2

per update, with storage independent of the stream length EtE_t3. This contrasts with traditional Newton-step unlearning, where direct Hessian inversion incurs EtE_t4 time and EtE_t5 memory, while per-sample gradient or curvature storage can grow as EtE_t6 or worse in long streams (Stewart, 13 Aug 2025). The stream-native framing therefore ties feasibility directly to bounded curvature memory rather than merely faster local updates.

The state-aware perspective of (Stewart, 17 May 2026) complicates this picture. For online L-BFGS, finite curvature memory does not imply finite influence in a naïve sense, because deleted data may continue to shape future curvature pairs through altered parameter trajectories. The paper introduces a memory-operator error

EtE_t7

a combined state error

EtE_t8

and an update-direction error

EtE_t9

These metrics show that streaming unlearning is a problem of aligning realizable optimizer states rather than only final parameters (Stewart, 17 May 2026). A plausible implication is that bounded-memory optimizers can make streaming unlearning computationally feasible while still leaving subtle long-range state contamination if memory is corrected only locally.

The same tension appears in hardware-oriented work. FiCABU integrates two specialized IPs—FIMD for Fisher-diagonal estimation and Dampening IP for layerwise parameter updates—into a streaming GEMM pipeline (Cho et al., 6 Nov 2025). The architecture is patch-level, with activations cached at checkpoints during a single forward pass and partial inference used to evaluate forget accuracy at intermediate layers. The paper reports that FIMD and Dampening together use 2,185 LUTs and 785 FFs on FPGA, and only 0.81 mW of system power, while the streaming unlearning path can achieve substantial MAC and energy reductions relative to SSD (Cho et al., 6 Nov 2025). The emphasis here is not on regret or counterfactual state but on co-design: unlearning becomes another streaming workload scheduled on the same inference pipeline.

A different type of memory offloading appears in "Ticketed Learning-Unlearning Schemes" (Ghazi et al., 2023). There, the learner stores only small central state θt\theta_t0 and returns small per-example tickets θt\theta_t1 to data owners. Unlearning later uses only the deleted examples, their tickets, and the central state to reconstruct exactly the predictor that would have been produced by retraining on survivors (Ghazi et al., 2023). This is not an online streaming-deletion model, but it is structurally related: central memory stays small by pushing part of the deletion-relevant state to the edge.

6. Continued training, drift, and persistent erasure

A central complication of streaming unlearning is that deletion may not remain effective once learning continues. This is explicit in federated and continual settings.

"Lethe: Adapter-Augmented Dual-Stream Update for Persistent Knowledge Erasure in Federated Unlearning" (Tan et al., 30 Jan 2026) identifies the failure mode of knowledge resurfacing: after unlearning in federated learning, continued FedAvg training on the remaining data can reactivate erased knowledge. The paper quantifies this with the Resurfacing Rate

θt\theta_t2

where θt\theta_t3, θt\theta_t4, and θt\theta_t5 are accuracies on the unlearning set before unlearning, after unlearning, and after continued training, respectively (Tan et al., 30 Jan 2026). Lethe addresses this by extracting a layerwise unlearning direction using a temporary adapter trained with gradient ascent on the unlearning data, then rectifying retained-data updates to subtract or oppose positively aligned components. Continued training is thus analyzed as a streaming process after the unlearning event, not as a separate postscript (Tan et al., 30 Jan 2026).

Continual-learning work provides additional context. In (Cossu et al., 2 Mar 2026), replay-based methods such as Experience Replay are shown to help under virtual drift but to become harmful under real drift because they keep reinforcing decision boundaries associated with obsolete concepts. The paper states that in the presence of real drift, replay preserves information that the drift has made obsolete (Cossu et al., 2 Mar 2026). This observation transfers directly to streaming unlearning: any mechanism that maintains or replays forgotten examples, or latent proxies for them, can undermine deletion goals during subsequent adaptation.

UIL (Wozniak et al., 15 Mar 2026) casts this issue in sliding-window terms. Once the active window is full, the oldest chunk should become irrelevant and the newest chunk relevant. The paper analyzes the deviation between the ideal sliding-window retrained model θt\theta_t6 and the unlearn-then-train approximation θt\theta_t7, assuming a per-step parameter discrepancy

θt\theta_t8

After a full window shift, this yields

θt\theta_t9

and, by θt=A(θt−1,Et),\theta_t = A(\theta_{t-1}, E_t),0-smoothness,

θt=A(θt−1,Et),\theta_t = A(\theta_{t-1}, E_t),1

(Wozniak et al., 15 Mar 2026). The interpretation is that continuous unlearning can approximate relevance-based forgetting, but approximation error accumulates with window size and with the quality of the unlearning operator.

The state-alignment analysis in (Stewart, 17 May 2026) offers a related persistence result at the optimizer level. Under contractive updates, the post-deletion state deviation satisfies

θt=A(θt−1,Et),\theta_t = A(\theta_{t-1}, E_t),2

leading to

θt=A(θt−1,Et),\theta_t = A(\theta_{t-1}, E_t),3

This means that misalignment decays geometrically only if new perturbations do not keep re-entering through future curvature updates (Stewart, 17 May 2026). In practical terms, continued learning can either wash out deletion error or re-inject it, depending on how the optimizer state is repaired.

7. Empirical evidence, limitations, and unresolved questions

Empirical evidence for streaming unlearning is spread across several distinct evaluation paradigms.

On image, tabular, and text benchmarks, SAFE is reported to remain close to retraining while being substantially faster than baseline unlearning methods. The paper evaluates RA, FA, TA, and MIA under 20-round and other multi-round deletion protocols and reports that SAFE consistently achieves performance closest to retrain while operating without revisiting the original dataset (Shen et al., 21 Jul 2025). These results support the claim that distribution-shift tracking can be an effective abstraction for streaming forgetting.

In (Stewart, 13 Aug 2025), preliminary experiments on MNIST treated as a stream compare Memory Pair with AdaGrad, SGD, and Online Newton Step. Memory Pair and AdaGrad achieve sublinear cumulative regret with vanishing instantaneous regret, whereas SGD and Online Newton Step do not show bounded regret in that setup (Stewart, 13 Aug 2025). The paper states that deletion-capacity experiments are forthcoming, so the empirical evidence is stronger on the online-learning side than on the full delete-interleaved setting.

FiCABU reports concrete edge metrics. On INT8 hardware with ResNet-18, system energy is reduced to 6.48 percent of the SSD baseline on CIFAR-20 and 0.13 percent on PinsFaceRecognition, while forget accuracy reaches random-guess level and retain preservation is improved (Cho et al., 6 Nov 2025). These experiments concern retraining-free on-device editing rather than replay-based or regret-based streaming with arbitrary delete times, but they establish that layer-adaptive unlearning can be embedded into a streaming accelerator without prohibitive overhead.

UIL evaluates unlearning-based sliding-window approximation under abrupt concept drift with ResNet-18. It is about 3× faster per batch than full sliding-window retraining in the reported setup and uses slightly less memory per batch, while recovery behavior varies by drift type: it is faster on MNIST noise drift but can lag on semantic drift in Fashion-MNIST (Wozniak et al., 15 Mar 2026). This indicates that approximate streaming forgetting may be easier when drift is low-level or covariate-like than when the label semantics themselves change.

Lethe evaluates persistent erasure under continued federated training. Across client-level, sample-level, and class-level unlearning, it maintains RR below 1% in most cases, while several baselines exhibit substantial resurfacing during follow-up rounds (Tan et al., 30 Jan 2026). This directly addresses a misconception common in static unlearning evaluations: low forget accuracy immediately after deletion does not imply persistent forgetting once learning resumes.

Several limitations recur across the literature. Convexity, strong convexity, smoothness, and bounded-Hessian assumptions are central to the strongest regret and state-alignment theorems (Stewart, 13 Aug 2025, Hu et al., 13 May 2025, Stewart, 17 May 2026). SAFE relaxes convexity but depends on latent Gaussian approximations and a single-step update from θt=A(θt−1,Et),\theta_t = A(\theta_{t-1}, E_t),4 (Shen et al., 21 Jul 2025). FiCABU targets practical random-guess forgetting and retain preservation rather than certified retrain equivalence (Cho et al., 6 Nov 2025). UIL explicitly treats its unlearning operator as approximate rather than strongly certified and notes sensitivity to hyperparameters and drift type (Wozniak et al., 15 Mar 2026). Ticketed exact unlearning delivers exactness for structured concept classes, but only in a one-shot deletion model rather than a fully dynamic stream (Ghazi et al., 2023).

Open questions therefore cluster around four themes. The first is extension to non-convex and large-scale models with meaningful guarantees, especially when the optimizer carries substantial internal state. The second is composition: how to support many deletion events without either repeated retraining or uncontrolled accumulation of noise and approximation error. The third is hybridization with continual-learning machinery such as replay, regularization, and modularity, especially under mixed real and virtual drift (Cossu et al., 2 Mar 2026). The fourth is systems co-design: how to place unlearning inside production pipelines, edge accelerators, or federated orchestration while preserving both efficiency and persistent erasure (Cho et al., 6 Nov 2025, Tan et al., 30 Jan 2026).

Streaming unlearning has thus emerged as a distinct research area rather than a mere variant of batch deletion. Its core problem is to align a live, stateful learning system with a deletion-edited counterfactual under continuous operation. Depending on the formulation, that alignment is expressed as regret against retraining, zCDP or Rènyi-indistinguishability, approximation to a sliding-window oracle, or full optimizer-state matching. What unifies these views is the insistence that forgetting must be compatible with ongoing learning rather than terminating it.

Topic to Video (Beta)

No one has generated a video about this topic yet.

Whiteboard

No one has generated a whiteboard explanation for this topic yet.

Follow Topic

Get notified by email when new papers are published related to Streaming Unlearning.