Form and Function: Machine Unlearning as a Problem of Misaligned States
Published 17 May 2026 in cs.LG and math.OC | (2605.17590v1)
Abstract: We formulate machine unlearning for online L-BFGS as a counterfactual state-alignment problem. Given an actual event stream and a deletion-edited counterfactual stream, the target of unlearning is the optimizer state that would have arisen had the deleted samples never been processed. We introduce state-aware metrics that separately measure parameter error, memory-operator error, combined state error, and update-direction error. The memory metric compares the inverse-Hessian actions induced by the o-L-BFGS memory, rather than treating curvature pairs as of finite influence. Under convexity assumptions, we derive a recursive bound on counterfactual state deviation. We then evaluate a state-aware benchmark of deletion interventions, including memory-only and parameter-only corrections, against an counterfactual oracle model. These results show that unlearning for online L-BFGS is not merely a parameter-correction problem: it requires alignment with a realizable counterfactual optimizer state.
The paper reframes machine unlearning for online L-BFGS as alignment of the full counterfactual state—parameters plus curvature memory—rather than parameter matching alone.
Theoretical bounds show that deletion effects can persist through indirect optimizer memory after directly contaminated curvature pairs leave the finite window, while experiments find parameter-only correction can produce substantially greater future state error than no intervention.
Strategically timed window replay is the most effective non-oracle approach, achieving exact recovery in 55.6% of 108 tested configurations, although performance declines on less stable ridge-logistic loss surfaces and remains untested for large-scale nonconvex models.
Overview
This paper reframes machine unlearning for stateful optimizers as a counterfactual state-alignment problem. Rather than asking whether an unlearned model's parameters or predictions resemble those of a retrained model, the author asks whether the full optimizer state — parameters together with the finite curvature memory of online L-BFGS (o-L-BFGS) — could have arisen from a deletion-edited event stream. The central claim is that unlearning is not merely a parameter-correction problem: interventions that correct only the parameter vector or only the curvature memory produce internally inconsistent hybrid states whose future optimization behavior diverges from the counterfactual, sometimes more than doing nothing at all.
Motivation and framing
The paper's starting point is an evaluation gap in the unlearning literature. Exact methods such as SISA guarantee removal but reduce utility and require retaining the training set, making them unsuitable for online or memory-constrained settings (Bourtoule et al., 2019). Approximate methods typically certify closeness to the retrained model via parameter distance or statistical indistinguishability (Sekhari et al., 2021, Suriyakumar et al., 2022), and recent work extends certified unlearning to online settings with regret guarantees (Qiao et al., 2024, Hu et al., 13 May 2025, Shen et al., 21 Jul 2025). However, prior work has documented residual information persisting after unlearning under both black-box and white-box evaluation (Chen et al., 2020, Bertran et al., 2024, Stewart, 24 Apr 2026). The paper argues this residue is structural: o-L-BFGS maintains τ curvature pairs (vi​,ri​) used to construct an inverse-Hessian approximation, and these pairs tangibly alter future update directions. A model's function (its predictions and loss) does not describe its form (its geometry), so performance-based evaluation is incomplete by construction.
Problem formulation
The formal setup defines an event stream et​=(ot​,it​,zt​) with insert/delete operations, a deletion-edited counterfactual history Ht−D​ that omits deleted samples while preserving temporal order, and learner states θt​=(wt​,Zt​) evolved by an update map Φt​. The unlearning target is the counterfactual state θt−U​ generated by applying the same optimizer to the edited history, and the state-alignment error is Δt​=dZ​(θt​,θt−U​).
A key contribution is the definition of (ϵ,δ)-state-certified unlearning, which connects state-deviation bounds to canonical certified unlearning: if the unlearned state deviates from the counterfactual by at most α with probability (vi​,ri​)0, then Gaussian noise injection with scale (vi​,ri​)1 yields (vi​,ri​)2 certification. This generalizes parameter-only noise calibration to the full optimizer state, including the memory-induced update operator.
Theoretical results
Under strong convexity, smoothness, bounded gradients, bounded inverse-Hessian approximations, and a contractive update map ((vi​,ri​)3, (vi​,ri​)4), the paper establishes a one-step recursion (vi​,ri​)5 and, via a discrete Grönwall inequality (Liao et al., 2018), the bound
(vi​,ri​)6
The two terms have distinct interpretations: the first captures initial deletion disruption; the sum captures indirect memory — information propagated through later parameters and later curvature pairs — decaying geometrically with contraction strength. Directly contaminated curvature pairs leave the window after (vi​,ri​)7 steps (a trivial proposition), but the theorem makes clear that direct clearance does not imply zero deviation. An accompanying certificate theorem calibrates Gaussian noise to (vi​,ri​)8, the worst-case state deviation over deletion sets of size (vi​,ri​)9, with exact unlearning obtained only when et​=(ot​,it​,zt​)0. These results are confined to convex objectives; nonconvex optimizers are explicitly deferred to future work.
Experimental findings
Experiments use synthetic drifting strongly convex quadratic losses and ridge-logistic streams (et​=(ot​,it​,zt​)1, et​=(ot​,it​,zt​)2, deletion sets of size 5), where oracle replay provides an exactly known counterfactual. Metrics include parameter error, memory-operator error measured through inverse-Hessian actions on probe vectors, combined state error, update-direction error, and future state-error AUC over a post-deletion horizon. Three patterns emerge consistently.
Parameter alignment understates persistence. Parameter-only correction produces the largest mean future state-error AUC among all methods, approximately et​=(ot​,it​,zt​)3 versus et​=(ot​,it​,zt​)4 for no-op deletion — a heavy-tailed failure mode also visible in the median ratio of 0.971 against a mean ratio of 16.985 relative to no-op. The corrected parameter paired with uncorrected curvature memory induces updates the counterfactual optimizer would never take, preserving a second-order trace of deleted data even after the first-order trace is removed. Drop-and-refill fares similarly poorly (mean AUC ratio 22.908), because indiscriminate erasure destroys useful information shared by actual and counterfactual histories.
Finite memory creates a pseudo-deletion horizon. Even after direct memory mass reaches zero, state error persists and decays slowly beyond the nominal et​=(ot​,it​,zt​)5 horizon. Notably, phase-specific decay-rate fits show et​=(ot​,it​,zt​)6 for intermediate memory depths immediately after deletion — the optimizer temporarily amplifies its state deviation — and this occurs even on convex quadratic objectives. Deeper memory improves optimization but slows counterfactual recovery, creating a learning–unlearning tradeoff in the choice of et​=(ot​,it​,zt​)7.
Replay succeeds when strategically timed. Window replay reconstructs the joint parameter-memory state from retained observations and is the only non-oracle family achieving exact recovery. Its effectiveness depends on whether the replay window covers the deletion-relevant history: for recent deletions, window replay achieves zero error exactly; for random and high-gradient deletions (where direct memory mass is already zero, isolating indirect trajectory dependence), only the et​=(ot​,it​,zt​)8 window reduces future state-error AUC meaningfully (ratios 0.694–0.996 versus ~1.0 for shorter windows). In the memory-horizon grid, et​=(ot​,it​,zt​)9 replay reaches exact recovery in all Ht−D​0 configurations, supporting the conclusion that the relevant recovery horizon includes enough preceding updates to regenerate the parameter trajectory that produced the current memory — not merely the formal curvature window. Across 108 configurations, Ht−D​1 window replay was best non-oracle in 27.8% of cases and better than no-op in 76.9%, with a 55.6% exact recovery rate.
An important caveat appears in the ridge-logistic regime: window replay performs markedly worse there than on quadratics, with local state adjustments comparatively stronger, which the author attributes to loss-surface instability under weaker effective convexity. Replay-based success is therefore not uniform across loss geometries.
Limitations
The paper concedes four constraints plainly. The theory requires convexity, smoothness, bounded gradients, bounded inverse-Hessian approximations, and contractive updates — assumptions that do not cover nonconvex training. Experiments are restricted to low-dimensional synthetic streams (Ht−D​2) with controlled counterfactual replay, leaving large-scale deployed models untested. Replay interventions assume access to sufficient event history, which may be infeasible for unbounded streams or memory-constrained environments such as IoT devices. Additionally, the contractivity assumption underlying the recursion is an idealization; the empirical observation of transient amplification (Ht−D​3) suggests real dynamics can violate it locally, and the gap between the theoretical geometric decay and observed phase-dependent behavior remains open.
Conclusion
The paper demonstrates, theoretically and empirically, that machine unlearning for online L-BFGS requires alignment with a realizable counterfactual optimizer state rather than correction of any single state component. Deleted information persists through indirect memory after direct clearance, parameter-only and memory-only interventions can be worse than no intervention, and only replay-based reconstruction of the joint parameter-memory state achieves reliable recovery within a finite horizon. The open questions left by the work are specific: whether analogous state-alignment certificates extend to nonconvex optimizers, and how replay-based reconstruction can be made feasible when event history cannot be retained.
“Emergent Mind helps me see which AI papers have caught fire online.”
Philip
Creator, AI Explained on YouTube
Sign up for free to explore the frontiers of research
Discover trending papers, chat with arXiv, and track the latest research shaping the future of science and technology.Discover trending papers, chat with arXiv, and more.