---
title: Safety-Preserving Adaptation (SPA)
url: https://www.emergentmind.com/topics/safety-preserving-adaptation-spa
type: topic
---

# Safety-Preserving Adaptation (SPA)

Searching arXiv for recent and foundational papers on Safety Preserving Adaptation and closely related formulations.
Safety Preserving Adaptation (SPA) denotes a family of methods that adapt a model, controller, or policy to new tasks, environments, or data distributions while preserving pre-existing safety properties. Across the literature, the term is used in multiple but structurally related senses. In robotics and control, SPA typically means wrapping adaptation inside a certified safety mechanism so that learning does not violate forward invariance of a safe set under uncertainty [1912.09095]. In large language models (LLMs), SPA usually refers to constraining fine-tuning, low-rank adaptation, or continual domain adaptation so that downstream capability gains do not erode refusal behavior or other safety alignment properties [2601.10141], [2604.08297], [2604.17691]. In reinforcement learning, SPA appears as safe policy updating under certified parameter-space or barrier-based constraints, so that downstream policy improvement preserves previously verified safety guarantees [2604.09452], [2310.08602]. Despite this diversity, the common principle is stable: adaptation is permitted only within a mechanism that either certifies safety-preserving updates, projects away safety-conflicting directions, freezes safety-critical components, or filters unsafe actions at execution time.

## 1. Conceptual scope and recurring formulation

SPA arises from a recurring failure mode: adaptation can improve utility while degrading safety. In the adaptive-control setting of robotic manipulation, standard adaptive control can produce inputs that are safe for the estimated model but unsafe for the true plant when parameters are uncertain [1912.09095]. In LLM post-training, fine-tuning can erode refusal behavior even on benign data, and small amounts of adversarial or unsafe data can sharply increase harmful compliance [2601.10141], [2506.18931], [2603.07445], [2605.30640]. In continual RL, downstream policy updates can catastrophically forget source-task safety unless adaptation is restricted to a certified safe parameter region [2604.09452].

This suggests a unifying view: SPA addresses the tension between **plasticity** and **stability**. The model or controller must change enough to acquire new task competence, but not along directions, states, parameters, or actions that would break safety [2602.07892], [2512.10150]. A plausible implication is that SPA is best understood not as a single algorithm, but as a design pattern for constrained adaptation.

Within that pattern, the literature repeatedly instantiates four mechanisms. One is **runtime shielding**, where the learned adaptive controller or policy proposes an action but a safety filter enforces a barrier or safe-set constraint before execution [1912.09095], [2310.08602]. Another is **update-space restriction**, where optimization steps are projected away from safety-sensitive subspaces or frozen safety-critical parameters [2602.07892], [2601.10141], [2604.08297]. A third is **post-hoc repair**, where an already trained adapter is pruned, translated, or geometrically corrected to restore safety with minimal loss of utility [2506.18931], [2605.04992], [2605.30640], [2508.15068]. A fourth is **continual-learning style preservation**, where safety alignment is treated as an earlier task that must not be forgotten during later fine-tuning [2602.07892], [2512.10150], [2604.17691].

## 2. Safe-set and barrier-based SPA in adaptive control and robotics

A foundational control-theoretic formulation appears in "Safe Adaptation with Multiplicative Uncertainties Using Robust Safe Set Algorithm" [1912.09095]. There, SPA is defined for a robot manipulation system with multiplicative or parametric uncertainty in control-affine dynamics
\[
\dot x = f(x) + g(x)u,
\]
where the true control effectiveness \(g(x)\) belongs to an uncertainty set \(\Sigma_g(x)\). The core idea is that an adaptive controller may estimate unknown parameters online, but every applied control must still satisfy a robust safety condition valid for all \(g(x)\in\Sigma_g(x)\) [1912.09095].

Safety is expressed through a scalar energy-like safety index \(\phi(x)\) augmented by a time-varying term \(\phi_\alpha(t)\), giving the composite index
\[
\phi_c(x,t)=\phi(x)+\phi_\alpha(t).
\]
The safe set is
\[
\mathcal X_S = \{x \mid \phi(x)+\phi_\alpha(t)\le 0\}.
\]
The robust safe-set condition requires that whenever the boundary is active,
\[
\phi(x(t))+\phi_\alpha(t)\ge 0 \Rightarrow \dot \phi(x(t))+\dot \phi_\alpha(t)\le -\eta,
\]
for some \(\eta>0\) [1912.09095]. Lemma 1 shows that this derivative condition yields forward invariance of the safe set, and Theorem 1 extends the guarantee to the true uncertain plant provided the true dynamics lie in the uncertainty family [1912.09095].

A central technical contribution is the robust admissible control set, denoted in the paper as something like \(\bar{\mathcal U}_S\), defined through the condition that there exists a control satisfying
\[
L_f \phi(x) + L_g \phi(x)\,u \le -\eta(t), \quad \forall g(x)\in \Sigma_g(x),
\]
when safety is active [1912.09095]. The optimization-based synthesis chooses the minimum-effort safe action by aligning control with the worst-case Lie-derivative direction. In the minimum-effort lemma, the control takes the form
\[
u = -c\,\frac{L_{g^*}\phi}{\|L_{g^*}\phi\|},
\]
with smallest admissible
\[
c = \frac{L_f\phi + \eta(t)}{\alpha^*}\,\|L_{g^*}\phi\|,
\qquad
\alpha^* := \min_{g(x)\in\Sigma_g(x)} L_{g^*}\phi \cdot L_g\phi.
\]
This yields the least corrective robustly safe input [1912.09095].

The adaptive component is a Slotine–Li controller,
\[
\tau_r = \hat M(\theta)\ddot \theta_r + \hat C(\theta,\dot\theta)\dot\theta_r - K_D s,
\]
with parameter update law
\[
\dot{\hat \xi} = \Gamma^{-1}Y^T(\theta,\dot\theta,\dot\theta_r,\ddot\theta_r),
\]
and
\[
\dot\theta_r = \dot\theta_d - \Lambda \tilde\theta,\qquad
\ddot\theta_r = \ddot\theta_d - \Lambda \dot{\tilde\theta},\qquad
s=\dot\theta-\dot\theta_r.
\]
SPA is realized by placing the robust safe-set filter on top of this adaptive tracker, so that inaccurate transient parameter estimates cannot directly generate unsafe inputs [1912.09095].

Related robotics work preserves safety under adaptation in different forms. "Safe Deep Policy Adaptation" [2310.08602] combines latent environment adaptation with a Control Barrier Function (CBF) safety filter. The deployment-time filter solves a quadratic program that keeps the executed action close to the adaptive policy while satisfying a discrete-time forward-invariance condition for a safe set \(\mathcal{C}^h = \{x:h(x)\ge 0\}\) [2310.08602]. "Domain Adaptation for Outdoor Robot Traversability Estimation from RGB data with Safety-Preserving Loss" [2009.07565] uses a different notion of SPA: unsupervised domain adaptation plus an asymmetric regression loss
\[
\mathcal{L}_s = \sum_{(\mathbf{I}, \mathbf{t})} \sum_{j=1}^k \left[ (\tilde{t}_j - t_j)^2 + \alpha \max(0,\tilde{t}_j - t_j)^2 \right] + \lambda \|\boldsymbol{\theta}\|_2^2,
\]
which penalizes dangerous overestimation of traversability more than conservative underestimation [2009.07565]. This shows that, in robotics, SPA can refer either to formal set invariance during adaptation or to loss shaping that biases adaptation toward safer errors.

A further generalization appears in "Generalizations of Backup Control Barrier Functions: Expansion and Adaptation for Input-Bounded Safety-Critical Control" [2603.18450]. There, the controller used to expand the implicit safe set is decoupled from the verified backup controller that certifies safety. The parameterized expansion controller can then be adapted online in an augmented-state formulation, while forward invariance is preserved by the generalized backup-CBF construction [2603.18450]. This suggests a broader control-theoretic interpretation of SPA: separate the mechanism that improves performance from the mechanism that certifies recoverability.

## 3. Gradient-, parameter-, and token-level SPA in LLM fine-tuning

In LLMs, SPA typically targets the safety degradation that accompanies downstream fine-tuning. One of the clearest geometric formulations is "Understanding and Preserving Safety in Fine-Tuned LLMs" [2601.10141], which introduces Safety-Preserving Fine-Tuning (SPF). The paper reports three empirical insights: safety gradients lie in a low-rank subspace; utility gradients span a broader space; and the dominant safety direction can be estimated from a single sample [2601.10141]. The update rule computes a utility gradient \(g_u\), a safety gradient \(g_s\), checks whether they conflict through \(\langle g_s,g_u\rangle<0\), and if so projects the utility update away from the safety subspace:
\[
G_t = G_u - U_s U_s^\top G_u.
\]
The parameters are then updated with the projected gradient [2601.10141]. The paper states that this preserves downstream utility while bounding safety drift, and empirically restores ASR close to the initial aligned model on Llama-3.1-8B-Instruct, Mistral-7B-Instruct-v0.3, and Qwen2.5-7B-Instruct [2601.10141].

A closely related continual-learning interpretation is developed in "Safety Alignment as Continual Learning: Mitigating the Alignment Tax via Orthogonal Gradient Projection" [2602.07892]. There, the alignment tax is defined as
\[
A_{\text{tax}} = \phi(\theta_{\text{pre}}; D_{\text{eval}}) - \phi(\theta_{\text{safe}}; D_{\text{eval}}),
\]
and attributed to heterogeneous continual learning, where sequential SFT and DPO overwrite pretrained capabilities [2602.07892]. OGPSA estimates a low-rank capability subspace from gradients on small reference datasets,
\[
S_{\text{gen}}(\theta) := \operatorname{span}\{g^{(1)}(\theta), \dots, g^{(M)}(\theta)\},
\]
and projects the safety gradient onto the orthogonal complement:
\[
\tilde g_{\text{safe}} = g_{\text{safe}} - U(U^\top g_{\text{safe}}).
\]
This constrains safety alignment updates not to move along directions important for general capability [2602.07892].

Other LLM SPA methods localize safety at different granularities. "Towards Identification and Intervention of Safety-Critical Parameters in Large Language Models" [2604.08297] introduces Expected Safety Impact (ESI),
\[
\text{ESI}(\theta_i) \triangleq |\sigma(\theta_i)\nabla_{\theta_i}\mathcal{S}(\theta)|,
\]
to identify safety-critical parameters. In SPA mode, the top-ranked safety-critical parameters are frozen during downstream task fine-tuning, while only non-critical parameters are updated [2604.08297]. The paper reports that this limits the safety degradation of aligned LLMs within \(1\%\) after a \(1{,}000\)-iteration instruction fine-tuning on different tasks [2604.08297].

At a finer granularity, "Few Tokens, Big Leverage: Preserving Safety Alignment by Constraining Safety Tokens during Fine-tuning" [2603.07445] argues that refusal behavior is concentrated in a small set of safety-related output tokens. It identifies a top-\(K\) safety token set \(\mathcal{S}_{\text{safety}}\) by discrepancy between aligned and base next-token distributions on harmful prompts and regularizes only the safety-token distribution with a weighted KL term:
\[
\mathcal{L} = \mathcal{L}_{\mathrm{CE}} + \lambda_{\mathrm{KL}} \mathcal{L}_{\mathrm{KL}}^{\text{safety}}.
\]
The method also calibrates the safety reference using a mixture of full-context and response-only reference logits to reduce harmful-prefix contamination [2603.07445]. This is a token-level SPA mechanism: preserve only the safety-relevant slice of the output distribution and leave the rest free for task adaptation.

A different structural decomposition appears in "A Guardrail for Safety Preservation: When Safety-Sensitive Subspace Meets Harmful-Resistant Null-Space" [2510.14301]. GuardSpace computes a covariance-preconditioned SVD of \(\mathbf W \mathbf C\) using harmful-prompt activations, freezes the large-singular-value safety-sensitive subspace, initializes low-rank adapters from the safety-irrelevant tail,
\[
\mathbf B = \mathbf U[:, -r:]\,\sqrt{\mathbf\Sigma[-r:]}, \quad
\mathbf A = \sqrt{\mathbf\Sigma[-r:]}\,(\mathbf V^\top \mathbf C^{-1})[-r:,:],
\]
and constrains the effective update through a null-space projector \(\mathbf P\) derived from the harmful-prompt covariance [2510.14301]. The key invariance relation
\[
(\mathbf W' + \mathbf B^*\mathbf A^* \cdot \mathbf P)\mathbf X = \mathbf W'\mathbf X, \quad \mathbf X \in \mathcal H
\]
is used to preserve the original behavior on harmful prompts throughout fine-tuning [2510.14301].

These methods differ in where they locate safety—gradient directions, parameter subsets, token distributions, or harmful-input null spaces—but they share the same adaptation logic: preserve a safety-relevant structure and route learning elsewhere.

## 4. Low-rank adaptation, post-hoc repair, and adapter-space SPA

A large subliterature studies SPA specifically for LoRA and related parameter-efficient fine-tuning. "SaLoRA: Safety-Alignment Preserved Low-Rank Adaptation" [2501.01765] is an early formulation that inserts a fixed safety module
\[
\mathbf{C}_{S} = \mathbf{I} - \mathbf{U}_C\mathbf{U}_C^\top
\]
to project trainable low-rank updates away from a harmful-feature subspace estimated from safety data [2501.01765]. It pairs this with task-specific initialization of the adapters using downstream task features. The reparameterized layer is
\[
\mathbf{W}' = \mathbf{W} - \mathbf{C}_S\mathbf{B}_S\mathbf{A}_S.
\]
The paper reports that SaLoRA sharply reduces harmful rate after Alpaca fine-tuning relative to LoRA, DoRA, and PiSSA while maintaining or improving utility on commonsense reasoning benchmarks [2501.01765].

Pruning-based variants remove parts of a trained LoRA update deemed most responsible for safety degradation. "Safe Pruning LoRA: Robust Distance-Guided Pruning for Safety Alignment in Adaptation of LLMs" [2506.18931] introduces Empirical-DIEM (E-DIEM) based on the discrepancy between a LoRA update \(\Delta\theta^i\) and its projection into an aligned subspace constructed from an aligned/unaligned model pair. Layers with large discrepancy are pruned entirely according to
\[
\mathcal{R}(\Delta \theta^i)=
\begin{cases}
\text { keep } \Delta \theta^i, & \text { if } u < t \\
\text { prune } \Delta \theta^i , & \text { if } u \ge t
\end{cases}
\]
[2506.18931]. The paper reports strong ASR reductions on Dialog Summary + PureBad, Alpaca + PureBad, and pure benign Alpaca settings, often with maintained or improved utility and reduced inference time [2506.18931].

"S3LoRA: Safe Spectral Sharpness-Guided Pruning in Adaptation of Agent Planner" [2508.15068] removes a different class of risky layers. It analyzes only the LoRA update \(\Delta W = AB\), performs Magnitude-Aware Spherically Normalized SVD, and computes the Spectral Sharpness Index
\[
\text{SSI} = \frac{\sigma_1^{\prime}}{\sum_{j=1}^h \sigma_j^{\prime}+\varepsilon}.
\]
Layers with top-\(\tau\) SSI are pruned post-hoc, without requiring base/aligned checkpoint pairs [2508.15068]. This suggests a more deployment-oriented SPA criterion: structurally sharp and concentrated updates are treated as potential safety risks even in the absence of a reference aligned subspace.

Another direction repairs rather than deletes unsafe update components. "CSULoRA: Closest Safe Update Low-Rank Adaptation" [2605.30640] estimates a safety-aligned subspace from the displacement between a safety-aligned checkpoint and its base checkpoint,
\[
V^i = W^i_{\mathrm{aligned}} - W^i_{\mathrm{base}},
\]
builds double-sided projectors \(P_L^i, P_R^i\), decomposes each LoRA update into four orthogonal blocks \(\Delta W_{LR}^i\), \(\Delta W_{L\bar R}^i\), \(\Delta W_{\bar L R}^i\), and \(\Delta W_{\bar L\bar R}^i\), and solves a closed-form penalized minimum-change problem [2605.30640]. The corrected update is
\[
\Delta W_\star^i = \Delta W_{LR}^i + \sum_{b \in \mathcal{B}\setminus\{LR\}} \gamma_b \Delta W_b^i,
\qquad
\gamma_b = \frac{1}{1+\lambda_b},
\]
with penalties set from relative block energy [2605.30640]. This is a geometric SPA: preserve the fully aligned part exactly and attenuate the rest rather than removing it.

The same post-hoc philosophy is extended beyond linear correction in "You Snooze, You Lose: Automatic Safety Alignment Restoration through Neural Weight Translation" [2605.04992]. NeWTral learns a non-linear parameter-space map from unsafe adapters to safe aligned adapters,
\[
\mathcal{T}(W_{dom}) = W_{cured} \approx W_{safe},
\]
using unsafe-to-safe adapter pairs and, in its main variant, a Mixture of Experts routing mechanism that interpolates between a safety-aggressive expert and a utility-preserving surgical expert [2605.04992]. The paper reports average ASR reduction from about \(70\%\) in unsafe experts to about \(13\%\) with the MoE translator while maintaining around \(90\%\) average knowledge fidelity [2605.04992].

The cumulative picture from these LoRA papers is that SPA in adapter space can be preventive, as in SaLoRA; selective, as in SPLoRA and S3LoRA; corrective, as in CSULoRA; or translational, as in NeWTral. This suggests that the adapter itself is treated as the main locus of safety erosion and therefore as the main intervention target.

## 5. Continual adaptation, forgetting, and multi-stage safety preservation

Several works explicitly frame SPA as a continual-learning problem. "Unforgotten Safety: Preserving Safety Alignment of Large Language Models with Continual Learning" [2512.10150] formalizes a two-stage pipeline: safety alignment first, user adaptation second. The goal is to keep
\[
\mathcal{L}_{\text{safe}}(\theta^\ast) \approx \mathcal{L}_{\text{safe}}(\theta_{\text{safe}})
\]
after downstream fine-tuning [2512.10150]. The paper evaluates regularization-based methods such as EWC and LwF, memory-based methods such as A-GEM and DER, and a merging method MagMax. Among these, DER is reported as the strongest overall method, drastically reducing ASR relative to standard fine-tuning while maintaining utility across GSM8K, SST2, and Code, in both benign and poisoned settings [2512.10150]. The paper’s central claim is that safety compromise is catastrophic forgetting of alignment rather than merely a one-off optimization artifact.

The same theme is sharpened in sequential multi-domain adaptation by "SafeAnchor: Preventing Cumulative Safety Erosion in Continual Domain Adaptation of Large Language Models" [2604.17691]. SafeAnchor has three components: Safety Subspace Identification (SSI) via empirical Fisher eigendecomposition in LoRA parameter space; Orthogonal Safety-Constrained Adaptation (OSCA), which projects each domain-task gradient away from the safety subspace,
\[
\tilde{g}_i^t = (I - V_i^{\text{safe}} (V_i^{\text{safe}})^\top) g_i^t;
\]
and Cumulative Safety Monitoring (CSM), which triggers corrective replay if the refusal rate falls below a threshold
\[
s_t < (1-\tau)s_0
\]
[2604.17691]. Across a Medical \(\rightarrow\) Legal \(\rightarrow\) Code pipeline on Llama-2-7B-Chat and Mistral-7B-Instruct, the paper reports retention of \(93.2\%\) and \(93.1\%\) of original safety alignment, respectively, while staying within about \(1.5\) points of unconstrained LoRA on domain tasks [2604.17691].

In continual RL, "SafeAdapt: Provably Safe Policy Updates in Deep Reinforcement Learning" [2604.09452] provides a parameter-space certificate rather than a replay-based or gradient-projection heuristic. It defines a Rashomon set
\[
\Theta_{\mathrm{source}}^{\mathrm{safe}}
=
\{\theta' : \theta_{\mathrm{source}}-\alpha^\star \le \theta' \le \theta_{\mathrm{source}}+\alpha^\star\},
\]
a center-symmetric orthotope in policy parameter space such that every policy inside it is certified safe on the source task [2604.09452]. Downstream policy updates are projected back into this region by element-wise clipping after each gradient step. The paper proves that if the certified lower bound on the surrogate critical-state safety rate exceeds a threshold, then source-task safety is preserved for every iteration of projected adaptation [2604.09452]. This is a stronger notion of SPA than most LLM work: safety is preserved a priori by parameter-space certification rather than statistically or geometrically encouraged.

This family of works indicates that SPA becomes more demanding when adaptation is repeated. A plausible implication is that one-shot safety-preserving fine-tuning methods may not suffice in long adaptation chains, because safety-relevant subspaces can drift and indirect erosion can accumulate across stages.

## 6. Self-triggered and preference-based uses of the term in LLM alignment

The term SPA is not used uniformly across all LLM alignment papers. In some works it denotes a broad category of safety-preserving adaptation rather than the name of the specific method. "Adaptive and Explicit safe: Triggering Latent Safety Awareness in Large Reasoning Models" [2606.16808] presents Safe Trigger, which the paper describes as a safety-preserving adaptation method for large reasoning models. Its key observation is **Latent Safety Awareness**: when the model reviews the original risky query together with its own reasoning trace, its Risk Identification Success Rate (RISR) is much higher than its direct-attack ASR, ranging from \(44.79\%\) to \(100\%\) [2606.16808]. The method uses Safe Trigger SFT with structured tags `<safe> ... </safe>` and Safe Trigger DPO on self-generated data. Across models, the aggregated results show Base / ST-S / ST-D harmful rates of \(16.22\%\), \(5.70\%\), and \(3.63\%\), jailbreak rates of \(27.21\%\), \(11.71\%\), and \(7.79\%\), and nearly unchanged general performance and over-refusal [2606.16808]. Here, SPA denotes selective activation of a safety-analysis module only on risky inputs, rather than geometric protection of safety parameters.

Another distinct usage appears in "SPA: Achieving Consensus in LLM Alignment via Self-Priority Optimization" [2511.06222], where SPA stands specifically for **Self-Priority Alignment** rather than Safety Preserving Adaptation. The method imposes a lexicographic trustworthy-before-helpful ordering,
\[
\min_{\theta} G_a(\theta)
\quad\text{followed by}\quad
\min_{\theta} G_b(\theta) \;\text{s.t.}\; G_a(\theta)\le G_a^\*,
\]
constructs lexicographically ordered preference pairs from self-generated samples, and optimizes an uncertainty-weighted SimPO-style loss [2511.06222]. Although the acronym overlaps, this is a distinct concept. The paper reports improved harmlessness or honesty together with improved helpfulness on Llama-3.1-8B-Instruct and Mistral-7B-Instruct across SafeRLHF, WildGuard, and HoneSet [2511.06222].

This lexical ambiguity matters. In the broader literature, “SPA” may refer to a method category—adapt safely while preserving existing behavior—or to the specific alignment paradigm of Self-Priority Alignment. A common misconception is that every “SPA” paper studies the same mechanism. The data instead indicate that the acronym spans at least two conceptually different traditions: safety-preserving model or controller adaptation, and trustworthy-before-helpful preference optimization.

## 7. Comparative structure, assumptions, and limitations

The SPA literature varies substantially in what safety means, what is preserved, and what guarantees are available.

| Regime | Preservation target | Main mechanism |
|---|---|---|
| Adaptive control / robotics | Forward invariance of safe set | Robust safe set, CBF, backup-set certification |
| LLM fine-tuning | Refusal behavior / safety alignment | Gradient projection, parameter freezing, token constraints |
| LoRA post-hoc repair | Safe behavior of adapted checkpoint | Pruning, subspace correction, parameter translation |
| Continual RL | Source-task certified safety | Projection into certified Rashomon set |

Control-theoretic SPA generally offers the strongest guarantees. The robust safe-set method in [1912.09095], the CBF-shielded SafeDPA in [2310.08602], the Rashomon-set projection of SafeAdapt [2604.09452], and the generalized backup-CBF adaptation framework [2603.18450] all provide formal forward-invariance or certified-safety statements under explicit assumptions. Those assumptions include bounded model error, Lipschitz continuity, nonempty admissible safe-control sets, or correct uncertainty families. The guarantees are therefore rigorous but model-dependent.

LLM SPA papers more often provide empirical safety-utility trade-offs with lighter theory. SPF proves utility convergence with bounded safety drift under a low-rank projection model [2601.10141]. OGPSA proves steepest feasible descent under first-order capability-preservation constraints [2602.07892]. GuardSpace provides a harmful-input invariance argument through its null-space projector [2510.14301]. Yet most LLM methods explicitly do not claim formal absolute safety. CSULoRA, for example, states that its safety subspace is only a proxy and “it is not a formal guarantee of safety” [2605.30640]. NeWTral likewise reduces ASR sharply but does not fully eliminate unsafe outputs [2605.04992].

The preservation target also differs. Some methods preserve **capability** during safety alignment, as in OGPSA [2602.07892]. Others preserve **safety** during capability tuning, as in SPF [2601.10141], ESI-based SPA [2604.08297], PACT [2603.07445], or GuardSpace [2510.14301]. LoRA repair methods preserve both the added specialization and the original safety alignment, but only relative to the chosen geometric or reference-model proxy [2506.18931], [2605.30640], [2605.04992].

Several limitations recur across papers. Many LLM methods depend on a reference aligned/base model pair or a safety calibration set, as in SPLoRA [2506.18931], SaLoRA [2501.01765], CSULoRA [2605.30640], and GuardSpace [2510.14301]. Others assume that safety is low-rank or localized, as in SPF [2601.10141], SafeAnchor [2604.17691], and ESI-based SPA [2604.08297]. Post-hoc methods can lose some utility when unsafe and useful directions overlap [2605.30640]. Continual methods must manage evolving safety geometry and cumulative drift [2604.17691]. These recurring assumptions suggest that the central open question is not whether safety can be preserved at all, but how robustly one can identify the correct safety structure under varying architectures, tasks, and adaptation regimes.

Overall, the literature portrays Safety Preserving Adaptation as a general solution strategy for a persistent systems problem: adaptation tends to move a model or controller into regions that improve immediate utility but destabilize previously aligned safe behavior. SPA methods differ in whether they act on actions, gradients, parameters, tokens, adapters, or certified parameter regions, but they converge on the same principle: learning is acceptable only when safety-preserving structure remains invariant, recoverable, or explicitly protected.

Source: https://www.emergentmind.com/topics/safety-preserving-adaptation-spa