---
title: 'BID-LoRA: Bi-Directional Low-Rank Adaptation'
url: https://www.emergentmind.com/topics/bi-directional-low-rank-adaptation-bid-lora
type: topic
---

# BID-LoRA: Bi-Directional Low-Rank Adaptation

Searching arXiv for the cited BID-LoRA and related LoRA variants to ground the article in the relevant papers.
arXiv search query: BID-LoRA 2604.12686 BoRA 2412.06441 Bi-LoRA 2508.19564
Bi-Directional Low-Rank Adaptation (BID-LoRA) is a parameter-efficient framework for Continual Learning Unlearning (CLU), a setting in which a model must repeatedly forget designated data or classes, learn new data or classes, and retain previously useful knowledge over multiple adaptation cycles. In its defining formulation, BID-LoRA uses three dedicated low-rank adapter pathways—retain, new, and unlearn—applied to attention layers, together with an escape unlearning objective that pushes forget-class embeddings to positions maximally distant from retained knowledge, while updating only about \(5\%\) of parameters [2604.12686]. The framework is motivated by the claim that naively combining existing continual learning (CL) and machine unlearning (MU) methods causes knowledge leakage, understood as a gradual degradation of foundational retained knowledge across repeated learn–forget cycles [2604.12686].

## 1. CLU as the problem setting

BID-LoRA is defined for a unified regime called Continual Learning Unlearning. In this regime, a pretrained model \(\mathcal{M}\) must support three simultaneous requirements at each step of adaptation: precise deletion of unwanted knowledge, efficient integration of new knowledge while preserving prior information, and minimizing knowledge leakage across cycles [2604.12686]. The motivating scenarios explicitly include identity management systems, privacy compliance under GDPR/CCPA, dynamic face recognition, and models that must adapt to changing users, policies, or classes [2604.12686].

The paper formalizes a pretrained model as a mapping \(f_{\mathcal{M}} : \mathcal{X}_D \to \mathcal{Y}_D\), with \(D_f\) denoting data to forget, \(D_r^{\text{full}}\) the full retain set, \(D_{\text{new}}\) new data to learn, and \(D_r \subset D_r^{\text{full}}\) a replay buffer. The assumptions are
\[
D = D_r^{\text{full}} \cup D_f,\qquad
D_r^{\text{full}} \cap D_f = \emptyset,\qquad
D_{\text{new}} \cap D = \emptyset.
\]
Before adaptation, the intended behavior is that the model maps \(\mathcal{X}_{D_f}\) to \(\mathcal{Y}_{D_f}\) and \(\mathcal{X}_{D_r}\) to \(\mathcal{Y}_{D_r}\), but not \(\mathcal{X}_{D_{\text{new}}}\) to \(\mathcal{Y}_{D_{\text{new}}}\). After adaptation, the desired behavior reverses this status for the forget and new sets: the model should no longer map \(\mathcal{X}_{D_f}\) to \(\mathcal{Y}_{D_f}\), should preserve mapping on \(\mathcal{X}_{D_r}\), and should now map \(\mathcal{X}_{D_{\text{new}}}\) to \(\mathcal{Y}_{D_{\text{new}}}\) [2604.12686].

For the continual setting over tasks \(t=1,\dots,T\), the update rule is written as
\[
\mathcal{M}^{(t)} =
\mathcal{F}\big(\mathcal{M}^{(t-1)}, D_f^{(t)}, D_r^{(t)}, D_{\text{new}}^{(t)}\big).
\]
Successful CLU is then defined through three operational criteria:
- forgetting: \(\text{Acc}(\mathcal{M}^{(t)}, D_f^{(i)}) \le \frac{1}{C}\),
- retention: \(\text{Acc}(\mathcal{M}^{(t)}, D_r^{(t)}) \approx \text{Acc}(\text{Oracle}, D_r^{(t)})\),
- learning: \(\text{Acc}(\mathcal{M}^{(t)}, D_{\text{new}}^{(t)}) \approx \text{Acc}(\text{Oracle}, D_{\text{new}}^{(t)})\),
where \(C\) is the number of classes [2604.12686].

This formulation places BID-LoRA in a distinct category from one-shot unlearning and standard class-incremental learning. The emphasis is sustained adaptation without cumulative drift. A plausible implication is that the method is intended less as an isolated PEFT variant and more as a systems-level mechanism for repeated state correction under retention constraints.

## 2. Tri-path low-rank architecture

The architectural point of departure is standard LoRA. In the paper’s notation, standard LoRA modifies a frozen weight matrix \(W\) by
\[
h = Wx + \mathcal{S}(BAx),
\]
with \(x \in \mathbb{R}^k\), \(h \in \mathbb{R}^d\), \(W \in \mathbb{R}^{d \times k}\) frozen, \(A \in \mathbb{R}^{r \times k}\), \(B \in \mathbb{R}^{d \times r}\), and \(\mathcal{S}\) a scaling factor [2604.12686]. BID-LoRA argues that, in CLU, a single adapter is overloaded because it must simultaneously support retention, learning, and forgetting; the consequence is gradient interference and eventually knowledge leakage [2604.12686].

BID-LoRA replaces that single update space with three dedicated low-rank pathways:
\[
W_{\text{ret}} = B_{\text{ret}}A_{\text{ret}},\qquad
W_{\text{new}} = B_{\text{new}}A_{\text{new}},\qquad
W_f = B_fA_f.
\]
These are applied jointly as
\[
h = Wx + \mathcal{S}\left(B_{\text{ret}}A_{\text{ret}} + B_{\text{new}}A_{\text{new}} + B_fA_f\right)x,
\]
while the backbone \(W\) remains frozen [2604.12686].

The separation is the central design principle. The retain adapter preserves prior knowledge, the new adapter learns new classes, and the forget adapter removes unwanted knowledge. Because they are trained with separate objectives and gradient masking, they do not directly interfere [2604.12686]. The paper applies these adapters to attention layers and classifier heads, preserving the PEFT premise that only small low-rank modules are updated.

This decomposition is called “bi-directional” in the paper’s title, but the operative mechanism is actually tri-path. The term therefore refers less to two opposing optimization directions than to coordinated directional control over retention and change within the CLU pipeline. This suggests that the distinctive contribution lies in pathway specialization rather than in a single algebraic modification of the LoRA parameterization.

## 3. Escape unlearning and the geometry of forgetting

The forget mechanism in BID-LoRA is escape unlearning. Its goal is to delete knowledge without damaging retained classes by moving forget-class embeddings toward a target that is maximally distant from retained knowledge [2604.12686].

The construction proceeds geometrically. For each class \(k\), the class centroid is
\[
c_k = \frac{1}{|D_k|}\sum_{x \in D_k} \mathrm{emb}(x),
\]
where \(\mathrm{emb}(x)\) is the learned embedding of sample \(x\) [2604.12686]. Using retain-class centroids \(\{c_{r_i}\}\), the paper defines the escape direction as
\[
\mathbf{d}^* = \arg\min_{\|\mathbf{d}\|=1} \max_i \left(\mathbf{d}^\top c_{r_i}\right).
\]
Here, \(\mathbf{d}^\top c_{r_i}\) measures alignment with retain centroid \(c_{r_i}\); the inner \(\max_i\) selects the retain centroid most aligned with \(\mathbf{d}\); and the outer minimization chooses the direction least aligned with all retain classes [2604.12686]. The interpretation given is that \(\mathbf{d}^*\) is the direction maximally distant from retained knowledge.

Because placing the escape point on the unit sphere can be unstable, the paper scales it as
\[
\mathbf{t}_{\text{escape}} = \lambda_{\text{esc}} \cdot \mathbf{d}^*,
\]
with \(\lambda_{\text{esc}}\) a scaling factor [2604.12686]. Forget samples are then driven toward this target through
\[
\mathcal{L}_f = \mathrm{MSE}\big(\mathrm{emb}(\mathbf{X}^f), \mathbf{t}_{\text{escape}}\big).
\]
The stated effect is that the model maps forget samples toward the same escape target, producing a many-to-one collapse that destroys class-discriminative structure [2604.12686].

The retain and new pathways use distinct objectives. Retention combines classification and embedding anchoring:
\[
\mathcal{L}_{\text{ret}} =
\lambda_{\text{ce}} \cdot \mathrm{CE}(\mathbf{z}_r, y_r) +
\lambda_{\text{emb}} \cdot \mathrm{MSE}(\mathbf{e}_r, \mathbf{e}_t),
\]
where \(\mathbf{z}_r\) are logits on retain samples, \(y_r\) retain labels, \(\mathbf{e}_r\) student embeddings, and \(\mathbf{e}_t\) frozen teacher embeddings from the initial model [2604.12686]. New knowledge uses the standard classification loss
\[
\mathcal{L}_{\text{new}} = \mathrm{CE}(\mathbf{z}_n, y_n).
\]

Taken together, these losses assign distinct geometric roles to the three pathways: anchoring for retention, discrimination for acquisition, and collapse toward a distant point for deletion. A plausible implication is that BID-LoRA treats unlearning not as simple error induction but as controlled relocation in embedding space.

## 4. Optimization protocol, isolation, and efficiency

Training is organized as isolated updates. The algorithm performs three stages per cycle:
1. retention update: freeze forget and new adapters, and update the retain adapter and retain classifier head using \(\mathcal{L}_{\text{ret}}\);
2. forget update: freeze retain and new adapters, and update the forget adapter and forget head using \(\mathcal{L}_f\);
3. new knowledge update: freeze retain and forget adapters, and update the new adapter and new head using \(\mathcal{L}_{\text{new}}\) [2604.12686].

At inference, the adapters are merged into the frozen backbone as
\[
W \gets W + \mathcal{S}\left(B_{\text{ret}}A_{\text{ret}} + B_{\text{new}}A_{\text{new}} - B_fA_f\right).
\]
The subtractive forget term appears in the merge expression exactly as written in the algorithm description [2604.12686].

The framework is explicitly parameter-efficient. BID-LoRA trains small low-rank adapters on attention layers and classifier heads while keeping the backbone frozen. The reported tunable-parameter ratio is about \(5.08\%\) on CIFAR-100 and \(5.00\%\) on CASIA-Face100 [2604.12686]. The paper presents this as practical for large pretrained ViTs and transformers, repeated adaptation cycles, settings where full retraining is too costly, and privacy-sensitive deployment where “surgical” updates are preferred [2604.12686].

The adapter configuration used in the main protocol assigns rank \(8\) to the retain adapter, rank \(8\) to the new adapter, and rank \(4\) to the forget adapter. The replay buffer \(D_r\) is \(10\%\) of the full retain set [2604.12686]. These design choices matter because the framework depends on both pathway separation and limited replay. The ablations report that performance saturates around rank \(8\), that more buffer helps retention, and that removing any one of the three pathways harms its corresponding objective [2604.12686].

## 5. Experimental protocol and empirical behavior

The empirical evaluation covers CIFAR-100 and CASIA-Face100, the latter described as a curated subset of \(100\) identities from CASIA-WebFace. The backbones are Data-efficient image transformer (DeiT) for classification and Face Transformer for face recognition [2604.12686]. The evaluation uses a six-task sliding window protocol:
- start from classes \(0\)–\(29\),
- task 1: \(10\)–\(39\),
- task 2: \(20\)–\(49\),
- task 3: \(30\)–\(59\),
- task 4: \(40\)–\(69\),
- task 5: \(50\)–\(79\),
- task 6: \(60\)–\(89\).

Each task retains \(20\) classes, forgets \(10\) classes, and learns \(10\) new classes [2604.12686]. The reported metrics are forget accuracy \((Acc_f)\), retain accuracy \((Acc_r)\), new accuracy \((Acc_n)\), overall accuracy \((Acc_o)\), MIA success rate, KL divergence to oracle, and tunable parameter ratio [2604.12686].

On CIFAR-100, BID-LoRA uses only \(5.08\%\) tunable parameters and achieves very low forget accuracy, roughly \(0.13\%\)–\(0.93\%\) across tasks; retain accuracy around \(70.83\%\)–\(76.00\%\); new accuracy up to \(83.20\%\); overall accuracy close to oracle; MIA near \(0.50\)–\(0.57\); and low KL divergence [2604.12686]. The examples given are task 1 overall accuracy \(76.03\) and task 6 overall accuracy \(73.51\), corresponding to only about a \(2.52\%\) drop across the six-task sequence [2604.12686].

On CASIA-Face100, BID-LoRA again uses about \(5.00\%\) tunable parameters and maintains forget accuracy near \(0\)–\(0.88\%\), high retain and new accuracy, overall accuracy around \(91\)–\(93\%\), MIA close to \(0.5\), and low KL divergence [2604.12686]. The paper reports task 1 overall accuracy \(93.20\) and task 6 overall accuracy \(91.22\), an overall drop of about \(1.98\%\) [2604.12686].

The baselines are LSF, CLPU-DER++, UniCLUN, UG-CLU, and UnCLe. The paper characterizes these as representing replay, distillation, hypernetworks, and gradient-based unified strategies [2604.12686]. The comparative interpretation is that BID-LoRA gives the best tradeoff between forgetting, retaining, learning new classes, and parameter efficiency, while several baselines either forget well at the cost of overall accuracy, show unstable retention, or exhibit inconsistent MIA or lower stability [2604.12686].

The ablation results sharpen this interpretation. Standard LoRA is reported to be inferior to BID-LoRA in CLU, supporting the claim that pathway separation matters. A larger escape scaling factor improves forgetting, with \(\lambda_{\text{esc}} = 10\) giving the best reported forgetting performance. The paper also presents t-SNE and 3D sphere plots showing forget embeddings migrating toward the escape point [2604.12686]. This suggests that the claimed deletion mechanism is not only metric-based but also geometrically observable in the learned representation space.

## 6. Relation to adjacent “bi-directional” LoRA variants, misconceptions, and limitations

The expression “Bi-Directional Low-Rank Adaptation” is potentially ambiguous because several LoRA variants use related nomenclature while solving different problems. The following comparison helps disambiguate the literature.

| Method | Defining mechanism | Primary setting |
|---|---|---|
| BID-LoRA | retain, new, and unlearn adapters plus escape unlearning | continual learning and unlearning |
| Bi-LoRA | primary LoRA and auxiliary LoRA for SAM-style perturbations | sharpness-aware fine-tuning |
| BoRA | row-wise and column-wise magnitude adaptation around a low-rank direction update | bi-dimensional weight decomposition |

BID-LoRA in the strict sense refers to the CLU framework with three dedicated pathways and escape unlearning [2604.12686]. By contrast, “Bi-LoRA: Efficient Sharpness-Aware Minimization for Fine-Tuning Large-Scale Models” introduces a dual-module design in which a primary LoRA module performs task adaptation via gradient descent and an auxiliary LoRA module models SAM-style perturbations via gradient ascent; only the primary branch is retained for inference [2508.19564]. That method addresses the mismatch that arises when SAM perturbations are forced into the same low-rank subspace used for adaptation, and it is framed around generalization and efficient sharpness-aware training rather than CLU [2508.19564].

BoRA, in turn, is “Bi-dimensional Weight-Decomposed Low-Rank Adaptation,” a symmetric extension of DoRA that learns both row-wise and column-wise magnitude information around a low-rank directional update. Its claim is symmetry across horizontal and vertical dimensions of the weight matrix, with trainable magnitude vectors \(m^r\) and \(m^c\) inserted into a two-stage normalization and scaling pipeline [2412.06441]. The paper explicitly states that BoRA can be interpreted as a bi-directional or bi-dimensional LoRA variant, but it does not use “BID-LoRA” as the method name [2412.06441].

A common misconception is therefore to treat BID-LoRA, Bi-LoRA, and BoRA as interchangeable. They are not. BID-LoRA is a unified continual learning plus machine unlearning framework; Bi-LoRA is a SAM-inspired fine-tuning method; and BoRA is a symmetric magnitude-direction decomposition for PEFT [2604.12686].

The limitations of BID-LoRA are also explicit. The framework relies on a replay buffer \(D_r\), assumed to be at least \(10\%\) of the retain set; it is validated mainly on vision classification and face recognition; it assumes a structured tri-partite setting with known retain, forget, and new partitions; and future work is needed to remove the buffer entirely and extend the method to other biometric modalities [2604.12686]. These constraints indicate that the framework is presently strongest in controlled, transformer-based identity or classification pipelines where repeated enrollment, revocation, and retention are all first-class requirements.

Source: https://www.emergentmind.com/topics/bi-directional-low-rank-adaptation-bid-lora