Local Support Learning

Local Support Learning reframes catastrophic forgetting as a geometric problem: conventional gradient updates are optimized on current data but modify the model's behavior on all future inputs. This paper proposes a solution that combines standard adapters with per-matrix gates that estimate where the current training distribution lies in activation space. By routing adapter updates only to inputs belonging to the learned support, the method achieves near-perfect retention of pretrained capabilities across multiple sequential learning phases while maintaining full adaptation capacity.
Script
Conventional gradient descent creates updates from current training examples, but those updates change the model's output for every input that shares any geometric overlap with the training activation. The result is a fundamental mismatch: local optimization objectives produce globally acting parameter changes.
Local Support Learning separates two functions that are usually entangled. Standard adapters learn the task-specific transformation with full representational capacity. Separately trained gates decide where that transformation should be applied, estimating the region of activation space occupied by the current learning phase.
The gate uses paired Gaussian mixture models. A positive model fits the current phase activations tightly. A negative model, trained once on roughly one million generic tokens, provides a broad reference density. The adapter activates only when the positive likelihood exceeds the negative, creating a conservative decision boundary that defaults to preserving the base model.
On Qwen 7B models trained sequentially on translation, chemistry, and cybersecurity, Local Support Learning preserves 98.8% of pretrained performance. Standard LoRA loses 44% after the first phase. The method decouples retention from adaptation: pretrained performance remains stable across changes in learning rate, adapter rank, and batch size that would normally force a trade off.
Gates are placed on individual weight matrices rather than using one global router. Since representations evolve across depth, the support of a task is matrix specific. Independent per matrix decisions confine routing errors locally and allow parallelization, though they prevent merging adapters into the base weights and make inference cost grow linearly with the number of learning phases.
The theoretical insight is asymmetric: a gate trained on current data directly observes errors inside the support but remains blind to erroneous activation outside it, since zero probability regions contribute nothing to the training loss. Density estimation solves this by imposing an inductive bias on off support behavior, where discriminative classifiers remain unconstrained. Visit EmergentMind.com to explore Local Support Learning further and create your own research videos.