Papers
Topics
Authors
Recent
Search
2000 character limit reached

Fora: From Weight-Space to Function-Space Protection in Capability-Preserving Fine-Tuning

Published 30 Jun 2026 in cs.LG | (2606.31092v1)

Abstract: Full fine-tuning adapts LLMs to new tasks but can erode capabilities they already possess. Existing remedies protect through proxies such as parameter distances, importance penalties, output matching, or dominant singular directions of the weights, but none directly asks which activation directions the preserved capability relies on. We argue that a capability is characterized more faithfully by the activation subspace it induces than by the singular geometry of the weight matrix, and develop function-space protection, instantiated as FORA (Function-space Orthogonal Residual Adaptation). From label-free calibration inputs, FORA estimates, per layer, the principal directions QQ of the input-activation covariance and forms a right projector PQ=IQQ<sup>TP_Q = I - QQ<sup>T. Paired with a left projector PUP_U from the weight SVD, the update is ΔW=PUMPQ+U2DδV2<sup>TΔW = P_U M P_Q + U_2 D_δ V_2<sup>T: a high-capacity branch structurally barred from reading capability-relevant function directions, plus a narrow spectral channel for controlled plasticity. The construction extends to parameter-efficient adaptation via M(α/r)BAM \to (α/r) BA. Across three settings on Qwen3-1.7B, including COGS and GSM8K learned while preserving translation and translation learned while preserving math, FORA consistently improves preservation over weight-space projection and standard regularization, with only a small new-task trade-off in the math-preservation setting. A controlled ablation isolating the projection source shows that the advantage comes not from projection itself, but from projecting onto capability-derived rather than weight-derived directions. Code is available at https://github.com/zrui239/FORA.

Authors (2)

Summary

No one has generated a summary of this paper yet.

Paper to Video (Beta)

No one has generated a video about this paper yet.

Whiteboard

No one has generated a whiteboard explanation for this paper yet.

Open Problems

We haven't generated a list of open problems mentioned in this paper yet.

Continue Learning

We haven't generated follow-up questions for this paper yet.