---
title: SFT+RL Protocol for LLM OOD Generalization
url: https://www.emergentmind.com/topics/sft-rl-protocol
type: topic
---

# SFT+RL Protocol for LLM OOD Generalization

A protocol combining Supervised Fine-Tuning (SFT) and Reinforcement Learning (RL)—hereafter "SFT+RL"—refers to post-training procedures for large language models (LLMs) that sequentially or jointly optimize both imitation and task reward objectives. The SFT+RL paradigm is now canonical in LLM post-training, especially for eliciting complex reasoning, arithmetic, and generalization capabilities beyond those feasible with either paradigm alone. The SFT+RL protocol in "RL Fine-Tuning Heals OOD Forgetting in SFT" provides a tightly characterized, empirically validated framework for understanding, implementing, and diagnosing this two-stage post-training pipeline [2509.12235].

## 1. Paradigm Definition and Formalization

The standard SFT+RL protocol consists of two distinct fine-tuning stages applied to a pretrained LLM with parameter vector $\theta_0$:

**(a) Supervised Fine-Tuning (SFT):** Minimize negative log-likelihood (cross-entropy) over a labeled dataset $\mathcal{D} = \{(x_i, y_i)\}$:
\[
L_{\text{SFT}}(\theta) = -\mathbb{E}_{(x, y)\sim \mathcal{D}}\, \log p_{\theta}(y|x).
\]
Empirically O($10^3$) SFT updates are made, producing a sequence of checkpoints $\theta^{(t)}$. Early SFT (0–50 updates) aligns to task format; intermediate SFT (50–140 updates) develops arithmetic/reasoning abilities.

**(b) RL Fine–Tuning:** Initialize from an SFT checkpoint $\theta_{\text{SFT}}$, apply PPO to maximize expected task reward $R(x)$:
\[
L_{\text{PPO}}(\theta) = \mathbb{E}_t\big[\min\big(r_t(\theta)\,A_t,\, \mathrm{clip}(r_t(\theta), 1-\epsilon, 1+\epsilon) A_t\big)\big],
\]
where $r_t(\theta) = p_{\theta}(a_t|s_t)/p_{\text{old}}(a_t|s_t)$ and $A_t$ is the advantage. Typical PPO convergence occurs within $\sim$10 checkpoints.

The protocol is operationalized by first training SFT for several hundred steps, then applying PPO on the best available SFT checkpoint, where “best” is not necessarily determined by in-distribution (ID) accuracy or loss but by early peaks in OOD generalization.

## 2. OOD Forgetting and Restoration Phenomena

### Out-of-Distribution (OOD) Metric
Generalization is measured in deliberately shifted evaluation conditions (e.g., a “GeneralPoints” card game with face cards J/Q/K remapped to 11/12/13). OOD accuracy is defined as:
\[
\mathrm{Acc}_{\rm OOD}(\theta) = \mathbb{E}_{(x, y) \sim Q_{\rm OOD}} \bigl[1\{\hat{y}_{\theta}(x) = y\}\bigr].
\]

### OOD Forgetting During SFT
Empirically, $\mathrm{Acc}_{\rm OOD}(\theta_{\text{SFT}})$ peaks extremely early (checkpoints 120–140), reaching $18$–$19\%$, after which it declines monotonically—termed "OOD forgetting"—even as ID metrics (loss, accuracy) continue to improve (ID accuracy $>80\%$ at end-SFT). Critically, the ID loss or accuracy do not warn about this drop; format errors plateau after $\sim$50 updates, and there is no sign of OOD capacity loss in standard validation diagnostics.

### RL-Mediated OOD Restoration
RL applied to SFT checkpoints recovers OOD accuracy up to within $<1\%$ of the early SFT OOD maximum, provided the SFT checkpoint is situated in a recovery window—empirically, SFT checkpoints with ID accuracy in $[40\%, 80\%]$ ($>420$ and $<1200$ updates for LLaMA-11B/Qwen-7B). RL does not deliver OOD generalization beyond the original SFT peak; its role is to restore previously lost capability rather than to synthesize fundamentally new generalization skills.

## 3. Protocol Boundaries, Empirical Guidance, and Failure Modes

The SFT+RL protocol is robust only within specific intervals of the SFT progression:

- **SFT too short (<420 updates):** The model's “base policy” is underfit, leading to reward hacking or RL collapse due to sparse positives.
- **SFT too long (>1200 updates):** Policy entropy collapses, causing RL advantage estimates to skew and PPO to stagnate, with no OOD recovery.

**Practical protocol:**
1. SFT for $20$–$50$ steps to learn format, continue to reach $40$–$60\%$ ID accuracy ($\sim200$–$400$ steps).
2. Switch to PPO: batch size $\sim$256, clip $\epsilon=0.1$–$0.2$, for $10$ rollouts or until OOD accuracy plateaus (typically within $0.5\%$–$1\%$ of SFT OOD peak).
3. Monitor positive-to-negative reward ratio between $0.4$ and $0.8$.

**Failures manifest as instability (RL from underfit SFT), reward hacking, or irrecoverable OOD loss (RL from overspecialized SFT).**

## 4. Mechanistic Insights: SVD Analysis and Singular Vector Rotation

Parameter matrices $W \in \mathbb{R}^{m \times n}$ are decomposed via SVD: $W = U\Sigma V^{\top}$. Contrary to prior assumption, the singular values $\Sigma$ remain nearly invariant through both SFT and RL ($\Delta\sigma_i \lesssim 5 \times 10^{-3}$). However, substantial rotation occurs in the left/right singular spaces $U, V$.

Quantitatively, principal angle spectra between $U_A, U_B$ (top-$k$ singular vector subspaces for $W_A, W_B$) are computed; larger mean rotation angles correspond to greater OOD forgetting. RL “restores” OOD by partially realigning these subspaces, reducing the angular deviation induced by over-specialized SFT.

Ablation of $U,V$ back to earlier SFT checkpoints can recover or erase OOD generalization, while manipulating $\Sigma$ is nearly inert.

## 5. Key Takeaways, Best Practices, and Theoretical Implications

- SFT rapidly aligns (“hard-specializes”) LLM parameters to ID data/modalities, producing early OOD generalization, then over-specializes, rotating singular vectors away from robust OOD-supporting modes.
- RL, given a “recoverable” SFT checkpoint, acts to softly undo these rotations, restoring previously accessible OOD behavior—but is fundamentally limited by the best OOD that SFT originally achieved.
- Monitoring “rotation magnitude” via principal angle metrics offers a direct diagnostic of OOD risk beyond standard loss curves.
- The optimal SFT+RL protocol avoids both insufficient (underfit) and excessive (overspecialized) SFT, emphasizing an intermediate “sweet spot” for policy handoff.

**Future work could include spectral-directional regularization during SFT to minimize harmful rotations, potentially obviating the need for RL-based restoration, and would likely further increase OOD maximality in a single-stage process.**

## 6. Quantitative Benchmarks

| Model           | SFT Peak OOD | SFT End OOD | RL End OOD | ID End Acc |
|-----------------|--------------|-------------|------------|------------|
| LLaMA-11B       | 18–19% (c.140) | ~10%         | 16–18%     | >80%       |
| Qwen-7B         | 18–19% (c.120) | ~10%         | 16–18%     | >80%       |

RL endpoints match—never exceed—the SFT-peak OOD accuracy. OOD cannot be reliably inferred from ID metrics. Proper protocol selection is essential for robust generalization.

---

**References:**
- "RL Fine-Tuning Heals OOD Forgetting in SFT" [2509.12235]

Source: https://www.emergentmind.com/topics/sft-rl-protocol