Papers
Topics
Authors
Recent
Search
2000 character limit reached

Output Embedding Centering (OEC)

Updated 6 January 2026
  • Output Embedding Centering (OEC) is a methodology that subtracts the global mean from embedding vectors to reveal relative variations.
  • It enhances spectral properties by eliminating the rank-one spike in uncentered data, leading to more accurate PCA/SVD outcomes.
  • OEC improves LLM training stability through methods like μ-centering and μ-loss, offering robustness with minimal computational overhead.

Output Embedding Centering (OEC) refers to a set of methodologies for subtracting the mean output embedding vector before performing downstream operations such as Singular Value Decomposition (SVD), Principal Component Analysis (PCA), or using logits in LLM training. The OEC paradigm ensures that the resulting embeddings capture relative variations rather than being dominated by global mean offsets. This centering process has both theoretical and practical consequences, including improved spectral properties, stable training dynamics, and robust mitigation of logit divergence, particularly in deep learning pipelines and LLM pretraining (Kim et al., 2023, Stollenwerk et al., 5 Jan 2026).

1. Mathematical Foundation of Output Embedding Centering

Let X∈Rn×pX \in \mathbb{R}^{n \times p} be a data matrix with rows xi∈Rpx_i \in \mathbb{R}^p. The mean embedding (row-mean vector) μ\mu is defined as:

μ=1nX⊤1n\mu = \frac{1}{n} X^\top 1_n

where 1n∈Rn1_n \in \mathbb{R}^n is the all-ones vector. The centered data matrix is

Xc=X−1nμ⊤X_c = X - 1_n \mu^\top

This guarantees each column of XcX_c has mean zero: 1n⊤Xc=01_n^\top X_c = 0. In the context of embedding-based models, OEC refers to subtracting this mean embedding from each output vector prior to further processing.

Centering is indispensable for PCA/SVD because the covariance matrix,

X⊤X=Xc⊤Xc+nμμ⊤X^\top X = X_c^\top X_c + n \mu \mu^\top

contains a rank-one "spike" nμμ⊤n \mu \mu^\top due to the mean, distorting the principal axes and eigen-spectrum (Kim et al., 2023).

2. OEC in PCA/SVD-Based Embedding Pipelines

SVD of the centered matrix yields xi∈Rpx_i \in \mathbb{R}^p0; for uncentered data, xi∈Rpx_i \in \mathbb{R}^p1 with respective spectral values. Absent centering, the first singular vector xi∈Rpx_i \in \mathbb{R}^p2 maximizes xi∈Rpx_i \in \mathbb{R}^p3 over xi∈Rpx_i \in \mathbb{R}^p4, aligning with the mean direction xi∈Rpx_i \in \mathbb{R}^p5. Proposition 1 ("parallel" condition) establishes that if the first centered singular vector xi∈Rpx_i \in \mathbb{R}^p6 is parallel to xi∈Rpx_i \in \mathbb{R}^p7, then xi∈Rpx_i \in \mathbb{R}^p8 (Kim et al., 2023).

If xi∈Rpx_i \in \mathbb{R}^p9, discarding the leading component from the uncentered embedding (μ\mu0) retrieves the centered μ\mu1-dimensional embedding up to sign: μ\mu2. More generally, if the span of the first μ\mu3 right singular vectors of μ\mu4 and μ\mu5 differ only by an orthogonal μ\mu6 change of basis, then their embeddings are related by an orthogonal transformation.

Spectrally, the eigenvalues of μ\mu7 interlace those of μ\mu8. The top eigenvalue μ\mu9 of μ=1nX⊤1n\mu = \frac{1}{n} X^\top 1_n0 includes the mean component μ=1nX⊤1n\mu = \frac{1}{n} X^\top 1_n1, "soaking up" variance and relegating relative structure to lower-order components. Thus, without centering, principal axes are prone to over-represent the mean, obscuring meaningful structure (Kim et al., 2023).

3. OEC for Stable LLM Pretraining: Geometric Diagnosis and Formalism

In LLMs, especially decoder-only architectures, output logits μ=1nX⊤1n\mu = \frac{1}{n} X^\top 1_n2 rely on output embeddings μ=1nX⊤1n\mu = \frac{1}{n} X^\top 1_n3 and hidden state μ=1nX⊤1n\mu = \frac{1}{n} X^\top 1_n4. At large learning rates, output-logit divergence occurs—some logits μ=1nX⊤1n\mu = \frac{1}{n} X^\top 1_n5 tend to μ=1nX⊤1n\mu = \frac{1}{n} X^\top 1_n6 or μ=1nX⊤1n\mu = \frac{1}{n} X^\top 1_n7, precipitating training collapse. Existing solutions such as z-loss regularization (μ=1nX⊤1n\mu = \frac{1}{n} X^\top 1_n8, μ=1nX⊤1n\mu = \frac{1}{n} X^\top 1_n9) suppress positive logit divergence but allow negative drift (Stollenwerk et al., 5 Jan 2026).

OEC directly targets the source: the anisotropic drift of mean output embedding 1n∈Rn1_n \in \mathbb{R}^n0. If 1n∈Rn1_n \in \mathbb{R}^n1 is uncontrolled, it induces an unbounded global shift in logits. Centering ensures 1n∈Rn1_n \in \mathbb{R}^n2, bounding logit values and preventing divergence.

There are two primary OEC variants:

  • μ-centering: Deterministically subtract 1n∈Rn1_n \in \mathbb{R}^n3 from every 1n∈Rn1_n \in \mathbb{R}^n4 after each optimization step:

1n∈Rn1_n \in \mathbb{R}^n5

Ensures 1n∈Rn1_n \in \mathbb{R}^n6 and does not alter loss or probabilities due to softmax invariance.

  • μ-loss: Regularize the mean norm by adding 1n∈Rn1_n \in \mathbb{R}^n7 to the standard negative log-likelihood:

1n∈Rn1_n \in \mathbb{R}^n8

Penalizes excessive drift in 1n∈Rn1_n \in \mathbb{R}^n9; Xc=X−1nμ⊤X_c = X - 1_n \mu^\top0 is typically Xc=X−1nμ⊤X_c = X - 1_n \mu^\top1.

4. Theoretical Guarantees and Algorithmic Implementation

Theorem 3.5 of (Stollenwerk et al., 5 Jan 2026) shows μ-centering provably tightens the bound on maximum logits. Assuming dot products Xc=X−1nμ⊤X_c = X - 1_n \mu^\top2 lie in Xc=X−1nμ⊤X_c = X - 1_n \mu^\top3, after centering, the effective spread is reduced by a ratio Xc=X−1nμ⊤X_c = X - 1_n \mu^\top4, enforcing Xc=X−1nμ⊤X_c = X - 1_n \mu^\top5.

For μ-loss, since unbounded Xc=X−1nμ⊤X_c = X - 1_n \mu^\top6 incurs infinite regularization, the optimizer keeps Xc=X−1nμ⊤X_c = X - 1_n \mu^\top7 close to zero, thereby bounding logits.

A generic PyTorch-style algorithm for OEC implementation is:

1n⊤Xc=01_n^\top X_c = 00 Computational cost is minimal: μ-centering adds Xc=X−1nμ⊤X_c = X - 1_n \mu^\top8 overhead; μ-loss is even less (<Xc=X−1nμ⊤X_c = X - 1_n \mu^\top9).

5. Best Practices, Hyperparameter Sensitivity, and Experimental Outcomes

OEC does not require changes to learning rate schedules. μ-centering is hyperparameter-free and stable across all tested learning rates. μ-loss recommends XcX_c0, insensitive to tuning in the range XcX_c1. By contrast, z-loss (XcX_c2) requires careful tuning; optimal XcX_c3 in practice.

Empirical results reported in (Stollenwerk et al., 5 Jan 2026) demonstrate:

  • Optimal test loss: All methods (baseline, z-loss, μ-loss, μ-centering) reach essentially identical minimum values.
  • Learning rate sensitivity (LRS): μ-loss and μ-centering yield the lowest LRS, indicating greatest stability. For the largest Transformer (221M params), LRS for μ-loss is 0.056, for μ-centering is 0.061, for z-loss is 0.109, and for baseline is 0.412.
  • Convergence under large learning rates: μ-based OEC variants remain stable at XcX_c4, while baseline and z-loss diverge for XcX_c5 and XcX_c6, respectively.
  • Mean embedding norm diagnostics: μ-centering enforces XcX_c7; μ-loss holds XcX_c8. Baseline and z-loss allow XcX_c9 to grow with learning rate.
  • Runtime overhead: μ-centering and μ-loss are competitive (<1\% overhead), outperforming z-loss (+0.8\%...6.4\%).
Method Learning Rate Sensitivity (LRS) Runtime Overhead
Baseline 0.412 0 %
z-loss (10⁻⁴) 0.109 +0.8 %
μ-loss (10⁻⁴) 0.056 +0.2 %
μ-centering 0.061 +0.3 %

6. Connections to Broader Embedding Theory and Robustness

OEC unifies several principles underlying dimensionality reduction and representation learning. In embedding-based pipelines (e.g., word embeddings, graph embeddings, deep representation vectors), failure to center results in the first axis being dominated by the global mean. Subtracting the mean ensures alignment of downstream principal axes with genuine structural directions, facilitates orthogonal invariance, and stabilizes spectral properties—removing a rank-one “mean spike” and restoring equivalence between covariance-based and SVD-based PCA (Kim et al., 2023).

When mean centering is impractical, discarding the first singular direction in uncentered SVD is an effective heuristic, closely matching centered PCA outcomes with an error bounded by the top centered singular value.

7. Implications, Limitations, and Prospects

OEC (both μ-centering and μ-loss) offers a theoretically grounded, low-overhead solution to output-logit divergence by addressing the root cause: anisotropic drift of output embeddings. A plausible implication is the broad applicability of OEC to other settings where mean drift affects representations, including vision models, multimodal architectures, and robust unsupervised learning. Hyperparameter robustness and ease of integration suggest immediate utility across modern LLM training frameworks, with minimal alterations to existing pipelines (Stollenwerk et al., 5 Jan 2026).

No significant controversies are identified around the necessity of centering for SVD/PCA or its stabilizing effect on output logits. Nevertheless, the choice between deterministic μ-centering and regularization-based μ-loss may depend on downstream compatibility and system constraints.

Output Embedding Centering constitutes a canonical best practice for modern embedding analysis and stable model training.

Definition Search Book Streamline Icon: https://streamlinehq.com
References (2)

Topic to Video (Beta)

No one has generated a video about this topic yet.

Whiteboard

No one has generated a whiteboard explanation for this topic yet.

Follow Topic

Get notified by email when new papers are published related to Output Embedding Centering (OEC).