---
title: Denoising Cosine Similarity (dCS)
url: https://www.emergentmind.com/topics/denoising-cosine-similarity-dcs
type: topic
---

# Denoising Cosine Similarity (dCS)

Denoising Cosine Similarity (dCS) refers to a suite of mathematically principled methods that correct the bias and reduce the noise intrinsic to empirical cosine similarity estimates, particularly when data are contaminated by sampling artifacts or stochastic noise. dCS approaches have recently been formalized both as spectral-cleaning operators for collaborative filtering and as self-supervised loss functions for robust representation learning. These procedures employ random matrix theory, eigenvalue shrinkage, explicit mean-correction, and statistical estimators to isolate true signal from noise-driven spurious similarity, substantially improving downstream performance in k-nearest neighbor (k-NN) systems and deep autoencoder frameworks [1905.07370][2304.09552].

## 1. Foundations of Cosine Similarity and Its Limitations

Cosine similarity is widely adopted for quantifying alignment between vectors in memory-based recommender systems and as an objective in representation learning. Given two vectors $u, v \in \mathbb{R}^D$, the standard cosine similarity and its negative loss form are
\[
\ell_{CS}(u, v) = - \frac{ \langle u, v \rangle }{ \| u \|_2 \| v \|_2 }.
\]
In collaborative filtering, the empirical cosine similarity matrix between $m$ items from an $n \times m$ user-item matrix $X$ (with $x_{ij}\in \mathbb{R}$) is constructed as $S_{cos} = D^{-1/2} X^\top X D^{-1/2}$, where $D = \mathrm{diag}( \|x_{:1}\|^2, ..., \|x_{:m}\|^2 )$. However, when $X$ is noisy or $n,m$ are comparable in size, empirical estimates can exhibit strong noise-induced eigenvalue spread and a systematic overestimation of the leading eigenvalues, especially due to nonzero mean effects [1905.07370][2304.09552].

## 2. Spectral Properties and Random Matrix Theory Analysis

Random Matrix Theory (RMT) provides the statistical underpinning for dCS corrections. For noise-only matrices ($X$ with i.i.d. zero-mean, unit variance entries), the empirical eigenvalue spectrum of $S_{cos}$ follows the Marčenko–Pastur law:
\[
\rho(\lambda) = \frac{1}{2\pi q \lambda} \sqrt{ (\lambda_{max} - \lambda)(\lambda - \lambda_{min}) },
\]
with $\lambda_{min} = (1-\sqrt{q})^2$, $\lambda_{max} = (1+\sqrt{q})^2$, $q = m/n$. This implies that in practical regimes ($q \sim O(1)$), most empirical eigenvalues arise from noise ('noise bulk'), and only those above $\lambda_{max}$ are potentially signal [1905.07370]. Notably, normalization by the column norms in $S_{cos}$ induces an inherent eigenvalue shrinkage relative to the Pearson estimator, but does not fully correct for mean-induced inflation of the largest eigenvalue.

## 3. Denoising Procedures: Eigenvalue Shrinkage, Clipping, and Bias Correction

dCS employs a two-step denoising procedure for similarity estimation:

1. **Eigenvalue Clipping (Noise Bulk Removal):**
   - Compute the SVD of the column-normalized data $X' = X \operatorname{diag}(1/\sigma_j)$ to obtain $S_{cos} \approx V_F \Sigma_F^2 V_F^\top$.
   - All empirical eigenvalues $\lambda_k$ below a threshold $\lambda_{cut} \approx \lambda_{max}$ (from MP law) are set to zero. This eliminates noise-dominated modes.

2. **Explicit Mean Correction (Top Eigenvalue Shrinkage):**
   - The largest empirical eigenvalue $\lambda_1$ is corrected by subtracting a rank-one bias term $\xi = \sum_j n(m_j/\sigma_j)^2$, where $m_j$ is the column mean of $X$ and $\sigma_j$ its norm.
   - Define the cleaned eigenvalues as:
     \[
     \lambda_k^{clean} = \begin{cases}
       \max(0, \lambda_1 - \xi), & k=1 \\
       \lambda_k, & 2 \le k \le F \wedge \lambda_k > \lambda_{cut} \\
       0, & \text{otherwise}
     \end{cases}
     \]
   - The denoised similarity matrix is $S_{denoised} = V_{clean} \Lambda_{clean} V_{clean}^\top$ with the cleaned eigenmodes.

This process yields "cleaned cosine similarity" (Editor’s term: dCS), which better reflects true correlation structure and avoids overfitting to noise [1905.07370].

## 4. dCS for Robust Representation Learning

dCS theory has been extended to self-supervised representation learning under noise [2304.09552]. Here, the problem is to learn representations $h_\theta$ from noisy samples $x = s + \epsilon$, where $s$ is the latent clean signal and $\epsilon$ is isotropic, zero-mean noise. The dCS loss is defined as:
\[
\ell_{dCS}(x; \theta) = \mathbb{E}_b \left[ \frac{ \ell_{CS}(b \odot x, b \odot h_\theta(\tilde x)) }
{ \hat k_{\|b\|_1, \sigma}( \| b \odot s \|_2 ) } \right],
\]
where $b$ is a random mask and $\hat k$ is a weight correcting bias introduced by normalization in the presence of noise. Under the specified noise model, dividing the masked cosine-similarity by $k_{D, \sigma}(t)$ yields an unbiased surrogate for the alignment of the underlying signals (Theorem 1 in [2304.09552]).

Practical estimation of the signal-to-noise ratio $c = \|s\|_2 / (\sigma\sqrt{D})$ uses two independent noisy views $(x, \tilde x)$ of the same $s$, with plug-in estimators:
\[
\hat{c} = \sqrt{ 2 \cdot \max\{ x^\top \tilde x, 0 \} / \| x - \tilde x \|^2 }.
\]
For large $D$, the correction factor $k_{D, \sigma}(\| s \|_2)$ can be directly approximated as $c / \sqrt{1 + c^2}$.

## 5. Implementation Guidelines and Computational Considerations

For memory-based recommenders, the denoised similarity estimator is implemented via the following steps [1905.07370]:

1. Compute column norms and means ($\sigma_j$, $m_j$).
2. Form column-normalized matrix $X'$.
3. Compute truncated SVD of $X'$ to obtain $F$ largest singular values.
4. Compute and correct the largest eigenvalue with $\xi$.
5. Clip all eigenvalues below $\lambda_{cut}$.
6. Construct $V_{clean}$ and $\Lambda_{clean}$ by retaining only nonzero modes.
7. The k-NN embedding $z_i = \sqrt{\lambda^{clean}} V_{clean}[i, :]$ allows fast similarity queries via dot-products.

For the dCS loss in deep networks, the practical algorithm samples masks, applies view-specific corruption (e.g., Blind-Spot Masking), and computes the unbiased loss using either a Monte Carlo or large-$D$ approximation for $k_{D,\sigma}$. Empirically, moderate mask rates ($\rho \approx 0.1$–0.3) and Monte Carlo sample sizes of $100$–$500$ are recommended.

## 6. Empirical Results and Performance Assessment

In k-NN recommendation, dCS produces similarity matrices whose spectra are closer (in spectral norm) to low-rank ground truth than both raw cosine and Pearson matrices, leading to higher accuracy and increased recommendation diversity [1905.07370].

In representation learning, dCS-regularized objectives have demonstrated superior or more stable downstream performance compared to standard CS, MSE, and recent self-supervised denoising baselines across vision and speech benchmarks. For example, in noisy MNIST, dCS yielded higher linear evaluation accuracy and degraded gracefully under increasing noise. On real-world datasets (MNIST, USPS, Pendigits, Fashion-MNIST), dCS matched or surpassed baselines. When used as a regularizer in SimSiam on CIFAR-100 and Tiny-ImageNet, dCS led to consistent gains over both CS-based and Noise2Void baselines, and in speech (ESC-50) with a Vision Transformer AE, dCS substantially outperformed other loss formulations [2304.09552].

## 7. Theoretical Guarantees, Limitations, and Future Directions

dCS is supported by theoretical guarantees on spectral denoising, bias correction, and statistical concentration. The estimator for the weight $k_{D, \sigma}$ is accurate up to $O(D^{-1/2})$ in the feature dimension under mild tail assumptions, and the surrogate loss tightly bounds the true signal alignment. However, dCS presumes isotropic, zero-mean noise and relies on masking schemes (e.g., Bernoulli masking); non-isotropic or structured noise remains an open area for extension. Potential avenues include adapting the correction for more general noise models, learning the masking distribution, and integrating dCS into end-to-end architectures, such as masked autoencoders in vision or speech [2304.09552].

---

**Key References:**

| Title                                                                     | arXiv ID     | Main Contribution                                |
|----------------------------------------------------------------------------|--------------|--------------------------------------------------|
| Cleaned Similarity for Better Memory-Based Recommenders                    | 1905.07370   | Spectral theory and practical estimator for dCS   |
| Denoising Cosine Similarity: A Theory-Driven Approach for Efficient Representation Learning | 2304.09552   | Bias-corrected loss for robust representation     |

These methodologies collectively define the modern denoising cosine similarity framework for robust similarity estimation and noise-aware representation learning.

Source: https://www.emergentmind.com/topics/denoising-cosine-similarity-dcs