---
title: Relativistic Critic in GANs
url: https://www.emergentmind.com/topics/relativistic-critic-discriminator
type: topic
---

# Relativistic Critic in GANs

A relativistic critic, also called a relativistic discriminator, is an architectural and objective reformulation of the discriminator module in adversarial frameworks—notably generative adversarial networks (GANs)—in which the discriminator does not estimate the solitary realism of a sample (“how real is $x$?”) but instead evaluates samples in a relative fashion (“how much more real is $x$ than $y$?”). This approach replaces the classical pointwise discrimination in standard GANs with a formulation that inherently couples the outputs for real and fake examples, producing well-characterized statistical divergences and leading to improved empirical stability, sample quality, and optimization characteristics in both generative modeling and modern large language model reinforcement learning paradigms [1807.00734, 2511.21667, 1901.02474, 2110.11293].

## 1. Mathematical Definition and Formulation

The canonical relativistic discriminator does not estimate $P(x \;\text{is real})$ directly. Instead, it estimates the probability that a real sample $x_r\sim P$ is more realistic than a generated (fake) sample $x_f\sim Q$:
\[
D_{\mathrm{rel}}(x_r, x_f) = \sigma(C(x_r) - C(x_f))
\]
where $C(\cdot) \in \mathbb{R}$ is the real-valued critic/logit and $\sigma(\cdot)$ is the sigmoid activation. This construction is extended to multiple variants:

- **Relativistic Standard GAN (RSGAN):**
  \[
  L_D^{\mathrm{RSGAN}} = -\mathbb{E}_{x_r, x_f}\left[\log \sigma(C(x_r) - C(x_f))\right]
  \]
  \[
  L_G^{\mathrm{RSGAN}} = -\mathbb{E}_{x_r, x_f}\left[\log \sigma(C(x_f) - C(x_r))\right]
  \]
  
- **Relativistic Average GAN (RaGAN):** Compares each sample’s score to the batch-mean of the opposite type.
  \[
  \overline{C}_f = \mathbb{E}_{x_f}[C(x_f)], \quad \overline{C}_r = \mathbb{E}_{x_r}[C(x_r)]
  \]
  For a real $x$:
  \[
  \bar{D}(x_r) = \sigma(C(x_r) - \overline{C}_f)
  \]
  Losses become:
  \[
  L_D^{\mathrm{RaSGAN}} = -\mathbb{E}_{x_r}\left[\log\bar{D}(x_r)\right] - \mathbb{E}_{x_f}\left[\log(1-\bar{D}(x_f))\right]
  \]
  
- **Extension to Arbitrary $f$-Divergences:** Any $f$-divergence GAN objective
  \[
  L_D = \mathbb{E}_{x_r}[f_1(C(x_r))] + \mathbb{E}_{x_f}[f_2(C(x_f))]
  \]
  is relativized by replacing arguments $C(x_r) \to C(x_r)-C(x_f)$ (or means), generalizing to paired and average relativistic $f$-divergence forms [1901.02474].

This construction ensures the discriminator's output is directly influenced by both real and fake batches, producing a margin-based, pairwise interaction rather than independent scalar values.

## 2. Motivation and Theoretical Benefits

The relativistic formulation is motivated by several deficiencies in classical SGAN learning:

- **Symmetry and Coupling:** In a mini-batch containing equal numbers of real and fake samples, making fake samples more realistic (raising their score) should naturally reduce the discriminator confidence in real samples. SGAN's pointwise setup lacks this coupling, potentially allowing for degenerate solutions.
  
- **Divergence Minimization:** Standard SGAN approximates the Jensen-Shannon divergence, where optimal convergence should force discriminator outputs for both real and fake to $0.5$. However, classical SGAN only forces $D(x_f) \to 1$, producing overconfident discriminators and vanishing gradients on the real side.

- **Gradient Dynamics:** In IPM-GANs, the critic's gradient always mixes $\mathbb{E}[\nabla C(x_r)]$ and $\mathbb{E}[\nabla C(x_f)]$, so both influence updates. In non-relativistic SGANs, perfect discrimination causes the real-sample gradient to vanish, leading to “critic stalling.” Relativistic discriminators keep both gradients active throughout training.

- **Statistical Divergence Properties:** For a broad class of concave $f$ functions, relativistic GAN objectives define bona fide statistical divergences, which are (topologically) strictly stronger than their classical counterparts, yet yield better optimization characteristics [1901.02474].

## 3. Variants and Generalizations

Multiple variants extend the core relativistic concept:

- **Relativistic Paired GAN (RpGAN):** Uses strictly pairwise comparisons:  
  \[
  D_f^{\mathrm{Rp}}(P, Q) = \sup_{C} 2\mathbb{E}_{x\sim P, y\sim Q}[f(C(x) - C(y))]
  \]

- **Relativistic Average (RaGAN):** Each sample is compared to the opposing batch mean.

- **Further Generalizations:**  
  - **RalfGAN:** One-sided relativization (only real batch compared to fake mean).
  - **RcGAN:** Centered scores using the mean of both distributions.

The choice of relativization (paired, batch mean, mixed mean) impacts both divergence strength and estimator bias/noise characteristics.

\[
\begin{aligned}
\text{Wasserstein GAN: } & D^W(P,Q) \\
\text{f-divergence GAN (SyGAN): } & D^{Sy}_f(P,Q) \\
\text{Relativistic Paired: } & D^{Rp}_f(P,Q) \\
\text{Relativistic Average: } & D^{Ra}_f(P,Q)
\end{aligned}
\]
with topological strength $D^W \prec D^{Sy}_f \prec D^{Rp}_f \prec D^{Ra}_f$ [1901.02474].

## 4. Algorithms and Implementation

The relativistic critic is implemented with minimal changes to canonical GAN architectures. The algorithmic steps are as follows:

- **For RSGAN and RpGAN:**  
  1. Sample batches of real ${x_r^i}$ and fake ${x_f^i}$.
  2. For $n_D$ discriminator steps, update $C$ by maximizing the paired loss over $i$.
  3. After $n_D$ steps, update the generator (or policy) using the reversed loss.

- **For RaGAN:**  
  1. Compute batch means $\mu_r$, $\mu_f$ of critic scores.
  2. For each $x$, use relativistic score differences with mean.
  3. Apply spectral normalization or gradient penalty for stability.
  4. Empirically, a single discriminator update per generator update suffices—no need for frequent critic steps.

- **In LLM Reasoning/RL (RARO):**  
  - Policy $\pi_\theta$ and critic $f_\phi(x, a)$ are trained adversarially.
  - Critic compares each expert answer $a_E$ to a policy (model) answer $a_G$ via $D_\phi(x, a_E, a_G) = \sigma(f_\phi(x, a_E) - f_\phi(x, a_G))$.
  - The policy is updated using PPO, with the critic's relativistic margin as reward.
  - Stabilization relies on two-time-scale updates, gradient penalties, and normalization [2511.21667].

## 5. Empirical Performance and Observed Effects

Extensive empirical studies demonstrate:

- **Stability:** Relativistic GANs (RSGAN, RaGAN, RaLSGAN, RMCosGAN) yield systematically lower FID variance and improved convergence as compared to non-relativistic GANs, across multiple datasets and architectures.
- **Sample Quality:** RaGAN with gradient penalty matches or surpasses WGAN-GP in FID with inexpensive, single-step D updates ($\sim$400% faster to reach SOTA). High-resolution image synthesis (e.g., 256x256 with $N=2011$) becomes tractable where SGAN, LSGAN, and even WGAN-GP collapse.
- **Comparative Performance:**

  | Loss         | FID (CIFAR-10, $n_D$=1) | FID (CAT, 64x64, min) |
  |--------------|------------------------|----------------------|
  | SGAN         | 40.64                  | 16.56                |
  | RSGAN        | 36.61                  | 19.03                |
  | RaSGAN       | 31.98                  | 15.38                |
  | WGAN-GP      | 83.89                  | >155 (at 256x256)    |
  | RSGAN-GP     | 25.60                  | —                    |
  | RaLSGAN      | —                      | **11.97**            |

  [1807.00734, 2110.11293]

- **Loss Function Variants:** Relativistic margin cosine loss (RMCosGAN) outperforms both CE and Ra-LS loss in FID and IS; e.g., CIFAR-10 FID = $31.34 \pm 0.13$ for RMCosGAN vs $33.82 \pm 0.03$ for CE [2110.11293].
- **Estimator Bias / Variance:** Minimum-variance unbiased estimators for judiciously relativistic divergences _do not_ improve sample quality—additional estimation noise regularly acts as implicit regularization, improving the generator [1901.02474].

## 6. Applications Beyond Image Synthesis

The relativistic critic is central not only in classic image-based GANs but in modern adversarial RL and imitation learning, especially in large language model training:

- **RARO (Relativistic Adversarial Reasoning Optimization):** Develops expert-level reasoning in LLMs where verifiers are absent, using a relativistic critic to deliver policy improvement signals [2511.21667].
- **RL Policy Updates:** The reward signal is the pairwise margin between model and expert response; this “relativizes" the imitation objective and prevents reward collapse or saturation.

## 7. Architectural and Hyperparameter Recommendations

- **Standard Practice:** DCGAN backbone, spectral normalization in $D$, batch norm in $G$, Adam optimizer ($2\times 10^{-4}$, $\beta_1=0.5$, $\beta_2=0.999$), batch size 64, $n_D=1$ sufficient for stability [1807.00734, 2110.11293].
- **RARO Critic:** Small MLP head (hidden size 512) atop a frozen LLM encoder; spectral norm or gradient penalty regularization; critic LR $3\times 10^{-5}$ vs policy LR $1\times 10^{-5}$; 2–5 critic steps per generator update; reward normalization/clipping [2511.21667].
- **Critical Hyperparameters:** For RMCosGAN, angular margin $m$ must be carefully tuned ($m \approx 0.15$ recommended), too large or too small leads to instability or feature collapse [2110.11293].

---

In sum, the relativistic critic fundamentally shifts adversarial learning from isolated sample scoring to paired, margin-based discrimination. This approach yields statistically principled divergences, improves optimization dynamics, and consistently enhances stability and generative performance across adversarial machine learning domains [1807.00734, 1901.02474, 2110.11293, 2511.21667].

Source: https://www.emergentmind.com/topics/relativistic-critic-discriminator