---
title: Granularity-Modulated Correlation (GMC)
url: https://www.emergentmind.com/topics/granularity-modulated-correlation-gmc
type: topic
---

# Granularity-Modulated Correlation (GMC)

Searching arXiv for the specified paper and closely related work to ground the article.
Search query: arXiv:2601.21738 GMC IQA correlation surface
Granularity-Modulated Correlation (GMC) is a general evaluation framework for Image Quality Assessment (IQA) that turns scalar correlation metrics into a structured, fine-grained analysis of IQA model behavior across two coupled assessment dimensions: absolute quality level (MOS) and pairwise quality granularity ($|\Delta \mathrm{MOS}|$). Proposed in "From Global to Granular: Revealing IQA Model Performance via Correlation Surface" [2601.21738], GMC augments the Generalized Correlation Coefficient (GCC) with a granularity modulator that applies Gaussian weights conditioned on MOS and $|\Delta \mathrm{MOS}|$ to probe local performance, and a distribution regulator that compensates for non-uniform test-set quality-score densities. The result is a 3D correlation surface over the joint $(\mathrm{MOS}, |\Delta \mathrm{MOS}|)$ space, plus a distribution-agnostic global score obtained by surface integration.

## 1. Motivation and problem setting

Evaluation of IQA models has long been dominated by global correlation metrics, such as Pearson Linear Correlation Coefficient (PLCC) and Spearman Rank-Order Correlation Coefficient (SRCC). While widely adopted, these metrics reduce performance to a single scalar, failing to capture how ranking consistency varies across the local quality spectrum. Two IQA models may achieve identical SRCC values, yet one ranks high-quality images, related to high Mean Opinion Score (MOS), more reliably, while the other better discriminates image pairs with small quality/MOS differences, related to $|\Delta \mathrm{MOS}|$. These complementary behaviors are invisible under global metrics [2601.21738].

The framework is motivated by two distinct evaluation axes. The first is the absolute quality scale, or MOS, which addresses how consistent predictions are at specific quality levels and corresponds to prediction accuracy. The second is pairwise granularity, or $|\Delta \mathrm{MOS}|$, which addresses how sensitive predictions are to small or large subjective differences and corresponds to discrimination capability. GMC reveals local performance variations by explicitly conditioning correlation computations on both axes.

A further motivation is distributional instability. Global correlations are sensitive to the MOS distribution: which quality levels are present and how densely they are sampled in the test set. Shifting the test-set distribution can flip the ranking of methods under PLCC and SRCC, even if their capabilities are unchanged. GMC further regularizes against distributional bias so that comparisons remain stable across test sets with different MOS histograms.

## 2. Mathematical foundation in generalized correlation

The framework adopts a pairwise formulation. For a dataset of $n$ images indexed by $i \in \{1,\ldots,n\}$, the ground-truth MOS for image $i$ is $M_i$, the model prediction is $y_i$, and the pairwise absolute MOS difference is
$$
|\Delta \mathrm{MOS}|_{ij} = |M_i - M_j|.
$$

Within Kendall–Gibbons’ generalized correlation (GCC), PLCC, SRCC, and KRCC reduce to pairwise differences:
$$
\Gamma =
\frac{\sum_{i,j} a_{ij} b_{ij}}
{\sqrt{\sum_{i,j} a_{ij}^2}\,\sqrt{\sum_{i,j} b_{ij}^2}}.
$$
The corresponding antisymmetric functions are:
- **PLCC**: $a_{ij} = y_i - y_j,\; b_{ij} = M_i - M_j$.
- **SRCC**: $a_{ij} = r(y_i) - r(y_j),\; b_{ij} = r(M_i) - r(M_j)$.
- **KRCC**: $a_{ij} = \operatorname{sgn}(y_i - y_j),\; b_{ij} = \operatorname{sgn}(M_i - M_j)$.

This representation makes explicit that GMC operates on pairwise relations, that is, ranking consistency is measured via differences $(M_i - M_j)$ versus $(y_i - y_j)$ or their ranks or signs. Conditioning then enters through pairwise weights depending on $M_i$, $M_j$, and $|M_i - M_j|$ [2601.21738].

The baseline scalar metrics remain the reference points. PLCC is defined between vectors $\mathbf{M}$ and $\mathbf{y}$ as
$$
\mathrm{PLCC}(\mathbf{M}, \mathbf{y})
=
\frac{\sum_i (M_i - \bar{M})(y_i - \bar{y})}
{\sqrt{\sum_i (M_i - \bar{M})^2}\,\sqrt{\sum_i (y_i - \bar{y})^2}},
$$
and SRCC is defined from ranks $r(M_i)$ and $r(y_i)$ as
$$
\mathrm{SRCC}(\mathbf{M}, \mathbf{y})
=
\frac{\sum_i (r(M_i) - \overline{r(M)})(r(y_i) - \overline{r(y)})}
{\sqrt{\sum_i (r(M_i) - \overline{r(M)})^2}\,\sqrt{\sum_i (r(y_i) - \overline{r(y)})^2}}.
$$
GMC does not discard these metrics; it instantiates them in weighted form inside the GCC framework.

## 3. Granularity modulation and distribution regulation

For sampling index $k$, GMC replaces uniform pairwise treatment with Gaussian-conditioned weighting centered at a query MOS level $Q_k^s$ and a query pairwise difference $Q_k^d$. The weight for pair $(i,j)$ is
$$
w_k^{ij} = P_k^s(i,j)\; P_k^d(i,j)\; P_k^t(i,j),
$$
where $P_k^s$ and $P_k^d$ are granularity weights and $P_k^t$ is the distribution regulator.

The absolute MOS conditioning term is
$$
P_k^s(i,j) =
\exp\!\left(
-\frac{(Q_k^s - M_i)^2}{2\sigma_i^2}
-\frac{(Q_k^s - M_j)^2}{2\sigma_j^2}
\right).
$$
Here $\sigma_i$ and $\sigma_j$ denote the standard deviations of subjective ratings for images $i$ and $j$, from the dataset if available; otherwise estimated via a parametric rating model, for example Beta regression. The stated rationale is that subjective ratings around MOS are approximately Gaussian, and the joint proximity of $M_i$ and $M_j$ to the local quality level $Q_k^s$ is captured by the product of Gaussians.

The pairwise $|\Delta \mathrm{MOS}|$ conditioning term is
$$
P_k^d(i,j) =
\exp\!\left(
-\frac{\big(Q_k^d - |M_i - M_j|\big)^2}{\sigma_i^2 + \sigma_j^2}
\right).
$$
The stated rationale is that, given Gaussian rating uncertainties, the difference $M_i - M_j$ has variance $\sigma_i^2 + \sigma_j^2$, so the kernel emphasizes pairs whose subjective separation aligns with the target granularity $Q_k^d$.

To decouple evaluation from the test-set MOS histogram, GMC introduces the distribution regulator
$$
P_k^{t}(i, j) = \frac{1}{\mathcal{D}(M_i)} \cdot \frac{1}{\mathcal{D}(M_j)},
$$
where $\mathcal{D}(M_i)$ is the kernel-smoothed density of MOS around $M_i$. It downweights overrepresented MOS regions and compensates underrepresented ones. If image-wise rating standard deviations $\sigma_u$ are available, the density is estimated as
$$
\mathcal{D}(M_i) = \frac{1}{n} \sum_{u=1}^{n}
\exp\!\left(-\frac{(M_u - M_i)^2}{2\sigma_u^2}\right).
$$
If $\sigma_u$ are unavailable, MOS is normalized to $[1,100]$, discretized into $Y$ equal-width bins with centers $c_b$ and frequencies $p_b$, and Gaussian kernel smoothing of the histogram is used:
$$
\mathcal{D}(M_i) \approx \sum_{b \in \mathcal{Y}}
\exp\!\left(-\frac{(M_i - c_b)^2}{2\sigma_h^2}\right)\, \frac{p_b}{n}.
$$
This avoids hard binning artifacts, interpolates empty bins, and prevents division by zero in $P_k^{t}$.

## 4. Correlation surface and global aggregation

At each query point $(Q_k^s, Q_k^d)$, GMC computes a localized weighted correlation $\Gamma_k$. The weighted instantiations of GCC are given explicitly. For PLCC,
$$
\Gamma_{k}(\mathrm{PLCC}) =
\frac{\sum_{i,j} w_k^{ij} (y_i - y_j)(M_i - M_j)}
{\sqrt{\sum_{i,j} w_k^{ij} (y_i - y_j)^2}\;\sqrt{\sum_{i,j} w_k^{ij} (M_i - M_j)^2}}.
$$
For SRCC,
$$
\Gamma_{k}(\mathrm{SRCC}) =
\frac{\sum_{i,j} w_k^{ij} \big(r(y_i) - r(y_j)\big)\big(r(M_i) - r(M_j)\big)}
{\sqrt{\sum_{i,j} w_k^{ij} \big(r(y_i) - r(y_j)\big)^2}\;\sqrt{\sum_{i,j} w_k^{ij} \big(r(M_i) - r(M_j)\big)^2}}.
$$
For KRCC,
$$
\Gamma_{k}(\mathrm{KRCC}) =
\frac{\sum_{i,j} w_k^{ij}\,\operatorname{sgn}(y_i - y_j)\operatorname{sgn}(M_i - M_j)}
{\sqrt{\sum_{i,j} w_k^{ij}\,\operatorname{sgn}(y_i - y_j)^2}\;\sqrt{\sum_{i,j} w_k^{ij}\,\operatorname{sgn}(M_i - M_j)^2}}.
$$
This joint conditioning simultaneously probes image-level accuracy at $\mathrm{MOS} = Q_k^s$ and ranking sensitivity at pairwise granularity $Q_k^d$.

The resulting GMC correlation surface is defined over the joint space of MOS and $|\Delta \mathrm{MOS}|$. Query points are sampled as centers $(Q_k^s, Q_k^d)$, each producing a localized weighted correlation $\Gamma_k$, and a continuous surface $\hat{\Gamma}(Q^s, Q^d)$ is fit by local linear kernel regression over the sampled points:
$$
\mathrm{GMC}(Q^s, Q^d) = \hat{\Gamma}(Q^s, Q^d)
\approx
\mathcal{F}\big(\{(Q_k^s, Q_k^d, \Gamma_k)\}_{k=1}^K\big),
$$
where $\mathcal{F}$ is a nonparametric 2D local linear smoother that reduces boundary bias [2601.21738].

To achieve uniform coverage with few samples, the query grid is sampled with 2D Latin Hypercube Sampling (LHS) over the feasible ranges:
$$
Q_k^s = \frac{\pi_x(k) - u_k}{K}\,\big(Q_{\max}^s - Q_{\min}^s\big) + Q_{\min}^s,\qquad
Q_k^d = \frac{\pi_y(k) - u_k}{K}\,\big(Q_{\max}^d - Q_{\min}^d\big) + Q_{\min}^d,
$$
with $\pi_x,\pi_y$ random permutations of $\{1,\ldots,K\}$ and $u_k \sim \mathcal{U}(0,1)$.

The global GMC score is obtained by integrating the surface over the full domain:
$$
\mathrm{GMC}_g =
\frac{1}{A}
\int_{Q_{\min}^s}^{Q_{\max}^s}
\int_{Q_{\min}^d}^{Q_{\max}^d}
\hat{\Gamma}(x, y)\, dx\, dy,
\qquad
A = (Q_{\max}^s - Q_{\min}^s)(Q_{\max}^d - Q_{\min}^d).
$$
Unlike PLCC and SRCC scalars, $\mathrm{GMC}_g$ reflects both absolute-quality alignment and fine-grained discrimination, and is stabilized by the distribution regulator.

## 5. Workflow, diagnostics, and computational profile

The algorithmic workflow is specified as a sequence of preprocessing, weighted correlation evaluation, surface fitting, and aggregation. The input consists of MOS $\{M_i\}$, predictions $\{y_i\}$, optional per-image $\sigma_i$, and a chosen base correlation among PLCC, SRCC, and KRCC.

**Preprocessing**: normalize MOS to a known range, for example $[1,100]$, to ease kernel parameterization. If $\sigma_i$ are missing, estimate rating uncertainty or set a reasonable constant, and build a KDE of MOS to obtain $\mathcal{D}(M_i)$.

**Sampling query centers**: use $K$-point LHS over $Q^s \in [Q_{\min}^s, Q_{\max}^s]$ and $Q^d \in [Q_{\min}^d, Q_{\max}^d]$. The recommended default is $K=100$; $Q_{\min}^d=0$ and $Q_{\max}^d$ is set to the maximum $|\Delta \mathrm{MOS}|$.

**Computing pairwise weights**: for all pairs $(i,j)$, compute $P_k^s(i,j)$, $P_k^d(i,j)$, and $P_k^t(i,j)$, then set $w_k^{ij} = P_k^s(i,j)P_k^d(i,j)P_k^t(i,j)$. To ensure stability, a minimum effective mass is enforced: if $\sum_{i,j} w_k^{ij}$ is below a threshold, skip $k$ or enlarge kernel bandwidths.

**Computing localized weighted correlation**: apply the corresponding weighted GCC instantiation. For SRCC, ranks $r(\cdot)$ are computed with average-tie handling.

**Surface fitting and global aggregation**: fit $\hat{\Gamma}(Q^s,Q^d)$ via local linear kernel regression on sampled points and numerically integrate $\hat{\Gamma}$ over the full domain to obtain $\mathrm{GMC}_g$, for example by a Riemann sum on a fine grid.

The framework also defines optional diagnostics. $\mathrm{GMC}_s$ is computed over specific MOS slabs and $\mathrm{GMC}_d$ over specific $|\Delta \mathrm{MOS}|$ slabs; these provide targeted insights into performance in LQ, MQ, and HQ regions, and in LD, MD, and HD regions.

The computational complexity is pairwise in the dataset size and linear in the number of query points, namely $O(n^2K)$. Vectorization of pairwise computations is recommended, and pairs may be subsampled for very large $n$ while maintaining coverage across MOS and $|\Delta \mathrm{MOS}|$. LHS drastically reduces $K$ needed for a stable surface, and $K \approx 100$ suffices in practice. Numerical stability requires guarding against tiny denominators in weighted correlation; small $\epsilon$ regularization may be applied if needed.

## 6. Empirical findings, practical interpretation, and scope

The reported experiments cover both full-reference and no-reference IQA settings [2601.21738].

| Setting | Benchmarks and models |
|---|---|
| FR-IQA | PSNR, SSIM, MS-SSIM, LPIPS, DISTS on KADID-10k and PIPAL |
| NR-IQA | NIQE, CLIP-IQA, CLIP-IQA+, QualiCLIP, MANIQA trained on KonIQ-10k and tested on LIVE-Challenge (LIVEC) and SPAQ |

Several findings are presented as behaviors exposed by GMC. On SPAQ, CLIP-IQA excels in high MOS and small $|\Delta \mathrm{MOS}|$, whereas NIQE is stronger at low MOS and large $|\Delta \mathrm{MOS}|$; GMC surfaces clearly display these behaviors, while PLCC and SRCC do not. On KADID-10k, MS-SSIM has high overall SRCC, yet DISTS outperforms MS-SSIM in low-quality MOS regions through $\mathrm{GMC}_s$. LPIPS is strong in low-difference regimes through $\mathrm{GMC}_d$ but can have lower overall PLCC than DISTS, again invisible to scalar metrics.

The reported robustness study concerns altered MOS distributions, including unimodal, bimodal, and trimodal subsets of PIPAL and SPAQ. Under these shifts, global SRCC and PLCC rankings can flip, whereas $\mathrm{GMC}_g$ remains stable and preserves model ordering. A practical implication is that distribution-agnostic aggregation is not only descriptive but also comparative.

The framework is further used for scenario-specific model selection. In high-quality retrieval, models selected by $\mathrm{GMC}_s$ rather than SRCC yield higher mean MOS for top-$k$ selections in HQ and LQ quartiles. In adversarial optimization experiments on PIPAL, models with higher $\mathrm{GMC}_d$ in LD regimes, such as MS-SSIM, produce perceptually superior results when used as target metrics under constraints, consistent with GMC analysis. GMC also reveals complementary pairs; on PIPAL, MS-SSIM complements LPIPS better than DISTS, and simple integration by polarity-aligned sum improves SRCC over naive choices suggested by global SRCC alone.

The paper places GMC in relation to weighted correlation, localized ranking analysis, and kernel smoothing. GMC generalizes weighted Pearson, Spearman, and Kendall within GCC to locally conditioned pairwise weights. By conditioning on $|\Delta \mathrm{MOS}|$, it quantifies discrimination capability in a data-driven kernelized manner. KDE-based density compensation is a standard density-estimation device adapted here to mitigate evaluation bias.

The stated limitations and trade-offs are equally explicit. Computational cost is quadratic in $n$ and linear in $K$, so large datasets may require pair subsampling or GPU-accelerated vectorization. If $\sigma_i$ are unknown and poorly estimated, kernels may misrepresent local neighborhoods; KDE helps but does not fully replace accurate uncertainty modeling. The Gaussian product assumes independence of rating errors across images, so correlations in ratings or content clusters could subtly bias weights. Proposed extensions include multi-attribute conditioning by adding axes for content semantics, distortion type, bitrate, or prompt correspondence for AIGC; alternative kernels, including adaptive or anisotropic kernels; and learned density-compensation terms.

Taken together, GMC transforms model evaluation from monolithic scalar correlation to a diagnostic correlation surface revealing where and how an IQA model performs. It decouples prediction accuracy along MOS from discrimination capability along $|\Delta \mathrm{MOS}|$, and regularizes against dataset distribution bias. This suggests a shift from single-number comparison toward regime-specific analysis, model selection, fusion, and deployment.

Source: https://www.emergentmind.com/topics/granularity-modulated-correlation-gmc