Papers
Topics
Authors
Recent
Search
2000 character limit reached

Conditional Fréchet Distance for Text-to-Image Evaluation

Updated 5 July 2026
  • Conditional Fréchet Distance (cFreD) is a metric that measures conditional distributions by comparing real and generated images based on given text prompts.
  • It evaluates both visual fidelity and text-prompt alignment, ensuring that realistic images also accurately reflect their conditioning text.
  • Empirical benchmarks show cFreD strongly correlates with human rankings, achieving metrics like ρ² up to 0.97 and Rank Accuracy of 91.1% on key datasets.

Conditional Fréchet Distance (cFreD) is a conditional variant of the classical Fréchet Distance / Fréchet Inception Distance designed for evaluating text-to-image generative models. It explicitly incorporates the conditioning text prompt xx into the comparison between the real-image distribution and the generated-image distribution, so that it simultaneously evaluates visual fidelity and text-prompt alignment. In its current form, cFreD adapts the conditional Fréchet Inception Distance formulation of Soloveitchik et al. to text-to-image generation and replaces the Inception backbone by a modern vision encoder, yielding a training-free metric based on pre-trained encoders only (Koo et al., 27 Mar 2025, Soloveitchik et al., 2021).

1. Definition and conceptual scope

cFreD is a distributional metric for conditional generation. The real object of comparison is not the marginal image distribution alone, but the family of conditional image distributions indexed by prompts. In the notation used by the text-to-image literature, xx denotes the prompt, yy a real image, and y^\hat y a generated image. The metric is constructed so that a model is penalized both when it generates unrealistic images and when it generates realistic images that are not properly related to their prompts (Koo et al., 27 Mar 2025).

This design addresses a specific deficiency of classical evaluation practice. Fréchet Inception Distance compares unconditional feature distributions of real and generated images and therefore ignores the prompt. CLIPScore measures image-text alignment for individual pairs but does not compare generated images to a reference image distribution. Human-preference-trained scoring models such as ImageReward, HPSv2, PickScore, and MPS can correlate strongly with human judgments, but they are explicitly trained on preference data and may require continual updating as models and domains change. cFreD is intended as a statistical alternative that uses only pre-trained encoders and conditional joint moments (Koo et al., 27 Mar 2025).

A standard motivating example is prompt permutation. If a system generates realistic “cat” images for “dog” prompts and realistic “dog” images for “cat” prompts, FID can remain small because the marginal image distribution is still correct, whereas cFreD increases because the text-image dependency structure is wrong (Koo et al., 27 Mar 2025).

2. Mathematical formulation

The immediate precursor of cFreD is Conditional Frechet Inception Distance (CFID), which defines a distance between conditional distributions by averaging a Wasserstein/Fréchet distance over the conditioning variable. In the Gaussian setting, one assumes conditional feature distributions

Qy∣x=N(μy∣x, Σyy∣x),Qy^∣x=N(μy^∣x, Σy^y^∣x),Q_{y|x} = \mathcal{N}(\mu_{y|x},\, \Sigma_{yy|x}), \qquad Q_{\hat{y}|x} = \mathcal{N}(\mu_{\hat{y}|x},\, \Sigma_{\hat{y}\hat{y}|x}),

and defines the conditional Fréchet distance as the expectation over xx of the Fréchet distance between these conditional Gaussians (Koo et al., 27 Mar 2025).

Equivalently, under a joint Gaussian model for (x,y,y^)(x,y,\hat y), CFID admits a closed form in terms of unconditional means, cross-covariances, and conditional covariances: CFID=∥my−my^∥2+Tr⁡[(Cyx−Cy^x)Cxx−1(Cxy−Cxy^)]+Tr⁡ ⁣(Cyy∣x+Cy^y^∣x−2(Cyy∣x1/2Cy^y^∣xCyy∣x1/2)1/2),\mathrm{CFID} = \|m_y - m_{\hat y}\|^2 + \operatorname{Tr}\Big[(C_{yx} - C_{\hat y x}) C_{xx}^{-1} (C_{xy} - C_{x\hat y}) \Big] + \operatorname{Tr}\!\left( C_{yy\mid x} + C_{\hat y\hat y\mid x} -2 \big( C_{yy\mid x}^{1/2} C_{\hat y\hat y\mid x} C_{yy\mid x}^{1/2} \big)^{1/2} \right), where

Cyy∣x=Cyy−CyxCxx−1Cxy,Cy^y^∣x=Cy^y^−Cy^xCxx−1Cxy^.C_{yy\mid x} = C_{yy} - C_{yx} C_{xx}^{-1} C_{xy}, \qquad C_{\hat y\hat y\mid x} = C_{\hat y\hat y} - C_{\hat y x} C_{xx}^{-1} C_{x\hat y}.

The first term measures unconditional mean mismatch, the second measures mismatch in the cross-covariance between inputs and outputs, and the third is a Fréchet term on the conditional covariances (Soloveitchik et al., 2021).

The 2025 cFreD formulation instantiates this conditional Fréchet distance in a joint image-text feature space. In the notation used there, the equivalent expression is written in terms of μy\mu_y, xx0, xx1, xx2, xx3, xx4, and xx5, with the same three-part interpretation: global realism, text-image correlation mismatch, and conditional variability (Koo et al., 27 Mar 2025).

For comparison, classical FID between two Gaussian feature distributions xx6 and xx7 is

xx8

which contains no conditioning term and therefore cannot distinguish prompt-agnostic generation from prompt-faithful generation (Koo et al., 27 Mar 2025).

3. Implementation in text-to-image evaluation

The recommended cFreD instantiation uses a DINOv2-G/14 Vision Transformer for image embeddings and the OpenCLIP ConvNeXt-B text encoder for text embeddings, with no fine-tuning applied. The metric is therefore training-free in the sense that it does not fit an evaluator to human preferences; it estimates means and covariances in a fixed joint feature space induced by pre-trained encoders (Koo et al., 27 Mar 2025).

Operationally, one computes text features xx9, real-image features yy0, and generated-image features yy1 for prompt-image pairs. From these one estimates yy2, yy3, yy4, yy5, yy6, yy7, yy8, and yy9, derives the conditional covariances using Gaussian conditioning identities, and evaluates the conditional Fréchet formula. If multiple generated samples are available per prompt, they simply contribute additional y^\hat y0 observations for the generated distribution (Koo et al., 27 Mar 2025).

The implementation choices are themselves empirically motivated. The study evaluates 43 image encoders and 43 text encoders, and reports that transformer-based image encoders outperform InceptionV3 as cFreD backbones. It also reports that correlation with human rankings depends on encoder pretraining scale, image resolution, model size, and feature dimensionality, with performance peaking around 1024 features in the reported ablations (Koo et al., 27 Mar 2025).

As a practical metric, cFreD is interpreted as a distance: lower is better. The reported usage is comparative rather than absolute, ranking models evaluated on the same prompt set with the same encoders (Koo et al., 27 Mar 2025).

4. Empirical behavior and benchmark results

The principal empirical claim is that cFreD correlates with human preferences more strongly than standard statistical metrics and, on some benchmarks, more strongly than metrics trained on human preference data. The evaluation protocol uses model-level rankings, the squared Spearman correlation coefficient y^\hat y1, and Rank Accuracy, defined as the probability that a random pair of models is ranked in the same order by the metric and by humans (Koo et al., 27 Mar 2025).

Benchmark cFreD result Reported comparison
HPDv2 y^\hat y2, Rank Accuracy y^\hat y3 Highest among all metrics, including human-trained
Parti-Prompts Correlation y^\hat y4 Best among statistical metrics
COCO arena Correlation y^\hat y5, Rank Accuracy y^\hat y6 Only statistical metric with clearly nontrivial alignment

On HPDv2, cFreD attains y^\hat y7 and Rank Accuracy y^\hat y8, outperforming FID, FDy^\hat y9, CLIPScore, CMMD, HPSv2, Aesthetic, ImageReward, and MPS in the reported table. On Parti-Prompts, cFreD is the strongest among purely statistical metrics, although HPSv2 and ImageReward are higher overall. On the COCO arena benchmark, cFreD again exceeds the other statistical metrics and is reported as the only one with clearly nontrivial alignment (Koo et al., 27 Mar 2025).

The paper also reports qualitative failure-mode sensitivity. In examples where one model preserves human ordering between best and worst generations for a prompt and competing statistical metrics do not, cFreD is presented as the only statistical metric that matches the human ordering. A plausible implication is that the cross-covariance and conditional-covariance terms are carrying evaluation signal that is absent from purely marginal metrics (Koo et al., 27 Mar 2025).

5. Relation to CFID and to other uses of “conditional Fréchet”

Historically, cFreD sits closest to CFID. CFID introduced the general problem of defining Wasserstein/Fréchet distances between conditional distributions, proved the inequalities

Qy∣x=N(μy∣x, Σyy∣x),Qy^∣x=N(μy^∣x, Σy^y^∣x),Q_{y|x} = \mathcal{N}(\mu_{y|x},\, \Sigma_{yy|x}), \qquad Q_{\hat{y}|x} = \mathcal{N}(\mu_{\hat{y}|x},\, \Sigma_{\hat{y}\hat{y}|x}),0

and derived the Gaussian closed form for conditional Fréchet distance in feature space. cFreD inherits that mathematical template but changes the application domain from general conditional generative modeling to text-to-image evaluation and substitutes the feature backbone configuration (Soloveitchik et al., 2021, Koo et al., 27 Mar 2025).

The term “conditional Fréchet” also appears in statistical Fréchet regression, but there it denotes something else: the target is the conditional Fréchet mean

Qy∣x=N(μy∣x, Σyy∣x),Qy^∣x=N(μy^∣x, Σy^y^∣x),Q_{y|x} = \mathcal{N}(\mu_{y|x},\, \Sigma_{yy|x}), \qquad Q_{\hat{y}|x} = \mathcal{N}(\mu_{\hat{y}|x},\, \Sigma_{\hat{y}\hat{y}|x}),1

not a distributional evaluator for generative models. In that literature, the conditional Fréchet functional is a regression risk over metric-space-valued responses, typically regularized by total variation or other smoothness priors (Lin et al., 2019).

In geometric and fine-grained algorithmics, “conditional Fréchet distance” is again used differently. There, the conditional aspect often refers to SETH-based conditional lower bounds or to input conditions such as Qy∣x=N(μy∣x, Σyy∣x),Qy^∣x=N(μy^∣x, Σy^y^∣x),Q_{y|x} = \mathcal{N}(\mu_{y|x},\, \Sigma_{yy|x}), \qquad Q_{\hat{y}|x} = \mathcal{N}(\mu_{\hat{y}|x},\, \Sigma_{\hat{y}\hat{y}|x}),2-packedness, rather than to a new statistical distance. One paper explicitly states that the phrase “Conditional Fréchet Distance (cFreD)” does not appear as a formal definition there and that there is no new distance measure called cFreD; the conditional aspect is about conditional complexity (Bringmann et al., 2014). Related work on approximating the Fréchet distance when only one curve is Qy∣x=N(μy∣x, Σyy∣x),Qy^∣x=N(μy^∣x, Σy^y^∣x),Q_{y|x} = \mathcal{N}(\mu_{y|x},\, \Sigma_{yy|x}), \qquad Q_{\hat{y}|x} = \mathcal{N}(\mu_{\hat{y}|x},\, \Sigma_{\hat{y}\hat{y}|x}),3-packed similarly treats “conditional” as a structural assumption on the input, not as a prompt-conditioned metric (Gudmundsson et al., 2024).

6. Limitations, robustness, and extensions

cFreD is training-free, but it is not model-free in the stronger sense of being representation-independent. Its behavior depends on the chosen image and text encoders, and the reported ablations show that backbone choice, pretraining scale, resolution, and feature dimensionality all affect metric–human correlation. This suggests that future encoder improvements can change cFreD performance without altering its underlying formula (Koo et al., 27 Mar 2025).

The method also inherits the Gaussian-moment assumption of FID-style metrics. The reported formulation uses means, covariances, and cross-covariances only, so it may miss non-Gaussian structure in conditional output distributions. The paper further notes that, like other global metrics, cFreD returns a single scalar per model and therefore does not diagnose which prompts or failure modes are responsible for a poor score (Koo et al., 27 Mar 2025).

At the same time, the conditional Fréchet framework is broader than the text-to-image setting in which cFreD was introduced. The 2025 study explicitly identifies extensions to text-to-video and to multi-conditional settings such as ControlNet, where conditioning may involve both text and structural images. It also proposes hybrid variants that keep the cFreD formula but replace the encoders with human-preference-trained backbones. This suggests a family of conditional Fréchet metrics rather than a single fixed evaluator (Koo et al., 27 Mar 2025).

In contemporary usage, therefore, cFreD refers most specifically to the 2025 text-to-image metric built from conditional Fréchet distance in a DINOv2/OpenCLIP feature space. More broadly, it belongs to a line of work that treats Fréchet-style geometry as a tool for comparing conditional distributions, with distinct interpretations in generative evaluation, metric-space regression, and fine-grained geometric algorithms (Koo et al., 27 Mar 2025, Soloveitchik et al., 2021).

Topic to Video (Beta)

No one has generated a video about this topic yet.

Whiteboard

No one has generated a whiteboard explanation for this topic yet.

Follow Topic

Get notified by email when new papers are published related to Conditional Fréchet Distance (cFreD).