---
title: 'FFHQ-Makeup: Synthetic Paired Makeup Dataset'
url: https://www.emergentmind.com/topics/ffhq-makeup
type: topic
---

# FFHQ-Makeup: Synthetic Paired Makeup Dataset

Searching arXiv for FFHQ-Makeup and closely related makeup-transfer papers to ground the article in current literature.
FFHQ-Makeup is a high-quality synthetic facial makeup dataset designed to provide paired bare–makeup facial images with preserved identity and expression across multiple styles. Built upon the FFHQ dataset, it transfers real-world makeup styles from existing datasets onto 18,000 identities, pairs each identity with 5 different makeup styles, and yields 90,000 high-quality bare–makeup image pairs at \(512 \times 512\) resolution. The dataset was introduced to address the scarcity of large-scale paired makeup data and to avoid the characteristic failure modes of earlier synthetic pipelines, particularly facial distortion in warping-based methods and identity or expression drift in text-to-image generation [2508.03241].

## 1. Definition and dataset profile

FFHQ-Makeup is centered on paired supervision: for a given identity, the bare image and the makeup image are intended to differ in makeup while remaining consistent in facial identity and expression. The paper describes this as a dataset in which each identity is associated with multiple makeup styles rather than a single one-to-one transformation, and it emphasizes manual curation and public release [2508.03241].

| Attribute | Value | Note |
|---|---|---|
| Base identities | 18,000 FFHQ faces | diverse target set |
| Styles per identity | 5 | random makeup styles |
| Total scale | 90,000 bare–makeup image pairs | multi-style pairing |
| Resolution | \(512 \times 512\) | high-resolution |
| Style sources | real-world makeup images from MT and LADN | source makeup corpus |

The dataset paper explicitly positions FFHQ-Makeup as a resource for beauty-related tasks such as virtual try-on, facial privacy protection, and facial aesthetics analysis. Its stated novelty is that it focuses specifically on constructing a makeup dataset rather than only introducing a transfer model [2508.03241].

## 2. Construction pipeline

The construction pipeline begins with two sets: source makeup styles \(\mathcal{S}\), drawn from real-world makeup images, and target identities \(\mathcal{T}\), drawn from FFHQ. For each makeup source image \(I^\mathcal{S}\), the pipeline computes a facial mask \(\mathcal{M}\), 2D landmarks \(\mathcal{L}\), and a 3D face reconstruction \(\mathcal{F}\) using a 3D Morphable Model, specifically FLAME. A bare-face estimate \(\hat{I}_b\) is reconstructed and blended with the original background, and the makeup residual is then defined as
\[
\mathcal{R} = I^\mathcal{S} - \hat{I}_b .
\]

This residual is treated as the makeup-bearing component. The pipeline then performs residual augmentation via 3D re-rendering. Vertex-wise color samples from \(\mathcal{R}\) are extracted using the source geometry \(\mathcal{F}\) and re-rendered onto a different target FFHQ 3D face \(\mathcal{F}^{\mathcal{T}}\). According to the dataset description, each source is augmented 100 times. Manual filtering removes failed cases such as segmentation errors and artifacts, leaving a final curated set of 2,257 high-fidelity makeup residuals. These residuals are then used to synthesize the paired data, with 5 random makeup styles applied to each selected FFHQ identity, followed by group-wise human quality control [2508.03241].

The resulting design is explicitly intended to separate makeup appearance from source-face geometry before transfer. This is the core technical mechanism behind the claim that the paired images preserve facial consistency while still exhibiting diverse makeup styles.

## 3. Generative mechanism and consistency preservation

The synthesis model used to generate FFHQ-Makeup is built upon Stable-Makeup and improved by FreeUV. Its architecture contains a Makeup Residual Detail Encoder, implemented with a frozen CLIP image encoder, and a Makeup Residual Learner with a FreeUV-inspired channel-attention mechanism. Structural control is provided by ControlNet, conditioned on the reconstructed bare face \(\hat{I}_b\) and facial landmarks \(\mathcal{L}\) [2508.03241].

In schematic form, the transfer stage is described as
\[
(I^{\mathcal{T}}, \tilde{\mathcal{R}}) \overset{\text{Diffusion + ControlNet}}{\longrightarrow} \text{makeup-applied image}.
\]

The dataset paper attributes facial consistency to three interacting components. First, FLAME-based 3D reconstruction is used to retain shape, expression, pose, and illumination cues. Second, the makeup residual \(\mathcal{R}\) is intended to discard intrinsic identity information from the makeup source while preserving makeup characteristics. Third, the re-rendering stage applies residuals to diverse target geometries, which the authors describe as enforcing invariance to facial geometry and expression. The paper also notes that, unlike many competing approaches, the synthesis model does not rely on paired ground-truth bare–makeup images for training, because the 3DMM-based pathway provides a self-supervised route to disentangle structure and cosmetics [2508.03241].

## 4. Empirical characteristics and reported comparisons

The empirical evaluation of FFHQ-Makeup emphasizes three properties: identity preservation, semantic consistency, and structural similarity. The reported metrics are ArcFace-based identity similarity (\(\text{Id}\)), DINO-I semantic consistency, and SSIM. In the dataset comparison table, FFHQ-Makeup is reported to achieve the highest \(\text{Id}\) and DINO-I scores among the compared synthetic datasets, while LADN-Syn has the highest SSIM [2508.03241].

| Dataset | Scores | Note |
|---|---|---|
| LADN-Syn | \(0.4973 / 0.9163 / 0.9173\) | Id / DINO-I / SSIM |
| BeautyBank | \(0.5034 / 0.9008 / 0.8060\) | Id / DINO-I / SSIM |
| FFHQ-Makeup | \(0.5888 / 0.9448 / 0.8371\) | Id / DINO-I / SSIM |

The paper further reports a GPT-4o-based visual preference study on 50 random sets, evaluating makeup realism and facial consistency. FFHQ-Makeup received the highest preference, with facial consistency reaching 92%. The ablation summary states that removing makeup residuals causes identity leakage from the source and that removing sampling and re-rendering augmentation leaves more structural artifacts and reduces facial diversity [2508.03241].

These results establish the dataset as a specifically engineered paired corpus rather than merely the by-product of an editing model. The evaluation is directed at dataset quality itself, not only at downstream transfer performance.

## 5. Naming ambiguity and relation to other FFHQ-derived resources

A persistent source of confusion is that the term “FFHQ-Makeup” is not used uniformly across the literature. In the EleGANt summary, “MT/FFHQ-Makeup” refers to the unpaired Makeup Transfer dataset inherited from BeautyGAN, comprising 1,115 non-makeup images and 2,719 makeup images, aligned and resized to \(256 \times 256\) [2207.09840]. In contrast, the 2025 FFHQ-Makeup paper defines a paired synthetic dataset built directly from FFHQ with 18,000 identities and 90,000 bare–makeup pairs [2508.03241].

A second nearby but distinct resource is RetouchingFFHQ, which is also FFHQ-derived but addresses retouching detection rather than paired makeup transfer. RetouchingFFHQ contains 58,158 original images and 652,568 retouched images, spans four retouching operations with four levels each, and is annotated for multi-type, multi-level retouching estimation [2307.10642]. Its target problem is therefore different from FFHQ-Makeup’s paired cosmetic transfer setting.

This suggests that the label “FFHQ-Makeup” should be interpreted in the immediate bibliographic context. In some papers it denotes a specific 2025 paired synthetic dataset; in others it is used more loosely for FFHQ-derived makeup corpora or even for the older MT benchmark lineage.

## 6. Position in subsequent research

Subsequent papers often treat FFHQ-Makeup as an important but imperfect precursor. “FLUX-Makeup” states explicitly that it does not use FFHQ-Makeup for training or testing; instead, it uses raw FFHQ as the base for generating a new paired dataset called HQMT. Its pipeline starts from 70,000 FFHQ images with 5 prompts each, producing 350,000 initial image pairs, and after filtering retains 56,900 high-quality pairs. In the reported comparison, Stable-Makeup has a pass rate of 6.8%, unfiltered HQMT 15.4%, and filtered HQMT 96.2%, with corresponding CLIP-I, SSIM, and L2-M improvements over the earlier synthetic baselines [2508.05069].

A later supervised diffusion paper likewise frames FFHQ-Makeup as a synthetic prior dataset with limitations. It describes FFHQ-Makeup as being created with LEDITS++ over FFHQ faces and notes alignment and style artifacts, then proposes a curated dataset of approximately 200,000 paired samples at \(512 \times 512\) resolution using a train–generate–filter–retrain strategy [2602.00729]. Taken together, these later works present high-quality paired makeup data as a central bottleneck in supervised makeup transfer and use FFHQ-Makeup as a comparison point for newer curation strategies.

Even where FFHQ-Makeup is not used directly, it remains methodologically influential. Later region-aware and controllable diffusion systems continue to rely on FFHQ-based synthetic generation, paired supervision, and explicit attempts to disentangle identity from makeup, all of which are central design themes of FFHQ-Makeup and its accompanying synthesis pipeline [2603.20012].

## 7. Research significance and applications

The FFHQ-Makeup paper presents paired bare–makeup facial images as essential for virtual try-on, facial privacy protection, and facial aesthetics analysis. Because the dataset provides multiple makeup styles per identity while aiming to keep identity and expression fixed, it is also relevant to makeup style transfer, makeup removal, robust face recognition under makeup, and controllable facial attribute manipulation [2508.03241].

Its main significance lies in the combination of scale, explicit pairing, and facial consistency. Earlier synthetic strategies were described as either distorting geometry through warping or drifting identity through text-to-image generation. FFHQ-Makeup addresses those issues through 3DMM reconstruction, residual extraction, residual augmentation, and diffusion-based synthesis under structural control. Later work that surpasses it in scale or filtering quality does not negate that role; rather, it indicates that FFHQ-Makeup helped crystallize paired synthetic makeup data as a distinct research object and a measurable component of system performance.

In that sense, FFHQ-Makeup occupies an intermediate position in the evolution of makeup-transfer datasets: more structured and consistency-oriented than earlier weakly paired or unpaired resources, yet also a clear antecedent to later large-scale curated datasets such as HQMT and other FFHQ-derived supervised corpora.

Source: https://www.emergentmind.com/topics/ffhq-makeup