Papers
Topics
Authors
Recent
Search
2000 character limit reached

FFHQ-Makeup: Synthetic Paired Makeup Dataset

Updated 18 July 2026
  • FFHQ-Makeup is a synthetic paired makeup dataset featuring 90,000 bare–makeup image pairs across 18,000 FFHQ faces with 5 varied makeup styles per identity.
  • The dataset’s pipeline employs 3D face reconstruction, makeup residual extraction, and ControlNet-guided augmentation to ensure consistent identity and expression.
  • Empirical evaluations show high ArcFace and DINO-I scores, underscoring its utility for virtual try-on, facial privacy protection, and facial aesthetics analysis.

Searching arXiv for FFHQ-Makeup and closely related makeup-transfer papers to ground the article in current literature. FFHQ-Makeup is a high-quality synthetic facial makeup dataset designed to provide paired bare–makeup facial images with preserved identity and expression across multiple styles. Built upon the FFHQ dataset, it transfers real-world makeup styles from existing datasets onto 18,000 identities, pairs each identity with 5 different makeup styles, and yields 90,000 high-quality bare–makeup image pairs at 512×512512 \times 512 resolution. The dataset was introduced to address the scarcity of large-scale paired makeup data and to avoid the characteristic failure modes of earlier synthetic pipelines, particularly facial distortion in warping-based methods and identity or expression drift in text-to-image generation (Yang et al., 5 Aug 2025).

1. Definition and dataset profile

FFHQ-Makeup is centered on paired supervision: for a given identity, the bare image and the makeup image are intended to differ in makeup while remaining consistent in facial identity and expression. The paper describes this as a dataset in which each identity is associated with multiple makeup styles rather than a single one-to-one transformation, and it emphasizes manual curation and public release (Yang et al., 5 Aug 2025).

Attribute Value Note
Base identities 18,000 FFHQ faces diverse target set
Styles per identity 5 random makeup styles
Total scale 90,000 bare–makeup image pairs multi-style pairing
Resolution 512×512512 \times 512 high-resolution
Style sources real-world makeup images from MT and LADN source makeup corpus

The dataset paper explicitly positions FFHQ-Makeup as a resource for beauty-related tasks such as virtual try-on, facial privacy protection, and facial aesthetics analysis. Its stated novelty is that it focuses specifically on constructing a makeup dataset rather than only introducing a transfer model (Yang et al., 5 Aug 2025).

2. Construction pipeline

The construction pipeline begins with two sets: source makeup styles S\mathcal{S}, drawn from real-world makeup images, and target identities T\mathcal{T}, drawn from FFHQ. For each makeup source image ISI^\mathcal{S}, the pipeline computes a facial mask M\mathcal{M}, 2D landmarks L\mathcal{L}, and a 3D face reconstruction F\mathcal{F} using a 3D Morphable Model, specifically FLAME. A bare-face estimate I^b\hat{I}_b is reconstructed and blended with the original background, and the makeup residual is then defined as

R=IS−I^b.\mathcal{R} = I^\mathcal{S} - \hat{I}_b .

This residual is treated as the makeup-bearing component. The pipeline then performs residual augmentation via 3D re-rendering. Vertex-wise color samples from 512×512512 \times 5120 are extracted using the source geometry 512×512512 \times 5121 and re-rendered onto a different target FFHQ 3D face 512×512512 \times 5122. According to the dataset description, each source is augmented 100 times. Manual filtering removes failed cases such as segmentation errors and artifacts, leaving a final curated set of 2,257 high-fidelity makeup residuals. These residuals are then used to synthesize the paired data, with 5 random makeup styles applied to each selected FFHQ identity, followed by group-wise human quality control (Yang et al., 5 Aug 2025).

The resulting design is explicitly intended to separate makeup appearance from source-face geometry before transfer. This is the core technical mechanism behind the claim that the paired images preserve facial consistency while still exhibiting diverse makeup styles.

3. Generative mechanism and consistency preservation

The synthesis model used to generate FFHQ-Makeup is built upon Stable-Makeup and improved by FreeUV. Its architecture contains a Makeup Residual Detail Encoder, implemented with a frozen CLIP image encoder, and a Makeup Residual Learner with a FreeUV-inspired channel-attention mechanism. Structural control is provided by ControlNet, conditioned on the reconstructed bare face 512×512512 \times 5123 and facial landmarks 512×512512 \times 5124 (Yang et al., 5 Aug 2025).

In schematic form, the transfer stage is described as

512×512512 \times 5125

The dataset paper attributes facial consistency to three interacting components. First, FLAME-based 3D reconstruction is used to retain shape, expression, pose, and illumination cues. Second, the makeup residual 512×512512 \times 5126 is intended to discard intrinsic identity information from the makeup source while preserving makeup characteristics. Third, the re-rendering stage applies residuals to diverse target geometries, which the authors describe as enforcing invariance to facial geometry and expression. The paper also notes that, unlike many competing approaches, the synthesis model does not rely on paired ground-truth bare–makeup images for training, because the 3DMM-based pathway provides a self-supervised route to disentangle structure and cosmetics (Yang et al., 5 Aug 2025).

4. Empirical characteristics and reported comparisons

The empirical evaluation of FFHQ-Makeup emphasizes three properties: identity preservation, semantic consistency, and structural similarity. The reported metrics are ArcFace-based identity similarity (512×512512 \times 5127), DINO-I semantic consistency, and SSIM. In the dataset comparison table, FFHQ-Makeup is reported to achieve the highest 512×512512 \times 5128 and DINO-I scores among the compared synthetic datasets, while LADN-Syn has the highest SSIM (Yang et al., 5 Aug 2025).

Dataset Scores Note
LADN-Syn 512×512512 \times 5129 Id / DINO-I / SSIM
BeautyBank S\mathcal{S}0 Id / DINO-I / SSIM
FFHQ-Makeup S\mathcal{S}1 Id / DINO-I / SSIM

The paper further reports a GPT-4o-based visual preference study on 50 random sets, evaluating makeup realism and facial consistency. FFHQ-Makeup received the highest preference, with facial consistency reaching 92%. The ablation summary states that removing makeup residuals causes identity leakage from the source and that removing sampling and re-rendering augmentation leaves more structural artifacts and reduces facial diversity (Yang et al., 5 Aug 2025).

These results establish the dataset as a specifically engineered paired corpus rather than merely the by-product of an editing model. The evaluation is directed at dataset quality itself, not only at downstream transfer performance.

5. Naming ambiguity and relation to other FFHQ-derived resources

A persistent source of confusion is that the term “FFHQ-Makeup” is not used uniformly across the literature. In the EleGANt summary, “MT/FFHQ-Makeup” refers to the unpaired Makeup Transfer dataset inherited from BeautyGAN, comprising 1,115 non-makeup images and 2,719 makeup images, aligned and resized to S\mathcal{S}2 (Yang et al., 2022). In contrast, the 2025 FFHQ-Makeup paper defines a paired synthetic dataset built directly from FFHQ with 18,000 identities and 90,000 bare–makeup pairs (Yang et al., 5 Aug 2025).

A second nearby but distinct resource is RetouchingFFHQ, which is also FFHQ-derived but addresses retouching detection rather than paired makeup transfer. RetouchingFFHQ contains 58,158 original images and 652,568 retouched images, spans four retouching operations with four levels each, and is annotated for multi-type, multi-level retouching estimation (Ying et al., 2023). Its target problem is therefore different from FFHQ-Makeup’s paired cosmetic transfer setting.

This suggests that the label “FFHQ-Makeup” should be interpreted in the immediate bibliographic context. In some papers it denotes a specific 2025 paired synthetic dataset; in others it is used more loosely for FFHQ-derived makeup corpora or even for the older MT benchmark lineage.

6. Position in subsequent research

Subsequent papers often treat FFHQ-Makeup as an important but imperfect precursor. “FLUX-Makeup” states explicitly that it does not use FFHQ-Makeup for training or testing; instead, it uses raw FFHQ as the base for generating a new paired dataset called HQMT. Its pipeline starts from 70,000 FFHQ images with 5 prompts each, producing 350,000 initial image pairs, and after filtering retains 56,900 high-quality pairs. In the reported comparison, Stable-Makeup has a pass rate of 6.8%, unfiltered HQMT 15.4%, and filtered HQMT 96.2%, with corresponding CLIP-I, SSIM, and L2-M improvements over the earlier synthetic baselines (Zhu et al., 7 Aug 2025).

A later supervised diffusion paper likewise frames FFHQ-Makeup as a synthetic prior dataset with limitations. It describes FFHQ-Makeup as being created with LEDITS++ over FFHQ faces and notes alignment and style artifacts, then proposes a curated dataset of approximately 200,000 paired samples at S\mathcal{S}3 resolution using a train–generate–filter–retrain strategy (Pan et al., 31 Jan 2026). Taken together, these later works present high-quality paired makeup data as a central bottleneck in supervised makeup transfer and use FFHQ-Makeup as a comparison point for newer curation strategies.

Even where FFHQ-Makeup is not used directly, it remains methodologically influential. Later region-aware and controllable diffusion systems continue to rely on FFHQ-based synthetic generation, paired supervision, and explicit attempts to disentangle identity from makeup, all of which are central design themes of FFHQ-Makeup and its accompanying synthesis pipeline (Gao et al., 20 Mar 2026).

7. Research significance and applications

The FFHQ-Makeup paper presents paired bare–makeup facial images as essential for virtual try-on, facial privacy protection, and facial aesthetics analysis. Because the dataset provides multiple makeup styles per identity while aiming to keep identity and expression fixed, it is also relevant to makeup style transfer, makeup removal, robust face recognition under makeup, and controllable facial attribute manipulation (Yang et al., 5 Aug 2025).

Its main significance lies in the combination of scale, explicit pairing, and facial consistency. Earlier synthetic strategies were described as either distorting geometry through warping or drifting identity through text-to-image generation. FFHQ-Makeup addresses those issues through 3DMM reconstruction, residual extraction, residual augmentation, and diffusion-based synthesis under structural control. Later work that surpasses it in scale or filtering quality does not negate that role; rather, it indicates that FFHQ-Makeup helped crystallize paired synthetic makeup data as a distinct research object and a measurable component of system performance.

In that sense, FFHQ-Makeup occupies an intermediate position in the evolution of makeup-transfer datasets: more structured and consistency-oriented than earlier weakly paired or unpaired resources, yet also a clear antecedent to later large-scale curated datasets such as HQMT and other FFHQ-derived supervised corpora.

Topic to Video (Beta)

No one has generated a video about this topic yet.

Whiteboard

No one has generated a whiteboard explanation for this topic yet.

Follow Topic

Get notified by email when new papers are published related to FFHQ-Makeup.