---
title: 'FaceGCD: Dynamic Face Identity Discovery'
url: https://www.emergentmind.com/topics/facegcd
type: topic
---

# FaceGCD: Dynamic Face Identity Discovery

FaceGCD is a face-specific generalized category discovery method introduced together with the task of **generalized face discovery (GFD)**, an open-world face recognition setting in which a model must recognize labeled known identities, exploit unlabeled samples from known identities, and discover previously unseen identities as distinct clusters [2507.22353]. The central claim is that face identity discovery is substantially harder than generic GCD because face datasets exhibit high cardinality, fine-grained discrimination, low inter-class margin, subtle identity-specific cues, and uneven sample distributions. To address this, FaceGCD uses a ViT-based architecture with **dynamic, instance-specific layer-wise prefixes** generated by a **HyperNetwork**, so that the feature extractor is modulated separately for each input image rather than relying on a static prompt or a single fixed feature encoder [2507.22353].

## 1. Task definition and relation to prior settings

FaceGCD is built around **generalized face discovery**, defined as a mixed setting with labeled and unlabeled data over both known and novel identities [2507.22353]. The model must simultaneously perform three functions: recognize known identities using labeled data, exploit unlabeled samples from those same identities, and discover entirely new identities from unlabeled or test data by clustering them into distinct unknown IDs.

This task differs from several adjacent formulations. Traditional face identification is a **closed-set** problem in which all test identities are assumed known during training, so no identity discovery is required [2507.22353]. Open-set or open-world face recognition allows unknown identities at test time, but typically treats them as a single rejection category or as outliers, rather than partitioning them into multiple newly discovered identities [2507.22353]. Generalized category discovery provides the closest conceptual analogue, but FaceGCD argues that direct transfer of generic GCD methods is ineffective because face IDs form a highly fine-grained and high-cardinality label space [2507.22353].

The method is therefore positioned as a face-specific response to the limitations of existing GCD systems on identity-centric data. A plausible implication is that the paper treats the representational demands of face recognition as qualitatively different from ordinary category discovery, even when the formal semi-supervised discovery structure is similar.

## 2. Formal problem setup and benchmark construction

The paper defines the labeled set as
$$
\mathcal{D}_{\mathcal{L}}=\{(x_i^L, y_i^L)\}_{i=1}^{N_L},
$$
with labels in the set of known identities $\mathcal{C}_{\mathcal{L}}$, and the unlabeled set as
$$
\mathcal{D}_{\mathcal{U}}=\{x_j^U\}_{j=1}^{N_U},
$$
where unlabeled samples may belong either to known identities in $\mathcal{C}_{\mathcal{L}}$ or to novel identities in $\mathcal{C}_{\mathcal{N}}$ [2507.22353]. Training therefore contains three semantic groups: labeled known identities, unlabeled known identities, and unlabeled novel identities.

Benchmark construction is based on **YouTube Faces (YTF)** and **CASIA-WebFaces** [2507.22353]. Subsets contain **500, 1000, or 2000 identities**, split **50/50 into known and unknown IDs**. For each known identity, **half of its training images are labeled** and half are treated as unlabeled, while unknown identities are **entirely unlabeled** during training. The train/test split per identity is **90%/10%** [2507.22353].

The six reported benchmarks are summarized below.

| Benchmark | Known/unknown split | Train / test size |
|---|---:|---:|
| YTF 500 | 250 / 250 | 48,089 / 11,779 |
| YTF 1000 | 500 / 500 | 96,002 / 23,523 |
| YTF 2000 | 1000 / 1000 | 190,248 / 46,615 |
| CASIA 500 | 250 / 250 | 46,991 / 11,999 |
| CASIA 1000 | 500 / 500 | 89,508 / 22,867 |
| CASIA 2000 | 1000 / 1000 | 184,432 / 47,114 |

The evaluation metric is **clustering accuracy** with **Hungarian assignment**, reported separately for **Known**, **Novel**, and **All** identities [2507.22353]. The paper also reports **Nearest-Neighbor Consistency (NNC)**,
$$
\text{NNC} = \frac{1}{N} \sum_{i=1}^{N} \frac{1}{k} \sum_{j \in \mathcal{N}_k(i)} \mathmybb{1}[\hat{y}_j = \hat{y}_i].
$$
An important assumption is that, for clustering, the number of clusters is set to match the ground-truth number of classes, i.e. $|\mathcal{C}_{\mathcal{L}}| + |\mathcal{C}_{\mathcal{N}}|$ [2507.22353].

## 3. Architecture: dynamic prefix generation with a HyperNetwork

FaceGCD uses a **pretrained ViT-B/16** backbone pretrained with **DINO** on **MS1MV3**, with patch size set to 8 for facial patch extraction and the **[CLS]** token used for downstream representation [2507.22353]. The method contains two transformer roles: a **static frozen feature extractor** used to produce conditioning features for the HyperNetwork, and the **main ViT backbone** that receives dynamically generated prefixes [2507.22353].

A pretrained **landmark CNN** based on **MobileNetV3**, following Part-FViT, detects facial landmarks and supports landmark-guided facial patch extraction [2507.22353]. This preprocessing is intended to make the tokenization more face-structure-aware.

The defining mechanism is the generation of **image-conditioned, layer-wise key/value prefixes**. For each transformer layer $L$, the frozen static ViT produces conditioning features $\mathbf{z}^{(L)}$, and a lightweight **2-layer MLP HyperNetwork** $\mathcal{H}(\cdot)$ maps those features to layer-specific parameters $\Phi_L$ that instantiate an input-dependent prefix generator [2507.22353]. The paper states the core generation relation as
$$
\mathbf{P}^{(L)}_{K}(x), \mathbf{P}^{(L)}_{V}(x) = G_x(\mathbf{P}_{\text{init}}; \mathcal{H}(\mathbf{z}^{(L)})),
$$
where $\mathbf{P}_{\text{init}}$ is a randomly initialized prefix.

The appendix formalizes the generated prefixes as
$$
\mathbf{P}^{(L)}_{K} = \mathcal{H}^{(L)}_{K}(\mathbf{P}_\tau) = \mathbf{U}^{(L)}_{K}(\text{ReLU}(\mathbf{D}^{(L)}_{K}(\mathbf{P}_\tau))),
$$
$$
\mathbf{P}^{(L)}_{V} = \mathcal{H}^{(L)}_{V} (\mathbf{P}_\tau) = \mathbf{U}^{(L)}_{V}(\text{ReLU}(\mathbf{D}^{(L)}_{V}(\mathbf{P}_\tau))),
$$
where $\mathbf{P}_\tau$ has shape $m \times h \times d$, $\mathbf{D}^{(L)}_{K/V} \in \mathbb{R}^{d \times b}$, $\mathbf{U}^{(L)}_{K/V} \in \mathbb{R}^{b \times h \times d}$, and $b \ll d$ [2507.22353]. The reported implementation implies 12 attention heads, 64 per-head dimensions, and a bottleneck dimension of 16 [2507.22353].

These prefixes are injected into self-attention by concatenating them to the key and value sequences while leaving queries unchanged:
$$
\tilde{K}^{(L)} = [\mathbf{P}^{(L)}_{K}; K^{(L)}], \quad \tilde{V}^{(L)} = [\mathbf{P}^{(L)}_{V}; V^{(L)}], \quad \tilde{Q}^{(L)} = Q^{(L)}.
$$
This design follows prompt-learning intuition in which the model adapts the context presented to attention without altering the original query stream [2507.22353].

The architecture therefore differs from both a **static prefix pool** and a **static prefix generator**. The former uses a fixed bank of prompts, and the latter uses one generator for all inputs; FaceGCD instead generates a different set of prefixes for each image and for each layer [2507.22353]. The paper’s main conceptual claim is that such per-instance modulation is better suited to the subtle identity differences that govern face discovery.

## 4. Training objective, optimization, and inference

FaceGCD is trained with a semi-supervised contrastive objective rather than with ArcFace loss [2507.22353]. The batch loss is
$$
\mathcal{L}_{B} = (1-\tau) \sum_{i\in B}\mathcal{L}^{u}_{i} + \lambda \sum_{i\in B \cap \mathcal{D}_{\mathcal{L}}}\mathcal{L}^{s}_{i},
$$
where $\mathcal{L}^{u}_{i}$ is an unsupervised contrastive loss, $\mathcal{L}^{s}_{i}$ is a supervised contrastive loss for labeled samples, $\tau$ is the temperature scaling hyperparameter, and $\lambda$ is the supervised loss weight [2507.22353].

The unsupervised term is
$$
\mathcal{L}^{u}_{i} = -\log \frac{\exp(\mathbf{z}_{i} \cdot \mathbf{z}'_{i}/\tau)} {\sum_{n} \mathmybb{1}_{[n \neq i]}\,\exp(\mathbf{z}_{i} \cdot \mathbf{z}_{n}/\tau)},
$$
and the supervised contrastive term is
$$
\mathcal{L}^{s}_{i} = - \frac{1}{\mathcal{N}(i)} \sum_{q\in\mathcal{N}(i)} \log \frac{\text{exp}(\mathbf{z}_{i} \cdot \mathbf{z}_{q}/\tau)} {\sum_{n} \mathmybb{1}_{[n \neq i]}\,\exp(\mathbf{z}_{i} \cdot \mathbf{z}_{n}\,/\,\tau)},
$$
where $\mathcal{N}(i)$ denotes samples in the batch sharing the same identity as $i$ [2507.22353]. All samples participate in the unsupervised contrastive term, while only labeled samples contribute to the supervised term [2507.22353]. The paper emphasizes that it does not introduce additional pseudo-label, entropy minimization, or self-distillation losses, so the architectural contribution is evaluated under the same general style of objective used in GCD [2507.22353].

During fine-tuning, FaceGCD trains only the **HyperNetwork**, the **final layer of the backbone ViT**, and the **DINO head** $H(\cdot)$, while most of the backbone ViT, the landmark CNN, and the static conditioning extractor remain frozen [2507.22353]. Pretraining is conducted for **50 epochs** using default hyperparameters of original DINO and ArcFace implementations, and GFD fine-tuning runs for **200 epochs** with **batch size 128**, **SGD**, **momentum 0.9**, **weight decay $2\times 10^{-5}$**, **base learning rate 0.1**, **warm-up learning rate $1\times 10^{-5}$**, **warm-up for the first 5 epochs**, and **cosine learning rate decay** [2507.22353]. The reported temperature is $\tau = 1.0$ and the supervised loss weight is $\lambda = 0.35$ [2507.22353].

Inference follows the same dynamic-prefix pipeline: landmark detection, frozen static ViT feature extraction, HyperNetwork-conditioned prefix generation, main ViT forward pass, and final embedding extraction [2507.22353]. Final assignments are then produced by clustering the embeddings with **semi-supervised k-means (SSK)** from GCD [2507.22353]. The paper explicitly states that prediction is not based on a classifier-plus-rejector architecture; instead, representations are clustered jointly and evaluated under a single Hungarian matching over all classes [2507.22353].

## 5. Empirical results and ablation evidence

FaceGCD reports state-of-the-art results on all six GFD benchmarks relative to the listed GCD baselines **GCD**, **SimGCD**, **PromptCAL**, and **CMS** [2507.22353]. On **YTF 500**, it achieves **81.2 / 93.8 / 68.6** for Known / Novel / All, improving the All score by **+6.0** over PromptCAL, **+7.1** over SimGCD, and **+10.7** over GCD [2507.22353]. On **YTF 1000**, it achieves **82.7 / 93.2 / 72.1**, with gains of **+7.7** All over PromptCAL, **+9.9** All over SimGCD, and **+12.3** All over GCD; the Novel gain is **+8.3** over PromptCAL and **+22.0** over SimGCD [2507.22353]. On **YTF 2000**, the result is **83.6 / 91.8 / 75.4**, including **+11.0** All over SimGCD, **+13.3** Novel over GCD, and **+18.7** Novel over PromptCAL [2507.22353].

On CASIA-based benchmarks, the method also leads consistently: **58.8 / 68.9 / 48.7** on **CASIA 500**, **52.2 / 52.3 / 52.2** on **CASIA 1000**, and **56.1 / 60.4 / 51.7** on **CASIA 2000** [2507.22353]. The most stable pattern across these tables is that FaceGCD wins on every benchmark in the **All** column, with particularly strong gains in **Novel** accuracy [2507.22353].

The method also outperforms a strong face recognition baseline under multiple clustering backends. On **YTF 1000** with **SSK**, FaceGCD achieves **82.7 / 93.2 / 72.1**, compared with **73.0 / 78.3 / 67.6** for ArcFace and **72.0 / 87.9 / 56.1** for ArcFace + GCD [2507.22353]. Its advantage persists under **K-Means**, **HAC**, and **DBSCAN**, which the paper interprets as evidence that the learned embeddings themselves are more discoverable rather than merely being better matched to one clustering algorithm [2507.22353].

Embedding quality is further supported by **NNC** on **YTF1000**: FaceGCD obtains **ACC 82.7, NNC 91.1**, compared with **ACC 70.4, NNC 83.6** for GCD and **ACC 73.0, NNC 82.7** for ArcFace [2507.22353]. This indicates improvements in both global clustering accuracy and local neighborhood consistency.

Ablation studies isolate the role of dynamic generation. A **Static Prefix Pool** yields **42.6 / 52.8 / 32.3** on YTF 1000, a major degradation relative to FaceGCD [2507.22353]. A **Static Prefix Generator** yields **76.9 / 81.7 / 72.1**, but requires **399.9M additional params**, **427.2M trainable params**, and **593.8M total params**, whereas FaceGCD uses **13.8M additional params (6.6%)**, **41.1M trainable params (19.8%)**, and **207.7M total params** [2507.22353]. The paper therefore argues that input-conditioned dynamic generation is both more accurate and more parameter-efficient than increasing static capacity.

Prefix-length ablations on YTF 1000 show **80.9 / 91.1 / 70.4** for prefix size 5, **82.9 / 92.9 / 72.8** for prefix size 10, and **82.7 / 93.2 / 72.1** for prefix size 20 [2507.22353]. The reported pattern is that prefix size 5 is inadequate, while performance saturates for moderate prefix sizes [2507.22353].

## 6. Interpretation, limitations, and broader context

The central contribution of FaceGCD is the claim that **open-world face recognition should be treated as adaptive per-instance representation learning** rather than as direct transfer of generic GCD or closed-set face identification [2507.22353]. In this view, faces are not merely another fine-grained category set; they require feature extractors that can adapt to subtle identity-specific cues on a sample-by-sample basis.

The paper also evaluates FaceGCD on generic GCD benchmarks including **CIFAR100**, **ImageNet100**, **CUB**, **Stanford Cars**, **FGVC Aircraft**, and **Herbarium19**, where it is competitive or state of the art on several fine-grained datasets, including **CUB: 64.5 All**, **Stanford Cars: 62.5 All, 67.6 Novel**, and **Herbarium19: 42.6 All, 40.8 Novel** [2507.22353]. This suggests that the mechanism is not strictly face-bound, although the paper emphasizes that its primary motivation is the particular difficulty of face identity discovery.

Several limitations are explicit or directly inferable from the reported setup. The clustering stage requires the **ground-truth number of classes** during evaluation, which weakens real-world open-world realism [2507.22353]. The final prediction protocol depends on **offline clustering** rather than online identity creation in a streaming environment [2507.22353]. The system also relies on substantial pretrained components, including a DINO-pretrained ViT on **MS1MV3** and a pretrained landmark detector [2507.22353]. In addition, the method is architecturally complex, involving a static extractor, a main backbone, hypernetwork-conditioned prefix generation, and facial landmark preprocessing [2507.22353].

The paper notes reproducibility ambiguities as well. The most salient is a **prefix size inconsistency**: the appendix says **prefix size = 20** in all experiments, while the ablation table reports the best YTF1000 result for **prefix size 10** [2507.22353]. The provided text also omits detailed augmentation recipes, even though the contrastive setup depends on augmented views $\mathbf{z}'_i$ [2507.22353]. These issues do not alter the methodological claim but do affect exact reproduction.

Within the broader literature, FaceGCD stands as a face-specific extension of generalized category discovery rather than a conventional face identification system. The related paper **"SGF-CDNet: A Consistency-Discrepancy Graph Network over Semantic-Geometric Fused Nodes for Face Forgery Detection"** describes its own approach as being in the same broad family as “FaceGCD-style methods” in the sense of prioritizing structural, component-level reasoning over manipulation-specific artifacts, though SGF-CDNet addresses face forgery detection rather than identity discovery [2607.03883]. This suggests that the label “FaceGCD” has begun to function, at least informally, as a reference point for structure-aware and generalization-oriented face analysis. A plausible implication is that FaceGCD’s longer-term significance lies not only in the specific GFD benchmark results, but also in establishing a face-domain template for open-world representation learning based on input-conditioned modulation [2507.22353].

Source: https://www.emergentmind.com/topics/facegcd