Prediction-Driven Top-K Jaccard Coefficient
- The paper introduces a dynamic refinement of Jaccard similarity by predicting optimal Top-K neighbors using neural models and a Sparse Differential Transformer.
- It employs a supervised prediction task to adaptively determine neighborhood sizes, which significantly enhances robustness and discriminative power in clustering.
- Empirical results on large-scale datasets demonstrate improved face clustering performance and generalization across multiple domains.
The prediction-driven Top-K Jaccard similarity coefficient is a dynamic, data-adaptive refinement of the traditional Jaccard approach for measuring pairwise relationships in face clustering graphs. Central to this methodology is the replacement of a static, globally fixed neighbor count with an individually predicted optimal Top-K for each node, determined through a supervised prediction task powered by neural models and further stabilized via a Sparse Differential Transformer (SDT). The resulting framework achieves increased discriminative power, improved robustness to noise, and state-of-the-art clustering performance on multiple large-scale datasets (Zhang et al., 27 Dec 2025).
1. Mathematical Foundations of the Top-K Jaccard Similarity
The classical Jaccard similarity for two sets and is given by:
In the context of face clustering, the sets and typically correspond to the K nearest neighbors of nodes and in the embedding space, measured, for instance, by cosine similarity:
In the prediction-driven Top-K extension, the neighbor count for each node is not fixed but predicted. If is the predicted number for node 0 (rounded to 1), then
2
3
with the intersection 4.
The prediction-driven Top-K Jaccard edge probability is then:
5
where 6 is a normalized pairwise similarity, obtained via distance-to-probability transformation:
7
8
9
This approach increases the reliability of similarity measurements by focusing on a purified, node-specific neighborhood (Zhang et al., 27 Dec 2025).
2. Data-Driven Prediction of Optimal Top-K
Rather than applying a fixed neighbor threshold, the optimal neighborhood size for each node is formalized as a supervised prediction problem. For node 0, the model considers the top-1 candidate neighbors (ranked by cosine similarity) and predicts a score 2 for each candidate 3, approximating the likelihood that 4 shares the same identity as 5.
During inference, the predicted Top-K is obtained by thresholding:
6
where 7 is typically set to 0.90.
The input to the predictor consists of the embedding sequence 8 and corresponding positional or relative distance encodings. These are projected into queries, keys, and values for Transformer-based processing.
Training employs a binary cross-entropy loss:
9
with 0 denoting the ground-truth same-identity indicator. Training samples primarily target a window around the true decision boundary to maintain class balance (Zhang et al., 27 Dec 2025).
3. Transformer-Based Predictors and the Sparse Differential Transformer (SDT)
The initial prediction framework employs a vanilla Transformer architecture (three layers, eight self-attention heads, hidden size 1024) with standard self-attention. However, the vanilla Transformer tends to overvalue irrelevant relationships, introducing noise into the neighborhood prediction.
Differential Self-Attention
The attention computation is refined by splitting queries and keys, and applying a subtractive update:
1
where 2 denotes softmax and 3 is learnable.
Sparse Differential Transformer (SDT)
Sparsity is imposed by masking all but the top-K entries per row:
4
with 5 denoting the Top-K mask. This structure masks weak key–query pairs, focusing the model on the most informative local relationships and improving resilience to noise.
Mixture-of-Experts SDT
MoE-SDT introduces a mixture over masks 6 (with 7) and combines their outputs using learnable weights 8:
9
This construct further mitigates prediction errors near the Top-K boundary (Zhang et al., 27 Dec 2025).
4. Integration Into Clustering Workflow
The integration of the prediction-driven Top-K Jaccard coefficient into face clustering follows these algorithmic steps:
- Extract face embeddings 0.
- For each 1, determine top-2 candidates via cosine similarity.
- Process 3 using the SDT predictor, yielding scores 4.
- Compute 5.
- Round to 6 (e.g., nearest multiple of 10 up to 7).
- Build refined neighborhoods 8 and compute 9.
- Threshold 0 to construct a sparse adjacency matrix.
- Apply the Map-Equation codec for community detection.
Pseudocode for the procedure is:
5. Empirical Results and Robustness
Evaluation on several benchmarks demonstrates the efficacy of the prediction-driven Top-K Jaccard approach.
MS-Celeb-1M Clustering:
| Method | 1 (584K) | 2 (584K) | 3 (5.21M) | 4 (5.21M) |
|---|---|---|---|---|
| FC-ESER | 95.28 | 93.85 | 89.40 | 88.80 |
| Diff-Cluster | 95.46 | 94.14 | 90.08 | 89.14 |
Sigmoid distance-transform yields +0.18% 5, +0.16% 6 versus exponential; adding Top-K filtering yields +0.35% 7, +0.46% 8.
Transformer Variant Ablations:
| Architecture | 9 | 0 |
|---|---|---|
| Vanilla Transformer | 94.25 | 92.73 |
| Vanilla+Top-K Mask | 94.78 | 93.25 |
| Differential (no mask) | 95.05 | 93.59 |
| Differential+Top-K (SDT) | 95.34 | 93.93 |
| MoE-SDT | 95.46 | 94.14 |
Noisy similarity matrix experiments simulate up to 40% noise: SDT maintains pairwise F-score, while vanilla Transformer degrades significantly.
Generalization evaluations confirm improvements across MSMT17 (person re-ID) and DeepFashion, as well as gains from substituting SDT into contemporary clustering pipelines (Zhang et al., 27 Dec 2025).
6. Impact and Significance
The prediction-driven Top-K Jaccard similarity coefficient constitutes an advancement in clustering methodology by dynamically adapting local connectivity based on learned data relationships rather than static heuristics. By defining neighborhood boundaries as a supervised prediction problem and employing SDT for robust selection, the approach produces neighborhood sets with enhanced purity, resulting in more reliable graph-based similarity estimation.
Empirical evidence establishes increases in overall clustering accuracy, robustness to random noise in similarity calculations, and improved performance generalization into domains beyond face clustering. A plausible implication is that prediction-driven Top-K filtering can be fruitfully adapted to related graph construction problems in metric learning and instance-level retrieval beyond the specific setting investigated.
In summary, this method advances the state of the art in graph-based face clustering with potential for transferability across adjacent domains (Zhang et al., 27 Dec 2025).