Papers
Topics
Authors
Recent
Search
2000 character limit reached

Prediction-Driven Top-K Jaccard Coefficient

Updated 3 January 2026
  • The paper introduces a dynamic refinement of Jaccard similarity by predicting optimal Top-K neighbors using neural models and a Sparse Differential Transformer.
  • It employs a supervised prediction task to adaptively determine neighborhood sizes, which significantly enhances robustness and discriminative power in clustering.
  • Empirical results on large-scale datasets demonstrate improved face clustering performance and generalization across multiple domains.

The prediction-driven Top-K Jaccard similarity coefficient is a dynamic, data-adaptive refinement of the traditional Jaccard approach for measuring pairwise relationships in face clustering graphs. Central to this methodology is the replacement of a static, globally fixed neighbor count with an individually predicted optimal Top-K for each node, determined through a supervised prediction task powered by neural models and further stabilized via a Sparse Differential Transformer (SDT). The resulting framework achieves increased discriminative power, improved robustness to noise, and state-of-the-art clustering performance on multiple large-scale datasets (Zhang et al., 27 Dec 2025).

1. Mathematical Foundations of the Top-K Jaccard Similarity

The classical Jaccard similarity for two sets AA and BB is given by:

J(A,B)=∣A∩B∣∣A∪B∣.J(A, B) = \frac{|A \cap B|}{|A \cup B|}.

In the context of face clustering, the sets NiN_i and NjN_j typically correspond to the K nearest neighbors of nodes ii and jj in the embedding space, measured, for instance, by cosine similarity:

J(Ni,Nj)=∣Ni∩Nj∣∣Ni∪Nj∣.J(N_i, N_j) = \frac{|N_i \cap N_j|}{|N_i \cup N_j|}.

In the prediction-driven Top-K extension, the neighbor count KK for each node is not fixed but predicted. If Top-Ki\mathrm{Top}\text{-}K_i is the predicted number for node BB0 (rounded to BB1), then

BB2

BB3

with the intersection BB4.

The prediction-driven Top-K Jaccard edge probability is then:

BB5

where BB6 is a normalized pairwise similarity, obtained via distance-to-probability transformation:

BB7

BB8

BB9

This approach increases the reliability of similarity measurements by focusing on a purified, node-specific neighborhood (Zhang et al., 27 Dec 2025).

2. Data-Driven Prediction of Optimal Top-K

Rather than applying a fixed neighbor threshold, the optimal neighborhood size for each node is formalized as a supervised prediction problem. For node J(A,B)=∣A∩B∣∣A∪B∣.J(A, B) = \frac{|A \cap B|}{|A \cup B|}.0, the model considers the top-J(A,B)=∣A∩B∣∣A∪B∣.J(A, B) = \frac{|A \cap B|}{|A \cup B|}.1 candidate neighbors (ranked by cosine similarity) and predicts a score J(A,B)=∣A∩B∣∣A∪B∣.J(A, B) = \frac{|A \cap B|}{|A \cup B|}.2 for each candidate J(A,B)=∣A∩B∣∣A∪B∣.J(A, B) = \frac{|A \cap B|}{|A \cup B|}.3, approximating the likelihood that J(A,B)=∣A∩B∣∣A∪B∣.J(A, B) = \frac{|A \cap B|}{|A \cup B|}.4 shares the same identity as J(A,B)=∣A∩B∣∣A∪B∣.J(A, B) = \frac{|A \cap B|}{|A \cup B|}.5.

During inference, the predicted Top-K is obtained by thresholding:

J(A,B)=∣A∩B∣∣A∪B∣.J(A, B) = \frac{|A \cap B|}{|A \cup B|}.6

where J(A,B)=∣A∩B∣∣A∪B∣.J(A, B) = \frac{|A \cap B|}{|A \cup B|}.7 is typically set to 0.90.

The input to the predictor consists of the embedding sequence J(A,B)=∣A∩B∣∣A∪B∣.J(A, B) = \frac{|A \cap B|}{|A \cup B|}.8 and corresponding positional or relative distance encodings. These are projected into queries, keys, and values for Transformer-based processing.

Training employs a binary cross-entropy loss:

J(A,B)=∣A∩B∣∣A∪B∣.J(A, B) = \frac{|A \cap B|}{|A \cup B|}.9

with NiN_i0 denoting the ground-truth same-identity indicator. Training samples primarily target a window around the true decision boundary to maintain class balance (Zhang et al., 27 Dec 2025).

3. Transformer-Based Predictors and the Sparse Differential Transformer (SDT)

The initial prediction framework employs a vanilla Transformer architecture (three layers, eight self-attention heads, hidden size 1024) with standard self-attention. However, the vanilla Transformer tends to overvalue irrelevant relationships, introducing noise into the neighborhood prediction.

Differential Self-Attention

The attention computation is refined by splitting queries and keys, and applying a subtractive update:

NiN_i1

where NiN_i2 denotes softmax and NiN_i3 is learnable.

Sparse Differential Transformer (SDT)

Sparsity is imposed by masking all but the top-K entries per row:

NiN_i4

with NiN_i5 denoting the Top-K mask. This structure masks weak key–query pairs, focusing the model on the most informative local relationships and improving resilience to noise.

Mixture-of-Experts SDT

MoE-SDT introduces a mixture over masks NiN_i6 (with NiN_i7) and combines their outputs using learnable weights NiN_i8:

NiN_i9

This construct further mitigates prediction errors near the Top-K boundary (Zhang et al., 27 Dec 2025).

4. Integration Into Clustering Workflow

The integration of the prediction-driven Top-K Jaccard coefficient into face clustering follows these algorithmic steps:

  1. Extract face embeddings NjN_j0.
  2. For each NjN_j1, determine top-NjN_j2 candidates via cosine similarity.
  3. Process NjN_j3 using the SDT predictor, yielding scores NjN_j4.
  4. Compute NjN_j5.
  5. Round to NjN_j6 (e.g., nearest multiple of 10 up to NjN_j7).
  6. Build refined neighborhoods NjN_j8 and compute NjN_j9.
  7. Threshold ii0 to construct a sparse adjacency matrix.
  8. Apply the Map-Equation codec for community detection.

Pseudocode for the procedure is:

jj1 (Zhang et al., 27 Dec 2025)

5. Empirical Results and Robustness

Evaluation on several benchmarks demonstrates the efficacy of the prediction-driven Top-K Jaccard approach.

MS-Celeb-1M Clustering:

Method ii1 (584K) ii2 (584K) ii3 (5.21M) ii4 (5.21M)
FC-ESER 95.28 93.85 89.40 88.80
Diff-Cluster 95.46 94.14 90.08 89.14

Sigmoid distance-transform yields +0.18% ii5, +0.16% ii6 versus exponential; adding Top-K filtering yields +0.35% ii7, +0.46% ii8.

Transformer Variant Ablations:

Architecture ii9 jj0
Vanilla Transformer 94.25 92.73
Vanilla+Top-K Mask 94.78 93.25
Differential (no mask) 95.05 93.59
Differential+Top-K (SDT) 95.34 93.93
MoE-SDT 95.46 94.14

Noisy similarity matrix experiments simulate up to 40% noise: SDT maintains pairwise F-score, while vanilla Transformer degrades significantly.

Generalization evaluations confirm improvements across MSMT17 (person re-ID) and DeepFashion, as well as gains from substituting SDT into contemporary clustering pipelines (Zhang et al., 27 Dec 2025).

6. Impact and Significance

The prediction-driven Top-K Jaccard similarity coefficient constitutes an advancement in clustering methodology by dynamically adapting local connectivity based on learned data relationships rather than static heuristics. By defining neighborhood boundaries as a supervised prediction problem and employing SDT for robust selection, the approach produces neighborhood sets with enhanced purity, resulting in more reliable graph-based similarity estimation.

Empirical evidence establishes increases in overall clustering accuracy, robustness to random noise in similarity calculations, and improved performance generalization into domains beyond face clustering. A plausible implication is that prediction-driven Top-K filtering can be fruitfully adapted to related graph construction problems in metric learning and instance-level retrieval beyond the specific setting investigated.

In summary, this method advances the state of the art in graph-based face clustering with potential for transferability across adjacent domains (Zhang et al., 27 Dec 2025).

Definition Search Book Streamline Icon: https://streamlinehq.com
References (1)

Topic to Video (Beta)

No one has generated a video about this topic yet.

Whiteboard

No one has generated a whiteboard explanation for this topic yet.

Follow Topic

Get notified by email when new papers are published related to Prediction-Driven Top-K Jaccard Similarity Coefficient.