Papers
Topics
Authors
Recent
Search
2000 character limit reached

Anchor-based Fair Clustering Framework

Updated 20 November 2025
  • The paper presents AFCF—a scalable algorithm that ensures exact per-cluster fairness by matching demographic proportions via novel anchor selection and constrained optimization.
  • It employs the FDAS mechanism to select representative anchors that maintain both spatial coverage and group balance, significantly reducing computational overhead.
  • The framework utilizes an ADMM-based solver to efficiently handle fairness-preserving label propagation, achieving linear scalability on large datasets.

The Anchor-based Fair Clustering Framework (AFCF) enables linear-time scalable fair clustering on large datasets, rigorously preserving demographic group fairness properties while drastically accelerating existing fair clustering algorithms. AFCF integrates novel fair sampling for anchor selection, a fairness-preserving label propagation mechanism grounded in constrained optimization, and an efficient ADMM solver, demonstrating consistent empirical efficacy across large benchmark datasets (Wei et al., 13 Nov 2025).

1. Fair Anchor Selection: FDAS Mechanism

AFCF achieves both spatial and demographic representativeness of a subset of anchors through the Fair Directly Alternate Sampling (FDAS) algorithm. Given a dataset X∈Rd×n\mathbf{X}\in\mathbb{R}^{d\times n}, a partition of the data into tt protected groups G={G1,…,Gt}\mathcal{G}=\{G_1, \dots, G_t\}, and group proportions ρr=∣Gr∣/n\rho_r=|G_r|/n, FDAS selects m≪nm\ll n anchors according to the following two-phase procedure:

(A) Quota Computation. Each group receives a quota qr=⌊m⋅ρr⌋q_r = \lfloor m \cdot \rho_r \rfloor, with the remainder Δ=m−∑r=1tqr\Delta = m - \sum_{r=1}^t q_r allocated iteratively to those groups underrepresented relative to mρrm\rho_r. This guarantees ∑rqr=m\sum_r q_r = m and for all rr, tt0.

(B) Within-Group Spatial Coverage. For each group tt1, points are scored via tt2, normalized to tt3. Iteratively, the highest scoring point within group tt4 is selected as an anchor. After each selection, scores are decayed as tt5 to promote spatial dispersion within the group. The process continues until tt6 anchors are chosen from every group.

The FDAS approach ensures that the selected anchors reflect both the global group proportions and spatial distribution, with computational complexity tt7, where tt8 is ambient dimensionality and tt9 (Wei et al., 13 Nov 2025).

2. Anchor Graph Construction and Fairness-Preserving Label Propagation

Post anchor selection, any fair clustering algorithm G={G1,…,Gt}\mathcal{G}=\{G_1, \dots, G_t\}0 is applied to the G={G1,…,Gt}\mathcal{G}=\{G_1, \dots, G_t\}1-anchor set, yielding cluster labels G={G1,…,Gt}\mathcal{G}=\{G_1, \dots, G_t\}2. The challenge is then to transfer these cluster assignments, preserving fairness, to the full dataset. This is mediated by constructing an G={G1,…,Gt}\mathcal{G}=\{G_1, \dots, G_t\}3 nonnegative affinity matrix G={G1,…,Gt}\mathcal{G}=\{G_1, \dots, G_t\}4 so that cluster label propagation maintains demographic parity.

The propagation problem is formalized as: G={G1,…,Gt}\mathcal{G}=\{G_1, \dots, G_t\}5 subject to

  • G={G1,…,Gt}\mathcal{G}=\{G_1, \dots, G_t\}6 (the G={G1,…,Gt}\mathcal{G}=\{G_1, \dots, G_t\}7-simplex for each G={G1,…,Gt}\mathcal{G}=\{G_1, \dots, G_t\}8),
  • For each cluster G={G1,…,Gt}\mathcal{G}=\{G_1, \dots, G_t\}9 and group ρr=∣Gr∣/n\rho_r=|G_r|/n0,

ρr=∣Gr∣/n\rho_r=|G_r|/n1

where ρr=∣Gr∣/n\rho_r=|G_r|/n2 is the anchor feature matrix, ρr=∣Gr∣/n\rho_r=|G_r|/n3 indexes anchors from cluster ρr=∣Gr∣/n\rho_r=|G_r|/n4, and ρr=∣Gr∣/n\rho_r=|G_r|/n5 indexes data for group ρr=∣Gr∣/n\rho_r=|G_r|/n6.

Fairness Preservation: The joint group-label constraint enforces that the final per-cluster group proportions on all ρr=∣Gr∣/n\rho_r=|G_r|/n7 data points match exactly those observed among the anchor assignments: ρr=∣Gr∣/n\rho_r=|G_r|/n8 where ρr=∣Gr∣/n\rho_r=|G_r|/n9 and balance is defined as m≪nm\ll n0.

Label propagation computes final soft assignments m≪nm\ll n1 (with m≪nm\ll n2 the anchor cluster one-hot matrix), and hard cluster labels by m≪nm\ll n3 (Wei et al., 13 Nov 2025).

3. ADMM-Based Optimization

To efficiently solve the constrained quadratic problem, AFCF employs an Alternating Direction Method of Multipliers (ADMM) framework. Introducing slack variable m≪nm\ll n4 and dual variable m≪nm\ll n5, the augmented Lagrangian is

m≪nm\ll n6

Iterative updates alternately minimize for m≪nm\ll n7 (simplex-constrained QPs), update m≪nm\ll n8 (closed form within each block to enforce the fairness constraint), and perform dual ascent on m≪nm\ll n9. Each ADMM iteration costs qr=⌊m⋅ρr⌋q_r = \lfloor m \cdot \rho_r \rfloor0, with qr=⌊m⋅ρr⌋q_r = \lfloor m \cdot \rho_r \rfloor1 (the number of anchors) typically qr=⌊m⋅ρr⌋q_r = \lfloor m \cdot \rho_r \rfloor2--qr=⌊m⋅ρr⌋q_r = \lfloor m \cdot \rho_r \rfloor3.

Convergence is measured via primal/dual residuals qr=⌊m⋅ρr⌋q_r = \lfloor m \cdot \rho_r \rfloor4 and qr=⌊m⋅ρr⌋q_r = \lfloor m \cdot \rho_r \rfloor5, which empirically decrease as the algorithm proceeds. Adaptive stepsize schemes (e.g., updating qr=⌊m⋅ρr⌋q_r = \lfloor m \cdot \rho_r \rfloor6 every 10 steps) are used to optimize convergence (Wei et al., 13 Nov 2025).

4. Theoretical Guarantees

AFCF provides two formal guarantees:

(a) Fairness Equivalence: Under the formulated group-label joint constraint, the final clustering of all qr=⌊m⋅ρr⌋q_r = \lfloor m \cdot \rho_r \rfloor7 points recovers the {\em exact} per-cluster demographic group proportions present in the anchor clustering. This implies preservation of standard fairness metrics, including balance and disparate impact.

(b) Linear-Time Scalability: For fixed qr=⌊m⋅ρr⌋q_r = \lfloor m \cdot \rho_r \rfloor8, qr=⌊m⋅ρr⌋q_r = \lfloor m \cdot \rho_r \rfloor9, and cluster count Δ=m−∑r=1tqr\Delta = m - \sum_{r=1}^t q_r0, the total computational complexity of AFCF is

Δ=m−∑r=1tqr\Delta = m - \sum_{r=1}^t q_r1

where Δ=m−∑r=1tqr\Delta = m - \sum_{r=1}^t q_r2 is the complexity of the fair clustering subroutine on Δ=m−∑r=1tqr\Delta = m - \sum_{r=1}^t q_r3 anchors. With Δ=m−∑r=1tqr\Delta = m - \sum_{r=1}^t q_r4, this yields overall linear scaling in the number of samples Δ=m−∑r=1tqr\Delta = m - \sum_{r=1}^t q_r5, a substantial reduction from the quadratic or higher costs of many existing fair clustering frameworks (Wei et al., 13 Nov 2025).

5. Empirical Evaluation

AFCF was benchmarked on five real-world datasets:

Dataset Size # Clusters Sensitive Attribute
Law School 18,000 2 Gender
Credit Card 29,000 5 Gender
Bank 41,000 2 Marital Status
Zafar 100,000 2 Binary Sensitive
Census II 2,460,000 5 Gender

Performance metrics included clustering quality (Accuracy, Normalized Mutual Information) and fairness (Balance, Minimal Normalized Conditional Entropy). Representative state-of-the-art methods—SpFC, VFC, FFC, FMSC, and fairletFC—were integrated into the AFCF pipeline.

Key empirical findings:

  • Computational Speedup: On Census II, VFC alone required ≈1,500s; VFC-AF (AFCF version) executed in ≈918s. SpFC could not complete within 30 minutes on Bank, whereas SpFC-AF finished in 35s. In general, AFCF enabled one to two orders of magnitude acceleration.
  • Clustering Quality and Fairness Preservation: Clustering accuracy and NMI varied by only a few percentage points; fairness metrics such as balance and MNCE were preserved within 1–2% of anchor clustering levels, consistent with the theoretical guarantee.
  • Ablation Analysis: Substituting FDAS for random or vanilla DAS anchor sampling resulted in degenerate clusters or substantial fairness loss. Excluding the group-label joint constraint in the graph update ("AC" ablation) degraded balance by up to 10% (Wei et al., 13 Nov 2025).

6. Significance and Implications

AFCF decouples scalability from the core fair clustering algorithm: any fair clustering routine applied to the anchor subset inherits AFCF’s linear-time scalability and exact fairness preservation when extended to the whole dataset. This modularity allows rapid experimentation and deployment across large-scale, high-stakes environments requiring fairness guarantees in unsupervised learning. The systematic empirical and theoretical analysis demonstrates AFCF’s ability to bridge the computational gap in fair clustering, establishing it as a plug-and-play, practical framework for scalable fair learning (Wei et al., 13 Nov 2025).

Definition Search Book Streamline Icon: https://streamlinehq.com
References (1)

Topic to Video (Beta)

No one has generated a video about this topic yet.

Whiteboard

No one has generated a whiteboard explanation for this topic yet.

Follow Topic

Get notified by email when new papers are published related to Anchor-based Fair Clustering Framework (AFCF).