Papers
Topics
Authors
Recent
Search
2000 character limit reached

Fair k-Clustering with Multiple Colors

Updated 17 November 2025
  • Fair k-clustering is a constrained clustering problem that ensures every cluster has equal representation from multiple protected groups under objectives like k-median, k-means, and k-center.
  • It employs a black-box reduction that transforms any α-approximation for standard clustering into an (α+2)-approximation for fair clustering, guaranteeing exact balance.
  • Experimental results show that the approach scales to large datasets while maintaining near-baseline costs and strict fairness without allowing additive violations.

A fair k-clustering problem is a constrained clustering formulation in which the assignments to clusters must respect the relative balance or representation of multiple “colors” (protected groups), with cost measured under classic objectives such as k-median, k-means, or k-center. The multi-color fair k-clustering problem, as studied in "Fair Clustering with Multiple Colors" (Böhm et al., 2020), formalizes fairness as a requirement that every cluster contain an identical count of each color, supporting arbitrary numbers of colors, clusters, and points. This problem poses unique structural and computational challenges beyond the two-color case, and remained open for true constant-factor approximation until the introduction of a black-box reduction from vanilla clustering objectives. The following sections survey formal definitions, objective functions, the reduction methodology and its theoretical guarantees, key analytic tools, computational complexity, and experimental highlights.

1. Formal Model: Fair k-Clustering with Multiple Colors

The input consists of a point set AA of size nrn r in Rd\mathbb R^d (or general metric space), partitioned by a coloring function c:A→[r]c: A \to [r] into rr color classes A(1),…,A(r)A^{(1)},\ldots,A^{(r)}, with ∣A(i)∣=n|A^{(i)}|=n for all ii. The clustering task is to select kk clusters, each assigned a center, and assign each point to a center.

Exact balance constraint: For all clusters CjC_j and colors nrn r0,

nrn r1

Consequently, each nrn r2 must be a multiple of nrn r3, and all nrn r4 must distribute evenly among clusters.

2. Clustering Objectives under Fairness

The fair k-clustering problem admits several center-based objectives, unified as:

nrn r5

with nrn r6 the set of cluster centers and nrn r7 the assignment. Special cases include:

  • k-median: nrn r8
  • k-means: nrn r9
  • k-center: Rd\mathbb R^d0, Rd\mathbb R^d1 arbitrary

All subject to the exact-balance constraint per cluster.

3. Black-Box Reduction from Unconstrained to Fair Clustering

The key contribution of (Böhm et al., 2020) is an algorithmic reduction that transforms any Rd\mathbb R^d2-approximation algorithm for the unconstrained (vanilla) Rd\mathbb R^d3-clustering problem into an Rd\mathbb R^d4-approximation for the fair Rd\mathbb R^d5-clustering problem with Rd\mathbb R^d6 colors. This is the first such reduction to achieve a true constant factor for all three objectives.

Algorithmic Framework (paraphrased from Algorithm 1):

For each color Rd\mathbb R^d7:

  1. For all Rd\mathbb R^d8 compute a min-cost perfect matching Rd\mathbb R^d9 under the c:A→[r]c: A \to [r]0 cost metric.
  2. Run the given c:A→[r]c: A \to [r]1-approximation algorithm for unconstrained c:A→[r]c: A \to [r]2 clustering on c:A→[r]c: A \to [r]3; obtain centers c:A→[r]c: A \to [r]4.
  3. For each color c:A→[r]c: A \to [r]5 and each c:A→[r]c: A \to [r]6, assign c:A→[r]c: A \to [r]7 to the closest center in c:A→[r]c: A \to [r]8 by matching via c:A→[r]c: A \to [r]9.
  4. Compute the total clustering cost for the combined assignment.

Return the best solution over all rr0.

Guarantee (Theorem 2.1): If the unconstrained problem has an rr1-approximation in time rr2, then fair rr3-clustering admits an rr4-approximation in time rr5, where rr6 is the time for min-cost perfect matching under the ground rr7 distance.

4. Theoretical Foundations and Proof Technique

Cost and feasibility analysis depend on several key tools:

  • Earth Mover's Distance (rr8): Used for matching between color blocks, rr9 is the minimum cost of perfect matching under A(1),…,A(r)A^{(1)},\ldots,A^{(r)}0 ground distance.
  • Color-block averaging argument: By averaging costs over pivot blocks A(1),…,A(r)A^{(1)},\ldots,A^{(r)}1, it follows that there exists a block for which the aggregate matching cost to other blocks plus the unconstrained cost is at most A(1),…,A(r)A^{(1)},\ldots,A^{(r)}2, where A(1),…,A(r)A^{(1)},\ldots,A^{(r)}3 is the optimal cost for fair A(1),…,A(r)A^{(1)},\ldots,A^{(r)}4-clustering.
  • Triangle inequality chain: The final cost for A(1),…,A(r)A^{(1)},\ldots,A^{(r)}5 assigned via the best pivot and matchings is bounded by

A(1),…,A(r)A^{(1)},\ldots,A^{(r)}6

  • The reduction is fully constructive due to exhaustive trial over all color classes.

5. Computational Complexity, Scope, and Generalization

The reduction involves A(1),…,A(r)A^{(1)},\ldots,A^{(r)}7 instances of min-cost perfect matching on A(1),…,A(r)A^{(1)},\ldots,A^{(r)}8 points, and A(1),…,A(r)A^{(1)},\ldots,A^{(r)}9 runs of the unconstrained clustering algorithm. The overall running time is thus ∣A(i)∣=n|A^{(i)}|=n0, with no restriction on ∣A(i)∣=n|A^{(i)}|=n1 or ∣A(i)∣=n|A^{(i)}|=n2.

This framework works for arbitrary finite metric spaces (using ∣A(i)∣=n|A^{(i)}|=n3 embedding if required), all center-based objectives, and arbitrary cluster counts and color numbers.

Special cases:

  • For k-center, a simple farthest-first traversal yields a 3-approximate fair solution in ∣A(i)∣=n|A^{(i)}|=n4 time.
  • When ∣A(i)∣=n|A^{(i)}|=n5 (i.e., every point is a center), a random-block-sampling routine gives a 2-approximate fair partition in ∣A(i)∣=n|A^{(i)}|=n6 time.

Hardness: Even the fair ∣A(i)∣=n|A^{(i)}|=n7-median or ∣A(i)∣=n|A^{(i)}|=n8-center problem is APX-hard for ∣A(i)∣=n|A^{(i)}|=n9; no PTAS is possible. Constant-factor approximation is therefore best possible in general.

6. Experimental Results

Implementation on six real benchmark datasets (Adults, Athletes, Bank, Diabetes, Credit cards, Census1990), with up to ii0 colors and ii1 points, demonstrates:

  • (α+2)-approximation algorithms are ii2–ii3 faster than previous bicriteria methods, with exact fairness.
  • Empirical costs lie ii4–ii5 above the unconstrained k-median baseline, closely matching theoretical bounds.
  • Fast special-case routines scale to hundreds of thousands of points in tens of seconds.

Empirical takeaways:

  • Fairness is achieved strictly (no violation), as opposed to prior algorithms permitting additive violation in cluster composition.
  • Solutions maintain high computational scalability for moderate to large datasets.

7. Extensions, Limitations, and Outlook

This reduction establishes a template for fair clustering applicable to any approximation algorithm for unconstrained clustering objectives, offering simplicity, modularity, and constant-factor guarantees in arbitrary finite metric spaces and arbitrary numbers of colors and clusters.

Extensions may target:

  • Generalizing the reduction to balance constraints beyond exact equality (e.g., range constraints, proportional, or disparate impact).
  • Adaptation to streaming or distributed environments via parallelization of the matching and unconstrained clustering phase.
  • Empirical study of trade-offs between cluster balance and utility for varied clustering criteria.

This framework resolves the multi-color fair k-clustering approximation barrier: for the first time, true constant-factor guarantees are established for fair k-median, k-means, and k-center under a strict balance model, with broad applicability and strong empirical validation (Böhm et al., 2020).

Definition Search Book Streamline Icon: https://streamlinehq.com
References (1)

Topic to Video (Beta)

No one has generated a video about this topic yet.

Whiteboard

No one has generated a whiteboard explanation for this topic yet.

Follow Topic

Get notified by email when new papers are published related to Fair k-Clustering Problem.