Papers
Topics
Authors
Recent
Search
2000 character limit reached

Stability-Validated K-Prototypes Clustering

Updated 10 January 2026
  • The paper presents a stability-based method that evaluates clustering reproducibility by balancing global (between-cluster) and local (within-cluster) stability.
  • It leverages the K-Prototypes algorithm to handle mixed numerical and categorical data through tailored perturbations and combined distance measures.
  • Empirical results show that the Stadion criterion outperforms traditional indices, ensuring robust model selection while mitigating subcluster overfitting.

Stability-Validated K-Prototypes Clustering is a model selection approach for clustering mixed-type data, rooted in the principle that an optimal clustering should be reproducible under data perturbations and show no stable sub-partitions within its clusters. This methodology, formalized in the Stadion criterion, provides a rigorous internal validation mechanism, extending stability-based cluster validation specifically to the mixed-data regime addressed by the K-Prototypes algorithm. The validity of a clustering configuration is assessed as a trade-off: maximizing stability of the clustering as a whole across perturbations, while ensuring that individual clusters do not admit reproducible internal structure under the same perturbations (Mourer et al., 2020).

1. Theoretical Principle and Formulation

The foundation of the Stability-Validated K-Prototypes framework is the assertion that “a good clustering is one which (a) is stable under small perturbations of the data, and (b) within each cluster no further stable partition exists.” This two-part definition encapsulates both global reproducibility and local indivisibility.

Let X={x1,...,xN}⊂Rpnum×{ categories}X = \{x_1, ..., x_N\} \subset \mathbb{R}^{p_{\mathrm{num}}} \times \{\,\text{categories}\} denote a dataset with mixed numerical and categorical attributes. For a given KK, K-Prototypes yields a partition CK={C1,...,CK}\mathcal{C}_K = \{C_1, ..., C_K\}, and s(C,C′)∈[0,1]s(\mathcal{C}, \mathcal{C}') \in [0, 1] quantifies the similarity between two labelings, typically using the Adjusted Rand Index (ARI).

Between-Cluster Stability

Between-cluster stability StabB\mathrm{Stab}_B is defined as the mean similarity between a reference clustering and clusterings of DD perturbed versions of the data: StabB(A,X,K)=1D∑d=1Ds(CK,CK(d))\mathrm{Stab}_B(\mathcal{A}, X, K) = \frac{1}{D} \sum_{d=1}^D s\bigl(\mathcal{C}_K, \mathcal{C}_K^{(d)}\bigr) where CK(d)=A(X(d),K)\mathcal{C}_K^{(d)} = \mathcal{A}(X^{(d)}, K), and each X(d)X^{(d)} is a perturbation of XX (additive noise to numericals, bootstrap or flip on categorical values).

Within-Cluster Stability

For a set KK0 of candidate subcluster counts (e.g., KK1), each cluster KK2 is re-clustered with KK3 both in the original data and under perturbations. For each KK4 and KK5, within-cluster stability is

KK6

Aggregate within-cluster stability KK7 is the cluster-size-weighted average of these stabilities over all KK8 and KK9.

Stadion Criterion

The Stadion index is the stability difference to be maximized: CK={C1,...,CK}\mathcal{C}_K = \{C_1, ..., C_K\}0 Optimal CK={C1,...,CK}\mathcal{C}_K = \{C_1, ..., C_K\}1 is the maximizer CK={C1,...,CK}\mathcal{C}_K = \{C_1, ..., C_K\}2.

2. Algorithmic Procedure

The procedure for Stability-Validated K-Prototypes is outlined as follows:

  • Input: data CK={C1,...,CK}\mathcal{C}_K = \{C_1, ..., C_K\}3; candidate cluster counts CK={C1,...,CK}\mathcal{C}_K = \{C_1, ..., C_K\}4; subcluster set CK={C1,...,CK}\mathcal{C}_K = \{C_1, ..., C_K\}5; number of perturbations CK={C1,...,CK}\mathcal{C}_K = \{C_1, ..., C_K\}6; noise levels grid CK={C1,...,CK}\mathcal{C}_K = \{C_1, ..., C_K\}7.
  • For each CK={C1,...,CK}\mathcal{C}_K = \{C_1, ..., C_K\}8, and each noise level CK={C1,...,CK}\mathcal{C}_K = \{C_1, ..., C_K\}9:
    • Generate s(C,C′)∈[0,1]s(\mathcal{C}, \mathcal{C}') \in [0, 1]0 perturbed realizations s(C,C′)∈[0,1]s(\mathcal{C}, \mathcal{C}') \in [0, 1]1.
    • Cluster each perturbed s(C,C′)∈[0,1]s(\mathcal{C}, \mathcal{C}') \in [0, 1]2 with K-Prototypes to obtain s(C,C′)∈[0,1]s(\mathcal{C}, \mathcal{C}') \in [0, 1]3.
    • Compute s(C,C′)∈[0,1]s(\mathcal{C}, \mathcal{C}') \in [0, 1]4 as mean similarity (ARI) between s(C,C′)∈[0,1]s(\mathcal{C}, \mathcal{C}') \in [0, 1]5 and s(C,C′)∈[0,1]s(\mathcal{C}, \mathcal{C}') \in [0, 1]6.
    • For every cluster s(C,C′)∈[0,1]s(\mathcal{C}, \mathcal{C}') \in [0, 1]7, and for each s(C,C′)∈[0,1]s(\mathcal{C}, \mathcal{C}') \in [0, 1]8:
    • Re-cluster s(C,C′)∈[0,1]s(\mathcal{C}, \mathcal{C}') \in [0, 1]9 and its perturbed versions StabB\mathrm{Stab}_B0 with StabB\mathrm{Stab}_B1.
    • Compute StabB\mathrm{Stab}_B2 and aggregate to obtain StabB\mathrm{Stab}_B3.
    • Compute the Stadion index per StabB\mathrm{Stab}_B4 and noise level.
  • Aggregate the Stadion index over noise levels (max or mean).
  • Return the StabB\mathrm{Stab}_B5 maximizing aggregated Stadion, along with full stability paths for diagnostic visualization.

Pseudocode is provided in the source for a step-by-step workflow, including cluster restriction, weighting, and perturbation strategies.

3. Adaptation to Mixed-Type Data and Perturbation Schemes

K-Prototypes clustering is designed for mixed numerical and categorical datasets, using a combined distance: StabB\mathrm{Stab}_B6 where StabB\mathrm{Stab}_B7 is a fixed weight for categorical mismatches.

Perturbation Strategy

  • Numerical features: Add uniform StabB\mathrm{Stab}_B8 or Gaussian StabB\mathrm{Stab}_B9 noise, after feature scaling to zero mean and unit variance.
  • Categorical features: Either bootstrap within-column (resampling with replacement), or independently flip values with small probability DD0, to jitter cluster boundaries.

Similarity is evaluated via ARI on the induced labelings. When comparing reference and perturbed clusterings, ARI is computed by transferring cluster labels from perturbed samples back to the original observations using sample correspondence.

Alternative Similarity Measures

A cluster co-occurrence matrix DD1, with

DD2

permits alternative definitions of stability, using within- and between-cluster co-occurrences.

4. Computational and Practical Considerations

The computational cost is dominated by repeated K-Prototypes clusterings across perturbations, noise levels, and cluster sizes: DD3 with DD4 the number of iterations per clustering. In typical practice, DD5 is 2–5, DD6–10, and DD7.

Empirical observations suggest DD8 or 2 already gives low-variance estimates, but DD9–10 is robust. The noise-level grid StabB(A,X,K)=1D∑d=1Ds(CK,CK(d))\mathrm{Stab}_B(\mathcal{A}, X, K) = \frac{1}{D} \sum_{d=1}^D s\bigl(\mathcal{C}_K, \mathcal{C}_K^{(d)}\bigr)0 should cover up to StabB(A,X,K)=1D∑d=1Ds(CK,CK(d))\mathrm{Stab}_B(\mathcal{A}, X, K) = \frac{1}{D} \sum_{d=1}^D s\bigl(\mathcal{C}_K, \mathcal{C}_K^{(d)}\bigr)1, so that at maximum noise all meaningful structure is dissolved and StabB(A,X,K)=1D∑d=1Ds(CK,CK(d))\mathrm{Stab}_B(\mathcal{A}, X, K) = \frac{1}{D} \sum_{d=1}^D s\bigl(\mathcal{C}_K, \mathcal{C}_K^{(d)}\bigr)2 emerges as optimal.

Initialization procedures (e.g., greedy, multiple restarts) should be consistent across runs to mitigate non-deterministic cluster assignments. Perturbations of the categorical component (bootstrapping or flipping) must be balanced: if category level proportions are highly imbalanced, bootstrapping is recommended.

5. Empirical Performance and Validation Approaches

Mourer et al. (Mourer et al., 2020) evaluated Stadion on 73 numerically-typed datasets (Gaussian mixtures, shapes, UCI benchmarks), comparing over 20 internal validation indices. Stadion-max with additive noise consistently ranked among the top two methods, outperforming conventional indices such as Silhouette, Dunn, Calinski-Harabasz (CH), Gap, and prior stability-based techniques.

For mixed-type evaluation, the prescribed protocol is:

  1. Assemble or simulate benchmark datasets with ground-truth clusters, including both synthetic (Gaussian numericals + Dirichlet-noise categorical prototypes) and real (e.g., UCI Adult, Heart Disease datasets).
  2. Apply Stability-Validated K-Prototypes as described above.
  3. Compare the selected StabB(A,X,K)=1D∑d=1Ds(CK,CK(d))\mathrm{Stab}_B(\mathcal{A}, X, K) = \frac{1}{D} \sum_{d=1}^D s\bigl(\mathcal{C}_K, \mathcal{C}_K^{(d)}\bigr)3 to true StabB(A,X,K)=1D∑d=1Ds(CK,CK(d))\mathrm{Stab}_B(\mathcal{A}, X, K) = \frac{1}{D} \sum_{d=1}^D s\bigl(\mathcal{C}_K, \mathcal{C}_K^{(d)}\bigr)4 (win counts), and compute StabB(A,X,K)=1D∑d=1Ds(CK,CK(d))\mathrm{Stab}_B(\mathcal{A}, X, K) = \frac{1}{D} \sum_{d=1}^D s\bigl(\mathcal{C}_K, \mathcal{C}_K^{(d)}\bigr)5.
  4. Use baselines tailored to mixed data: adaptations of Silhouette with Gower distance, WB-index with Gower, categorical CH, EM-based BIC, mixed-data indices (e.g., E–CV index), or mixed-consensus clustering.
  5. Inspect stability paths plotted against noise to confirm the expected dissolution of subclusters and cluster merges as perturbation increases.

The expected empirical outcome is that Stadion-max remains a top-performing internal criterion for mixed data, contingent on an appropriate balance between numerical noise and categorical perturbation.

6. Connections, Limitations, and Perspectives

Stability-based validation circumvents the ill-defined objectives of unsupervised clustering by appealing to reproducibility under controlled perturbations. The Stadion criterion subsumes previous notions of stability but overcomes the inadequacy of classical stability for detecting underestimation of StabB(A,X,K)=1D∑d=1Ds(CK,CK(d))\mathrm{Stab}_B(\mathcal{A}, X, K) = \frac{1}{D} \sum_{d=1}^D s\bigl(\mathcal{C}_K, \mathcal{C}_K^{(d)}\bigr)6 by penalizing within-cluster stability. This framework is model-agnostic, but its adaptation to K-Prototypes is central for practical handling of heterogeneous attributes.

A plausible implication is that the explicit modeling of both global and local stability can extend to other mixed-type clustering algorithms, provided suitable perturbation and similarity definitions. However, the computational overhead is substantial, and the performance is governed by the adequacy of perturbation schemes and similarity measures for both data types. The method is inherently internal and does not leverage possible external or semi-supervised information.

The neutral point is: empirical results demonstrate top-tier performance for Stadion-max as an internal criterion, but its generalization to non-K-Prototypes algorithms or very high-dimensional mixed data remains to be explored (Mourer et al., 2020).

Definition Search Book Streamline Icon: https://streamlinehq.com
References (1)

Topic to Video (Beta)

No one has generated a video about this topic yet.

Whiteboard

No one has generated a whiteboard explanation for this topic yet.

Follow Topic

Get notified by email when new papers are published related to Stability-Validated K-Prototypes Clustering.