Papers
Topics
Authors
Recent
Search
2000 character limit reached

Progressive k-Annealing: Adaptive Clustering

Updated 6 February 2026
  • Progressive k-Annealing is a data-adaptive learning paradigm that dynamically adjusts prototypes and model complexity through a temperature-based annealing process.
  • It utilizes a temperature-parameterized free-energy formulation with Gibbs soft assignments and stochastic approximation to ensure robustness and convergence.
  • The method automatically grows clusters via bifurcation phenomena, is robust to initialization, and extends to hierarchical and reinforcement learning frameworks.

Progressive k-Annealing is a data-adaptive learning paradigm in which both the number of prototypes (k) and their locations are dynamically adapted online as a function of a decreasing “annealing” parameter, generally the temperature TT, or its inverse, β\beta. This approach extends classical deterministic annealing for clustering and classification, enabling automatic model complexity growth through bifurcation phenomena while mitigating sensitivity to initialization and poor local minima. The methodology is grounded in free-energy minimization, stochastic approximation, and Bregman divergence regularization, yielding an interpretable, robust, and complexity-adaptive framework for unsupervised and supervised learning (Mavridis et al., 2021, Mavridis et al., 2022, Mavridis et al., 2022).

1. Free-Energy Formulation and Annealing Principle

At its core, Progressive k-Annealing replaces the non-convex hard-clustering objective with a temperature-parameterized free-energy functional: FT[M,p]=D[M,p]−T H[p]F_T[M,p] = D[M,p] - T\,H[p] where D[M,p]=EX[∑i=1kp(μi∣X)d(X,μi)]D[M, p] = \mathbb{E}_X \bigl[\sum_{i=1}^k p(\mu_i|X) d(X, \mu_i)\bigr] is the expected divergence between data points and prototypes, and H[p]=−EX[∑i=1kp(μi∣X)log⁡p(μi∣X)]H[p] = -\mathbb{E}_X\bigl[\sum_{i=1}^k p(\mu_i|X) \log p(\mu_i|X)\bigr] is the Shannon entropy of the assignment probabilities. Here p(μi∣x)p(\mu_i|x) are soft assignments, and d(⋅,⋅)d(\cdot,\cdot) is a user-selected Bregman divergence (e.g., Euclidean distance, KL divergence) (Mavridis et al., 2021). As T→0T\to0, the entropy term vanishes, recapitulating hard assignments as in Lloyd’s algorithm; as T→∞T\to\infty, the assignments become uniform and complexity collapses to k=1k=1.

The soft-assignment takes the Gibbs form: β\beta0 which ensures the differentiability and tractability of the optimization path as β\beta1 is lowered. The cluster centroids update in closed-form as weighted means: β\beta2 for all members of the Bregman divergence class (Mavridis et al., 2021).

2. Bifurcation Phenomenon and Automatic Model Complexity Growth

Unlike fixed-β\beta3 clustering, the annealing process in Progressive k-Annealing enables the number of clusters to increase naturally via bifurcations as β\beta4 decreases. The critical temperature β\beta5 at which an existing prototype bifurcates is characterized by a loss in local stability of the free-energy minimum, most precisely by the criterion: β\beta6 where β\beta7 is the Hessian of the Bregman generator β\beta8, and β\beta9 is the local covariance matrix in the dual space. In the canonical Euclidean case, the condition simplifies to FT[M,p]=D[M,p]−T H[p]F_T[M,p] = D[M,p] - T\,H[p]0 (Mavridis et al., 2021, Mavridis et al., 2022, Mavridis et al., 2022). In practice, split detection can be implemented via the “virtual split” heuristic: after convergence at a given FT[M,p]=D[M,p]−T H[p]F_T[M,p] = D[M,p] - T\,H[p]1, perturb each prototype by FT[M,p]=D[M,p]−T H[p]F_T[M,p] = D[M,p] - T\,H[p]2, update, and observe if the perturbed prototypes diverge (indicating a true bifurcation and increment in FT[M,p]=D[M,p]−T H[p]F_T[M,p] = D[M,p] - T\,H[p]3) or coalesce (no split) (Mavridis et al., 2021).

3. Stochastic Approximation and Online Prototype Updates

Progressive k-Annealing is realized via online, gradient-free stochastic approximation algorithms. For each prototype FT[M,p]=D[M,p]−T H[p]F_T[M,p] = D[M,p] - T\,H[p]4, running estimates of cluster responsibility FT[M,p]=D[M,p]−T H[p]F_T[M,p] = D[M,p] - T\,H[p]5 and assigned data-weighted sum FT[M,p]=D[M,p]−T H[p]F_T[M,p] = D[M,p] - T\,H[p]6 are tracked: FT[M,p]=D[M,p]−T H[p]F_T[M,p] = D[M,p] - T\,H[p]7 where weights FT[M,p]=D[M,p]−T H[p]F_T[M,p] = D[M,p] - T\,H[p]8 are derived from the Gibbs assignment, and FT[M,p]=D[M,p]−T H[p]F_T[M,p] = D[M,p] - T\,H[p]9 is the step size (D[M,p]=EX[∑i=1kp(μi∣X)d(X,μi)]D[M, p] = \mathbb{E}_X \bigl[\sum_{i=1}^k p(\mu_i|X) d(X, \mu_i)\bigr]0). For settings involving local parametric models within clusters, a two-timescale stochastic approximation is employed: cluster prototypes evolve on a slow timescale D[M,p]=EX[∑i=1kp(μi∣X)d(X,μi)]D[M, p] = \mathbb{E}_X \bigl[\sum_{i=1}^k p(\mu_i|X) d(X, \mu_i)\bigr]1, and local model parameters are updated on a faster timescale D[M,p]=EX[∑i=1kp(μi∣X)d(X,μi)]D[M, p] = \mathbb{E}_X \bigl[\sum_{i=1}^k p(\mu_i|X) d(X, \mu_i)\bigr]2, with D[M,p]=EX[∑i=1kp(μi∣X)d(X,μi)]D[M, p] = \mathbb{E}_X \bigl[\sum_{i=1}^k p(\mu_i|X) d(X, \mu_i)\bigr]3 (Mavridis et al., 2022).

Convergence of these updates to a local minimizer of D[M,p]=EX[∑i=1kp(μi∣X)d(X,μi)]D[M, p] = \mathbb{E}_X \bigl[\sum_{i=1}^k p(\mu_i|X) d(X, \mu_i)\bigr]4 is guaranteed under standard regularity and step-size conditions (Mavridis et al., 2021, Mavridis et al., 2022, Mavridis et al., 2022).

4. Algorithmic Workflow and Hyper-Parameter Control

The progressive k-annealing workflow can be summarized as follows (Mavridis et al., 2021, Mavridis et al., 2022):

Initialization:

  • Select Bregman divergence, temperature schedule D[M,p]=EX[∑i=1kp(μi∣X)d(X,μi)]D[M, p] = \mathbb{E}_X \bigl[\sum_{i=1}^k p(\mu_i|X) d(X, \mu_i)\bigr]5 (typically D[M,p]=EX[∑i=1kp(μi∣X)d(X,μi)]D[M, p] = \mathbb{E}_X \bigl[\sum_{i=1}^k p(\mu_i|X) d(X, \mu_i)\bigr]6 with D[M,p]=EX[∑i=1kp(μi∣X)d(X,μi)]D[M, p] = \mathbb{E}_X \bigl[\sum_{i=1}^k p(\mu_i|X) d(X, \mu_i)\bigr]7), perturbation D[M,p]=EX[∑i=1kp(μi∣X)d(X,μi)]D[M, p] = \mathbb{E}_X \bigl[\sum_{i=1}^k p(\mu_i|X) d(X, \mu_i)\bigr]8, tolerances D[M,p]=EX[∑i=1kp(μi∣X)d(X,μi)]D[M, p] = \mathbb{E}_X \bigl[\sum_{i=1}^k p(\mu_i|X) d(X, \mu_i)\bigr]9, H[p]=−EX[∑i=1kp(μi∣X)log⁡p(μi∣X)]H[p] = -\mathbb{E}_X\bigl[\sum_{i=1}^k p(\mu_i|X) \log p(\mu_i|X)\bigr]0, H[p]=−EX[∑i=1kp(μi∣X)log⁡p(μi∣X)]H[p] = -\mathbb{E}_X\bigl[\sum_{i=1}^k p(\mu_i|X) \log p(\mu_i|X)\bigr]1, and maximum H[p]=−EX[∑i=1kp(μi∣X)log⁡p(μi∣X)]H[p] = -\mathbb{E}_X\bigl[\sum_{i=1}^k p(\mu_i|X) \log p(\mu_i|X)\bigr]2.
  • Start with H[p]=−EX[∑i=1kp(μi∣X)log⁡p(μi∣X)]H[p] = -\mathbb{E}_X\bigl[\sum_{i=1}^k p(\mu_i|X) \log p(\mu_i|X)\bigr]3 prototype.

Annealing Loop:

  • For each temperature H[p]=−EX[∑i=1kp(μi∣X)log⁡p(μi∣X)]H[p] = -\mathbb{E}_X\bigl[\sum_{i=1}^k p(\mu_i|X) \log p(\mu_i|X)\bigr]4:
    • For each prototype, attempt a virtual split by H[p]=−EX[∑i=1kp(μi∣X)log⁡p(μi∣X)]H[p] = -\mathbb{E}_X\bigl[\sum_{i=1}^k p(\mu_i|X) \log p(\mu_i|X)\bigr]5.
    • Iterate stochastic approximation updates until convergence.
    • Prune coalesced or idle prototypes (small H[p]=−EX[∑i=1kp(μi∣X)log⁡p(μi∣X)]H[p] = -\mathbb{E}_X\bigl[\sum_{i=1}^k p(\mu_i|X) \log p(\mu_i|X)\bigr]6).
    • Increment H[p]=−EX[∑i=1kp(μi∣X)log⁡p(μi∣X)]H[p] = -\mathbb{E}_X\bigl[\sum_{i=1}^k p(\mu_i|X) \log p(\mu_i|X)\bigr]7 when a genuine bifurcation is detected.
    • Decrease H[p]=−EX[∑i=1kp(μi∣X)log⁡p(μi∣X)]H[p] = -\mathbb{E}_X\bigl[\sum_{i=1}^k p(\mu_i|X) \log p(\mu_i|X)\bigr]8 for the next stage.

Termination:

  • Optionally fine-tune assignments with hard-clustering at H[p]=−EX[∑i=1kp(μi∣X)log⁡p(μi∣X)]H[p] = -\mathbb{E}_X\bigl[\sum_{i=1}^k p(\mu_i|X) \log p(\mu_i|X)\bigr]9.

Hyper-parameter roles:

  • p(μi∣x)p(\mu_i|x)0 determines initial smoothness (single cluster stability).
  • p(μi∣x)p(\mu_i|x)1 sets the annealing rate (proximal steps vs computational effort).
  • p(μi∣x)p(\mu_i|x)2 and merge thresholds (p(μi∣x)p(\mu_i|x)3, p(μi∣x)p(\mu_i|x)4) regularize the minimal cluster "resolution" and suppress redundant splits.

The table below summarizes key functional aspects:

Element Description Typical Range
p(μi∣x)p(\mu_i|x)5 Initial temperature (smoothness) p(μi∣x)p(\mu_i|x)6
p(μi∣x)p(\mu_i|x)7 Decay factor for temperature p(μi∣x)p(\mu_i|x)8--p(μi∣x)p(\mu_i|x)9
d(⋅,⋅)d(\cdot,\cdot)0 Split perturbation size d(⋅,⋅)d(\cdot,\cdot)1
d(⋅,⋅)d(\cdot,\cdot)2 Convergence tolerance for updates Small (d(⋅,⋅)d(\cdot,\cdot)3)
d(⋅,⋅)d(\cdot,\cdot)4 Prototype merge threshold Moderate (d(⋅,⋅)d(\cdot,\cdot)5 data scale)
d(⋅,⋅)d(\cdot,\cdot)6 Pruning threshold for idle prototypes Small

5. Application Domains and Hierarchical Extensions

Progressive k-Annealing provides a general-purpose, online, robust clustering and classification methodology. It extends seamlessly to hierarchical and multi-resolution settings as realized in Multi-Resolution Online Deterministic Annealing (MRODA), whereby codebooks are organized as a tree structure, and ODA is invoked recursively within each node. This hierarchical organization localizes search, preserves computational tractability (d(⋅,⋅)d(\cdot,\cdot)7 for depth d(⋅,⋅)d(\cdot,\cdot)8), and naturally exploits data locality and variable-rate partitioning akin to deep architectures. Within each resolution, bifurcations drive local complexity growth, and the process yields adaptive, interpretable variable-depth partitioning (Mavridis et al., 2022).

In reinforcement learning, the two-timescale stochastic approximation of k-annealing integrates with Q-learning via joint updates: fast-timescale temporal difference (TD) learning for Q-values and slower prototype updates, yielding an adaptive state-action aggregation scheme (Mavridis et al., 2022).

6. Convergence Guarantees and Empirical Performance

Convergence of Progressive k-Annealing is established via the ODE method for stochastic approximation. Provided the step-sizes satisfy d(⋅,⋅)d(\cdot,\cdot)9 and standard martingale-difference conditions, the iterates converge almost surely to locally stable equilibria of the corresponding ODE in the parameter space. In the two-timescale case, joint convergence of prototype and local-model parameters is guaranteed (Mavridis et al., 2021, Mavridis et al., 2022, Mavridis et al., 2022).

Empirical evaluations show that Progressive k-Annealing:

  • Automatically discovers the effective number of clusters T→0T\to00, with graceful complexity-performace trade-offs as measured by distortion or classification error.
  • Matches or surpasses the online convergence rate and accuracy of batch deterministic annealing, k-means, linear SVMs, and approaches shallow neural network or random forest accuracy, often with greater interpretability (Mavridis et al., 2021).
  • Is robust to initialization and avoids poor local minima due to annealing from high-T→0T\to01 solutions.
  • Enables online “dial-in” of desired model complexity by stopping annealing early, providing flexible computational and representational control.

7. Relation to Classical Deterministic Annealing and Extensions

Progressive k-Annealing is a direct online, gradient-free stochastic approximation realization of classical deterministic annealing frameworks (e.g., Rose ’98), where the number T→0T\to02 emerges at eigenvalue-driven bifurcation points as the temperature is reduced (Mavridis et al., 2021, Mavridis et al., 2022). The methodology extends to hierarchical architectures, two-timescale estimation, variable-resolution clustering/classification, and function approximation in each partition. In contrast to offline deterministic annealing, progressive k-annealing operates in a single pass, with the capacity to grow T→0T\to03 as data and complexity demands, and provides online adaptivity and interpretability (Mavridis et al., 2022).

Topic to Video (Beta)

No one has generated a video about this topic yet.

Whiteboard

No one has generated a whiteboard explanation for this topic yet.

Follow Topic

Get notified by email when new papers are published related to Progressive k-Annealing.