Papers
Topics
Authors
Recent
Search
2000 character limit reached

Fast Greedy Evidence Maximization Algorithm

Updated 16 November 2025
  • The paper presents a fast greedy algorithm that maximizes the determinant of a Gram matrix by iteratively selecting vectors based on residual norms.
  • It achieves a composable coreset approximation guarantee through rigorous local swap-optimality bounds, significantly improving previous theoretical limits.
  • Optimizations such as incremental QR updates and lazy residual computations enable scalable performance in distributed and large-scale data settings.

The Fast Greedy Evidence Maximization Algorithm addresses the determinant maximization problem, fundamental to MAP inference for determinantal point processes (DPPs). Given a ground set of nn vectors in Rd\mathbb{R}^d, the task is to select kk vectors whose span maximizes the volume, equivalent to maximizing the determinant of the k×kk \times k Gram matrix of the selected vectors. Determinant maximization has practical importance in diversity modeling and large-scale data applications, making scalable and composable algorithms crucial. The Fast Greedy Evidence Maximization Algorithm, referred to hereafter as "Greedy", provides a scalable approach with strong theoretical and empirical guarantees in the composable coreset setting, where data is partitioned across multiple machines or shards.

1. Algorithmic Structure: Greedy Determinant Maximization

The Greedy procedure constructs a coreset C⊆PC \subseteq P of size kk by iteratively selecting vectors that maximize the incremental volume. The key selection criterion is the squared distance from the candidate vector pp to the span of CC, computed as r(p)2=∥(I−QQT)p∥22r(p)^2 = \|(I - QQ^T)p\|_2^2, where QQ is the orthonormal basis for the current coreset. The volume increment when adding Rd\mathbb{R}^d0 is given by Rd\mathbb{R}^d1, with Rd\mathbb{R}^d2 as the Gram matrix of Rd\mathbb{R}^d3. Each iteration orthonormalizes the chosen vector against Rd\mathbb{R}^d4 (e.g., via Gram–Schmidt) and updates Rd\mathbb{R}^d5.

Pseudocode: CC2 Tracking and updating residuals Rd\mathbb{R}^d6 and their squared norms enables computational efficiency.

2. Composable Coreset Guarantee: Approximation Bounds

In distributed or partitioned contexts, each subset Rd\mathbb{R}^d7 (e.g., data on one machine or shard) independently runs Greedy to produce a coreset Rd\mathbb{R}^d8 of size Rd\mathbb{R}^d9. The union kk0 (of size kk1, with kk2 subdivisions) serves as the candidate pool for a final Greedy round selecting kk3 vectors for the global coreset.

The algorithm achieves a composable-coreset approximation guarantee: kk4 where OPT is the maximum determinant achievable by any kk5-vector subset of kk6. In big-kk7 notation, this is kk8—a significant improvement over previous guarantees of kk9, aligning closely with local-search bounds k×kk \times k0 from earlier work.

This result follows directly by demonstrating k×kk \times k1-local optimality for Greedy, allowing application of general coreset approximation theorems for locally optimal sets.

3. Local-Optimality via Single-Swap Lemma

Central to the improved analysis is a swap (local optimality) lemma which asserts that exchanging any chosen point k×kk \times k2 in the Greedy solution k×kk \times k3 with any non-selected k×kk \times k4 yields an increase in volume by at most k×kk \times k5: k×kk \times k6 with volume defined as k×kk \times k7. The proof leverages orthogonalization (Gram–Schmidt), matrix determinant lemmas, and bounds on rank-one perturbations. This upper bound is tight up to the additive "+1"—exemplified by choices such as k×kk \times k8 and k×kk \times k9 for C⊆PC \subseteq P0, where swapping produces a volume increase by C⊆PC \subseteq P1.

4. Computational Complexity and Scalability

The naïve implementation incurs C⊆PC \subseteq P2 cost per round for recomputing all candidate residuals. Practical optimizations include:

  • Maintaining per-vector residuals C⊆PC \subseteq P3 and their squared norms.
  • Upon adding a new orthonormal direction C⊆PC \subseteq P4, each residual is updated by one dot-product and subtraction:

C⊆PC \subseteq P5

  • Initialization requires C⊆PC \subseteq P6; subsequent updates per round are C⊆PC \subseteq P7; maintaining C⊆PC \subseteq P8 is C⊆PC \subseteq P9 via modified Gram–Schmidt or Cholesky updates.
  • For kk0, kernel or low-rank optimizations yield per-iteration complexity kk1.
  • In very large-scale settings, sub-sampling or "lazy Greedy" prioritization heuristics (priority-week by residual norms) avoid full scans, reducing dot-products by factors of 10–50×.

Empirical runs process hundreds of thousands of points and kk2 in the hundreds within seconds on a single machine.

5. Empirical Evidence and Performance Assessment

Mahabadi et al. tested Greedy’s empirical swap-optimality on MNIST kk3 and GENES kk4, sampling workstreams of size kk5–kk6 and varying kk7 from 1 up to 300. The observed swap-factor kk8 was consistently below 1.5 even at kk9; for pp0, it remained under 1.4 on benchmarked data and ~1.2 on random sphere points. Variation in pp1 across orders of magnitude produced negligible impact on the worst-case swap factor. This suggests that, in practical applications, Greedy is not just provably pp2-locally-optimal but regularly outperforms theoretical bounds.

6. Implementation Recommendations and Optimization Strategies

Effective deployment of Greedy entails:

  • Regularly updating the orthonormal basis pp3 via incremental QR (e.g., modified Gram–Schmidt).
  • Storing and iteratively updating residuals pp4 and norms; after initial pp5 computation, next-point selection proceeds in pp6 per round.
  • For very large pp7, implement lazy updating via priority queues on candidate residual norms, recalculating only when necessary.
  • In cases where pp8 or the feature space is approximately low-rank, use kernel DPP approaches: maintain pp9 Gram matrices with rank-one Cholesky updates.
  • In composable or streaming-data settings, run Greedy independently on all shards/servers, then assemble the CC0-sized coresets by a final Greedy pass.

These techniques yield a highly scalable and near-optimal Greedy solver for DPP MAP inference; the composable-coreset quality approaches the CC1 theoretical bound, performing substantially better in empirical evaluations. A plausible implication is that, for practical deployment in distributed or streaming environments, Greedy offers an efficient balance of solution quality and computational tractability.

Topic to Video (Beta)

No one has generated a video about this topic yet.

Whiteboard

No one has generated a whiteboard explanation for this topic yet.

Follow Topic

Get notified by email when new papers are published related to Fast Greedy Evidence Maximization Algorithm.