Papers
Topics
Authors
Recent
Search
2000 character limit reached

Streaming L_p-Sampling Algorithms

Updated 16 January 2026
  • Streaming L_p-sampling algorithms are one-pass, space-efficient protocols that sample vector indices with probability proportional to |x_i|^p.
  • They employ randomized scaling, CountSketch recovery, and norm estimation to balance error rates and update times across various p regimes.
  • These methods enable practical applications like heavy hitters, duplicates detection, distributed monitoring, and regression coresets in data streams.

A streaming LpL_p-sampling algorithm is a one-pass, small-space protocol that, given turnstile or insertion-only updates to a vector x∈Rnx \in \mathbb{R}^n, returns index ii with probability exactly ∣xi∣p/∥x∥pp|x_i|^p/\|x\|_p^p, or with specified relative error, while failing rarely. These primitives are central to randomized tracking of "frequency moments" in streaming models, supporting tasks such as heavy hitters, duplicates, distributed monitoring, and online learning. The field has developed tight space, error, and time bounds for all parameter regimes of pp, leveraging probabilistic scaling, sparse recovery, linear sketches, precise order-statistics, and sampling/rejection frameworks.

1. Formal Problem Statement and Distributional Guarantees

Given an underlying vector x∈Rnx \in \mathbb{R}^n in a turnstile or insertion-only data stream, the LpL_p-distribution on [n][n] is

Pr⁡[i]=∣xi∣p∥x∥pp,∥x∥p=(∑i=1n∣xi∣p)1/p.\Pr[i] = \frac{|x_i|^p}{\|x\|_p^p}, \quad \|x\|_p = \left( \sum_{i=1}^n |x_i|^p \right)^{1/p}.

An LpL_p-sampling algorithm outputs x∈Rnx \in \mathbb{R}^n0 so that, conditioned on success,

x∈Rnx \in \mathbb{R}^n1

for any x∈Rnx \in \mathbb{R}^n2, failure probability at most x∈Rnx \in \mathbb{R}^n3 (Jowhari et al., 2010). For x∈Rnx \in \mathbb{R}^n4, the algorithm samples uniformly from x∈Rnx \in \mathbb{R}^n5. Algorithms may be required to be "perfect" (exactly matching x∈Rnx \in \mathbb{R}^n6 distribution up to x∈Rnx \in \mathbb{R}^n7 additive error), or "approximate" (allowing small multiplicative error).

2. Algorithmic Frameworks for x∈Rnx \in \mathbb{R}^n8

Space-optimal x∈Rnx \in \mathbb{R}^n9-sampling for ii0 is achieved via randomized scaling and heavy-hitter detection, typically combining:

A canonical method initializes random scaling parameters, runs CountSketches and norm estimators for recovery, and employs a statistical test to avoid ambiguous outputs. For ii6, optimal space is ii7 bits, perfect sampling is achievable in ii8 update time (Swartworth et al., 29 Nov 2025).

ii9 regime Space (bits) Update time Reference
∣xi∣p/∥x∥pp|x_i|^p/\|x\|_p^p0 ∣xi∣p/∥x∥pp|x_i|^p/\|x\|_p^p1 ∣xi∣p/∥x∥pp|x_i|^p/\|x\|_p^p2 (Jowhari et al., 2010, Swartworth et al., 29 Nov 2025)
∣xi∣p/∥x∥pp|x_i|^p/\|x\|_p^p3 ∣xi∣p/∥x∥pp|x_i|^p/\|x\|_p^p4 ∣xi∣p/∥x∥pp|x_i|^p/\|x\|_p^p5 (Jowhari et al., 2010)
∣xi∣p/∥x∥pp|x_i|^p/\|x\|_p^p6 ∣xi∣p/∥x∥pp|x_i|^p/\|x\|_p^p7 ∣xi∣p/∥x∥pp|x_i|^p/\|x\|_p^p8 (Jowhari et al., 2010)
∣xi∣p/∥x∥pp|x_i|^p/\|x\|_p^p9 pp0 pp1 (Jowhari et al., 2010)

For pp2, a logarithmic factor in space is currently unavoidable under present techniques (Swartworth et al., 29 Nov 2025).

3. Extensions: pp3 and General pp4-Sampling

For pp5, perfect pp6 sampling requires fundamentally more space. Recent advances employ a sampling-and-rejection method:

  • Instantiate pp7 independent perfect pp8 samplers.
  • For each pp9 sample, estimate x∈Rnx \in \mathbb{R}^n0 via CountSketch and approximate rejection step.
  • Accept with probability proportional to x∈Rnx \in \mathbb{R}^n1 (Woodruff et al., 9 Apr 2025).

Generalizations to non–scale-invariant x∈Rnx \in \mathbb{R}^n2 functions use similar sampling and rejection logic, extending to polynomials, logarithms, or capped powers, with perfect sampling in x∈Rnx \in \mathbb{R}^n3 bits for polynomials and x∈Rnx \in \mathbb{R}^n4 bits for x∈Rnx \in \mathbb{R}^n5 and x∈Rnx \in \mathbb{R}^n6 (Woodruff et al., 9 Apr 2025).

4. Derandomization, Update Complexity, and Lower Bounds

  • Derandomization: PRGs such as Nisan’s and Gopalan–Kane–Meka enable deterministic sampling with negligible increase in space (Jayaram et al., 2018, Swartworth et al., 29 Nov 2025).
  • Update time: Recent algorithms attain x∈Rnx \in \mathbb{R}^n7 per-update cost for x∈Rnx \in \mathbb{R}^n8 by simulating exponentials and their rescalings efficiently, using Fourier inversion and rapid quadrature (Swartworth et al., 29 Nov 2025).
  • Lower bounds: One-pass streaming algorithms for x∈Rnx \in \mathbb{R}^n9 sampling on LpL_p0 vectors require LpL_p1 bits, matching optimal algorithms. Heavy hitters and duplicates detection admit matching lower bounds LpL_p2 bits (Jowhari et al., 2010).
  • Truly perfect sampling: In the general turnstile model, achieving truly perfect sampling (LpL_p3) requires LpL_p4 space, but insertion-only or sliding-window models allow LpL_p5 bits for LpL_p6 (Jayaram et al., 2021).

5. Practical Implications and Continuous Sampling

In practical data streaming, LpL_p7-sampling algorithms enable:

  • Duplicates detection: LpL_p8 sampling gives LpL_p9 bits for one-pass detection, improving previous bounds.
  • Heavy hitters: Matching [n][n]0 upper and lower bounds in turnstile models.
  • Continuous sampling: Algorithms maintain a valid sample at each time, supporting applications in distributed monitoring and online statistics (Lin et al., 9 Aug 2025).
  • Regression and coresets: In turnstile matrix streams, [n][n]1 leverage score and row sampling enables streaming coreset constructions for regression and neural learning tasks (Munteanu et al., 2024).
  • Distributed monitoring: Perfect [n][n]2 sampling over [n][n]3 servers is resolved for all [n][n]4, with optimal [n][n]5 communication (Lin et al., 26 Oct 2025).

6. Open Questions and Future Directions

Major unresolved issues include:

  • Eliminating the remaining [n][n]6 factor for [n][n]7.
  • Achieving truly perfect sampling (zero additive error) in general turnstile or adversarial streams without linear space (Jayaram et al., 2021).
  • Reducing the polylogarithmic dependence on [n][n]8 in approximate sampling.
  • Extensions to more general non-linear update patterns, adaptive data streams, and privacy settings (Lin et al., 9 Aug 2025, Jayaram et al., 2021).

7. References and Historical Context

Initial upper bounds were given by Monemizadeh and Woodruff (2010) for approximate [n][n]9 sampling (Jowhari et al., 2010). Subsequent work by Andoni, Krauthgamer, and Onak improved the space-time complexity (Jowhari et al., 2010). The first perfect Pr⁡[i]=∣xi∣p∥x∥pp,∥x∥p=(∑i=1n∣xi∣p)1/p.\Pr[i] = \frac{|x_i|^p}{\|x\|_p^p}, \quad \|x\|_p = \left( \sum_{i=1}^n |x_i|^p \right)^{1/p}.0 samplers for Pr⁡[i]=∣xi∣p∥x∥pp,∥x∥p=(∑i=1n∣xi∣p)1/p.\Pr[i] = \frac{|x_i|^p}{\|x\|_p^p}, \quad \|x\|_p = \left( \sum_{i=1}^n |x_i|^p \right)^{1/p}.1 were constructed by Jayaram and Woodruff, with later derandomization and efficient update-time results (Jayaram et al., 2018, Swartworth et al., 29 Nov 2025). For Pr⁡[i]=∣xi∣p∥x∥pp,∥x∥p=(∑i=1n∣xi∣p)1/p.\Pr[i] = \frac{|x_i|^p}{\|x\|_p^p}, \quad \|x\|_p = \left( \sum_{i=1}^n |x_i|^p \right)^{1/p}.2, Woodruff, Xie, and Zhou established new tight bounds for perfect and polynomial samplers (Woodruff et al., 9 Apr 2025). Practical frameworks for regression were developed with Pr⁡[i]=∣xi∣p∥x∥pp,∥x∥p=(∑i=1n∣xi∣p)1/p.\Pr[i] = \frac{|x_i|^p}{\|x\|_p^p}, \quad \|x\|_p = \left( \sum_{i=1}^n |x_i|^p \right)^{1/p}.3 leverage-score sampling in turnstile matrix models, yielding the first streaming coresets for logistic regression (Munteanu et al., 2024). Distributed monitoring with adversarial robustness is now solved up to log factors for all Pr⁡[i]=∣xi∣p∥x∥pp,∥x∥p=(∑i=1n∣xi∣p)1/p.\Pr[i] = \frac{|x_i|^p}{\|x\|_p^p}, \quad \|x\|_p = \left( \sum_{i=1}^n |x_i|^p \right)^{1/p}.4 (Lin et al., 26 Oct 2025).

Streaming Pr⁡[i]=∣xi∣p∥x∥pp,∥x∥p=(∑i=1n∣xi∣p)1/p.\Pr[i] = \frac{|x_i|^p}{\|x\|_p^p}, \quad \|x\|_p = \left( \sum_{i=1}^n |x_i|^p \right)^{1/p}.5-sampling has been central in closing the gap between theoretical lower bounds and real-time data analytics, delivering near-optimal algorithms for a spectrum of streaming and monitoring settings.

Topic to Video (Beta)

No one has generated a video about this topic yet.

Whiteboard

No one has generated a whiteboard explanation for this topic yet.

Follow Topic

Get notified by email when new papers are published related to Streaming $L_p$-Sampling Algorithms.