Papers
Topics
Authors
Recent
Search
2000 character limit reached

Probabilistic PROTES Method

Updated 2 February 2026
  • The Probabilistic PROTES method is a black-box optimization framework that leverages tensor-train representations to efficiently explore vast discrete spaces.
  • It models the search distribution using low-parametric TT formats, enabling tractable sampling and gradient updates without exhaustively enumerating the combinatorial grid.
  • Empirical results show that PROTES outperforms classical discrete optimizers on challenging benchmarks like QUBO and binary control problems, effectively reducing the curse of dimensionality.

Probabilistic PROTES Method

The Probabilistic PROTES method (PROTES: Probabilistic Optimization with Tensor Sampling) is a black-box optimization approach targeting extremely high-dimensional discrete spaces by leveraging probabilistic sampling from low-parametric tensor-train (TT) representations. PROTES is specifically designed for minimizing an objective function defined on a Cartesian product grid, efficiently handling settings with up to 21002^{100} candidates, such as binary optimization and discretized control problems. The core innovation is expressing and manipulating the search distribution in TT format to bypass the curse of dimensionality, enabling effective exploration and exploitation in combinatorial and control domains (Batsheva et al., 2023).

1. Optimization Problem and Motivation

PROTES addresses the black-box minimization problem: min⁡xf(x),x=(n1,...,nd),ni∈{1,...,Ni}\min_{x} f(x), \quad x=(n_1, ..., n_d), \quad n_i \in \{1, ..., N_i\} where ff is an expensive black-box function, and the search space forms a dd-dimensional grid of size N1×⋯×NdN_1 \times \cdots \times N_d. As this product explodes combinatorially for large dd or NiN_i, brute-force search is infeasible. Existing heuristics such as evolutionary algorithms, PSO, or CMA-ES become ineffective or inapplicable due to the extreme dimensionality and discrete structure. PROTES overcomes these limitations by parametrizing an adaptable probability distribution P(x)P(x) via a compact TT decomposition, enabling tractable sampling and distribution updates even when ∏iNi\prod_i N_i is astronomically large (Batsheva et al., 2023).

2. Tensor-Train Representation of Discrete Distributions

The search distribution P∈R+N1×⋯×NdP \in \mathbb{R}_+^{N_1 \times \cdots \times N_d} is modeled in the TT format: min⁡xf(x),x=(n1,...,nd),ni∈{1,...,Ni}\min_{x} f(x), \quad x=(n_1, ..., n_d), \quad n_i \in \{1, ..., N_i\}0 where each core tensor min⁡xf(x),x=(n1,...,nd),ni∈{1,...,Ni}\min_{x} f(x), \quad x=(n_1, ..., n_d), \quad n_i \in \{1, ..., N_i\}1 controls the min⁡xf(x),x=(n1,...,nd),ni∈{1,...,Ni}\min_{x} f(x), \quad x=(n_1, ..., n_d), \quad n_i \in \{1, ..., N_i\}2th dimension, and min⁡xf(x),x=(n1,...,nd),ni∈{1,...,Ni}\min_{x} f(x), \quad x=(n_1, ..., n_d), \quad n_i \in \{1, ..., N_i\}3 are TT ranks with min⁡xf(x),x=(n1,...,nd),ni∈{1,...,Ni}\min_{x} f(x), \quad x=(n_1, ..., n_d), \quad n_i \in \{1, ..., N_i\}4. The number of parameters grows only as min⁡xf(x),x=(n1,...,nd),ni∈{1,...,Ni}\min_{x} f(x), \quad x=(n_1, ..., n_d), \quad n_i \in \{1, ..., N_i\}5 for uniform ranks min⁡xf(x),x=(n1,...,nd),ni∈{1,...,Ni}\min_{x} f(x), \quad x=(n_1, ..., n_d), \quad n_i \in \{1, ..., N_i\}6. This format enables efficient storage and scalable manipulation of min⁡xf(x),x=(n1,...,nd),ni∈{1,...,Ni}\min_{x} f(x), \quad x=(n_1, ..., n_d), \quad n_i \in \{1, ..., N_i\}7, which otherwise would be intractable for large min⁡xf(x),x=(n1,...,nd),ni∈{1,...,Ni}\min_{x} f(x), \quad x=(n_1, ..., n_d), \quad n_i \in \{1, ..., N_i\}8. In practice, sampling and updates operate with this TT factorization, sidestepping explicit enumeration over hyper-exponential cardinality (Batsheva et al., 2023).

3. Sampling and Update Algorithm

PROTES repeatedly draws samples from the current TT-modeled distribution, using a sequential conditional algorithm adapted from Dolgov & Savostyanov (2020):

  • Forward/backward message computation: For each TT-core, partial contraction ("messages") min⁡xf(x),x=(n1,...,nd),ni∈{1,...,Ni}\min_{x} f(x), \quad x=(n_1, ..., n_d), \quad n_i \in \{1, ..., N_i\}9 (forward) and ff0 (backward) are evaluated to obtain marginals and conditional probabilities for efficient sampling of each coordinate sequentially.
  • Sequential sampling: Each ff1 is sampled conditional on the previously chosen coordinates, with explicit distributions computed from the TT structure and current ff2, ff3 values.
  • Batch sampling: ff4 independent samples ff5 are generated per iteration. The computational cost is ff6, with ff7 and ff8 the cost of categorical sampling over ff9 values.

After batch evaluation,

  • The top-dd0 sample indices with lowest dd1 are selected as the elite set.
  • The TT parameters dd2 are updated via dd3 steps of Adam (or any gradient optimizer) on the loss:

dd4

This corresponds to a REINFORCE-style policy gradient, weighted by elite selection (Batsheva et al., 2023).

Complete Iterative Scheme (Pseudocode)

P(x)P(x)2 This process requires only black-box function evaluations and can enforce additional constraints by constraining the support of the initial TT cores (Batsheva et al., 2023).

4. Computational Complexity and Scalability

Each PROTES iteration consists of:

  • Sampling: dd5
  • Gradient steps: dd6 per step, dd7 per iteration
  • Total cost over dd8 function evaluations: dd9

This scaling is essentially linear in dimension N1×⋯×NdN_1 \times \cdots \times N_d0 (assuming fixed N1×⋯×NdN_1 \times \cdots \times N_d1), and polynomial in N1×⋯×NdN_1 \times \cdots \times N_d2, a dramatic reduction from exponential growth in naive discrete optimization. For moderate to large N1×⋯×NdN_1 \times \cdots \times N_d3, PROTES remains computationally tractable where other discrete optimizers are not (Batsheva et al., 2023).

5. Theoretical Foundations and Relation to Policy-Gradient Methods

The update rule of PROTES can be derived from maximizing the expected reward N1×⋯×NdN_1 \times \cdots \times N_d4, with N1×⋯×NdN_1 \times \cdots \times N_d5 a Fermi–Dirac (sharpened) function,

N1×⋯×NdN_1 \times \cdots \times N_d6

In the zero-temperature limit (N1×⋯×NdN_1 \times \cdots \times N_d7), N1×⋯×NdN_1 \times \cdots \times N_d8 selects only samples close to the empirical minimum, yielding the empirical top-N1×⋯×NdN_1 \times \cdots \times N_d9 aggregation. The gradient update thus becomes a hard-selection analogue of the REINFORCE estimator, concentrating probability mass on promising regions while maintaining exploration. This perspective clarifies why the TT-parameterized search distribution is suitable for black-box optimization with no derivative information about dd0 (Batsheva et al., 2023).

6. Empirical Results and Performance

In comprehensive experiments, PROTES was evaluated on:

  • Analytic 7D benchmark functions (Ackley, Rastrigin, Schwefel) on grids up to dd1
  • Four dd2-bit QUBO instances (Max-Cut, Vertex Cover, Knapsack)
  • Binary optimal control problems for dd3 (search space up to dd4)
  • Constrained binary control (e.g., "at least three ones," encoded via an indicator TT)

With hyperparameters dd5, dd6, dd7, dd8, dd9, NiN_i0, PROTES found the minimal known value in NiN_i1 out of NiN_i2 cases, consistently outperforming both TT-based (TTOpt, Optima-TT) and classical discrete optimizers in the nevergrad suite (PSO, CMA-ES, Differential Evolution, NoisyBandit, Portfolio). Convergence was typically faster in terms of NiN_i3 versus number of objective evaluations (Batsheva et al., 2023).

7. Strengths, Limitations, and Application Domains

Strengths:

  • Bypasses the curse of dimensionality for large discrete domains by TT factorization
  • No need for objective gradients; only black-box evaluations required
  • Structured constraints (e.g., combinatorial restrictions) easily incorporated by modifying initial TT support
  • Demonstrated robust performance on combinatorial, control, and synthetic problems

Limitations:

  • TT-rank NiN_i4, sample size NiN_i5, and elite size NiN_i6 may need tuning per problem
  • Sampling and autodifferentiation through TT is nontrivial for very large NiN_i7 and/or NiN_i8; more advanced manifold optimization (e.g., Riemannian methods) may be preferable for NiN_i9
  • The method is currently more expensive per iteration than simpler heuristics if P(x)P(x)0 or P(x)P(x)1 grow large

Applications:

PROTES is applicable to high-dimensional combinatorial optimization (QUBO, graph partitioning), black-box parameter tuning in machine learning, discrete control, resource allocation, and any setting with Cartesian product structure and latent low-rank solution geometry (Batsheva et al., 2023).

Definition Search Book Streamline Icon: https://streamlinehq.com
References (1)

Topic to Video (Beta)

No one has generated a video about this topic yet.

Whiteboard

No one has generated a whiteboard explanation for this topic yet.

Follow Topic

Get notified by email when new papers are published related to Probabilistic PROTES Method.