---
title: 'RDIS Algorithm: Optimization, Depth & Imputation'
url: https://www.emergentmind.com/topics/rdis-algorithm
type: topic
---

# RDIS Algorithm: Optimization, Depth & Imputation

The acronym RDIS refers to three distinct algorithms and frameworks in the literature, each addressing a separate core problem: (1) Recursive Decomposition for Nonconvex Optimization; (2) Relative Depth in Stereo for monocular depth estimation pretraining; (3) Random Drop Imputation with Self-Training for incomplete time series imputation. The following article focuses on each framework in turn, providing precise definitions, core methodology, and empirical findings.

## 1. Recursive Decomposition for Nonconvex Optimization

### 1.1. Problem Formulation

RDIS solves the global optimization problem $\min_{x \in \mathbb{R}^n} f(x)$ where $f$ is continuously differentiable and possesses at least one global minimizer $x^*$ with finite $f^* = f(x^*) > -\infty$. The variable indices are denoted $I = \{1,2,\ldots,n\}$. For any subset $C \subseteq I$, the variable block $x$ is partitioned into $x_C$ and $x_U$ with $U = I \setminus C$. A partial assignment $\rho_C$ fixes $x_C$; $f|_{\rho_C}(x_U)$ denotes the function with $x_C$ fixed [1611.02755].

### 1.2. Recursive Decomposition Strategy

RDIS alternates two phases:
- **Value Selection:** Select a cutset $C$ of variables (via hypergraph partitioning of the factorized $f$), then optimize $f$ over $C$ (holding $U$ fixed) to obtain an assignment $\rho_C$.
- **Decomposition:** Simplify $f|_{\rho_C}(x_U)$ by omitting or approximating negligible terms. Identify $k > 1$ independent subfunctions $\{f_i(x_{U_i})\}$ on disjoint $U_i$, then recurse on each.

This approach exploits local separability after key variable assignments, similar to DPLL-style SAT solvers and recursive conditioning in inference [1611.02755].

### 1.3. Variable Selection via Hypergraph Partitioning

For $f(x) = \sum_{i=1}^m f_i(x)$, define a hypergraph $H = (V,E)$ with one vertex per term and one hyperedge per variable, connecting all terms involving a variable. A $k$-way partition minimizes the number of cut hyperedges under balance, yielding a small $C$ whose assignment decomposes the residual function. Tools such as PaToH are used for hypergraph partitioning [1611.02755].

### 1.4. Pseudocode and Algorithmic Structure

The main RDIS pseudocode proceeds as follows:
1. Choose cutset $C$ via hypergraph cut.
2. For each restart, partition $x^*$, optimize $C$ via a user-chosen nonconvex subspace optimizer $S$, yielding assignment $\rho_C$.
3. Simplify $f|_{\rho_C}(x_U)$ using a tolerance $\epsilon$.
4. Decompose the simplified function into independent components and recurse on each.
5. Update the global record if a new best function value is found.
6. Terminate according to a preset criterion, e.g., fixed outer restarts or full variable assignment [1611.02755].

### 1.5. Theoretical Guarantees

Assuming at every recursion a decomposition into $k > 1$ subproblems and cut-block size $d=|C|$, let $\xi(d)$ be the number of subspace optimizer calls required on $d$ variables. Recurrence analysis yields:
$$
T(n) = O\big((k \cdot \xi(d))^{\log_k(n/d)}\big) = O\left(\frac{n}{d}\, \xi(d)^{\log_k(n/d)}\right)
$$
This result demonstrates exponential speedups over grid search or random restart descent, under mild technical conditions on $S$ for global convergence. For $\epsilon = 0$ and $S$ satisfying Armijo/gradient-norm decrease, all limit points are stationary and global optimality is achieved with high probability under random restarts [1611.02755].

### 1.6. Use of Standard Optimizers

RDIS is agnostic to the choice of local optimizer $S$ at recursion leaves. $S$ may be gradient descent with restart, Levenberg–Marquardt, or similar; the optimizer focuses on the chosen cutset block, treating other variables as fixed. If the remaining variable set is empty, RDIS directly applies $S$ to the full initial problem [1611.02755].

### 1.7. Empirical Evaluation

Empirical results demonstrate RDIS's advantage in several domains:
- **Structure from Motion:** On bundle adjustment (up to 23,000 variables), RDIS consistently finds lower reprojection error than Levenberg–Marquardt (LM) and block-coordinate LM, with advantages increasing at scale.
- **Highly Multimodal Synthetic Functions:** RDIS outperforms conjugate gradient and block variants by orders of magnitude in objective value and time.
- **Protein Sidechain Placement:** On 21 proteins (up to 943 variables), RDIS attains lower energy than CGD and BCD-CGD, with simplification tolerance $\epsilon$ trading speed and final energy.

The combination of intelligent cutset selection, function simplification, and recursive decomposition yields exponential gains over standard multistart or block-coordinate techniques [1611.02755].

## 2. RDIS Dataset and Method for Monocular Depth Estimation

### 2.1. Dataset Construction and Label Semantics

The RDIS dataset is built from 70 rectified 3D movies, yielding 97,652 stereo keyframes. Semi-Global Matching (SGM) is used to compute dense disparity maps, with subsequent boundary correction and quality control. Relative depth ground-truth is encoded as ordinal relationships on point pairs: "closer" ($r=+1$), "farther" ($r=-1$), or "equal" ($r=0$), based on a thresholded disparity difference [1806.00585].

### 2.2. Network Architecture

The approach employs a "wide" ResNet (seven units) with pre-activation BatchNorm–ReLU style, leveraging ImageNet+Places365 pretraining. The network head is configured for regression ($C_{out}=1$) in pretraining and per-pixel classification ($C_{out}=B$) in finetuning [1806.00585].

### 2.3. Pretraining on Ordinal Depth

Pretraining utilizes sampled ordinal pairs with a ranking loss:
$$
E_k=
\begin{cases}
\log(1+e^{-z_{i_k}+z_{j_k}}), & r_k=+1,\\
\log(1+e^{z_{i_k}-z_{j_k}}), & r_k=-1,\\
(z_{i_k}-z_{j_k})^2, & r_k=0
\end{cases}
$$
Empirically, $K=1000$ pairs per image optimizes transfer performance [1806.00585].

### 2.4. Depth as Classification and Information Gain Loss

Finetuning discretizes depth into $B$ bins (log-space), with network outputs as per-pixel logits. The multinomial logistic loss is modulated by an information gain matrix:
$$
L_{\log} = -\frac1N\sum_{i=1}^N \sum_{d=1}^B H(D^*_i,d)\log P(d|z_i)
$$
with $H(p,q)=\exp[-\alpha(p-q)^2]$. This allows near-correct predictions to be weighted, improving gradient signal for ambiguous cases [1806.00585].

### 2.5. Evaluation and Ablation

Ablation studies confirm benefits of RDIS pretraining, network width, and the information gain loss. The full method achieves state-of-the-art metrics on NYU v2 and KITTI, outperforming prior methods on root-mean-squared error, relative error, log error, and accuracy within thresholds. The pretraining enables generalization to relative-depth benchmarks (DIW test: WHDR $0.197$, improving on prior best of $0.215$) [1806.00585].

## 3. Random Drop Imputation with Self-Training (Time Series)

### 3.1. Problem Context and Main Steps

Given incomplete multivariate time series $X \in \mathbb{R}^{T \times D}$ with original mask $M \in \{0,1\}^{T \times D}$, RDIS trains an imputation model $F(X;\theta)$ to estimate the unobserved entries. The methodology consists of:
- **Random-Drop Imputation (RDI):** Mask a random subset of observed entries, optimize $F$ to recover them using explicit loss.
- **Self-training:** Train an ensemble of $N$ models. The ensemble is used to generate pseudo-labels for the original missing entries, filtered by prediction-variance, and used for further fine-tuning [2010.10075].

### 3.2. Explicit Imputation and Self-training Losses

The loss on each random-dropped instance is:
$$
L_{\mathrm{impute}}(\tilde{X},\tilde{M}) = \|X \odot \Delta M - F(\tilde{X};\theta)\odot \Delta M\|_2 + \|\tilde{X} \odot \tilde{M} - F(\tilde{X};\theta)\odot \tilde{M}\|_2
$$
where $\Delta M$ indicates the newly dropped entries. The self-training loss incorporates pseudo-value targets $\hat X$ at confident unobserved positions (variance threshold $\tau$):
$$
L_{\mathrm{self},k} = \|\hat{X}\odot U - F_k(\tilde{X}_k;\theta_k) \odot U\|_2 + \|X\odot M - F_k(\tilde{X}_k;\theta_k)\odot M\|_2
$$
with $U = 1-M$ [2010.10075].

### 3.3. Pseudocode and Model-Agnosticity

The RDIS framework supports any $F$ (e.g., GRU, Bi-GRU, Transformer, TCN, GAN). The pseudocode details alternating RDI training (explicit dropout and recovery) and periodic self-training cycles (ensemble pseudo-label generation, entropy filtering, model update). Ensemble size, drop probability, entropy threshold, and update-frequency are principal hyperparameters [2010.10075].

### 3.4. Empirical Validation

Empirical comparisons on the Air Quality (11 variables, 48 time steps) and Gas Sensor (19 variables) datasets show that RDIS (with Bi-GRU) delivers minimal mean squared error among baselines (e.g., at 50% missing, BRITS: $0.0338$ vs. RDIS(Bi-GRU): $0.0277$). Ablative comparisons indicate that ensemble RDI and the self-training stage both confer performance advantages; the largest gains accrue at elevated missing rates ($\geq 50\%$) [2010.10075].

### 3.5. Practical Considerations

Utilizing an ensemble increases computational cost and necessitates careful tuning of drop rate, entropy threshold, and update frequency. Pseudo-label quality hinges on ensemble variance estimation, and RDIS only imputes point estimates; extensions for full predictive distributions are left to further work [2010.10075].

## 4. Comparison of RDIS Methodologies

| Context                                | RDIS Meaning                                              | Core Principle            |
|-----------------------------------------|-----------------------------------------------------------|---------------------------|
| Nonconvex Optimization                  | Recursive Decomposition for Nonconvex Optimization        | Divide-and-conquer opt.   |
| Depth Estimation                        | Relative Depth in Stereo Dataset and Pretraining          | Ordinal pretraining       |
| Time Series Imputation                  | Random Drop Imputation with Self-Training                 | Explicit drop + ensembling|

All RDIS variants emphasize decomposition (explicit or statistical) and leverage auxiliary structures (graph partitioning, ordinal structure, ensemble consensus) to enhance model performance in the respective domains. Each achieves state-of-the-art or competitive empirical results in its target application [1611.02755, 1806.00585, 2010.10075].

## 5. Significance and Research Impact

RDIS as recursive decomposition for nonconvex optimization has advanced scalable global optimization by importing principles from combinatorial problem solving and outperforming prior continuous methods across vision and molecular modeling tasks [1611.02755]. As a depth estimation pretraining method, RDIS democratizes dense ordinal-depth supervision, bridging the scarcity of metric ground truth and producing robust monocular predictors [1806.00585]. In imputation, RDIS delivers explicit supervision and confidence-calibrated pseudo-labels for incomplete time series, outperforming strong baselines at high missingness [2010.10075]. The common thread is the augmentation of training with problem-informed structure, be it decompositional, ordinal, or variance-filtered.

Source: https://www.emergentmind.com/topics/rdis-algorithm