---
title: 'Oxytrees: Efficient Proxy-Based Biclustering Trees'
url: https://www.emergentmind.com/topics/oxytrees
type: topic
---

# Oxytrees: Efficient Proxy-Based Biclustering Trees

Oxytrees are proxy-based biclustering model trees designed for efficient and interpretable bipartite learning—predicting binary interactions for pairs from two feature domains. They address scale, inductive generalization, and interpretability in settings where interactions are organized as sparse matrices, such as drug–target, RNA–disease, or regulatory networks. Oxytrees improve upon previous biclustering forests by compressing the interaction matrix into row and column proxies, enabling fast construction and inference with shallower trees whose leaves host expressive linear models parameterized by Kronecker product kernels [2511.12713].

## 1. Bipartite Learning Setting and Motivation

Bipartite learning seeks to infer a function $Y^{ij} = f(x_1^i, x_2^j) \in \{0,1\}$ given samples $X_1 = \{x_1^i\}_{i=1}^{n_1} \subseteq \mathbb{R}^{m_1}$ and $X_2 = \{x_2^j\}_{j=1}^{n_2} \subseteq \mathbb{R}^{m_2}$, forming a sparse bipartite interaction matrix $Y \in \{0,1\}^{n_1 \times n_2}$. Challenges arise from:

- **Scalability**: $n_1$ and $n_2$ may both reach thousands.
- **Generalization**: Models must predict for unseen $x_1, x_2$ instances (inductive setting).
- **Interpretability**: The goal is a two-dimensional partitioning—biclustering—that remains comprehensible and locally predictive.

Previous state-of-the-art biclustering forests such as BICTR scale as $\Theta(m n^2 \log n)$ and utilize constant-value leaf predictions, limiting both efficiency and model expressivity. Oxytrees integrate proxy-based acceleration, efficient batch inference, and model tree structure with linear (kernelized) leaves to address these limitations [2511.12713].

## 2. Proxy-based Biclustering and Split Scoring

Oxytrees exploit the property that impurity measures on interaction submatrices can be reformulated in a proxy-friendly form:
$$ I(Y_\text{node}) = \rho\left( \sum_{i,j} \mu(Y^{ij}_\text{node}) \right) $$
For variance impurity, $\mu(z) = (1, z, z^2)$ and $\rho(u_1, u_2, u_3) = u_3/u_1 - (u_2/u_1)^2$. This enables computation of two proxy matrices per node:
- **Row proxy:** $\tilde Y_1^i = \sum_j \mu(Y^{ij}_\text{node})$
- **Column proxy:** $\tilde Y_2^j = \sum_i \mu(Y^{ij}_\text{node})$

Horizontal splits use prefix sums of $\tilde Y_1$, while vertical splits use $\tilde Y_2$, ensuring that all split candidates per feature are evaluated in $O(1)$ per threshold. The computational complexity is reduced to $\Theta(n_1^2(\log n_1 + m))$, a theoretical and observed $\log(n)$–$m$ fold speedup relative to traditional biclustering trees [2511.12713].

## 3. Model Tree Construction and Leaf Modeling

Oxytrees construct a binary tree top-down. At each node, $m$ features are randomly sampled from $X_1$ or $X_2$, akin to extremely randomized trees. Each candidate split is evaluated for impurity reduction $\Delta I$, computed expediently via the proxy mechanism. Splits with the maximum $\Delta I$ are chosen, and growth halts if $n_\text{node}$ falls below a minimum or $\Delta I \leq 0$.

Distinctly, Oxytrees assign each leaf a local response matrix $Y_\text{leaf}$ and fit a linear model parameterized by a Kronecker product kernel. This approach improves expressivity over mean-value leaves and yields shallower, more efficient trees [2511.12713].

## 4. Efficient Batch Leaf Assignment and Inference

For prediction over Cartesian products $X_{1,\text{test}} \times X_{2,\text{test}}$, Oxytrees avoid redundant traversals. The inference procedure is:

- Horizontal splits partition $X_{1,\text{test}}$, passing all of $X_{2,\text{test}}$ downstream.
- Vertical splits partition $X_{2,\text{test}}$, passing all of $X_{1,\text{test}}$ downstream.
- At each leaf, all predictions for $(X_{1,\text{leaf}}, X_{2,\text{leaf}})$ are computed in batch via matrix multiplication.

This routine requires each test instance be visited only once per tree depth, with empirical complexity $\Theta(n_\text{test}^2)$, compared to $\Theta(n_\text{test}^2 \log n_\text{train})$ for previous state-of-the-art methods [2511.12713].

## 5. Kronecker Product Kernel Models in Leaves

Each leaf utilizes Regularized Least Squares regression with a Kronecker product kernel $K((x_1, x_2), (x_1', x_2')) = K_1(x_1, x_1') \otimes K_2(x_2, x_2')$. Given training Gram matrices $\Phi_1$ and $\Phi_2$ for leaf instances, eigendecompositions $\Phi_{1} = U_1 \operatorname{diag}(\lambda_{1}) U_1^T$ and $\Phi_{2} = U_2 \operatorname{diag}(\lambda_{2}) U_2^T$ produce the optimal weight matrix:
$$
W = U_2 \left[ \Lambda \odot (U_2^T Y U_1) \right] U_1^T, \quad \Lambda_{ij} = \frac{1}{\alpha + \lambda_{2,i} \lambda_{1,j}}
$$
Predictions for test batch similarity matrices $\Phi_{1,\text{test}}$, $\Phi_{2,\text{test}}$ are:
$$
\widehat Y_{\text{test}} = \Phi_{2,\text{test}} W \Phi_{1,\text{test}}^T
$$
This structure allows efficient, batch-wise computation for any leaf-level domain sizes [2511.12713].

## 6. Empirical Results and Performance Benchmarks

Oxytrees were evaluated on 15 bipartite datasets, including drug–nuclear receptor, kinase inhibitors, and lncRNA–disease, with sizes ranging from $30 \times 50$ to $2000 \times 400$ and densities 1–20%. Baselines included MLPs on concatenated features, local models (RLS-avg, BLMNII, WkNNIR), RLS-Kron, NRLMF, and BICTR.

- **Validation** encompassed instance-wise + dyad-wise splits yielding transductive (TD), semi-inductive (LT/TL), and fully inductive (TT) test sets; positive-unlabeled (PU) masking at 0–75%.
- **Metrics** included AUROC, AUPRC, and Friedman + Nemenyi tests.

Key outcomes:

- Predictive performance: Oxytrees match or outperform BICTR and RLS-Kron, notably in inductive scenarios.
- Efficiency: Oxytrees train $\approx$35$\times$ faster and predict $\approx$10$\times$ faster than BICTR (p < $10^{-5}$).
- Model parsimony: Oxytrees reach 98% accuracy with $\sim$42% fewer trees due to expressive leaves.
- Proxy mechanism and leaf model ablation confirm that these components are crucial; removing leaf models degrades performance.
- Robustness: Maintained accuracy under PU-masking and with increased minimum leaf size [2511.12713].

## 7. Software Implementation and Reproducibility

A Python package "bipartite_learn" provides the Oxytrees framework, including:

- `OxytreeClassifier` class with `fit(X1, X2, Y)`, `predict(X1, X2)`, `predict_proba()`, and `score()`
- hyperparameters for number of trees, minimum leaf size, feature subsample size, impurity, regularization ($\lambda$), and kernel functions $K_1$, $K_2$
- efficient batch inference
- access to all 15 preprocessed datasets and evaluation/PU-masking utilities

All results are reproducible via open-source code, with tutorials and Jupyter notebooks available at [https://github.com/pedroilidio/oxytrees2025] [2511.12713].

| Component          | Characteristic                                 | Complexity/Efficiency    |
|--------------------|------------------------------------------------|--------------------------|
| Split Scoring      | Proxy row/col matrices ($\tilde{Y}_{1,2}$)     | $O(n_1^2(\log n_1 + m))$ |
| Leaf Model         | Kronecker-kernel RLS                           | Efficient per-leaf SVD   |
| Inference          | Batch, single-pass per test set                | $O(n_{\text{test}}^2)$   |

Oxytrees thereby constitute an interpretable, fast, and accurate methodology for inductive bipartite learning, combining proxy-based biclustering, expressive kernelized leaf modeling, and scalable batch inference [2511.12713].

Source: https://www.emergentmind.com/topics/oxytrees