Sharp dimension-dependent rates in unbalanced nonlinear latent factor models

Characterize the sharp dependence on the respective dimensions of the left and right factors in the estimation error rates for generalized latent factor models with nonlinear links and missing observations when the matrix dimensions are unbalanced.

Background

The paper’s theory reaches the fixed-rank degrees-of-freedom sampling scale up to logarithmic factors in the balanced regime, where the two matrix dimensions are comparable. The authors point out that existing singular-subspace perturbation theory for linear models suggests different error-rate scalings for the left and right factors when one dimension is much larger than the other.

The unresolved problem is to establish the corresponding sharp dimension dependence in the present nonlinear setting, specifically for regimes in which n is much larger than p or p is much larger than n.

References

Several paths are left open for further work: (i) such nonlinearity-encoded factor structure is also natural for symmetric network data as well as tensor data, thus it would be interesting to see whether we can have computationally tractable and statistically optimal procedures for such models; (ii) While our analysis reaches the fixed-rank degrees-of-freedom sampling scale up to logarithmic factors in the balanced regime $(n \approx p)$, the singular subspace perturbation theory for the linear models \citep{cai2018rate,zhang2022heteroskedastic,cai2021subspace} suggests that, in the unbalanced regime ($n \gg p$ or $n \ll p$), the error rates for the left and right factors should scale according to their respective dimensions. Characterizing the sharp dependence on these dimensions in the present nonlinear setting is an important question for future study.

From Good Starts to Optimal Inference: Generalized Latent Factor Models with Missingness and Implicit Regularization  (2609.11740 - Huang et al., 10 Sep 2026) in Section 7, Conclusion and Discussion