Papers
Topics
Authors
Recent
Search
2000 character limit reached

Parameter estimation and application in two types of uncertain single-index models

Published 6 Jul 2026 in stat.ME and stat.AP | (2607.04699v1)

Abstract: Uncertain data often arises in complex environments because of frequency instability and subjective judgment. This paper establishes two types of uncertain single-index models to capture the inherent properties of such data. Based on the semiparametric least-squares principle, the Nadaraya-Watson kernel and B-spline methods are used to estimate the unknown coefficients in various scenarios with both crisp and imprecise explanatory variables. Residual analysis and hypothesis testing under uncertainty assess the fit of the proposed models. Furthermore, simulation studies verify the models' validity, and a real-data application demonstrates their effectiveness in practical settings.

Authors (2)

Summary

  • The paper introduces two uncertain single-index models (USIC and USIU) that extend traditional SIMs to imprecise data using Liu's uncertainty theory.
  • It employs a profile least-squares method with kernel smoothing and B-spline estimation to infer nonparametric link functions and coefficient vectors.
  • Simulation studies and a real weather data application demonstrate robust parameter recovery and improved model fit over conventional techniques.

Parameter Estimation and Application in Two Types of Uncertain Single-Index Models

Introduction

The paper addresses the challenge of modeling and inference in high-dimensional settings where data are characterized by imprecision due to frequency instability or subjectivity. Traditional single-index models (SIMs) are well-established within a probabilistic framework for crisp data, but their theoretical and applied advantages do not extend directly to settings where observations are uncertain rather than random. To this end, the authors develop two distinctive types of uncertain single-index models (USIC and USIU) grounded in Liu's uncertainty theory, which replaces probability with uncertain measure to systematically handle imprecise data.

Uncertain Single-Index Models: Formulation

The foundation is built upon Liu's uncertainty theory, utilizing uncertain variables and their distributions to formalize imprecise observations. Two types of models are introduced:

  • USIC Model (Uncertain Single-Index Model with Crisp Covariates):

y~i=g(βTxi)+ϵi\tilde{y}_i = g(\bm{\beta}^T \bm{x}_i) + \epsilon_i

where xi\bm{x}_i is a crisp predictor, y~i\tilde{y}_i and ϵi\epsilon_i are uncertain variables.

  • USIU Model (Uncertain Single-Index Model with Uncertain Covariates):

y~i=g(βTx~i)+ϵi\tilde{y}_i = g(\bm{\beta}^T \tilde{\bm{x}}_i) + \epsilon_i

where both y~i\tilde{y}_i and x~i\tilde{\bm{x}}_i are uncertain variables.

Both models utilize a semiparametric structure, encapsulating high-dimensional complexity within a nonparametric link function gg and a low-dimensional linear index characterized by β\bm{\beta}.

Estimation Methodology

Profile Least Squares for USIC

Parameter estimation is formulated via a profile least-squares criterion in the space of uncertain variables, reducing ultimately to minimization over the expected responses:

(g^,β^)=argming,βi=1n(E[y~i]g(βTxi))2(\hat{g}, \hat{\bm{\beta}}) = \arg\min_{g, \bm{\beta}} \sum_{i=1}^n (E[\tilde{y}_i] - g(\bm{\beta}^T \bm{x}_i))^2

The estimation of xi\bm{x}_i0 is addressed via an uncertain version of the Nadaraya–Watson (N–W) kernel estimator:

xi\bm{x}_i1

The estimation proceeds through a two-step iterative procedure: alternately updating xi\bm{x}_i2 and xi\bm{x}_i3 until convergence.

For kernel and bandwidth selection, the derivative of the standard uncertain normal distribution is used as kernel, and bandwidth is optimized via V-fold cross-validation and the Fibonacci search algorithm.

B-spline Estimation for USIU

For the USIU model, with imprecise covariates, explicit monotonicity constraints on xi\bm{x}_i4 are imposed. The function xi\bm{x}_i5 is approximated using B-splines:

xi\bm{x}_i6

Parameter estimation is performed by minimizing:

xi\bm{x}_i7

with monotonicity enforced on the derivative of the spline approximation. Cross-validation is similarly used for selection of spline basis number.

Residual Analysis and Hypothesis Testing

For both models, estimation of the expectation and variance of residuals is developed—enabling both model fit assessment and the construction of uncertain hypothesis tests on disturbance terms or coefficient significance, using uncertainty theory analogues of classical techniques.

Simulation Studies

The authors conduct extensive simulation experiments under both the USIC (crisp covariate) and USIU (all variables uncertain) settings. Results include:

  • USIC Simulation: Estimated coefficients (xi\bm{x}_i8, xi\bm{x}_i9, y~i\tilde{y}_i0) closely match true values. The optimal kernel bandwidth is y~i\tilde{y}_i1. Residual expectation and variance are y~i\tilde{y}_i2, y~i\tilde{y}_i3. The model passes the uncertainty-based residual distribution hypothesis test with only 4/500 residuals in the rejection region.
  • USIU Simulation: For cubic B-splines and optimal number of basis functions y~i\tilde{y}_i4, coefficient estimates (y~i\tilde{y}_i5, y~i\tilde{y}_i6, y~i\tilde{y}_i7) and spline coefficients are presented. Residual analysis yields y~i\tilde{y}_i8, y~i\tilde{y}_i9, and the model again passes the uncertainty residual test.

Simulated confidence intervals for predictions are constructed using the derived uncertainty distributions of the predicted outcomes.

Real-Data Application

Application to a weather dataset (Lagos, Nigeria) demonstrates empirical utility. The daily temperature range is modeled as a function of precipitation, windspeed, humidity, sea-level pressure, and cloud cover. Key findings:

  • The main driver is humidity (ϵi\epsilon_i0).
  • Secondary factors include sea-level pressure (ϵi\epsilon_i1) and cloud cover (ϵi\epsilon_i2).
  • The estimated link function ϵi\epsilon_i3 displays a marked threshold between radiation-dominated and convection-dominated meteorological regimes, consistent with physical understanding of tropical climate dynamics.
  • The model fits exhibit a residual variance ϵi\epsilon_i4, outperforming parametric alternatives for ϵi\epsilon_i5 (e.g., quadratic, exponential).
  • Residual-based uncertain hypothesis tests confirm the adequacy of the model and validity of residual distributional assumptions.

An uncertain significance test is employed to evaluate whether small estimated coefficients (specifically ϵi\epsilon_i6) are statistically nonzero. The corresponding model passes the acceptance region criteria under both null and alternative hypotheses per uncertainty theory-based decision criteria.

Theoretical and Practical Implications

The proposed methodology extends the SIM framework to settings where data are inherently imprecise or subjective, providing estimation and inference tools that respect the underlying epistemic uncertainty. The flexibility of the nonparametric link function avoids model mis-specification and accommodates complex nonlinear relationships under uncertainty, as evidenced by enhanced fit relative to fully parametric specifications.

Theoretical innovations include:

  • The adaptation of kernel and spline techniques for parameter estimation in uncertain settings.
  • Derivation of explicit forms for inverse uncertainty distributions for imprecise functional arguments.
  • Uncertainty theory-based hypothesis testing mechanisms for both model fit and coefficient significance.

Practically, the methodology enables robust modeling for high-dimensional, nonlinear systems where conventional statistical methods are inadequate due to data imprecision—relevant in fields such as environmental modeling, biomedicine, and finance.

Future Developments

Possible next steps include extending estimation and inference theory to uncertain partially linear single-index models and uncertain varying-coefficient single-index models, enhancing flexibility for complex, heterogeneous systems. Theoretical developments on robustness, finite-sample properties, and optimization under different types of imprecise data (interval, fuzzy, hybrid) remain open directions. Methodological integration with scalable algorithms (for large-scale or streaming uncertain data) is also warranted to meet the demands of contemporary AI and data-intensive applications.

Conclusion

This work establishes a comprehensive inference framework for single-index modeling under uncertainty theory, accommodating both crisp and uncertain data. By employing kernel smoothing and B-spline methods within a profile least-squares architecture—coupled with systematic residual analysis and hypothesis testing—the proposed uncertain SIMs offer a flexible, robust, and theoretically principled tool for modeling complex systems where imprecision is intrinsic to the data. Empirical validation on both simulated and real-world data reinforces the model's applicability and practical effectiveness.


Reference: "Parameter estimation and application in two types of uncertain single-index models" (2607.04699)

Paper to Video (Beta)

No one has generated a video about this paper yet.

Whiteboard

No one has generated a whiteboard explanation for this paper yet.

Open Problems

We haven't generated a list of open problems mentioned in this paper yet.