---
title: Two-Step Sieve ML Estimator
url: https://www.emergentmind.com/topics/two-step-sieve-ml-estimator
type: topic
---

# Two-Step Sieve ML Estimator

to=arxiv_search  天天中彩票软件  天天中彩票不中返json
{"query":"\"two-step sieve\" estimator maximum likelihood semiparametric copula control function arXiv","max_results":10,"sort_by":"relevance","sort_order":"descending"}【อ่านข้อความเต็มjson to=arxiv_search 
to=arxiv_search ,最新高清无码专区 经彩票json
{"query":"2401.17334", "max_results": 5, "sort_by":"relevance", "sort_order":"descending"}
A two-step sieve ML estimator is a semiparametric estimation procedure in which an infinite-dimensional nuisance object is approximated by a finite-dimensional sieve, while estimation itself is split into two stages: a preliminary estimator or nuisance fit is obtained first, and a second-stage likelihood, pseudo-likelihood, or empirical-risk problem is then solved over the sieve space. In recent arXiv work, this architecture is used to estimate marginal parameters in multivariate models with unknown dependence [2401.17334], average treatment effects under single-index semiparametric models [2207.03281], and quantile regression coefficients in a triangular system with endogeneity and measurement error [2605.20601]. A closely related line of work combines a first-stage machine-learning regression with a second-stage series projection onto a data-adaptive sieve, yielding efficient plug-in estimators for pathwise differentiable functionals [2003.01856].

## 1. Core two-step architecture

Across these formulations, the defining feature is the separation between nuisance recovery and final estimation. In the multivariate semiparametric model of marginal distributions and unknown copula density, the first step ignores dependence and computes the quasi-MLE
\[
\hat\theta^{(0)}=\arg\max_{\theta\in\Theta}\ell_{\mathrm{marg}}(\theta),
\]
whereas the second step jointly maximizes a pseudo-log-likelihood over the marginal parameter and sieve copula weights [2401.17334]. In the single-index treatment-effect model, the first step estimates the propensity-score and outcome-regression nuisance functions by sieve-M estimation, and the second step plugs those fits into a Robinson-type least-squares estimator for the effect parameter \(\alpha\) [2207.03281]. In endogenous quantile regression with measurement error, the first step constructs a control function \(\hat V\), either parametrically or by series regression, and the second step maximizes a sieve likelihood that integrates the generated control through copula weights [2605.20601].

| Setting | Step 1 | Step 2 |
|---|---|---|
| Multivariate marginals with unknown copula | QMLE from \(\ell_{\mathrm{marg}}(\theta)\) | Joint sieve MLE over \((\theta,w)\) |
| Single-index ATE model | Sieve-ML for \((\beta,g_1)\) and \((\gamma,g_2)\) | Plug-in least-squares estimator of \(\alpha\) |
| Endogenous quantile regression with measurement error | Estimate control function \(\hat V\) | Sieve likelihood maximization over \((b,\sigma,\gamma)\) |
| ML + series plug-in framework | Flexible ML fit \(\theta_n^0\) | Empirical-risk projection onto a data-adaptive sieve |

This recurring architecture reflects a common statistical objective: retain flexibility in the nuisance structure while recovering root-\(n\) or efficiency properties for a finite-dimensional target. A plausible implication is that the “two-step” designation refers less to a single algorithm than to a design pattern for semiparametric estimation.

## 2. Sieve construction and parameterization

The sieve component differs across models, but in each case it converts an infinite-dimensional problem into a constrained finite-dimensional one. In the multivariate copula model, the unknown copula density \(c(u)\) is approximated by a Bernstein–Kantorovich polynomial copula,
\[
c_K(u;w)=\sum_{k_1=0}^K\cdots\sum_{k_d=0}^K w_{k_1\cdots k_d}\prod_{j=1}^d B_{k_j,K}(u_j),
\]
with nonnegativity, summation, and uniform-marginal constraints on the weight tensor \(w\). The resulting sieve copula has \((K+1)^d-d(K+1)+d-1\) free parameters [2401.17334].

In the treatment-effect model, the nuisance link functions \(g_{1,0}\) and \(g_{2,0}\) lie in \(L^2(\mathbb{R},\pi)\), with \(\pi(t)=\exp(-t^2/2)\), and are expanded in the Hermite basis. Truncating the expansion at \(k_n-1\) yields finite-dimensional sieve spaces
\[
S_{\ell,k_n}=\mathrm{span}\{h_0,\ldots,h_{k_n-1}\}, \qquad \ell=1,2.
\]
Identifiability is enforced through \(\|\beta\|=1\), \(\|\gamma\|=1\), a nonnegative first coordinate, and a bound \(\|C\|\le M\) on the sieve coefficient vector [2207.03281].

In endogenous quantile regression with measurement error, each quantile coefficient function \(\beta_k(u)\) is approximated by a spline of order \(r\) with \(J_n\) knots on \([0,1]\), using B-spline basis functions. The second-stage parameter is \(\theta=(b,\sigma,\gamma)\), where \(b\) indexes the sieve approximation, \(\sigma\) parameterizes the measurement-error density, and \(\gamma\) parameterizes the copula \(f_c(u\mid v;\gamma)\). The parameter space \(\Theta_n\) includes a monotonicity restriction requiring \(x'\beta(u)\) to be increasing in \(u\) for all \(x\) [2605.20601].

The ML + series framework generalizes the same logic by defining a data-adaptive sieve
\[
\Theta_n=\mathrm{Span}\{\psi_1,\ldots,\psi_K\}\circ \theta_n^0,
\]
where \(\theta_n^0\) is a first-stage ML fit and the basis is applied to the fitted values rather than directly to covariates. This is still a sieve, but one whose geometry is induced by the initial estimator [2003.01856].

## 3. Estimation criteria in the second step

The second-step criterion is model-specific, but it always exploits the sieve approximation to recover information discarded or inaccessible in the first step. In the multivariate marginal model, the pseudo-log-likelihood is
\[
\ell_n(\theta,w)=\sum_{i=1}^n\left[\log c_K\!\bigl(F_1(Y_{i1};\theta_1),\ldots,F_d(Y_{id};\theta_d);w\bigr)+\sum_{j=1}^d\log f_j(Y_{ij};\theta_j)\right].
\]
Starting at \(\hat\theta^{(0)}\), the estimator solves
\[
(\hat\theta,\hat w)=\arg\max_{\theta\in\Theta,\;w\in W_K}\ell_n(\theta,w),
\]
with \(W_K\) the convex polytope of admissible copula weights [2401.17334].

In the single-index ATE model, the first-stage nuisance fits are obtained from two separate likelihood-type objectives: a logistic criterion for the propensity score and a least-squares criterion for the outcome regression. The second-stage estimator then uses
\[
\hat\alpha=
\frac{\sum_{i=1}^n\bigl[Y_i-\hat g_{2,k_n}(X_i'\hat\gamma)\bigr]\bigl[D_i-\hat e(X_i)\bigr]}
{\sum_{i=1}^n\bigl[D_i-\hat e(X_i)\bigr]^2},
\]
which is the simple least-squares estimator arising after partialling out the estimated nuisance components [2207.03281].

In endogenous quantile regression with measurement error, the second-stage pseudo-likelihood for observation \(i\) is
\[
\ell_i(\theta;\hat V_i)=\log\int_0^1 f_\epsilon(Y_i-X_i' b S(u);\sigma)\,f_c(u\mid \hat V_i;\gamma)\,du,
\]
and the estimator maximizes the empirical average over \(\Theta_n\) [2605.20601]. In the ML + series framework, the analogous second step is empirical-risk minimization over the data-adaptive sieve, followed by the plug-in estimator \(\hat\Psi_n=\Psi(\theta_n^*)\) [2003.01856].

These constructions show that “ML” in the term need not always denote the same object. In some formulations it is genuine maximum likelihood or pseudo-likelihood; in others it appears as “likelihood-type” nuisance estimation followed by a plug-in target estimator.

## 4. Asymptotic theory and efficiency

The central theoretical motivation for two-step sieve ML methods is that they can preserve semiparametric flexibility without forfeiting asymptotic efficiency. In the multivariate copula model, the first-step QMLE is consistent but inefficient, whereas the sieve MLE is \(\sqrt n\)-consistent and asymptotically normal with covariance \(\Sigma_{\mathrm{eff}}=I_{\mathrm{eff}}^{-1}\), and this covariance attains the semiparametric efficiency bound [2401.17334]. The efficient score takes the form
\[
S_{\mathrm{eff}}(Y)=s_{\mathrm{marg}}(Y)+\frac{1}{c_0(u)}g^*(u),
\]
where \(g^*\) solves a least-squares projection problem in the nuisance tangent space.

In the single-index ATE model, standard sieve-M-estimation theory yields
\[
\|\hat\beta-\beta_0\|+\|\hat\gamma-\gamma_0\|=O_p(k_n^{-m}+\sqrt{k_n/n}),
\]
and similarly for the \(L^2\) error of the estimated link functions. Under conditions C1–C6, the effect estimator satisfies
\[
\sqrt n(\hat\alpha-\alpha)\to_d N(0,V), \qquad V=\sigma^2\Psi^{-1},
\]
with influence function \(\psi_i=((D_i-e(X_i))\epsilon_i)/\Psi\) [2207.03281].

The ML + series framework makes the efficiency statement especially explicit. Under Conditions A1–A3 and C1–C4, the plug-in estimator admits the expansion
\[
\hat\Psi_n-\Psi(\theta_0)=\frac{1}{n}\sum_{i=1}^n IF(V_i)+o_p(n^{-1/2}),
\]
and the influence function coincides with the canonical gradient under the nonparametric model, implying asymptotic efficiency [2003.01856].

In endogenous quantile regression with measurement error, the asymptotic normalization is \(\sqrt{n\kappa_{J_n}}(\hat\theta_n-\theta_0)\), reflecting the ill-posedness encoded by the minimal eigenvalue \(\kappa_{J_n}\). Consistency requires \(J_n\to\infty\) and \(J_n/n\to 0\); asymptotic normality requires additional growth conditions linking \(J_n\), \(\kappa_{J_n}\), and the ordinary-smoothness order of the measurement error [2605.20601].

## 5. Empirical behavior and applications

The empirical results in the cited literature emphasize both efficiency gains and robustness. In simulations for semiparametric multivariate models with bivariate exponential marginals and Gaussian, Clayton, Frank, or Plackett copulas, asymptotic relative efficiency satisfies \( \mathrm{ARE}\approx 1 \) under independence. With strong negative dependence, \(\mathrm{ARE}(\mathrm{FMLE}/\mathrm{QMLE})\) is reported up to \(5\)–\(6\), while \(\mathrm{ARE}(\mathrm{SMLE}/\mathrm{QMLE})\) is up to \(1.5\)–\(2\); in three dimensions, similar patterns hold, although computational cost grows roughly as \(O(K^d)\) [2401.17334].

The same paper reports two substantive applications. In an insurance context with one censored variable, the sieve MLE produces tighter parameter estimates. In weekly \(5\%\) Value-at-Risk forecasting for Bank of America stock returns, with models involving trading-volume change, realized volatility, or both, exceedance rates are controlled at approximately \(5\%\) for all methods, but the sieve MLE achieves higher average censored log-scores than both QMLE and parametric FMLE, with mean gain \(0.007\) and \(t\)-statistic \(>5\) in the bivariate case and mean gain \(0.002\) and \(t\)-statistic \(>2\) in the trivariate case [2401.17334].

For the single-index ATE model, the finite-sample performance is evaluated through simulation studies and an empirical example [2207.03281]. For endogenous quantile regression with measurement error, Monte Carlo simulations show that the proposed estimator markedly reduces bias relative to existing methods, and the bootstrap is proposed for inference [2605.20601].

A recurring pattern is that the second step is most valuable when the first step deliberately sacrifices structure for robustness or tractability. This suggests why gains are minimal under independence but can be substantial when dependence, endogeneity, or latent nuisance structure is pronounced.

## 6. Computation, tuning, and methodological scope

The computational burden of two-step sieve ML methods is driven by basis dimension, constraints, and the need to propagate first-stage uncertainty. In the multivariate copula setting, \(K\) may be selected by AIC, BIC, or cross-validation, with theory requiring \(K_n\to\infty\) and \(K_n^d/n\to 0\). Initialization may use \(\hat\theta^{(0)}\) and either uniform weights or empirical copula histograms, while optimization may proceed via block-coordinate ascent or general-purpose constrained optimization such as interior-point or sequential quadratic programming [2401.17334].

In the single-index ATE model, practical optimization uses Lagrange-multiplier formulations, Newton–Raphson or quasi-Newton solvers, projection of the index vector onto the unit sphere, and alternating updates for \(\gamma\) and the least-squares coefficients [2207.03281]. In the ML + series framework, the sieve dimension \(K\) may be chosen by standard \(k\)-fold cross-validation on the series risk, and under an additional balanced-approximation condition the cross-validated choice still yields an efficient estimator [2003.01856]. In endogenous quantile regression with measurement error, the second stage uses Gauss–Legendre quadrature for the latent-\(u\) integral, gradient-based optimizers such as interior-point or trust-region methods, and either analytic sandwich variance estimation or nonparametric pairs bootstrap or weighted bootstrap [2605.20601].

Two misconceptions are addressed directly by the cited work. First, a two-step procedure is not necessarily statistically inefficient: several of these estimators are root-\(n\) consistent and asymptotically efficient under their stated conditions [2401.17334; 2003.01856]. Second, full parametric likelihood is not automatically preferable: in the copula model, full MLE is efficient only under a correct copula specification and may be biased if the copula is misspecified, whereas the sieve MLE is introduced precisely to improve over QMLE without that drawback [2401.17334].

A related but distinct methodology is the multi-step MLE process for ergodic diffusion, where a preliminary estimator from a short learning interval is followed by one-step and two-step score corrections. That construction is asymptotically efficient, but it is based on repeated Le Cam one-step updates rather than on sieve approximation [1504.01869]. Its relevance is conceptual: it shows that the broader logic of a rough first step followed by a refined second step is not confined to sieve-based estimation.

Source: https://www.emergentmind.com/topics/two-step-sieve-ml-estimator