---
title: Sieve Maximum Pseudo-Likelihood Estimation
url: https://www.emergentmind.com/topics/sieve-maximum-pseudo-likelihood-estimation-smple
type: topic
---

# Sieve Maximum Pseudo-Likelihood Estimation

Searching arXiv for the cited SMPLE papers and closely related terminology to ground the article in current arXiv records.
arxiv_search(query="\"Finite sample analysis of profile M-estimation in the Single Index model\" OR \"Recurrent Event Analysis with Ordinary Differential Equations\" OR \"Sieve Maximum Pseudo-Likelihood Estimation\"", max_results=10, sort_by="relevance")
Sieve Maximum Pseudo-Likelihood Estimation (SMPLE) is a sieve-based M-estimation methodology in which a pseudo-likelihood or quasi-likelihood is optimized jointly over a finite-dimensional parameter of interest and a growing finite-dimensional approximation to an infinite-dimensional nuisance component. In the arXiv literature represented here, SMPLE appears in two technically distinct but structurally related settings: sieve profile quasi maximum likelihood estimation for the single-index model with a nonparametric link approximated by \(C^3\)-Daubechies wavelets [1406.4052], and sieve maximum pseudo-likelihood for recurrent event data in which the conditional mean function is modeled through an ordinary differential equation and the nuisance functions are approximated by B-splines under an NHPP working model [2507.20396]. Across these formulations, the common principle is to replace an infinite-dimensional nuisance problem by a sequence of finite-dimensional sieve spaces, optimize an empirical pseudo-likelihood, and derive finite-sample or asymptotic guarantees for the target finite-dimensional parameter.

## 1. Conceptual definition and scope

SMPLE denotes estimation by maximizing a pseudo-likelihood over a sieve parameterization of nuisance functions. In the recurrent-event formulation, the method is introduced explicitly as a “Sieve Maximum Pseudo-Likelihood Estimation (SMPLE) method, employing the NHPP as a working model” [2507.20396]. In the single-index formulation, the paper describes a “sieve profile quasi maximum likelihood estimator” based on least squares, and states that “the paper’s ‘sieve profile quasi-MLE’ is an SMPLE” because it maximizes a quasi-likelihood over a sieve for the nonparametric link and profiles out that link to estimate the index parameter [1406.4052].

The “pseudo-likelihood” terminology is essential. In the single-index model, the criterion is least squares and “This is quasi-MLE because \(\epsilon\) need not be Gaussian; as M-estimation it plays the role of a pseudo-likelihood” [1406.4052]. In the recurrent-event model, the pseudo-likelihood is built from the NHPP log-likelihood even when the true process may be non-Poisson, so the estimator targets the parameters that best fit the conditional mean ODE “in the pseudo-likelihood sense (M-estimation)” [2507.20396].

This suggests that SMPLE is not tied to a single stochastic model class. Rather, it is a general estimation strategy for semiparametric problems in which the parameter of interest is finite-dimensional, the nuisance structure is infinite-dimensional, and a working likelihood is available but may be misspecified.

## 2. SMPLE in the single-index model

In the single-index setting, the starting point is the regression model
\[
Y_i = f(X_i^\top \theta) + \epsilon_i,
\]
with \(\theta \in S_1^{p,+} = \{\theta \in \mathbb{R}^p: \|\theta\| = 1, \theta_1 > 0\}\) and i.i.d., centered, sub-Gaussian errors independent of \(X_i\) [1406.4052]. The paper defines the link through conditional expectation, \(f(t) := E[g(X)\mid X^\top \theta = t]\), and uses the “true index” \(\theta^*\) to characterize the target [1406.4052].

The sieve approximation uses a \(C^3\)-Daubechies wavelet basis with compact support on \([ -s_X, s_X]\). Writing \(\{\psi_k\}_{k\ge 1}\) for the induced orthonormal basis on \(L^2([ -s_X, s_X])\), the link is approximated by
\[
f_\eta(t) := \sum_{k=1}^{m} \eta_k \psi_k(t),
\]
with sieve dimension \(m\) and resolution level \(J\) satisfying \(m \asymp 2^J\) [1406.4052]. The empirical criterion is
\[
L_m(\theta,\eta) := - \sum_{i:\|X_i\|\le s_X} |Y_i - f_\eta(X_i^\top \theta)|^2 / 2.
\]
The joint sieve maximizer is
\[
\tilde \upsilon_m := \arg\max_{(\theta,\eta)\in \Upsilon_m} L_m(\theta,\eta), \qquad \tilde \theta_m := \Pi_\theta \tilde \upsilon_m,
\]
where \(\Upsilon_m := S_1^{p,+} \times B_{r^{\circ}}^m\) and the Euclidean ball constraint on \(\eta\) is used to “avoid boundary issues” [1406.4052]. The profiled objective is
\[
M_n(\theta) := \sup_{\eta \in \mathbb{R}^m:\|\eta\|\le r^{\circ}} L_m(\theta,\eta),
\]
with
\[
\hat \theta := \arg\max_{\theta \in S_1^{p,+}} M_n(\theta), \qquad \hat f := \arg\max_{\eta:\|\eta\|\le r^{\circ}} L_m(\hat \theta,\eta).
\]

The theoretical targets distinguish the infinite-dimensional oracle
\[
\upsilon^* := \arg\max_{(\theta,\eta)\in S_1^{p,+}\times l^2} E L(\theta,\eta)
\]
from the finite-sieve target
\[
\upsilon_m^* := \arg\max_{(\theta,\eta)\in S_1^{p,+}\times \mathbb{R}^m} E[L_m(\theta,\eta)].
\]
This separation makes explicit the two sources of deviation: stochastic estimation error and sieve approximation bias [1406.4052].

## 3. SMPLE for recurrent event analysis via ODEs

In the recurrent-event formulation, SMPLE is built around the conditional mean function
\[
\mu_x(t) = E[N(t)\mid X=x],
\]
modeled as the solution to the ODE
\[
\begin{cases}
\mu'_x(t) = f\{t,\mu_x(t),x\},\\
\mu_x(0)=0,
\end{cases}
\]
with focus on the class
\[
f\{t,\mu_x(t),x\} = \alpha(t)\exp(x^\top \beta)q(\mu_x(t)),
\]
or equivalently
\[
\mu'_x(t) = \exp\big(x^\top \beta + \gamma(t) + g(\mu_x(t))\big),
\]
where \(\gamma(t)=\log \alpha(t)\) and \(g(u)=\log q(u)\) [2507.20396].

Under the NHPP working model, the conditional intensity is set equal to the mean rate,
\[
\lambda_x(t) = \mu'_x(t) = \exp\big(x^\top \beta + \gamma(t) + g(\mu_x(t))\big),
\]
and the pseudo-log-likelihood becomes
\[
\ell_n(\beta,\gamma,g)
=
\frac{1}{n}\sum_{i=1}^n
\left\{
\int_0^{C_i}
\big[X_i^\top \beta + \gamma(t) + g(\mu_i(t))\big]\,dN_i(t)
-
\mu_i(C_i;\beta,\gamma,g)
\right\}.
\]
Here \(\mu_i(t;\beta,\gamma,g)\) solves the subject-specific ODE with \(X=X_i\) [2507.20396].

The sieve spaces are spline-based. For \(\gamma(t)\) on \([0,\tau]\), the paper chooses a polynomial spline space \(S_n(T_{K_n},K_n,p_1)\) with B-spline basis \(\{A_j(t)\}_{j=1}^{q_n^1}\) and representation
\[
\hat \gamma(t)=\sum_{j=1}^{q_n^1} a_j A_j(t).
\]
For \(g(u)\) on \([0,\mu]\), it uses \(S_n(T_{L_n},L_n,p_2)\) with basis \(\{B_j(u)\}_{j=1}^{q_n^2}\) and
\[
\hat g(u)=\sum_{j=1}^{q_n^2} b_j B_j(u)
\]
[2507.20396].

The sieve parameter is \(\theta=(\beta,a,b)\), subject to the identifiability constraints \(\beta_1=1\) and \(\gamma(t_0)=0\), equivalently
\[
\sum_{j=1}^{q_n^1} a_j A_j(t_0)=0.
\]
With these expansions, the objective becomes
\[
\ell_n(\beta,a,b)
=
\frac{1}{n}\sum_{i=1}^n
\left[
\int_0^{C_i}
\left\{
X_i^\top \beta + \sum_{j=1}^{q_n^1} a_j A_j(t) + \sum_{j=1}^{q_n^2} b_j B_j\big(\mu_i(t;\beta,a,b)\big)
\right\}dN_i(t)
-
\mu_i(C_i;\beta,a,b)
\right].
\]
No explicit penalization is required; instead, identifiability constraints are enforced directly [2507.20396].

## 4. Sieve construction, profiling, and identifiability

The two arXiv formulations differ in basis choice and data-generating structure, but they share a common architecture: the nuisance function is approximated by a finite-dimensional sieve, and the target finite-dimensional component is extracted either by profiling or by joint optimization.

In the single-index paper, smoothness of the true wavelet coefficients is imposed through
\[
\sum_{k=1}^\infty k^{2\alpha}(\eta_k^*)^2 \le C_{\|\eta^*\|}^2 < \infty, \qquad \alpha>2,
\]
together with bounded derivatives \(\|f'\|_\infty < \infty\) and \(\|f''\|_\infty < \infty\) [1406.4052]. Identifiability is enforced through the unit half-sphere restriction on \(\theta\), the lower bound on the design density \(p_X\), the conditional variance lower bound for orthogonal projections, and a salience condition requiring a region on which \(|f'(X^\top \theta^*)| > c_{f'} > 0\) [1406.4052]. The paper also allows model bias through
\[
\|E[g(X)\mid X^\top \theta^*]-g(X)\| \le C_{\mathrm{bias}} \ge 0,
\]
supplemented by curvature conditions away from \(\theta^*\) [1406.4052].

In the recurrent-event paper, identifiability arises because \(\alpha(\cdot)\) and \(q(\cdot)\) are both unspecified. The imposed normalizations \(\beta_1=1\) and \(\gamma(t_0)=0\) anchor the scale of the parametric and nonparametric components [2507.20396]. Regularity assumptions include compact covariate support, covariate density bounded below, non-singularity of \(P(XX^\top)\), independent censoring given \(X\), Hölder smoothness for \(\gamma_0\) and \(g_0\), and the technical conditions (C1)–(C9), including a lower bound on \(\mathrm{Var}(X\mid U(t),V)\), absence of strong collinearity between functional derivatives with respect to \(\gamma\) and \(g\), existence of least favorable directions, and non-singularity of the information-type matrix \(A\) [2507.20396].

A useful contrast is that the single-index paper develops the SMPLE within a profile M-estimation framework using the conditions of Andresen A. and Spokoiny V., including AssId, \((LL_0)\), \((CS\,DF_0)\), \((CS\,DF_1)\), \((CS\,r)\), and \((cc(L)\,r)\) [1406.4052]. The recurrent-event paper instead formulates the nuisance structure through ODE-constrained splines and develops semiparametric efficiency through least favorable directions and orthogonalized scores [2507.20396].

## 5. Finite-sample and asymptotic theory

For the single-index SMPLE, the paper derives nonasymptotic Wilks and Fisher expansions. Let \(D:=p+m\) and define the profiled information matrix
\[
D_{P,m}^{-2} := \Pi_\theta DF_m(\upsilon_m^*)^{-2}\Pi_\theta^\top,
\]
together with the profile score
\[
\xi_m := \nabla_\theta \zeta(\upsilon_m^*) - A_m H_m^{-2}\nabla_\eta \zeta(\upsilon_m^*),
\]
where \(\zeta:=L-E_\epsilon L\) [1406.4052]. Under the stated assumptions and sample-size constraints, the paper proves that, for \(n\) large enough and with probability at least
\[
1 - 12 e^{-x} - \exp\{-m^3 x\} - \exp\{-n c_{(Q)}/4\},
\]
\[
\big| 2L_r(\tilde \theta_m,\theta_m^*) - \|\xi_m\|^2 \big|
\le
C\big[\sigma\sqrt{p+x} + \mathfrak{E}(x)\big]\mathfrak{E}(x),
\]
and
\[
\|D_{P,m}(\upsilon_m^*)(\tilde \theta_m-\theta_m^*)-\xi_m\|
\le
\mathfrak{E}(x),
\]
with
\[
\mathfrak{E}(x):=
\mathrm{CONST}_{\diamond}(D^{5/2}+C_{\mathrm{bias}}D^{7/2}+x)/\sqrt{n}
\]
[1406.4052]. The corresponding sieve bias bound is
\[
\alpha(m) \le C\sqrt{n}\big[m^{-(\alpha+1/2)} + C_{\mathrm{bias}}m^{-(\alpha-1)}\big].
\]

When \(C_{\mathrm{bias}}=0\) and \(D^{5/2}/\sqrt{n}\to 0\), these finite-sample results imply the asymptotic limits
\[
D_P(\tilde \theta_m-\theta^*) \Rightarrow N(0,\sigma^2 I_p), \qquad
2L_r(\tilde \theta_m,\theta^*) \Rightarrow \chi_p^2,
\]
and, in the correctly specified Gaussian regression case,
\[
\mathrm{Var}\big(\sqrt{n}(\tilde \theta_m-\theta^*)\big)=\sigma^2 D_P^{-2},
\]
which the paper states is efficient in the considered setting [1406.4052].

For the recurrent-event SMPLE, the asymptotic theory is organized differently. Under (C1)–(C7),
\[
d(\hat \theta_n,\theta_0)\xrightarrow{P}0,
\]
where \(d(\cdot,\cdot)\) is defined through the Euclidean norm for \(\beta\), \(L^2\)-distance for \(\gamma\), and an \(L^2\)-distance involving \(\zeta(t,x,\beta,\gamma)=g(\mu(t,x,\beta,\gamma,g))\) [2507.20396]. Under the conditions of Theorem 2 together with (C8)–(C9),
\[
\sqrt{n}(\hat \beta_n-\beta_0)
=
\sqrt{n}A^{-1}P_n\Xi^*(\beta_0,\gamma_0,\zeta_0;W) + o_p(1)
\Rightarrow
N(0,A^{-1}BA^{-T}),
\]
where
\[
\Xi^* = \ell'_\beta(\theta_0;W) - \ell'_\gamma(\theta_0;W)[\mathbf{v}^*] - \ell'_\zeta(\theta_0;W)[\mathbf{h}^*(\cdot,\beta_0,\gamma_0)]
\]
is the efficient influence function after orthogonalization with respect to the nuisance tangent directions [2507.20396].

Under the NHPP working model, the paper states that \(A=B=I(\beta_0)\), so
\[
\sqrt{n}(\hat \beta_n-\beta_0)\Rightarrow N\big(0,I^{-1}(\beta_0)\big),
\]
and the estimator “achieves the semiparametric efficiency bound” [2507.20396]. For the sieve components, the overall convergence rate satisfies
\[
d(\hat \theta_n,\theta_0)=O_p\big(n^{-\min\{p_1\nu_1,\;p_2\nu_2,\;[1-\max\{\nu_1,\nu_2\}]/2\}}\big),
\]
and in the balanced case \(p_1=p_2=p\), \(\nu_1=\nu_2=\nu=1/(1+2p)\),
\[
d(\hat \theta_n,\theta_0)=O_p\big(n^{-p/(1+2p)}\big)
\]
[2507.20396].

## 6. Algorithms, computation, and covariance estimation

The single-index paper analyzes an alternating maximization procedure. Initialization is performed on a coarse grid \(G_N=\{\theta_l\}_{l=1}^N\subset S_1^{p,+}\), with closed-form update
\[
\tilde \eta_l^{(0)}
=
\arg\max_\eta L_m(\theta_l,\eta)
=
\left[
\frac{1}{n}\sum_{i=1}^n \psi(X_i^\top \theta_l)\psi(X_i^\top \theta_l)^\top
\right]^{-1}
\left[
\frac{1}{n}\sum_{i=1}^n Y_i\psi(X_i^\top \theta_l)
\right],
\]
followed by iterative \(\eta\)-updates and \(\theta\)-updates [1406.4052]. The \(\eta\)-updates are linear least squares in wavelet coefficients, whereas the \(\theta\)-updates are nonconvex [1406.4052]. Under the stated conditions, the paper proves high-probability deviation bounds for the alternating sequence and a convergence-to-global-maximizer result via Theorem 2.4 of AASPalternating [1406.4052]. It also reports that each \(\eta\)-update solves an \(m\times m\) linear system with cost “\(\sim O(nm^2+m^3)\)” [1406.4052].

The recurrent-event paper uses block coordinate ascent together with ODE-based gradients. Given
\[
f(t;\mu,\theta)=\exp\!\left(X^\top \beta + \sum a_j A_j(t) + \sum b_j B_j(\mu(t))\right),
\]
the gradient requires the sensitivity ODE
\[
\begin{cases}
G'(t)=\dfrac{\partial f}{\partial \theta}(t;\mu,\theta)+\dfrac{\partial f}{\partial \mu}(t;\mu,\theta)G(t),\\
G(0)=0,
\end{cases}
\]
where \(G(t)=\partial \mu(t;\theta)/\partial \theta\) [2507.20396]. A central computational device is the time transformation
\[
\tilde t = e^{X^\top \beta}\int_0^t \exp\!\left(\sum_{j=1}^{q_n^1} a_j A_j(s)\right)\,ds,
\qquad
\tilde \mu(\tilde t)=\mu(t),
\]
which yields
\[
\frac{d\tilde \mu(\tilde t)}{d\tilde t}
=
\exp\!\left(\sum_{j=1}^{q_n^2} b_j B_j(\tilde \mu(\tilde t))\right),
\]
“independent of \(X\), enabling a single ODE solve” [2507.20396]. The optimization alternates between updating \((a,b)\) subject to \(\gamma(t_0)=0\) and updating \(\beta\) subject to \(\beta_1=1\), using a gradient-based optimizer [2507.20396].

Covariance estimation is also treated differently in the two settings. The recurrent-event paper develops a resampling procedure to estimate the asymptotic covariance matrix without closed-form least favorable directions. The core approximation is
\[
n^{-1/2}U_n(\hat \theta_n + n^{-1/2}Z) \approx A Z,
\]
with \(Z\) generated from a mean-zero distribution, so that \(A\) can be estimated row-wise by regression of perturbed scores on the perturbations [2507.20396]. This yields \(\hat \Sigma = \hat A^{-1}\hat B(\hat A^{-T})\) [2507.20396].

## 7. Applications, interpretation, and limitations

The single-index paper situates SMPLE within projection pursuit. After estimating one direction and link, it updates residuals by
\[
Y_i^{(l+1)} := Y_i^{(l)} - f_{\tilde \eta_{(l)}}(X_i^\top \tilde \theta_{(l)}),
\qquad
Y_i^{(1)}=Y_i,
\]
and derives the bound
\[
\sup_{x\in B_{s_X}(0)}
\left|
\sum_{l=1}^M f_{\tilde \eta_{(l)}}(x^\top \tilde \theta_{(l)})
-
\sum_{l=1}^M f_{\eta^*_{(l)}}(x^\top \theta^*_{(l)})
\right|
\le
C M \sqrt{m}\left[\frac{D^{7/2}+x}{n} + \frac{\sqrt{D+x}}{\sqrt{n}}\right]
\]
under the stated conditions [1406.4052]. The paper remarks that if
\[
b(M):=\left\|g-\sum_{l=1}^M f_{\eta^*_{(l)}}(\cdot^\top \theta^*_{(l)})\right\|_\infty
\]
decays suitably, then the total approximation error improves over the full multivariate nonparametric rate \(n^{-\alpha/(2\alpha+p)}\) [1406.4052].

The recurrent-event paper evaluates SMPLE through simulations and an ICU readmission application. It reports NHPP scenarios, gamma-frailty non-Poisson scenarios, and a real-data analysis of the MIMIC-III ICU readmission frequency dataset with “18,252 patients, 20,360 events” and “20 covariates after preprocessing” [2507.20396]. In that application, the fitted general linear transformation model with quartic B-splines yields estimated \(\log \alpha(t)\) and \(\log q(u)\) that “are not constant nor linear in \(\log t\), indicating that Cox- or AFT-type models are inadequate” [2507.20396].

Several limitations are explicit. In the single-index setting, smoothness conditions require \(\alpha>2\), and “\(\alpha > 14/3\) if \(C_{\mathrm{bias}} > 0\) is needed to force \(\alpha(m)\to 0\) sufficiently fast” [1406.4052]. The design assumptions include Lipschitz \(p_X\), bounded support, and conditional variance lower bounds [1406.4052]. In the recurrent-event setting, the pseudo-likelihood relies on the NHPP working model for construction, and under “heavy-tailed or highly dependent processes, variance estimation may require larger resampling \(B\)” [2507.20396]. Joint identifiability of \(\alpha\) and \(q\) requires constraints, and spline-dimension selection may require BIC or AIC tuning [2507.20396].

A common misconception is that pseudo-likelihood necessarily implies inefficiency. The two papers jointly show a more nuanced picture. In the single-index model, the profile quasi-MLE admits Wilks and Fisher expansions and, in the correctly specified Gaussian regression case, achieves the stated efficient asymptotic variance [1406.4052]. In the recurrent-event model, SMPLE is robust under non-Poisson misspecification through the sandwich covariance \(A^{-1}BA^{-T}\), yet becomes semiparametrically efficient when the NHPP working model is valid [2507.20396]. This suggests that the “pseudo” qualifier refers to the construction of the criterion, not to an inevitable loss of first-order optimality.

A second misconception is that sieve methods are purely asymptotic devices. The single-index paper directly counters this by establishing finite-sample Wilks and Fisher results, nonasymptotic confidence sets, and convergence guarantees for alternating maximization [1406.4052]. The recurrent-event paper complements this with a scalable implementation based on a shared ODE solve, sensitivity equations, and resampling-based covariance estimation [2507.20396]. Taken together, these works place SMPLE at the intersection of semiparametric efficiency theory, sieve approximation, and computationally tractable pseudo-likelihood optimization.

Source: https://www.emergentmind.com/topics/sieve-maximum-pseudo-likelihood-estimation-smple