---
title: Split-Panel Jackknife Bias Correction
url: https://www.emergentmind.com/topics/split-panel-jackknife-bias-correction
type: topic
---

# Split-Panel Jackknife Bias Correction

Split-panel jackknife bias correction is a jackknife-type procedure for panel estimators whose leading asymptotic bias scales with panel length, cross-sectional size, or both. In long-panel fixed-effects settings, the incidental-parameter bias typically has the form \(B/T+D/N\); in time-split settings it is often \(b/T\). The basic device is to recompute the estimator on large subpanels—most commonly half-panels—so that the leading bias term is doubled on the subsample, and then to combine full-sample and subpanel estimators so that the \(O(1/T)\) and, where relevant, \(O(1/N)\) components cancel. Across the literatures on fixed-effects likelihood, kernel smoothing, quantile regression, local projections, common correlated effects, and factor-augmented regressions, SPJ is treated as a non-parametric or “automatic” alternative to analytical bias correction, with the central attraction that it does not require direct estimation of model-specific bias constants [1709.08980][1904.04217].

## 1. Incidental-parameter bias as the organizing principle

The immediate motivation for split-panel jackknifing is the incidental-parameter problem. In semiparametric panel models with unobserved individual and time effects,
\[
y_{it}\mid x_{it},\alpha_i,\gamma_t\sim f\bigl(y\mid x_{it}'\beta+\alpha_i+\gamma_t\bigr),
\]
the fixed-effect estimator of \(\beta\) is contaminated by the estimation noise in \(\{\widehat\alpha_i\}\) and \(\{\widehat\gamma_t\}\). Under large-\(N\), large-\(T\) asymptotics, the canonical expansion is
\[
E[\widehat\beta-\beta_0]\approx \frac{B}{T}+\frac{D}{N},
\qquad
\mathrm{Var}(\widehat\beta)=O\!\bigl((NT)^{-1}\bigr),
\]
so the estimator remains offset by a bias of order \(O(1/T+1/N)\) unless that term is removed [1709.08980].

Fernández-Val and Weidner’s long-panel framework, summarized in the review of fixed-effect estimation, also presents a unifying \(p/n\) heuristic: in a cross-sectional MLE with \(p\) parameters and sample size \(n\),
\[
E[\widehat\theta-\theta_0]=O(p/n).
\]
In panels, \(p=N+T+\dim(\beta)\) and \(n=NT\), giving the \(B/T+D/N\) structure directly [1709.08980]. In binary-choice panels with individual and time effects, the same logic is written as
\[
E[\widehat\beta]=\beta+b_N^\beta+b_T^\beta+o(1/N+1/T),
\]
with \(b_N^\beta=B^\beta/T\) and \(b_T^\beta=C^\beta/N\) [1904.04217].

The same leading-bias phenomenon appears outside fixed-effects likelihood. In smoothed quantile panel models,
\[
\widehat\beta(\tau)=\beta(\tau)+\frac{b}{T}+O_P\bigl((NT)^{-1/2}\bigr),
\]
so the fixed-effects bias directly contaminates inference on the slope coefficient [1911.04729]. In panel local projections, Nickell bias arises for all regressors in the fixed-effect estimator, even if lagged dependent variables are absent in the regression, because the predictive specification violates strict exogeneity at horizon \(h\ge 1\) [2302.13455]. In heterogeneous-dynamics panels and kernel density estimation of unit-level objects, the first-stage estimation noise generates \(O(1/T)\) and \(O(1/(Th^2))\) biases that play the same structural role [1803.09452][1802.08825].

## 2. Core constructions and bias-cancellation algebra

The split-panel jackknife descends from the classical observation of Quenouille and later jackknife theory: if an estimator’s bias scales like \(1/n\), then a suitable linear combination of full-sample and subsample estimates can eliminate the first-order term. In panel data, the relevant “sample size” is directional—\(N\) in the cross-sectional dimension and \(T\) in the time dimension—so the estimator is split along \(i\), along \(t\), or along both dimensions [1904.04217].

| Scheme | Panel partition | Bias-corrected estimator |
|---|---|---|
| Half-panel jackknife | \(t=1,\dots,T/2\) and \(t=T/2+1,\dots,T\) | \(\hat\theta_{SSJK}=2\hat\theta-\tfrac12(\hat\theta_1+\hat\theta_2)\) |
| SPJ1 | separate half-splits in \(i\) and \(t\) | \(\tilde\beta_{spj1}=2\hat\beta-\hat\beta^{N}-\hat\beta^{T}\) |
| SPJ2 | four quadrants from simultaneous half-splits in \(i\) and \(t\) | \(\tilde\beta_{spj2}=\hat\beta-\hat b_{spj2}^\beta,\ \hat b_{spj2}^\beta=\hat\beta^{NT}-\hat\beta\) |

For two-way fixed-effects models, the most transparent derivation uses the expansions
\[
E[\widehat\beta_{N,T}]=\beta_0+\frac{B}{T}+\frac{D}{N},\quad
E[\widehat\beta_{N/2,T}]=\beta_0+\frac{B}{T}+\frac{2D}{N},\quad
E[\widehat\beta_{N,T/2}]=\beta_0+\frac{2B}{T}+\frac{D}{N}.
\]
Choosing weights \((w_0,w_1,w_2)=(3,-1,-1)\) gives
\[
\widehat\beta_{SPJ}
=
3\,\widehat\beta_{N,T}
-
\widehat\beta_{N/2,T}
-
\widehat\beta_{N,T/2},
\]
and removes both leading components simultaneously:
\[
E[\widehat\beta_{SPJ}]
=
\beta_0+o(T^{-1}+N^{-1}) .
\]
A simpler time-split version,
\[
\widehat\beta_{SPJ}^{(\alpha)}
=
2\,\widehat\beta_{N,T}-\widehat\beta_{N,T/2},
\]
is used when only the \(B/T\) component is targeted [1709.08980].

The binary-choice literature makes the same cancellation explicit. If
\[
E[\widehat\beta_{\mathrm{full}}]=\beta+b_N+b_T+o(1/N+1/T),
\]
then
\[
E[\widehat\beta^N]=\beta+2b_N+b_T+o(1/N+1/T),
\qquad
E[\widehat\beta^T]=\beta+b_N+2b_T+o(1/N+1/T).
\]
Hence
\[
E[\widehat b_{spj1}^\beta]
=
E[\widehat\beta^N]+E[\widehat\beta^T]-2E[\widehat\beta_{\mathrm{full}}]
=
b_N+b_T,
\]
so
\[
E[\tilde\beta_{spj1}]
=
\beta+o(1/N+1/T),
\]
with the remaining bias typically of order \(O(1/N^2+1/T^2)\) [1904.04217].

This same half-panel subtraction,
\[
2\widehat\theta-\tfrac12(\widehat\theta_a+\widehat\theta_b),
\]
reappears in smoothed quantile panels, panel local projections, heterogeneous-dynamics estimation, and factor extraction problems. The common algebra is that each half-sample estimator has approximately twice the full-sample first-order bias, so averaging the halves yields an empirical bias estimate that can be subtracted from the full estimator [1911.04729][2302.13455][1803.09452].

## 3. Assumptions and regimes of validity

SPJ is not assumption-free. In fixed-effects binary-choice models with individual and time effects, the key long-panel conditions are that \(N,T\to\infty\) with \(N/T\to\kappa^2\in(0,\infty)\), errors are i.i.d. or weakly dependent across \(i\) and \(t\), unobserved \(\alpha_i,\gamma_t\) enter additively and are arbitrary, and covariates satisfy conditional exogeneity [1904.04217]. For SPJ specifically, unconditional homogeneity is required: there should be no deterministic trends or structural breaks in \(x_{it}\), so that subsample-based MLEs have the same target \(\beta\) [1904.04217].

In the panel local projection framework, the high-level conditions are phrased differently but play the same role. The vector of innovations is assumed to be a strictly stationary martingale-difference sequence, the roots of the companion VAR\((p)\) lie outside the unit circle, the within-demeaned Gram matrices converge to a positive-definite limit \(Q\), and a joint CLT for the score sums and their half-sample analogues yields a limit variance matrix \(R\). The resulting limit theory requires \(N/T^3\to0\) after horizon adjustment [2302.13455].

Kernel and nonparametric settings impose additional rate restrictions because the first-stage estimation noise interacts with smoothing. For kernel density estimation of heterogeneous means, autocovariances, and autocorrelations, the required rates are
\[
Nh\to\infty,\quad Nh^5\to C,\quad Th^2\to\infty,\quad N/(T^3h^5)\to0,
\]
under i.i.d. units, stationarity, mixing, moment, and smoothness conditions [1802.08825]. In the local-linear smoothed quantile estimator for nonparametric panel quantile regression, SPJ is developed under
\[
N/(Th^d)\to\kappa^2,\qquad Th^{d+1}\to\infty,\qquad Th^{d+3}\to0,
\]
together with further restrictions on the smoothing bandwidth \(b\) and the order \(m\) of the smoothing kernel [1911.01824].

Balanced panels are the cleanest environment. With unbalanced data, the review of fixed-effect estimation allows analogous formulas provided no “commodity rows/columns vanish,” and suggests splitting by ranking units \(i\) by \(N_i\) and time periods \(t\) by \(T_t\), or taking random halves of roughly equal sample size [1709.08980]. The fixed-effects binary-choice analysis is more cautionary: in unbalanced panels one typically “ignores” the missing-at-random mechanism in the splits, and this can lead to very uneven subsample sizes [1904.04217]. A plausible implication is that the practical validity of SPJ depends not only on asymptotic divisibility of the panel, but on the empirical comparability of the halves.

## 4. Major variants across econometric literatures

In fixed-effects binary-choice models, SPJ is presented as a non-parametric correction for the leading \(1/N\) and \(1/T\) asymptotic bias of the fixed-effects binary-choice estimator. Czarnowske and Stammann discuss two implementations: SPJ1, based on separate cross-sectional and time splits, and SPJ2, based on a four-way \(NT\) split. After forming the corrected slope estimate, one re-plugs the debiased coefficient into the incidental-parameter problem to obtain debiased fixed effects and hence partial-effect estimates [1904.04217].

In kernel estimation for heterogeneous dynamics, Okui and Yanagi use the half-panel jackknife to remove both the incidental-parameter bias \(A_{\xi,1}(x)/T\) and the second-order nonlinearity bias \(A_{\xi,2}(x)/(Th^2)\) in
\[
\hat f_{\hat\xi}(x)
=
\frac1{Nh}\sum_{i=1}^N
K\!\Bigl(\frac{x-\hat\xi_i}{h}\Bigr).
\]
The corrected estimator is
\[
\hat f_{\hat\xi}^H(x)=2\hat f_{\hat\xi}(x)-\bar f(x),
\]
and under the same double-asymptotic conditions as the uncorrected estimator,
\[
\sqrt{Nh}\Bigl(\hat f^H_{\hat\xi}(x)-f_\xi(x)-h^2\tfrac{\kappa_1f_\xi''(x)}2\Bigr)\dto \mathcal N(0,\kappa_2f_\xi(x)).
\]
Because the smoothing bias remains, the same paper proposes a robust bias-corrected \(t\)-statistic and pointwise confidence intervals [1802.08825].

In quantile panel data models using smoothed quantile regressions, the estimator is constructed by a within-type first step followed by a smoothed QR objective. Splitting the time dimension into two halves and recomputing the same two-step estimator on each half yields
\[
\widehat\beta_{spj}(\tau)
=
2\,\widehat\beta(\tau)-\tfrac12\bigl[\widehat\beta_1(\tau)+\widehat\beta_2(\tau)\bigr].
\]
Since
\[
\widehat\beta_j(\tau)=\beta(\tau)+\frac{2b}{T}+o_P(T^{-1}),
\]
the \(b/T\) term cancels exactly, and
\[
\sqrt{NT}\,\bigl(\widehat\beta_{spj}(\tau)-\beta(\tau)\bigr)
\overset{d}{\longrightarrow}
\mathcal N\bigl(0,\Sigma^{-1}\Omega\Sigma^{-1}\bigr)
\]
under \(N/T\to\kappa^2>0\) and the stated smooth-kernel conditions [1911.04729].

In nonparametric quantile regressions for panel data, the local-linear smoothed quantile regression estimator admits an explicit incidental-parameter bias term \(B_{IP}=(Th^d)^{-1}B^{(2)}\) at boundary points. Chen’s SPJ correction,
\[
\check\beta_\tau^{bc}(x)
=
2\check\beta_\tau(x)-\tfrac12\bigl[\check\beta_{\tau,1}(x)+\check\beta_{\tau,2}(x)\bigr],
\]
is designed to cancel both the local-linear bias \(hB^{(1)}\) and the incidental-parameter bias \(B_{IP}\) [1911.01824].

In panel local projections, the split-panel jackknife is a direct remedy for the Nickell bias of the fixed-effect LP estimator. Denoting the full-sample and two half-sample fixed-effect estimators by \(\widehat\theta^{(h)}_{fe}\), \(\widehat\theta^{(h)}_a\), and \(\widehat\theta^{(h)}_b\), the SPJ estimator is
\[
\widehat\theta^{(h)}_{spj}
=
2\,\widehat\theta^{(h)}_{fe}
-\tfrac12\bigl(\widehat\theta^{(h)}_a+\widehat\theta^{(h)}_b\bigr),
\]
and satisfies
\[
\sqrt{N(T-h)}\bigl(\widehat\theta^{(h)}_{spj}-\theta^{(h)}\bigr)\dto N(0,Q^{-1}RQ^{-1})
\]
when \((N,T-h)\to\infty\) jointly with \(N/T^3\to0\) [2302.13455].

In nonlinear panel data models with interactive fixed effects estimated by common correlated effects, Chen and Zhang adapt the two-way split formula:
\[
\hat\beta_{SPJ}
=
3\,\hat\beta
-\tfrac12\bigl(\hat\beta_{N/2,T}^1+\hat\beta_{N/2,T}^2\bigr)
-\tfrac12\bigl(\hat\beta_{N,T/2}^1+\hat\beta_{N,T/2}^2\bigr).
\]
Given a Bahadur expansion with bias vectors \(b/T\) and \(d/N\), the linear combination removes both leading terms and preserves the \(\sqrt{NT}\) normal limit [2304.13199].

In factor-augmented regressions with weak factors, Jiang, Uematsu, and Yamagata use a cross-sectional split rather than a time split. After randomly partitioning the \(N\) series into two half-panels, extracting principal components on each half, and aligning subpanel factors to the full-sample factors, they define
\[
\hat\delta_{bcjk}
=
2\,\hat\delta-\tfrac12(\hat\delta_1+\hat\delta_2).
\]
Theorem 3.4 shows that when \(\alpha_r=1\) the leading bias vanishes exactly; for weak factors \((\alpha_r<1)\), the correction shrinks rather than fully annihilates the leading bias [2509.02066].

## 5. Finite-sample behavior, limitations, and controversies

The strongest positive evidence for SPJ comes from balanced-panel simulations and settings where the dominant bias clearly scales with the halved dimension. In fixed-effects binary-choice simulations with balanced panels \((N,T)=(200,15),(200,20),(200,25)\), SPJ1 and SPJ2 both shrink bias almost completely to zero, with relative bias below \(1\%\) for \(\beta\) and coverage around \(0.90\)–\(0.93\) for \(95\%\) confidence intervals [1904.04217]. In kernel density estimation for heterogeneous dynamics, both the half-panel jackknife and the third-order jackknife remove the dominant \(O(1/T)\) and \(O(1/(Th^2))\) biases, producing bias reductions of up to \(90\%\) and improving coverage from roughly \(50\%\)–\(60\%\) to \(90\%\)–\(95\%\) even when \(T=12\) [1802.08825]. In panel local projections with \((N,T)=(50,120)\), the FE impulse response is severely attenuated toward zero at medium horizons and \(95\%\) confidence intervals cover only \(30\%\)–\(40\%\) of the time, whereas the SPJ impulse response lies on top of the true \(\beta^{(h)}\) and empirical coverage is virtually \(95\%\) [2302.13455].

The main caveat is that favorable first-order bias removal does not imply uniformly superior finite-sample performance. The fixed-effects binary-choice simulations report that analytical bias corrections have even better coverage, around \(0.94\)–\(0.95\), and slightly smaller MSE in balanced panels [1904.04217]. In unbalanced panels, the same study finds that analytical bias correction remains virtually unaffected by the missing-data pattern, while SPJ1 can deteriorate severely: for \(T=15\), the relative bias of the lag coefficient is \(-32\%\) with coverage around \(11\%\) under one missing-data pattern, and bias is \(-20\%\) with coverage around \(44\%\) under another; SPJ2 shows similar sensitivity, sometimes slightly less extreme [1904.04217].

A second limitation is variance inflation beyond first order. The higher-order analysis of efficient bias correction shows that split-sample jackknife bias estimates are not \(\sqrt n\)-consistent in the i.i.d. setting considered there, and the higher-order variance term of the split-sample jackknife is twice as large as that of the leave-one-out jackknife:
\[
V^{(2)}_{SSJK}=2\cdot V^{(2)}_{LOOJK}.
\]
This does not contradict first-order equivalence: the split-sample jackknife shares the same leading variance \(I^{-1}\), but pays a higher-order variance cost [2207.09943]. The same paper therefore recommends avoiding split-sample jackknife in i.i.d. cross-sectional or panel settings when analytical, bootstrap, or leave-one-out corrections are available, while also noting that in dependent data—especially time series or serially correlated panels—leave-one-out corrections typically fail to remove bias and split-sample jackknife remains valid [2207.09943].

A common misconception is that SPJ is uniformly “automatic” once one can split a panel in half. The literature does not support that generalization. The binary-choice evidence ties poor performance to unbalancedness and heterogeneous splits [1904.04217]; the review literature stresses that SPJ requires homogeneity of the leading bias terms across subpanels [1709.08980]; and the CCE literature notes that if the data exhibit strong regime shifts between the first and second halves, SPJ may not entirely remove the bias [2304.13199].

## 6. Implementation, inference, and practical use

Implementation is conceptually simple but computationally nontrivial: the nonlinear fixed-effects or two-step estimator must be recomputed on each subpanel. In the two-way FE binary-choice case, SPJ1 requires four auxiliary fits—two \(N\)-splits and two \(T\)-splits—and SPJ2 requires four quadrant regressions; if average partial effects are needed, the debiased \(\tilde\beta\) is then held fixed while \(\alpha_i,\gamma_t\) are re-estimated by an offset algorithm [1904.04217]. For large \(N\) and \(T\), the same literature recommends high-dimensional FE solvers such as **reghdfe** in Stata, **lfe** in R, and **pyhdfe** in Python to demean and avoid direct estimation of \(\alpha_i,\gamma_t\) [1904.04217].

Inference is model-specific. In heterogeneous-dynamics panels, Okui and Yanagi pair the half-panel jackknife with cross-sectional bootstrap inference by resampling the unit-level triplets \((\hat\theta_i,\hat\theta_i^{(1)},\hat\theta_i^{(2)})\), and show that the bootstrap consistently estimates the asymptotic variance of the jackknife estimator under the stated conditions [1803.09452]. In kernel density estimation, the half-panel jackknife is combined with a robust bias-corrected variance estimator and \(t\)-statistic to account for the remaining \(O(h^2)\) smoothing bias [1802.08825]. In panel local projections, standard errors are obtained by recalculating residuals from \(\widehat\theta_{spj}\) and using the usual cluster-robust variance, or two-way clustering if needed [2302.13455]. In nonlinear CCE estimation, SPJ requires no explicit estimation of the bias vectors \(b\) and \(d\), and no kernel or HAC bandwidth; its only tuning choice is how to split the panel, with halves recommended so that each subpanel remains large enough for the asymptotics [2304.13199].

The practical recommendations in the literature are cautious. In fixed-effects binary-choice models, SPJ is described as conceptually simple and easy to code, but it should be treated as a robustness check alongside analytical correction, and standard errors should account for the extra variability from overlapping subpanels [1904.04217]. In smoothed quantile panel regressions, SPJ is extremely simple to implement once the two-step estimator is coded, but when \(T\) is very small it can inflate variance substantially, so analytical bias correction may be preferred if density estimates are stable [1911.04729]. In factor-augmented regressions with weak factors, the recommended practice is to randomize the ordering of cross-sectional units before splitting, use at least \(R\ge 50\) random splits and average the corrected estimators, and align subpanel factors to the full-sample factors by maximizing pairwise correlations [2509.02066].

As a method class, split-panel jackknife bias correction occupies an intermediate position between derivative-heavy analytical correction and computationally intensive leave-one-out procedures. It is most persuasive when the leading bias is known to scale mechanically with the panel dimension being halved, when subpanels remain homogeneous and large, and when dependence structures make leave-one-out arguments unattractive. It is least persuasive when unbalancedness, regime shifts, or higher-order variance costs dominate the first-order bias reduction [1709.08980][2207.09943].

Source: https://www.emergentmind.com/topics/split-panel-jackknife-bias-correction