---
title: 'TWICE Methods: Dual Penalization & Wage Inference'
url: https://www.emergentmind.com/topics/twice
type: topic
---

# TWICE Methods: Dual Penalization & Wage Inference

TWICE denotes two distinct methodological constructions in recent arXiv literature. In one usage, TWICE is the acronym for **Tree-based Wage Inference with Clustering and Estimation**, a framework for modeling conditional wage functions from observables with gradient-boosted trees, observable-anchored worker and firm partitions, and a variance decomposition that includes sorting and non-additive interactions [2601.00776]. In another usage, “twice” refers to a **twice-penalized P-spline approach** for handling overlapping asymmetric datasets, where a flexible regression on a smaller cohort is regularized both by a customary P-spline roughness penalty and by a second penalty that aligns horizontal-cohort marginals with information learned from a larger cohort [2311.10489]. The two methods are unrelated in application domain, but both are organized around a dual-regularization or dual-structure idea: one combines tree ensembles with clustering and estimation, while the other combines smoothing with marginal alignment.

## 1. Terminological scope and problem classes

The two uses of TWICE arise from different statistical problems. McTeer et al. study **overlapping asymmetric datasets**, motivated by healthcare settings in which “a small amount of information” may be available for a larger number of patients, while “a small number of patients may have had extensive further testing” [2311.10489]. Their objective is to model the smaller cohort against a response while considering the larger cohort, without relying on missing data imputation when the smaller cohort is “significantly different in scale to the larger sample.”

Bakirov, Del Prato, and Zacchia define TWICE explicitly as **Tree-based Wage Inference with Clustering and Estimation** [2601.00776]. Their setting is matched worker–firm panel data with log wages \(Y_{it}=\ln w_{it}\), worker covariates \(Z_{it}\), and firm covariates \(X_j\). The framework is proposed against a background in which standard latent fixed-effects approaches depend on worker mobility, enforce additivity, and may produce decompositions with limited interpretability.

These two constructions therefore address different inferential objects. The twice-penalized P-spline targets a predictive function on an asymmetric, partially overlapping data structure. TWICE in wage inference targets the conditional-expectation function \(m_0(X_{J(i,t)},Z_{it})\) and then projects that function onto observable worker and firm partitions. A plausible implication is that the commonality in naming is methodological rather than substantive: both use a “two-part” architecture to stabilize estimation under structural asymmetry.

## 2. Twice-penalized P-spline methodology for overlapping asymmetric datasets

The twice-penalized P-spline approach is formulated around a B-spline basis and a penalized objective on the smaller “horizontal” cohort \(H\) [2311.10489]. Let \(x\) and \(z\) lie in \([a,b]\), with an ordered knot sequence \(t_1\le t_2\le\cdots\le t_{m+q}\), where \(q\) is the B-spline order and \(m\) is the number of interior segments. The Cox–de Boor recursion defines B-splines \(B_j^{(q)}(x)\), and the design matrix \(D_H\) is constructed either additively, with separate bases in \(x\) and \(z\), or as a tensor-product basis \(B_j(x_i)B_k(z_i)\) when interactions are allowed.

The central objective is to estimate \(\beta\) so that \(\hat y=D_H\beta\) fits \(y\) in the small cohort while enforcing two regularization criteria. The first uses the customary P-spline roughness matrix \(\Omega\), for example \(\Omega=C'C\) with a finite-difference operator of order \(d=2\). The second uses a matrix \(W\) that maps \(\beta\) to estimated horizontal-cohort marginals on a grid \(x_{\text{test}}\), together with \(\theta_V(x_{\text{test}})\), the vector of marginal fits from the large “vertical” cohort \(V\). The criterion is
\[
J(\beta)=\sum_{i\in H}\bigl(y_i-[D_H\beta]_i\bigr)^2+\lambda_1\beta^\top\Omega\beta+\lambda_2(W\beta-\theta_V)^\top(W\beta-\theta_V).
\]

For continuous responses, the estimator has the closed form
\[
\hat\beta=(D_H'D_H+\lambda_1\Omega+\lambda_2W'W)^{-1}(D_H'y+\lambda_2W'\theta_V),
\]
with \(\hat y=D_H\hat\beta\). For binary responses, the method maximizes a penalized log-likelihood with \(\theta_i=\expit(d_i'\beta)\) and fits the model by penalized iteratively weighted least squares, specifically Newton–Raphson or Fisher scoring. The gradient is
\[
\nabla \ell = D_H'(y-\theta)-2\lambda_1\Omega\beta-2\lambda_2W'(W\beta-\theta_V),
\]
and the Hessian includes the usual GLM curvature term together with the two penalty contributions.

The second penalty is the distinctive component. It is created “through discrepancies in the marginal value of covariates that exist in both the smaller and larger cohorts” [2311.10489]. This makes the approach different from a once-penalized P-spline: the estimator is not only smoothed, but also pulled toward consistency with a marginal structure learned from the larger cohort.

## 3. Tuning, identifiability, and empirical behavior of the twice-penalized P-spline

The two tuning parameters have distinct roles. \(\lambda_1\) is chosen by K-fold cross-validation in the horizontal cohort \(H\), cycling through a grid of \(\lambda_1\ge 0\) and selecting the value that optimizes one of several held-out criteria: sum of squares, penalized log-likelihood, or AUC for binary \(y\) [2311.10489]. \(\lambda_2\) is then selected over a modest grid \(\lambda_2\ge 0\) by requiring either a prescribed reduction in the marginal discrepancy \(\lVert W\beta-\theta_V\rVert_2^2\), such as 50%, or the absolute best marginal fit. In simulated or real NAFLD data, the paper reports \(\lambda_2\approx O(1\text{–}20)\) when \(\lambda_1\approx O(1\text{–}10)\).

Theoretical properties are stated in terms of the bias–variance trade-off, identifiability, and convergence. \(\lambda_1\) controls smoothness of \(\theta(x,z)\), while \(\lambda_2\) controls bias–variance in the marginal alignment. Two separate penalties “ensure one does not collapse into the other,” and \(\Omega\) is full rank after adding a small ridge if necessary. Convergence is immediate in the continuous case because of the closed-form estimator, while local convergence for the binary case is guaranteed under standard GLM regularity conditions.

The empirical results reported in the paper are substantial. In simulations with \(N_H=100\text{–}400\), noise \(\sigma=0.2\text{–}1.0\), with or without \(x\)–\(z\) interaction, and \(p_x=p_z=8\text{–}18\) knots, the twice-penalized fit (“Fit 2”) reduced the sum of squares of the marginal curve by 40–90% over the single-penalty P-spline (“Fit 1”) and by even more relative to an unpenalized B-spline linear model (“Fit 0”) [2311.10489]. Applied to NAFLD registry data with \(V: 6024\) patients and \(H: 1456\) patients, a binary “At-Risk NASH” outcome, 19 routine covariates in \(V\), and 37 genomic covariates in \(H\), the method achieved \(>50\%\) reduction in marginal-SS across linear predictor, PCA, and tSNE dimension-reduction schemes, and \(>65\%\) improvement in AUC versus standard P-splines. The authors characterize this as a flexible and numerically stable way to borrow “vertical” cohort accuracy into a smoother “horizontal” cohort fit, without imputation.

## 4. TWICE as Tree-based Wage Inference with Clustering and Estimation

In the wage-inequality setting, TWICE begins from the conditional-expectation model
\[
Y_{it}=m_0\bigl(X_{J(i,t)},Z_{it}\bigr)+u_{it},\qquad \mathbb E\bigl[u_{it}\mid X_{J(i,t)},Z_{it}\bigr]=0
\]
and estimates \(m_0\) with a gradient-boosted tree ensemble [2601.00776]. The fitted predictor is
\[
\widehat f(x,z)=\sum_{b=1}^B T_b(x,z),
\]
where each tree partitions the feature space into leaves and assigns each leaf a constant score. Estimation minimizes a regularized squared-error objective
\[
L(\{T_b\}_{b=1}^B)=\sum_{i,t}\bigl(Y_{it}-\widehat f(X_{J(i,t)},Z_{it})\bigr)^2+\sum_{b=1}^B\Omega(T_b),
\]
with
\[
\Omega(T)=\gamma L_T+\tfrac12\lambda\sum_{\ell=1}^{L_T}w_\ell^2.
\]
Here \(\gamma\) penalizes the number of leaves and \(\lambda\) shrinks leaf weights toward zero; early stopping, learning-rate control, maximum depth, and minimum leaf-size constraints are used for regularization.

The framework then constructs **observable-anchored partitions** rather than latent worker and firm effects. For firms, TWICE computes each firm-year’s mean log wage \(\bar Y_{jt}\) and fits a single regression tree of depth \(K\) to predict \(\bar Y_{jt}\) from \(X_j\). The \(K\) terminal leaves define firm-cells \(h:\mathcal X\to\{1,\dots,K\}\). For workers, it aggregates each worker to mean log wage \(\bar Y_i\) and mean covariates \(\bar Z_i\), then fits a depth-\(L\) tree to predict \(\bar Y_i\) from \(\bar Z_i\). The \(L\) leaves define worker-cells \(g:\mathcal Z\to\{1,\dots,L\}\). At each observation \((i,t)\), assignment uses the period-specific \(Z_{it}\), so workers may move between leaves over time.

With \(g\) and \(h\) fixed, the final predictor is estimated via LightGBM using raw covariates \((X,Z)\) together with one-hot indicators of worker-cell and firm-cell membership. The procedure is therefore a hybrid of nonparametric prediction and structured discretization: tree ensembles recover nonlinearities in the conditional wage function, while the worker and firm partitions create an interpretable representation of two-sided heterogeneity.

## 5. Cross-fitting, decomposition, and relation to AKM in TWICE

A defining estimation feature of TWICE is **two-way ID-blocked cross-fitting** [2601.00776]. Worker IDs are partitioned into \(B\) blocks \(\{\mathcal I_a\}\), firm IDs into \(B\) blocks \(\{\mathcal J_b\}\), and for each pair \((a,b)\) the validation set is
\[
S_{ab}=\{(i,t): i\in\mathcal I_a,\; J(i,t)\in\mathcal J_b\}.
\]
Training excludes any data with the same worker or firm as those in \(S_{ab}\), and the blocked MSE
\[
\mathcal L(K,L)=\frac1{B^2}\sum_{a,b}\frac1{|S_{ab}|}\sum_{(i,t)\in S_{ab}}\bigl(Y_{it}-\widehat f^{(-ab)}(X_{J(i,t)},Z_{it})\bigr)^2
\]
is minimized over a grid in \((K,L)\). The final model is retrained on all training data with the selected \((K^*,L^*)\) and evaluated on an external firm-level holdout set.

The fitted ensemble induces predicted cell means
\[
\hat\mu_{\ell k}=\frac{1}{N_{\ell k}}\sum_{(i,t):g(Z_{it})=\ell,\;h(X_{J(i,t)})=k}\widehat f\bigl(X_{J(i,t)},Z_{it}\bigr),
\]
which are decomposed as
\[
\hat\mu_{\ell k}=\bar Y+\hat\alpha_\ell+\hat\psi_k+\hat\kappa_{\ell k}.
\]
The pair \((\hat\alpha,\hat\psi)\) solves a weighted two-way least-squares projection using cell probabilities \(\pi_{\ell k}\), and the interaction term \(\hat\kappa_{\ell k}\) is orthogonal to both additive components by construction. Defining
\[
\alpha_{it}=\hat\alpha_{g(Z_{it})},\qquad \psi_{it}=\hat\psi_{h(X_{J(i,t)})},\qquad \xi_{it}=Y_{it}-\hat\mu_{g(Z_{it}),h(X_{J(i,t)})},
\]
TWICE yields the exact variance decomposition
\[
\Var(Y_{it})=
\underbrace{\Var(\alpha_{it})}_{\text{worker}}+
\underbrace{\Var(\psi_{it})}_{\text{firm}}+
\underbrace{2\,\Cov(\alpha_{it},\psi_{it})}_{\text{sorting}}+
\underbrace{\Var\bigl(\kappa_{g(Z_{it}),h(X_{J(i,t)})}\bigr)}_{\text{interaction}}+
\underbrace{\Var(\xi_{it})}_{\text{residual}}.
\]

The comparison benchmark is the canonical AKM model
\[
Y_{it}=\alpha_i+\psi_{J(i,t)}+X_{it}^\top\beta+\varepsilon_{it}.
\]
Bakirov, Del Prato, and Zacchia emphasize two AKM limitations: **additivity**, which rules out complementarities, and **limited-mobility bias**, under which sparse job-switching networks inflate \(\Var(\widehat\psi_j)\) and attenuate estimated sorting \(\Cov(\widehat\alpha_i,\widehat\psi_j)\) [2601.00776]. TWICE addresses these by replacing latent effects with observable partitions, allowing explicit non-additivity \(\kappa_{\ell k}\), and accepting a trade-off: it gives up some ability to capture purely idiosyncratic unobservables in exchange for robustness to sampling noise and out-of-sample portability.

## 6. Empirical results and interpretability diagnostics in TWICE

The empirical application uses Portuguese administrative data, combining Quadros de Pessoal and SCIE over 2012–2019, restricted to full-time workers aged 20–65 at firms with at least five employees and to the largest mobility-connected set [2601.00776]. The sample contains approximately 3.4 million worker-years, 750,000 workers, and 96,000 firms. On an external firm-holdout test, TWICE with two-way cross-fitted LightGBM and worker-plus-firm cells achieves **Test MSE \(\approx 0.092\)** and **Test \(R^2\approx 0.493\)**, whereas a flexible OLS benchmark with age polynomials attains **Test \(R^2\lesssim 0.41\)**.

The variance decomposition shares are reported as follows:

| Component | TWICE share | Comparison values |
|---|---:|---:|
| Worker | 27.6% | AKM: 59%; Bonhomme and Manresa (2019): 50% |
| Firm | 8.7% | AKM: 20%; Bonhomme and Manresa (2019): 5% |
| Sorting | 11.6% | AKM: 7%; Bonhomme and Manresa (2019): 20% |
| Interaction | 7.3% | AKM: none; Bonhomme and Manresa (2019): none |
| Residual | 44.8% | AKM: 14%; Bonhomme and Manresa (2019): 25% |

These figures support the paper’s central substantive claim that sorting is quantitatively important and that non-additive interactions are nonzero [2601.00776]. The interaction share is “modest (7.3%)” but present, while sorting is larger than in AKM on the same sample.

Interpretability is provided through Partial Dependence Plots and Accumulated Local Effects. The paper reports that age–wage PDPs by worker qualification and education show standard concave profiles, gender gaps, and steeper slopes for managers; tenure ALEs reveal a brief probationary dip at very low tenure followed by sustained firm-specific returns; and firm-size PDPs conditional on productivity are flat, suggesting that the canonical size premium largely reflects sorting on productivity and workforce composition rather than size per se [2601.00776]. The framework also compares observable partitions to AKM fixed effects using \(\eta^2\), the fraction of AKM effect-variance explained by TWICE cell membership, finding \(\eta^2\approx 0.40\) for workers and \(\eta^2\approx 0.25\) for firms on an observation-weighted basis.

## 7. Comparative significance and future directions

The two TWICE-related methods occupy different positions in contemporary statistical methodology. The twice-penalized P-spline approach is framed as, to the authors’ knowledge, “the first work to propose additional marginal penalties in a flexible regression” for asymmetric datasets, and it is explicitly motivated by avoiding missing data imputation while incorporating information from a larger cohort [2311.10489]. Its future work includes adaptation “to not require dimensionality reduction” and extension to “parametric modelling methods.” This suggests a broader agenda in which the second penalty might be embedded in richer semiparametric or multivariate structures.

TWICE in wage inference is positioned as an alternative to latent fixed-effects decompositions. Its contribution is to model the wage CEF directly from observables, to cluster workers and firms into interpretable cells, and to decompose wage dispersion into worker, firm, sorting, interaction, and residual components [2601.00776]. The authors emphasize that this replaces latent effects with observable-anchored partitions and thereby trades off the ability to capture idiosyncratic unobservables for robustness to sampling noise and out-of-sample generalization.

Taken together, these two usages illustrate different meanings of “TWICE” in current quantitative research. In one case, the term denotes a **dual-penalty smoothing architecture** for overlapping asymmetric datasets. In the other, it denotes a **tree-based inferential framework** for wage inequality decomposition. The shared conceptual thread is the deliberate introduction of a second structural layer—marginal alignment in one case, clustered observable heterogeneity in the other—to address limitations of simpler one-stage estimators.

Source: https://www.emergentmind.com/topics/twice