Papers
Topics
Authors
Recent
Search
2000 character limit reached

Focused Weighted-Average Least Squares (FWALS)

Updated 5 July 2026
  • Focused Weighted-Average Least Squares (FWALS) is a targeted model-averaging method in linear regression that simplifies uncertainty by semi-orthogonalizing auxiliary regressors.
  • It employs a computational strategy that reduces an exponential submodel selection problem to managing at most k₂ regressor-wise weights through a structured orthogonalization process.
  • FWALS delivers a data-driven, focused estimator by minimizing the asymptotic mean squared error for a specific scalar target under a local-to-zero framework.

Focused Weighted-Average Least Squares (FWALS) is a focused model-averaging estimator for linear regression with model uncertainty over auxiliary regressors. It is designed for settings in which the inferential target is not the full parameter vector, nor generic prediction risk, but a specific smooth scalar function of the coefficients on regressors that are always included. Its defining feature is a semi-orthogonalization of the auxiliary regressors that reduces the weighting problem from 2k22^{k_2} candidate submodels to at most k2k_2 regressor-wise weights, yielding a tractable but sub-optimal alternative to full focused model averaging (Yin, 3 Mar 2026).

1. Target of inference and model-uncertainty setting

FWALS is formulated for the linear regression

yi=xiβ+ϵi=x1iβ1+x2iβ2+ϵi,i=1,,N,y_i=x_i^\top\beta+\epsilon_i=x_{1i}^\top\beta_1+x_{2i}^\top\beta_2+\epsilon_i,\qquad i=1,\dots,N,

or in matrix form

y=Xβ+ϵ=X1β1+X2β2+ϵ.y=X\beta+\epsilon=X_1\beta_1+X_2\beta_2+\epsilon.

Here X1X_1 contains the primary or core regressors, which are always included, whereas X2X_2 contains auxiliary regressors whose inclusion is uncertain (Yin, 3 Mar 2026).

The focused parameter is a smooth scalar function

μ(β1),\mu(\beta_1),

so the objective is explicitly target-specific. In the full focused model-averaging formulation, one considers all

M=2k2M=2^{k_2}

submodels generated by including or excluding each of the k2k_2 auxiliary regressors. For submodel mm, with selection matrix k2k_20, the least-squares estimator is denoted k2k_21, and the focused estimate is

k2k_22

Traditional focused averaging then forms

k2k_23

FWALS keeps the same focused target but replaces this exponentially large weighting problem by a lower-dimensional construction (Yin, 3 Mar 2026).

A recurring misconception is that FWALS is simply ordinary model averaging with a new label. It is instead a focused estimator in the strict sense that the weights are chosen for k2k_24, not for overall fit, prediction error, or the full coefficient vector. A second misconception is that FWALS averages focused submodel estimates in the same order as standard focused model averaging. It does not; its aggregation acts on an orthogonalized coefficient representation before the focused function is applied (Yin, 3 Mar 2026).

2. Semi-orthogonalization and estimator construction

FWALS builds on the WALS orthogonalization idea of Magnus et al. and De Luca et al. by transforming the auxiliary regressors into a semi-orthogonalized representation

k2k_25

where

k2k_26

and

k2k_27

This residualizes k2k_28 with respect to k2k_29, rescales it, and rotates it so that the transformed auxiliary directions are orthogonal in the residualized space (Yin, 3 Mar 2026).

In this representation, the submodel estimator for yi=xiβ+ϵi=x1iβ1+x2iβ2+ϵi,i=1,,N,y_i=x_i^\top\beta+\epsilon_i=x_{1i}^\top\beta_1+x_{2i}^\top\beta_2+\epsilon_i,\qquad i=1,\dots,N,0 can be written as

yi=xiβ+ϵi=x1iβ1+x2iβ2+ϵi,i=1,,N,y_i=x_i^\top\beta+\epsilon_i=x_{1i}^\top\beta_1+x_{2i}^\top\beta_2+\epsilon_i,\qquad i=1,\dots,N,1

with

yi=xiβ+ϵi=x1iβ1+x2iβ2+ϵi,i=1,,N,y_i=x_i^\top\beta+\epsilon_i=x_{1i}^\top\beta_1+x_{2i}^\top\beta_2+\epsilon_i,\qquad i=1,\dots,N,2

A WALS-type average over submodels yields

yi=xiβ+ϵi=x1iβ1+x2iβ2+ϵi,i=1,,N,y_i=x_i^\top\beta+\epsilon_i=x_{1i}^\top\beta_1+x_{2i}^\top\beta_2+\epsilon_i,\qquad i=1,\dots,N,3

where

yi=xiβ+ϵi=x1iβ1+x2iβ2+ϵi,i=1,,N,y_i=x_i^\top\beta+\epsilon_i=x_{1i}^\top\beta_1+x_{2i}^\top\beta_2+\epsilon_i,\qquad i=1,\dots,N,4

Because of the semi-orthogonalized structure, yi=xiβ+ϵi=x1iβ1+x2iβ2+ϵi,i=1,,N,y_i=x_i^\top\beta+\epsilon_i=x_{1i}^\top\beta_1+x_{2i}^\top\beta_2+\epsilon_i,\qquad i=1,\dots,N,5 is diagonal: yi=xiβ+ϵi=x1iβ1+x2iβ2+ϵi,i=1,,N,y_i=x_i^\top\beta+\epsilon_i=x_{1i}^\top\beta_1+x_{2i}^\top\beta_2+\epsilon_i,\qquad i=1,\dots,N,6 The crucial reduction is that the original yi=xiβ+ϵi=x1iβ1+x2iβ2+ϵi,i=1,,N,y_i=x_i^\top\beta+\epsilon_i=x_{1i}^\top\beta_1+x_{2i}^\top\beta_2+\epsilon_i,\qquad i=1,\dots,N,7-dimensional model-weight vector enters only through the yi=xiβ+ϵi=x1iβ1+x2iβ2+ϵi,i=1,,N,y_i=x_i^\top\beta+\epsilon_i=x_{1i}^\top\beta_1+x_{2i}^\top\beta_2+\epsilon_i,\qquad i=1,\dots,N,8 partial-sum weights yi=xiβ+ϵi=x1iβ1+x2iβ2+ϵi,i=1,,N,y_i=x_i^\top\beta+\epsilon_i=x_{1i}^\top\beta_1+x_{2i}^\top\beta_2+\epsilon_i,\qquad i=1,\dots,N,9. These weights lie in

y=Xβ+ϵ=X1β1+X2β2+ϵ.y=X\beta+\epsilon=X_1\beta_1+X_2\beta_2+\epsilon.0

and, unlike the original model weights, they do not need to sum to one (Yin, 3 Mar 2026).

The FWALS estimator is then

y=Xβ+ϵ=X1β1+X2β2+ϵ.y=X\beta+\epsilon=X_1\beta_1+X_2\beta_2+\epsilon.1

This ordering is essential: FWALS first aggregates the transformed coefficient contributions and only then applies the focused function y=Xβ+ϵ=X1β1+X2β2+ϵ.y=X\beta+\epsilon=X_1\beta_1+X_2\beta_2+\epsilon.2 (Yin, 3 Mar 2026).

3. Asymptotic theory under local misspecification

FWALS is analyzed under a local-to-zero framework for the auxiliary coefficients,

y=Xβ+ϵ=X1β1+X2β2+ϵ.y=X\beta+\epsilon=X_1\beta_1+X_2\beta_2+\epsilon.3

which makes the inclusion bias and estimation variance comparable asymptotically (Yin, 3 Mar 2026).

Let

y=Xβ+ϵ=X1β1+X2β2+ϵ.y=X\beta+\epsilon=X_1\beta_1+X_2\beta_2+\epsilon.4

and define

y=Xβ+ϵ=X1β1+X2β2+ϵ.y=X\beta+\epsilon=X_1\beta_1+X_2\beta_2+\epsilon.5

Then, for nonstochastic y=Xβ+ϵ=X1β1+X2β2+ϵ.y=X\beta+\epsilon=X_1\beta_1+X_2\beta_2+\epsilon.6,

y=Xβ+ϵ=X1β1+X2β2+ϵ.y=X\beta+\epsilon=X_1\beta_1+X_2\beta_2+\epsilon.7

where

y=Xβ+ϵ=X1β1+X2β2+ϵ.y=X\beta+\epsilon=X_1\beta_1+X_2\beta_2+\epsilon.8

The first term is the asymptotic bias from partial omission of auxiliary effects; the second term is the asymptotic variance component (Yin, 3 Mar 2026).

For a smooth focused function, with gradient

y=Xβ+ϵ=X1β1+X2β2+ϵ.y=X\beta+\epsilon=X_1\beta_1+X_2\beta_2+\epsilon.9

the delta method gives

X1X_10

Hence the asymptotic bias of FWALS depends jointly on the local signal X1X_11, the correlation structure X1X_12, the orthogonalization geometry X1X_13, and the sensitivity of the focused target through X1X_14 (Yin, 3 Mar 2026).

This suggests a sharp distinction between FWALS and global shrinkage methods. Two weighting rules that look similar in coefficient space can behave differently once the target function changes, because X1X_15 changes the relevant bias-variance tradeoff.

4. Plug-in AMSE criterion and data-driven weights

The infeasible optimal FWALS weights minimize the asymptotic mean squared error of the focused estimator over

X1X_16

The resulting AMSE is a quadratic form in X1X_17 involving the bias term X1X_18, the covariance matrix X1X_19, the focused gradient X2X_20, and the transformed design objects X2X_21, X2X_22, and X2X_23 (Yin, 3 Mar 2026).

Operationally, the paper proposes plug-in AMSE minimization. The unknown local-bias component is estimated through

X2X_24

but because X2X_25 is not unbiased for X2X_26, it is corrected as

X2X_27

With this correction, the data-driven weights are chosen by

X2X_28

and the final estimator is

X2X_29

The plug-in criterion is shown to be asymptotically unbiased for the infeasible AMSE, and the estimated weights converge in distribution to the AMSE-optimal limit (Yin, 3 Mar 2026).

A practical implication is that FWALS is not a fixed-weight procedure. Its weights are data-driven and target-driven, and the focus enters through the gradient of μ(β1),\mu(\beta_1),0, not merely through model selection over μ(β1),\mu(\beta_1),1.

5. Computational reduction, sub-optimality, and empirical behavior

The defining computational advantage of FWALS is that it replaces an optimization over

μ(β1),\mu(\beta_1),2

submodels by an optimization over at most μ(β1),\mu(\beta_1),3 regressor-wise weights. This reduction is exact at the level of the transformed WALS representation but not at the level of full focused model averaging, because different model-weight vectors can produce the same diagonal partial-sum matrix μ(β1),\mu(\beta_1),4. FWALS is therefore explicitly described as a tractable sub-optimal procedure rather than a full solution to the original focused model-averaging problem (Yin, 3 Mar 2026).

The paper also explains why conventional focused averaging does not admit the same collapse. Even after orthogonalization, the quantities μ(β1),\mu(\beta_1),5 remain correlated through the common narrow-model component, so the full μ(β1),\mu(\beta_1),6-dimensional weight structure is not eliminated (Yin, 3 Mar 2026).

Simulation evidence is reported for a baseline linear-regression design with μ(β1),\mu(\beta_1),7, μ(β1),\mu(\beta_1),8, pairwise regressor correlation μ(β1),\mu(\beta_1),9, M=2k2M=2^{k_2}0, varying M=2k2M=2^{k_2}1, and focused parameter

M=2k2M=2^{k_2}2

Across 1000 Monte Carlo replications, FWALS, FIC, and mMSE are often similar; FWALS generally dominates SAIC and SBIC except sometimes when M=2k2M=2^{k_2}3 is very small and M=2k2M=2^{k_2}4 is large; prior-based WALS can be competitive when signals are weak but tends to perform worse when M=2k2M=2^{k_2}5 is high because it is not designed for the focused parameter; and the differences between FWALS and FIC are usually small (Yin, 3 Mar 2026).

A second design uses impulse response functions as the focus: M=2k2M=2^{k_2}6 For horizons M=2k2M=2^{k_2}7, FWALS and FIC remain very close, whereas mMSE becomes unstable across horizons, especially when the focused function is nonlinear. This is the basis for the claim that FWALS delivers stable performance when the focused function is designed for impulse response function (Yin, 3 Mar 2026).

The computational savings are substantial. For M=2k2M=2^{k_2}8, reported timings are as follows.

M=2k2M=2^{k_2}9 FWALS FIC mMSE
8 0.012 sec 0.851 sec 0.130 sec
9 0.015 sec 10.330 sec 0.992 sec
10 0.016 sec 137.098 sec 9.804 sec
11 0.017 sec 1407.540 sec 97.055 sec

These timings illustrate the central design goal of FWALS: focused inference with substantially lower computational burden than full focused model averaging (Yin, 3 Mar 2026).

6. Position within the broader weighted least-squares literature

FWALS belongs to a broader family of weighted least-squares methods, but its notion of focus is specific. In optimal weighted least-squares approximation, the weights and sampling law are coupled through Christoffel-type leverage scores so that stability and near-best approximation are obtained with sample size scaling linearly in model dimension up to a logarithmic factor (Cohen et al., 2016). On irregular domains, a related theory constructs weighted empirical least-squares estimators from surrogate discrete orthonormal bases while preserving stability and quasi-optimality (Migliorati, 2019). These methods are weighted and often leverage-focused, but their target is approximation in k2k_20, not a smooth focused function of a core parameter vector.

Another related strand concerns estimation under model or state-space pathologies. In the subcritical Heston process, a weighted least-squares estimator with weight

k2k_21

is introduced to regularize the singular region near k2k_22; the weighting is “focused” on the boundary problem rather than on a user-chosen target parameter (Chaumaray, 2015). In sufficient dimension reduction, Slice Weighted Average Regression constructs

k2k_23

from slice-wise least-squares slopes, with influence-based weights used to improve robustness (Masioti et al., 2022). These are close methodological relatives in the sense that they combine least-squares objects through structured weights, but they do not formulate the focused target as k2k_24.

A distinct but conceptually adjacent result studies weighted least squares for estimating a target linear projection of a misspecified regression function under a chosen weighting measure. There, weighted least squares with importance weights is required because ordinary least squares under a non-target design is inconsistent for the desired projection (Azriel, 2021). This suggests a broader interpretation: “focus” in weighted least squares can refer either to a chosen inferential target, as in FWALS, or to a chosen target measure, as in projection-based design problems.

Within this landscape, FWALS is characterized by three simultaneous commitments: a focused scalar target, model uncertainty restricted to auxiliary regressors, and computational reduction through semi-orthogonalization. What it does not claim is full optimality over the original k2k_25 model-weight simplex. Its contribution is instead to provide a practical and robust alternative whose weights are directly tailored to the focused parameter and whose computational cost remains manageable even when the number of auxiliary regressors is moderately large (Yin, 3 Mar 2026).

Topic to Video (Beta)

No one has generated a video about this topic yet.

Whiteboard

No one has generated a whiteboard explanation for this topic yet.

Follow Topic

Get notified by email when new papers are published related to Focused Weighted-Average Least Squares (FWALS).