Series HAR Two-Sample t-Tests
- Series HAR two-sample t-tests are robust procedures that compare means in time series data exhibiting heteroskedasticity and serial dependence.
- The method employs orthonormal basis projections to estimate long-run variances, replacing conventional sample-variance formulas.
- It integrates a Welch-type t adjustment and a series-based wild bootstrap to enhance finite-sample performance in diverse applications.
Searching arXiv for the specified paper to ground the article in the current record. Series HAR two-sample t-tests are procedures for testing equality of means across two univariate time series when the data may exhibit heteroskedasticity and serial dependence. They are developed for the setting
where the two series are independent of each other, while the innovations may be heteroskedastic and serially dependent. The target hypothesis is
and the central methodological feature is a heteroskedasticity-and-autocorrelation-robust standardization based on orthonormal basis projections rather than conventional sample-variance formulas. The framework is presented as accommodating structural breaks, treatment-control comparisons, and group-averaged panel data, with a Welch-type approximation and a series-based HAR wild bootstrap providing finite-sample refinements under long-run variance heterogeneity (Hounyo et al., 12 Dec 2025).
1. Formal setting and inferential objective
The formal setup considers two univariate time series,
with independent series and potentially heteroskedastic, serially dependent innovations. The inferential objective is the two-sample mean comparison under the null hypothesis
Under mild regularity, expressed through a functional CLT, the sample means satisfy
where
is the long-run variance of series (Hounyo et al., 12 Dec 2025).
This formulation places the problem outside the scope of classical iid two-sample procedures. A plausible implication is that the relevant uncertainty is governed not by one-period innovation variance alone, but by the long-run variance induced by temporal dependence. That distinction motivates the series HAR construction.
2. Series-HAR standardization through orthonormal projections
The standardization step uses a mean-zero orthonormal basis on 0, denoted 1 for 2, satisfying
3
A convenient choice is the trigonometric system
4
For each series, the demeaned observations are projected onto the basis through
5
and the 6th partial long-run variance estimator is defined as
7
The series-HAR long-run variance estimator is then the simple average
8
The stated intuition is that each 9 picks out a frequency-band projection of the series, and averaging over 0 mimics a nonparametric fixed-1 estimator of the long-run variance that is robust to general heteroskedasticity and autocorrelation (Hounyo et al., 12 Dec 2025). This suggests that the method achieves robustness through basis-domain aggregation rather than direct blockwise time-domain smoothing.
3. Test statistic and Welch-type degrees-of-freedom adjustment
The unequal-long-run-variance statistic is defined by
2
When 3 and 4, the statistic satisfies
5
However, finite-sample size distortions may arise under serial dependence.
To address that issue, a Welch-type 6 approximation is constructed by matching the first two moments of the denominator with a scaled 7. Let 8. Under 9 and fixed 0,
1
has mean approximately
2
and variance approximately
3
Matching to 4 yields the adjusted degrees of freedom
5
with feasible implementation replacing 6 by 7 and 8 by 9. The resulting rule compares 0 to 1 critical values (Hounyo et al., 12 Dec 2025).
This correction is explicitly motivated by long-run variance heterogeneity across the two series. In contrast to classical Welch adjustments based on sample variances under independence, the present version is built around HAR long-run variance estimators.
4. Series-based HAR wild bootstrap
The series-based HAR wild bootstrap, denoted SHAR-WB, is designed to replicate both heteroskedasticity and serial dependence. Its algorithm proceeds as follows.
First, residuals are computed as
2
Second, the null is imposed through the common mean
3
Third, dependent wild multipliers 4 are generated via a second orthonormal basis 5, 6, for example
7
With iid draws 8 for 9, the multipliers are set to
0
These satisfy
1
and
2
which mimics a fixed-3 Daniell kernel.
Fourth, bootstrap errors and bootstrap data are formed as
4
Fifth, one recomputes 5, residuals 6, series-HAR long-run variances 7, and
8
Finally, after 9 repetitions, the 0 and 1 quantiles of the bootstrap distribution of 2 are used, and the null is rejected when the original 3 lies outside those quantiles (Hounyo et al., 12 Dec 2025).
A central feature of SHAR-WB is that it avoids resampling blocks of observations. The paper characterizes this as an extension of traditional wild bootstrap methods to the time-series setting.
5. Assumptions and asymptotic properties
The assumptions are organized around basis regularity, weak convergence, higher-order dependence control, and properties of the external wild variables. The basis functions 4 and 5 are assumed to be piecewise-smooth, mean-zero for 6, orthonormal, and uniformly bounded. The functional CLT is stated as
7
For each series, fourth-order cumulant summability is imposed:
8
and
9
The external wild variables satisfy the previously stated conditional moment and covariance conditions (Hounyo et al., 12 Dec 2025).
Under these assumptions, several asymptotic results are reported. In the equal-long-run-variance case with fixed 0, under 1 and 2,
3
In the unequal-long-run-variance case with 4 and 5,
6
For the bootstrap, large-7 validity is expressed as
8
in probability. The Welch-approximation 9 is stated to have asymptotically correct size to a higher order than the plain normal approximation (Hounyo et al., 12 Dec 2025).
These results separate the equal- and unequal-long-run-variance cases in a way analogous to classical pooled and Welch testing, but with long-run variance playing the role ordinarily occupied by one-sample variance.
6. Finite-sample behavior and empirical use
The reported finite-sample evidence is based on Monte Carlo designs with AR(1) errors, 0, normal or 1 innovations, equal versus unequal 2, and 3. Within these experiments, classical 4 and Welch 5 massively overreject when 6. The series HAR normal test, meaning 7 with a normal cutoff, improves performance but remains slightly oversized in small samples under strong dependence. The series HAR 8 approximation using 9 is reported to have much better size control. SHAR-WB delivers the best size control across all settings, including strong dependence and unequal 0, with power only moderately below the infeasible oracle (Hounyo et al., 12 Dec 2025).
Empirical illustrations include WFH productivity and pre/post structural breaks in U.S. macro series. In those examples, classical 1-tests reject differences that fail to survive the serial-dependence-robust series HAR and bootstrap tests. A plausible implication is that methods ignoring serial dependence can attribute significance to mean differences that are better interpreted as consequences of underestimated uncertainty.
7. Scope, implementation, and relation to conventional practice
The framework is described as accommodating a wide range of applications, including structural breaks, treatment-control comparisons, and group-averaged panel data (Hounyo et al., 12 Dec 2025). In all such settings, the unifying issue is inference on mean differences when serial dependence and heteroskedasticity make classical two-sample variance formulas unreliable.
Implementation is summarized as straightforward in the sense that one specifies two integers, 2 and 3. The first controls the number of orthonormal projections entering the series-HAR long-run variance estimator, and the second governs the multiplier construction in SHAR-WB. This suggests a modular architecture: basis-projection standardization for the statistic itself, and basis-driven multiplier dependence for bootstrap calibration.
A common misconception is that robustness to serial dependence in two-sample problems necessarily requires block bootstrap resampling or explicit parametric modeling of the autocovariance structure. The series HAR approach provides a different route: the long-run variance is estimated through orthonormal series projections, and the bootstrap analogue reproduces dependence through dependent wild multipliers rather than blocks. Within the reported evidence, this combination is associated with valid inference under heterogeneity and nonparametric dependence structures, together with superior finite-sample performance for the bootstrap variant (Hounyo et al., 12 Dec 2025).