Papers
Topics
Authors
Recent
Search
2000 character limit reached

Mirror Statistics in High-Dimensional Analysis

Updated 14 July 2026
  • Mirror Statistic is a family of quantitative constructs that use mirror symmetry to evaluate sign consistency in high-dimensional regression, cosmology, and image analysis.
  • In high-dimensional regression, it enables p-value-free false discovery rate control by comparing two independent coefficient estimates via mirror symmetry.
  • The approach extends to cosmology and computer vision, quantifying mirror-parity in the cosmic microwave background and scoring image symmetry with measurable precision.

Mirror statistic denotes several distinct quantitative constructions whose common element is a mirror relation: sign symmetry around zero, geometric reflection through a plane, or reflectional self-consistency under left–right reversal. In high-dimensional regression, it is a p-value-free device for false discovery rate (FDR) control built from two independent coefficient estimates; in cosmology, it denotes estimators for approximate mirror-parity of the cosmic microwave background (CMB); in image analysis, it denotes a scalar score measuring how mirror-symmetric an image is about a specified axis (Molinari et al., 2024, Ben-David et al., 2014, Woehrer, 9 Jul 2026).

1. Mirror principle and conceptual scope

Across these literatures, the defining idea is not a single canonical formula but a shared use of symmetry under a mirror operation. In high-dimensional inference, the relevant symmetry is distributional: null statistics are required to be approximately symmetric around $0$, so large negative values can estimate the number of false positives among large positive values. In CMB analysis, the mirror operation is literal spatial reflection through a plane on the sphere. In image symmetry scoring, the mirror operation is left–right reflection relative to a candidate axis, and the statistic measures agreement between an image and its reflected counterpart (Molinari et al., 2024, Ben-David et al., 2014, Woehrer, 9 Jul 2026).

This suggests a family resemblance rather than a single standardized object. In all cases, the statistic is constructed so that a specific form of mirror agreement or mirror symmetry becomes operationally testable. What differs is the object being mirrored: coefficient estimates, sky directions, or image representations.

2. Mirror statistics for FDR control in high-dimensional regression

In high-dimensional regression, the Mirror Statistic is a method to control FDR without relying on p-values. The construction requires two independent estimates of each regression coefficient. For predictor XjX_j, the statistic is defined as

Mj=sign ⁣(β^j(1)β^j(2))f ⁣(β^j(1),  β^j(2)),M_j = \operatorname{sign}\!\left(\hat{\beta}^{(1)}_j \hat{\beta}^{(2)}_j\right)\, f\!\left(|\hat{\beta}^{(1)}_j|,\; |\hat{\beta}^{(2)}_j|\right),

where f(,)f(\cdot,\cdot) is any non-negative, monotonically increasing function. In the simulations and application of the 2024 paper, the choice is

f(a,b)=a+b,f(a,b)=a+b,

so that

Mj=sign ⁣(β^j(1)β^j(2))(β^j(1)+β^j(2)).M_j = \operatorname{sign}\!\left(\hat{\beta}^{(1)}_j \hat{\beta}^{(2)}_j\right)\left(|\hat{\beta}^{(1)}_j|+|\hat{\beta}^{(2)}_j|\right).

If the two estimates have the same sign, then Mj>0M_j>0; if they disagree in sign, then Mj<0M_j<0. Large positive values suggest a stable non-null effect, whereas negative values are treated as evidence of null behavior or instability (Molinari et al., 2024).

The FDR argument is based on null symmetry. If, for a null feature jS0j\in S_0, the sampling distribution of at least one of the two coefficient estimates is symmetric around zero, then MjM_j is also symmetric around zero. This yields

XjX_j0

The threshold is chosen by

XjX_j1

and the selected set is

XjX_j2

The method therefore selects the smallest threshold at which the ratio of negative to positive mirror statistics is at most the target FDR level XjX_j3 (Molinari et al., 2024).

Independence of the two coefficient estimates is essential. In the original data-splitting formulation, the dataset is divided into two independent subsets, one used to estimate XjX_j4 and the other to estimate XjX_j5. The paper emphasizes that this is straightforward and universally valid, but it reduces the sample size available to each step and therefore hurts power and stability (Molinari et al., 2024).

A related frequentist construction reviewed in the Bayesian paper is the Gaussian Mirrors form

XjX_j6

Here XjX_j7 measures signal and XjX_j8 measures noise or variability. Under the null, the statistic centers near XjX_j9; under the alternative, the sum term tends to dominate. This establishes that, even within FDR control, “Mirror Statistic” refers to a method family rather than a single algebraic form (Molinari et al., 1 Oct 2025).

3. Randomisation and Bayesian reformulations

A central limitation of the classical Mirror Statistic is the need for two independent coefficient estimates. One 2024 approach replaces data splitting with outcome randomisation. Instead of partitioning the sample, it generates two pseudo-outcomes

Mj=sign ⁣(β^j(1)β^j(2))f ⁣(β^j(1),  β^j(2)),M_j = \operatorname{sign}\!\left(\hat{\beta}^{(1)}_j \hat{\beta}^{(2)}_j\right)\, f\!\left(|\hat{\beta}^{(1)}_j|,\; |\hat{\beta}^{(2)}_j|\right),0

Mj=sign ⁣(β^j(1)β^j(2))f ⁣(β^j(1),  β^j(2)),M_j = \operatorname{sign}\!\left(\hat{\beta}^{(1)}_j \hat{\beta}^{(2)}_j\right)\, f\!\left(|\hat{\beta}^{(1)}_j|,\; |\hat{\beta}^{(2)}_j|\right),1

where Mj=sign ⁣(β^j(1)β^j(2))f ⁣(β^j(1),  β^j(2)),M_j = \operatorname{sign}\!\left(\hat{\beta}^{(1)}_j \hat{\beta}^{(2)}_j\right)\, f\!\left(|\hat{\beta}^{(1)}_j|,\; |\hat{\beta}^{(2)}_j|\right),2. The paper notes that

Mj=sign ⁣(β^j(1)β^j(2))f ⁣(β^j(1),  β^j(2)),M_j = \operatorname{sign}\!\left(\hat{\beta}^{(1)}_j \hat{\beta}^{(2)}_j\right)\, f\!\left(|\hat{\beta}^{(1)}_j|,\; |\hat{\beta}^{(2)}_j|\right),3

and that these two outcomes are constructed to be independent. Coefficients estimated from Mj=sign ⁣(β^j(1)β^j(2))f ⁣(β^j(1),  β^j(2)),M_j = \operatorname{sign}\!\left(\hat{\beta}^{(1)}_j \hat{\beta}^{(2)}_j\right)\, f\!\left(|\hat{\beta}^{(1)}_j|,\; |\hat{\beta}^{(2)}_j|\right),4 and Mj=sign ⁣(β^j(1)β^j(2))f ⁣(β^j(1),  β^j(2)),M_j = \operatorname{sign}\!\left(\hat{\beta}^{(1)}_j \hat{\beta}^{(2)}_j\right)\, f\!\left(|\hat{\beta}^{(1)}_j|,\; |\hat{\beta}^{(2)}_j|\right),5 then replace the two split-sample estimates in the Mirror Statistic. The resulting method, RandMS, is reported to increase testing power in highly correlated covariates and high percentages of active variables, while retaining a low memory footprint and requiring only a single run on the full dataset. In the reported benchmark, with Mj=sign ⁣(β^j(1)β^j(2))f ⁣(β^j(1),  β^j(2)),M_j = \operatorname{sign}\!\left(\hat{\beta}^{(1)}_j \hat{\beta}^{(2)}_j\right)\, f\!\left(|\hat{\beta}^{(1)}_j|,\; |\hat{\beta}^{(2)}_j|\right),6, the method took about Mj=sign ⁣(β^j(1)β^j(2))f ⁣(β^j(1),  β^j(2)),M_j = \operatorname{sign}\!\left(\hat{\beta}^{(1)}_j \hat{\beta}^{(2)}_j\right)\, f\!\left(|\hat{\beta}^{(1)}_j|,\; |\hat{\beta}^{(2)}_j|\right),7 seconds and used about Mj=sign ⁣(β^j(1)β^j(2))f ⁣(β^j(1),  β^j(2)),M_j = \operatorname{sign}\!\left(\hat{\beta}^{(1)}_j \hat{\beta}^{(2)}_j\right)\, f\!\left(|\hat{\beta}^{(1)}_j|,\; |\hat{\beta}^{(2)}_j|\right),8 GB of memory (Molinari et al., 2024).

A Bayesian reformulation removes data splitting in a different way. Rather than constructing two estimates from two datasets or two pseudo-outcomes, it takes two independent posterior draws: Mj=sign ⁣(β^j(1)β^j(2))f ⁣(β^j(1),  β^j(2)),M_j = \operatorname{sign}\!\left(\hat{\beta}^{(1)}_j \hat{\beta}^{(2)}_j\right)\, f\!\left(|\hat{\beta}^{(1)}_j|,\; |\hat{\beta}^{(2)}_j|\right),9 and defines

f(,)f(\cdot,\cdot)0

Repeating this f(,)f(\cdot,\cdot)1 times gives a Monte Carlo sample f(,)f(\cdot,\cdot)2. After fixing a mirror threshold

f(,)f(\cdot,\cdot)3

the method defines posterior mirror-inclusion probabilities

f(,)f(\cdot,\cdot)4

and then uses a second threshold

f(,)f(\cdot,\cdot)5

The selected set is f(,)f(\cdot,\cdot)6 (Molinari et al., 1 Oct 2025).

The Bayesian version is designed for continuous and discrete outcomes and more complex predictors, such as mixed models. To remain scalable in high dimensions, it relies on Automatic Differentiation Variational Inference with mean-field Gaussian variational approximations. The paper is explicit that the FDR control is currently supported mainly by simulation evidence and that a full formal proof is still open (Molinari et al., 1 Oct 2025).

4. Mirror statistics for confounder selection

In observational studies, mirror statistics have been adapted to confounder selection under the union-set and minimal-set criteria. With treatment f(,)f(\cdot,\cdot)7, outcome f(,)f(\cdot,\cdot)8, and covariates f(,)f(\cdot,\cdot)9, the relevant sets are

f(a,b)=a+b,f(a,b)=a+b,0

with

f(a,b)=a+b,f(a,b)=a+b,1

The corresponding false discovery proportions are defined target-specifically. For the union-set approach,

f(a,b)=a+b,f(a,b)=a+b,2

and for the minimal-set approach,

f(a,b)=a+b,f(a,b)=a+b,3

The method remains p-value-free and uses symmetry of the mirror statistic under the relevant null to calibrate selection (Harada et al., 2024).

Two constructions are developed. In paired mirrors, separate mirror statistics f(a,b)=a+b,f(a,b)=a+b,4 and f(a,b)=a+b,f(a,b)=a+b,5 are built for the outcome and treatment models. The basic form is

f(a,b)=a+b,f(a,b)=a+b,6

with f(a,b)=a+b,f(a,b)=a+b,7 nonnegative, exchangeable, and monotone, and f(a,b)=a+b,f(a,b)=a+b,8 used as the default original mirror choice. In unified mirrors, the paper combines two-dimensional standardized statistics

f(a,b)=a+b,f(a,b)=a+b,9

expressed in polar coordinates Mj=sign ⁣(β^j(1)β^j(2))(β^j(1)+β^j(2)).M_j = \operatorname{sign}\!\left(\hat{\beta}^{(1)}_j \hat{\beta}^{(2)}_j\right)\left(|\hat{\beta}^{(1)}_j|+|\hat{\beta}^{(2)}_j|\right).0. For the union-set approach, the unified mirror has the form

Mj=sign ⁣(β^j(1)β^j(2))(β^j(1)+β^j(2)).M_j = \operatorname{sign}\!\left(\hat{\beta}^{(1)}_j \hat{\beta}^{(2)}_j\right)\left(|\hat{\beta}^{(1)}_j|+|\hat{\beta}^{(2)}_j|\right).1

where Mj=sign ⁣(β^j(1)β^j(2))(β^j(1)+β^j(2)).M_j = \operatorname{sign}\!\left(\hat{\beta}^{(1)}_j \hat{\beta}^{(2)}_j\right)\left(|\hat{\beta}^{(1)}_j|+|\hat{\beta}^{(2)}_j|\right).2 can be based on Mj=sign ⁣(β^j(1)β^j(2))(β^j(1)+β^j(2)).M_j = \operatorname{sign}\!\left(\hat{\beta}^{(1)}_j \hat{\beta}^{(2)}_j\right)\left(|\hat{\beta}^{(1)}_j|+|\hat{\beta}^{(2)}_j|\right).3 or its sign, and Mj=sign ⁣(β^j(1)β^j(2))(β^j(1)+β^j(2)).M_j = \operatorname{sign}\!\left(\hat{\beta}^{(1)}_j \hat{\beta}^{(2)}_j\right)\left(|\hat{\beta}^{(1)}_j|+|\hat{\beta}^{(2)}_j|\right).4 can be built from sums or products of radii or coordinatewise maxima. For the minimal-set approach, the statistic adds a third factor,

Mj=sign ⁣(β^j(1)β^j(2))(β^j(1)+β^j(2)).M_j = \operatorname{sign}\!\left(\hat{\beta}^{(1)}_j \hat{\beta}^{(2)}_j\right)\left(|\hat{\beta}^{(1)}_j|+|\hat{\beta}^{(2)}_j|\right).5

with

Mj=sign ⁣(β^j(1)β^j(2))(β^j(1)+β^j(2)).M_j = \operatorname{sign}\!\left(\hat{\beta}^{(1)}_j \hat{\beta}^{(2)}_j\right)\left(|\hat{\beta}^{(1)}_j|+|\hat{\beta}^{(2)}_j|\right).6

because the null geometry for Mj=sign ⁣(β^j(1)β^j(2))(β^j(1)+β^j(2)).M_j = \operatorname{sign}\!\left(\hat{\beta}^{(1)}_j \hat{\beta}^{(2)}_j\right)\left(|\hat{\beta}^{(1)}_j|+|\hat{\beta}^{(2)}_j|\right).7 is mixed (Harada et al., 2024).

The thresholding rule is the usual mirror-counting criterion,

Mj=sign ⁣(β^j(1)β^j(2))(β^j(1)+β^j(2)).M_j = \operatorname{sign}\!\left(\hat{\beta}^{(1)}_j \hat{\beta}^{(2)}_j\right)\left(|\hat{\beta}^{(1)}_j|+|\hat{\beta}^{(2)}_j|\right).8

with OR or AND combinations used for paired mirrors. The theoretical claims are asymptotic or approximate rather than finite-sample exact. The paper emphasizes that the approach is compatible with MLE, cross-fitting, lasso, debiased lasso, and multiple data splitting. In high-dimensional simulations, the proposed methods generally controlled FDR and often outperformed p-value-based procedures, especially with lasso-based estimation. In the right-heart-catheterization application, union-set approaches selected many variables, minimal-set approaches selected a sparser set, and the selected variables included disease category, SUPPORT score, APACHE score, oxygenation, and transfer status (Harada et al., 2024).

5. Mirror-parity statistics in cosmology

In CMB analysis, mirror statistics are estimators for approximate reflection symmetry about a plane in the sky. If Mj=sign ⁣(β^j(1)β^j(2))(β^j(1)+β^j(2)).M_j = \operatorname{sign}\!\left(\hat{\beta}^{(1)}_j \hat{\beta}^{(2)}_j\right)\left(|\hat{\beta}^{(1)}_j|+|\hat{\beta}^{(2)}_j|\right).9 is the normal to a candidate mirror plane, then a sky direction Mj>0M_j>00 is reflected as

Mj>0M_j>01

A temperature field Mj>0M_j>02 has even mirror-parity when Mj>0M_j>03 and odd mirror-parity when Mj>0M_j>04. In harmonic space, after rotating the candidate axis to the Mj>0M_j>05-axis, spherical harmonics transform as

Mj>0M_j>06

so modes with Mj>0M_j>07 even contribute to even mirror-parity and those with Mj>0M_j>08 odd contribute to odd mirror-parity (Ben-David et al., 2014).

The paper studies both pixel-space and harmonic-space estimators. The pixel-space statistic is

Mj>0M_j>09

where low Mj<0M_j<00 indicates strong even parity and low Mj<0M_j<01 indicates strong odd parity. The harmonic estimator is constructed from rotated coefficients Mj<0M_j<02, the sign factor Mj<0M_j<03, and the observed power spectrum

Mj<0M_j<04

with a subtraction of Mj<0M_j<05 as a normalization convention to make the ensemble mean zero. In this convention, high Mj<0M_j<06 means even parity and low Mj<0M_j<07 means odd parity (Ben-David et al., 2014).

For each sky map, the authors scan all directions and define a best-direction statistic Mj<0M_j<08 by taking the appropriate directional extremum: Mj<0M_j<09 They also define

jS0j\in S_00

where jS0j\in S_01 and jS0j\in S_02 are the mean and standard deviation of the full score map jS0j\in S_03 over all directions. The normalized statistic measures how exceptional the best direction is relative to the full directional score map on the same sky (Ben-David et al., 2014).

The paper’s main methodological warning is that significance depends strongly on Galactic masks and on a-posteriori choices. In pixel space, masking introduces anisotropic bias, especially with large masks such as U73. In harmonic space, the low-jS0j\in S_04 coefficients on a cut sky must be reconstructed by a maximum-likelihood method, and the authors conservatively restrict masked-map harmonic analysis to jS0j\in S_05. Using the large U73 mask and raw jS0j\in S_06, pixel-based odd-parity directions can appear significant at about the jS0j\in S_07 level, with reported values jS0j\in S_08 for SMICA, jS0j\in S_09 for LGMCA, and MjM_j0 for NILC. With MjM_j1, these fall to about MjM_j2–MjM_j3. Harmonic-space results are more stable but lower, reaching up to MjM_j4 with raw MjM_j5 and about MjM_j6 with MjM_j7. The conclusion is cautious: there is a mild tendency toward odd mirror-parity, especially around MjM_j8, but the signal “poses no real challenge to the concordance model” (Ben-David et al., 2014).

6. Mirror statistics in image symmetry scoring

In computer vision, a mirror statistic is a scalar quantitative measure of how mirror-symmetric an image is about a specified axis. The benchmark paper distinguishes detection from scoring: detection asks where the symmetry axis is, while scoring asks how symmetric the image is about a given axis. A scorer maps an image and a candidate reflection axis to a scalar that should increase with the strength of mirror symmetry about that axis (Woehrer, 9 Jul 2026).

The paper unifies 13 methods through the template

MjM_j9

where XjX_j00 is the left–right mirror operator, XjX_j01 is an image representation, XjX_j02 is a comparison between the image and its mirror, and XjX_j03 aggregates the comparison to a scalar. Evaluation is not based on absolute calibration across images, but on whether the statistic ranks the true axis above perturbed wrong axes. The benchmark metric is the chance-anchored discrimination skill

XjX_j04

with AUC computed within each image by taking the fraction of wrong axes that the true axis outranks, ties counting as XjX_j05, and then averaging equally across images (Woehrer, 9 Jul 2026).

The benchmark spans 13 scorers: 11 hand-crafted methods and 2 frozen deep-feature readouts. The methods are PixCorr, SlideWin, EROS, WBS, Grad, MS-Grad, DCT, Gabor, HOG, PHOG, PatchNN, AlexNet-C2, and DeepFeat. HOG is the best classical descriptor, with tuned configuration cell 16, 18 bins, unsigned gradients. DeepFeat is the best frozen deep readout, with default benchmark backbone MambaOut-Base stage-1 (Woehrer, 9 Jul 2026).

The evaluation uses a reflection-exact harness. Each image-axis pair is warped so the axis becomes the exact vertical centerline; the crop has even width and half-pixel offset; flipping is exact, with no resampling asymmetry; and out-of-image regions use constant border fill rather than reflection padding. The datasets cover four single-axis and five multi-axis benchmarks, totaling about 3950 axis units, with multi-axis contributing about 3300. The main negatives are perturbed axes obtained by perpendicular shifts XjX_j06 of crop size and in-plane rotations XjX_j07, with the main results using the coarser half, namely shifts XjX_j08 and rotations XjX_j09 (Woehrer, 9 Jul 2026).

The reported mean single-axis skills are 0.83 for DeepFeat, 0.81 for AlexNet-C2, 0.80 for HOG, and 0.71 for Gabor, followed by lower scores for the remaining methods. The paper reports DeepFeat > HOG by XjX_j10 skill with Holm-adjusted XjX_j11, HOG > Gabor by XjX_j12 with adjusted XjX_j13, and no significant difference between DeepFeat and AlexNet-C2 or between AlexNet-C2 and HOG. HOG is about 0.5 ms/crop, DeepFeat about 170 ms/crop, so HOG is roughly 300× faster on CPU. The interpretive conclusion is that discrimination concentrates in mid-scale oriented features: deep backbones peak at a low or mid stage, HOG peaks at a mid cell size, and unsigned gradients outperform signed gradients (Woehrer, 9 Jul 2026).

Related statistical literature uses mirror constructions for hypothesis testing even when the phrase “mirror statistic” is not the primary label. One example is the Mirror Bootstrap for testing hypotheses of one population mean. The method addresses the contradiction between requiring the sample to be representative of the population and simultaneously imposing the null XjX_j14 when the sample mean differs from XjX_j15. Given a sample XjX_j16, it constructs a symmetric pseudo-population by reflecting each observation around XjX_j17: XjX_j18 so that the constructed population is

XjX_j19

Bootstrap samples of size XjX_j20 are then drawn with replacement from this XjX_j21-point population, and the p-value is approximated by

XjX_j22

Simulations reported in the paper show that the method is slightly conservative for very small samples, but its validity and power quickly approach those of the XjX_j23-test for moderate sample sizes (Varvak, 2012).

Another symmetry-testing construction is based on cumulative past and residual extropy of record values. The population quantity is

XjX_j24

with symmetry characterized by XjX_j25. The implemented test specializes to

XjX_j26

and rejects for large XjX_j27. The paper emphasizes that the procedure does not require estimating the centre of symmetry, proves consistency under XjX_j28, XjX_j29, and XjX_j30, and calibrates the test by Monte Carlo because the asymptotic distribution is analytically intractable (Chaudhary et al., 2022).

These related constructions do not define the modern high-dimensional Mirror Statistic directly, but they show that reflection and symmetry arguments have been used in several distinct inferential settings. A plausible implication is that “mirror” methods are best understood as a broader methodological motif in which a reflected or sign-symmetrized object is used to enforce a null structure while retaining observed information (Varvak, 2012, Chaudhary et al., 2022).

Topic to Video (Beta)

No one has generated a video about this topic yet.

Whiteboard

No one has generated a whiteboard explanation for this topic yet.

Follow Topic

Get notified by email when new papers are published related to Mirror Statistic.