Papers
Topics
Authors
Recent
Search
2000 character limit reached

Tilted Knockoff Method

Updated 9 July 2026
  • Tilted Knockoff method is a variable-selection procedure that redefines the covariate law by incorporating selection probabilities to restore knockoff exchangeability.
  • It constructs knockoffs under a tilted distribution, ensuring valid conditional independence tests and rigorous FDR control even with biased sampling.
  • The method applies standard knockoff filtering with modified importance scores, demonstrating improved error control over naive applications in simulations and real genetic studies.

The Tilted Knockoff method is a variable-selection procedure for secondary-phenotype analysis under biased sampling, developed for settings in which observations are not selected uniformly at random from the population of interest. Its central purpose is to test population-level conditional null hypotheses of the form H0j:Xj ⁣ ⁣ ⁣YXjH_{0j}: X_j \perp\!\!\!\perp Y \mid X_{-j} while controlling the False Discovery Rate (FDR), even when the observed data arise from a selection mechanism that distorts the joint law of (X,Y)(X,Y). The method modifies the model-X knockoff framework by replacing the population covariate law with a tilted law that incorporates the selection probability, and then constructing knockoff variables relative to that tilted distribution rather than the unweighted population law (Zhao et al., 25 Aug 2025).

1. Problem formulation under biased sampling

The starting point is a population distribution

P(X,Y),(XRp,YR),P(X,Y), \quad (X \in \mathbb{R}^p, Y \in \mathbb{R}),

with scientific interest focused on testing the conditional nulls

H0j:Xj ⁣ ⁣ ⁣YXjH_{0j}: X_j \perp\!\!\!\perp Y \mid X_{-j}

under PP. In the setting considered by the method, the data are not observed as i.i.d. draws from PP. Instead, the sample is selected by a mechanism s(X,Y)[0,1]s(X,Y) \in [0,1], so that

(Xi,Yi)Pobs,Pobs(x,y)s(x,y)P(x,y).(X_i, Y_i) \sim P_{\rm obs}, \quad P_{\rm obs}(x,y) \propto s(x,y)\,P(x,y).

A classical example is case-control sampling, where YY is a secondary phenotype and selection depends on a disease indicator DD:

(X,Y)(X,Y)0

These definitions frame the central difficulty: analyses performed directly on the observed sample target the wrong distribution if the selection mechanism is ignored (Zhao et al., 25 Aug 2025).

The motivating pathology is collider bias. In a case-control study, one may identify a spurious association between an exposure and a secondary phenotype when both affect the case-control status. The method is designed for the high-dimensional regime where tests of independence under biased sampling are available in principle but typically do not apply when the number of variables is large. A common misconception is that standard model-X knockoffs remain valid after biased sampling if one simply applies them to the observed sample; the method explicitly shows that naive application fails to control FDR in this setting.

2. Tilted population law

Because knockoffs require knowledge of the marginal law of (X,Y)(X,Y)1 in the sample, the method introduces a tilted version of the population covariate law:

(X,Y)(X,Y)2

Equivalently,

(X,Y)(X,Y)3

In a regression or case-control setting, the paper also works with the conditional tilt

(X,Y)(X,Y)4

so that for any fixed (X,Y)(X,Y)5, (X,Y)(X,Y)6 has the same (X,Y)(X,Y)7 conditionals as the true sampled law of (X,Y)(X,Y)8 given (X,Y)(X,Y)9 (Zhao et al., 25 Aug 2025).

This construction is the conceptual core of the method. Rather than attempting to “debias” the response model directly, it redefines the covariate distribution used for knockoff generation so that the exchangeability argument underlying model-X knockoffs is restored under the sampled law. This suggests that the relevant object for valid selective inference is not the original population law of P(X,Y),(XRp,YR),P(X,Y), \quad (X \in \mathbb{R}^p, Y \in \mathbb{R}),0 alone, but the covariate law induced after integrating the selection mechanism into the sampling process.

3. Knockoff construction under the tilt

Under the tilted law, the method imposes the usual pairwise exchangeability condition

P(X,Y),(XRp,YR),P(X,Y), \quad (X \in \mathbb{R}^p, Y \in \mathbb{R}),1

or, more explicitly for each null index P(X,Y),(XRp,YR),P(X,Y), \quad (X \in \mathbb{R}^p, Y \in \mathbb{R}),2,

P(X,Y),(XRp,YR),P(X,Y), \quad (X \in \mathbb{R}^p, Y \in \mathbb{R}),3

so that P(X,Y),(XRp,YR),P(X,Y), \quad (X \in \mathbb{R}^p, Y \in \mathbb{R}),4 and its knockoff P(X,Y),(XRp,YR),P(X,Y), \quad (X \in \mathbb{R}^p, Y \in \mathbb{R}),5 are swap-invariant under P(X,Y),(XRp,YR),P(X,Y), \quad (X \in \mathbb{R}^p, Y \in \mathbb{R}),6 (Zhao et al., 25 Aug 2025).

When P(X,Y),(XRp,YR),P(X,Y), \quad (X \in \mathbb{R}^p, Y \in \mathbb{R}),7 is known in closed form, the method allows exact model-X sampling. The paper lists exact SDP, Gaussian mixture knockoffs, and metropolized knockoffs as admissible samplers for generating P(X,Y),(XRp,YR),P(X,Y), \quad (X \in \mathbb{R}^p, Y \in \mathbb{R}),8. When only the first two moments of P(X,Y),(XRp,YR),P(X,Y), \quad (X \in \mathbb{R}^p, Y \in \mathbb{R}),9 are available, the method uses a second-order Gaussian approximation

H0j:Xj ⁣ ⁣ ⁣YXjH_{0j}: X_j \perp\!\!\!\perp Y \mid X_{-j}0

Second-order knockoffs are then built by choosing a H0j:Xj ⁣ ⁣ ⁣YXjH_{0j}: X_j \perp\!\!\!\perp Y \mid X_{-j}1 matrix H0j:Xj ⁣ ⁣ ⁣YXjH_{0j}: X_j \perp\!\!\!\perp Y \mid X_{-j}2 such that

H0j:Xj ⁣ ⁣ ⁣YXjH_{0j}: X_j \perp\!\!\!\perp Y \mid X_{-j}3

and setting

H0j:Xj ⁣ ⁣ ⁣YXjH_{0j}: X_j \perp\!\!\!\perp Y \mid X_{-j}4

The significance of this construction is methodological rather than merely computational. The method does not introduce a new knockoff statistic; instead, it alters the law relative to which knockoffs are constructed. A plausible implication is that much of the existing model-X knockoff machinery can be reused, provided the sampler is calibrated to the tilted distribution rather than the uncorrected population or observed covariate law.

4. Test statistics and the knockoff filter

After generating knockoffs, the procedure follows the standard antisymmetric knockoff workflow. One fits any predictive model of H0j:Xj ⁣ ⁣ ⁣YXjH_{0j}: X_j \perp\!\!\!\perp Y \mid X_{-j}5 from H0j:Xj ⁣ ⁣ ⁣YXjH_{0j}: X_j \perp\!\!\!\perp Y \mid X_{-j}6—the paper gives Lasso as an example—and records symmetric importance scores H0j:Xj ⁣ ⁣ ⁣YXjH_{0j}: X_j \perp\!\!\!\perp Y \mid X_{-j}7 for H0j:Xj ⁣ ⁣ ⁣YXjH_{0j}: X_j \perp\!\!\!\perp Y \mid X_{-j}8 and H0j:Xj ⁣ ⁣ ⁣YXjH_{0j}: X_j \perp\!\!\!\perp Y \mid X_{-j}9 for PP0. One then forms

PP1

so that under exchangeability a null feature’s PP2 is equally likely to be positive or negative (Zhao et al., 25 Aug 2025).

For target FDR level PP3, the knockoff filter threshold is

PP4

and the selection set is PP5.

This stage of the method is intentionally conservative in design: it preserves the standard knockoff+ thresholding rule, so the inferential novelty lies entirely in how PP6 is generated under biased sampling. A common misunderstanding would be to attribute validity to the choice of predictive model or feature-importance statistic. In fact, the method’s guarantee is tied to exchangeability under the tilt; the importance model enters through power rather than through the FDR argument itself.

5. Theoretical guarantee

The principal theoretical statement is that, under the only assumption that for every null PP7 the sampler PP8 is pairwise exchangeable with respect to PP9 satisfying

PP0

the knockoff+ procedure controls FDR at level PP1 (Zhao et al., 25 Aug 2025).

The proof sketch follows the standard model-X logic. By pairwise exchangeability of PP2 under PP3 for null PP4, the signs PP5 are i.i.d. Rademacher for nulls. Standard arguments then yield

PP6

with all expectations taken under the tilted law, which coincides with the conditional law of PP7 given PP8 under the sampled distribution.

The paper also states that power depends on how closely PP9 and the chosen s(X,Y)[0,1]s(X,Y) \in [0,1]0 match the true law of s(X,Y)[0,1]s(X,Y) \in [0,1]1 in the selected sample. This sharply separates validity from efficiency: exact or well-approximated tilting is necessary for honest FDR control, whereas the choice between exact and second-order samplers affects the strength of the resulting discoveries. This suggests that misspecification in the selection-aware covariate model primarily manifests as power loss or invalid calibration, not as a benign nuisance.

6. Algorithmic workflow and empirical behavior

The paper gives the following pseudocode for Tilted Knockoff. The input is a biased sample s(X,Y)[0,1]s(X,Y) \in [0,1]2, target level s(X,Y)[0,1]s(X,Y) \in [0,1]3, and known or estimated s(X,Y)[0,1]s(X,Y) \in [0,1]4:

  1. (Optional) Pre-estimate s(X,Y)[0,1]s(X,Y) \in [0,1]5 or its moments from an external reference.
  2. For each observed s(X,Y)[0,1]s(X,Y) \in [0,1]6 (or bin/discretization) compute tilt weight s(X,Y)[0,1]s(X,Y) \in [0,1]7.
  3. Estimate tilted moments

s(X,Y)[0,1]s(X,Y) \in [0,1]8

s(X,Y)[0,1]s(X,Y) \in [0,1]9

  1. Generate knockoffs (Xi,Yi)Pobs,Pobs(x,y)s(x,y)P(x,y).(X_i, Y_i) \sim P_{\rm obs}, \quad P_{\rm obs}(x,y) \propto s(x,y)\,P(x,y).0 of (Xi,Yi)Pobs,Pobs(x,y)s(x,y)P(x,y).(X_i, Y_i) \sim P_{\rm obs}, \quad P_{\rm obs}(x,y) \propto s(x,y)\,P(x,y).1 under (Xi,Yi)Pobs,Pobs(x,y)s(x,y)P(x,y).(X_i, Y_i) \sim P_{\rm obs}, \quad P_{\rm obs}(x,y) \propto s(x,y)\,P(x,y).2 (second-order) or any exchangeable sampler for (Xi,Yi)Pobs,Pobs(x,y)s(x,y)P(x,y).(X_i, Y_i) \sim P_{\rm obs}, \quad P_{\rm obs}(x,y) \propto s(x,y)\,P(x,y).3.
  2. Fit importance model to (Xi,Yi)Pobs,Pobs(x,y)s(x,y)P(x,y).(X_i, Y_i) \sim P_{\rm obs}, \quad P_{\rm obs}(x,y) \propto s(x,y)\,P(x,y).4, obtain (Xi,Yi)Pobs,Pobs(x,y)s(x,y)P(x,y).(X_i, Y_i) \sim P_{\rm obs}, \quad P_{\rm obs}(x,y) \propto s(x,y)\,P(x,y).5.
  3. Compute threshold (Xi,Yi)Pobs,Pobs(x,y)s(x,y)P(x,y).(X_i, Y_i) \sim P_{\rm obs}, \quad P_{\rm obs}(x,y) \propto s(x,y)\,P(x,y).6 and output selections (Xi,Yi)Pobs,Pobs(x,y)s(x,y)P(x,y).(X_i, Y_i) \sim P_{\rm obs}, \quad P_{\rm obs}(x,y) \propto s(x,y)\,P(x,y).7 (Zhao et al., 25 Aug 2025).

The empirical evaluation comprises both simulation and real-data analysis. In simulation with a case-control design and a continuous secondary outcome, the setting used (Xi,Yi)Pobs,Pobs(x,y)s(x,y)P(x,y).(X_i, Y_i) \sim P_{\rm obs}, \quad P_{\rm obs}(x,y) \propto s(x,y)\,P(x,y).8 covariates, with 10% true signals (Xi,Yi)Pobs,Pobs(x,y)s(x,y)P(x,y).(X_i, Y_i) \sim P_{\rm obs}, \quad P_{\rm obs}(x,y) \propto s(x,y)\,P(x,y).9, and selection probability

YY0

Under this design, standard knockoffs ignoring YY1 showed median FDP YY2, whereas tilted knockoffs, whether exact or second-order, had FDP YY3. Their power was slightly below that of naive knockoffs but with honest FDR.

In real data, the method was applied to genetic mapping of neurocognitive endophenotypes using a case-control sample with YY4 for severe mental illness, YY5 candidate SNPs, and four secondary phenotypes (accuracy/speed in two tasks). The population law YY6 was approximated by weighted Gaussian moments from cases/controls and registry incidence, and

YY7

was estimated by logistic regression with intercept correction. The analysis generated tilted second-order knockoffs for each YY8, with multiple-knockoff construction for stability. In this application, naive knockoffs selected more SNPs (likely spuriously), whereas tilted knockoffs made fewer selections consistent with inverse-probability-weighted follow-up analysis.

7. Interpretation and relation to secondary-phenotype inference

The method is designed specifically for using samples collected for one scientific purpose to investigate secondary questions without losing finite-sample selective error control. Its formal contribution is to show that the standard model-X knockoff requirement—knowledge of the covariate law relevant to the sample—must be interpreted through the selection mechanism when the sample is biased. In this sense, the method is not merely a correction for case-control studies; it is a selection-aware generalization of model-X knockoffs for secondary-phenotype analysis (Zhao et al., 25 Aug 2025).

Its scope is also delimited clearly. The guarantee depends on the availability of a known or estimated selection mechanism and on the ability to construct a pairwise-exchangeable sampler with respect to the corresponding tilt. The method does not claim that any biased sample can be analyzed without assumptions; rather, it restores FDR control by re-weighting the presumed population covariate law by the selection mechanism and then applying a model-X knockoff sampler to that tilted law.

The principal practical lesson is therefore negative as well as positive. Ignoring the sampling design can induce spurious discoveries through collider bias and can invalidate naive knockoff inference. Conversely, when the selection mechanism is incorporated into the knockoff construction, variable selection for secondary phenotypes can retain replicability guarantees at the target FDR level.

Definition Search Book Streamline Icon: https://streamlinehq.com
References (1)

Topic to Video (Beta)

No one has generated a video about this topic yet.

Whiteboard

No one has generated a whiteboard explanation for this topic yet.

Follow Topic

Get notified by email when new papers are published related to Tilted Knockoff Method.