Papers
Topics
Authors
Recent
Search
2000 character limit reached

Weighted and Penalized mRMR

Updated 4 March 2026
  • Weighted and Penalized mRMR is a continuous, weight-based feature selection method that balances relevance and redundancy through a quadratic objective.
  • It incorporates nonconvex penalties such as SCAD and MCP to enforce sparsity and accurately recover active features in high-dimensional settings.
  • The approach employs a two-stage knockoff+ procedure to control the false discovery rate and ensure oracle screening properties.

Weighted and penalized minimum Redundancy Maximum Relevance (mRMR), formalized in the SmRMR framework, denotes a class of feature selection methodologies that unifies continuous, weight-based redundancy-aware variable screening with sparsity-inducing nonconvex penalization, and formal false discovery rate (FDR) control through a multi-stage knockoff procedure. SmRMR is motivated by the need for scalable, model-free selection of relevant features in ultra-high-dimensional datasets, particularly where classical mRMR is computationally prohibitive and unable to provide explicit statistical error control (Naylor et al., 26 Aug 2025).

1. Continuous Weighted mRMR Objective

SmRMR generalizes the classic discrete mRMR approach by introducing a continuous optimization framework with feature-wise nonnegative weights. Let X1,…,XpX_1,\ldots,X_p denote pp features and YY the target. Define an association measure D(⋅,⋅)D(\cdot,\cdot) that is nonnegative and zero iff the arguments are independent (e.g., Hilbert-Schmidt Independence Criterion (HSIC) or Projection Correlation). The model constructs two matrices:

  • RyX∈RpR_{yX} \in \mathbb{R}^p with (RyX)j=D(Xj,Y)(R_{yX})_j = D(X_j, Y) (feature-response associations)
  • RXX∈Rp×pR_{XX} \in \mathbb{R}^{p \times p} with (RXX)j,k=D(Xj,Xk)(R_{XX})_{j,k} = D(X_j, X_k) (feature-feature associations)

The continuous, unpenalized mRMR objective is

max⁡w≥0Q(w):=w⊤RyX−12w⊤RXXw,\max_{w \geq 0} Q(w) := w^\top R_{yX} - \frac{1}{2} w^\top R_{XX} w,

which corresponds to minimizing

L0(w):=−w⊤RyX+12w⊤RXXw.\mathcal{L}_0(w) := -w^\top R_{yX} + \frac{1}{2} w^\top R_{XX} w.

Feature weights pp0 directly capture each variable's overall redundancy-aware utility.

2. Nonconvex Penalization and Sparse Optimization

To induce sparsity (i.e., automatic exclusion of inactive features), SmRMR augments pp1 with a componentwise penalty:

pp2

where pp3 is a nonconvex regularizer such as:

  • SCAD (Smoothly Clipped Absolute Deviation; param pp4)
  • MCP (Minimax Concave Penalty; param pp5)

The explicit forms are:

Penalty pp6 Derivative pp7
SCAD pp8, pp9<br>YY0, YY1<br>YY2, YY3 YY4
MCP YY5, YY6<br>YY7, YY8 YY9

Asymptotically, D(⋅,⋅)D(\cdot,\cdot)0 or D(⋅,⋅)D(\cdot,\cdot)1 recovers LASSO. Nonconvex penalties sharpen zeroing of truly inactive coefficients, with the sparsity pattern in D(⋅,⋅)D(\cdot,\cdot)2 providing the support of relevant features.

3. Numerical Algorithm: Local Linear Approximation

SmRMR leverages the Local Linear Approximation (LLA) algorithm to address the nonconvexity of SCAD/MCP:

  1. Initialize D(⋅,⋅)D(\cdot,\cdot)3 by solving the nonnegative LASSO (i.e., D(⋅,⋅)D(\cdot,\cdot)4).
  2. Iterate for D(⋅,⋅)D(\cdot,\cdot)5:
    • (a) Compute D(⋅,⋅)D(\cdot,\cdot)6,
    • (b) Solve the convex weighted-LASSO problem D(⋅,⋅)D(\cdot,\cdot)7 (via coordinate descent or QP),
    • (c) If D(⋅,⋅)D(\cdot,\cdot)8, terminate.
  3. Return D(⋅,⋅)D(\cdot,\cdot)9.

Typically, RyX∈RpR_{yX} \in \mathbb{R}^p0 suffices for convergence. This iterative reweighting adapts the penalty based on the current coefficient estimates, promoting sharper recovery of the feature support.

4. Multi-Stage Knockoff+ Filter for FDR Control

SmRMR integrates a two-stage knockoff pipeline to ensure explicit control of the false discovery rate at a user-specified level RyX∈RpR_{yX} \in \mathbb{R}^p1: Stage 1: Screening.

  • Randomly split the data RyX∈RpR_{yX} \in \mathbb{R}^p2.
  • Apply penalized mRMR to RyX∈RpR_{yX} \in \mathbb{R}^p3 samples; retain a working set RyX∈RpR_{yX} \in \mathbb{R}^p4 of size RyX∈RpR_{yX} \in \mathbb{R}^p5 with RyX∈RpR_{yX} \in \mathbb{R}^p6.

Stage 2: Knockoff+ Selection.

  • Generate model-X knockoff variables RyX∈RpR_{yX} \in \mathbb{R}^p7 for RyX∈RpR_{yX} \in \mathbb{R}^p8 (e.g., using the equi-correlated construction).
  • On RyX∈RpR_{yX} \in \mathbb{R}^p9 samples, solve penalized mRMR over all (RyX)j=D(Xj,Y)(R_{yX})_j = D(X_j, Y)0 and (RyX)j=D(Xj,Y)(R_{yX})_j = D(X_j, Y)1, yielding weights (RyX)j=D(Xj,Y)(R_{yX})_j = D(X_j, Y)2, (RyX)j=D(Xj,Y)(R_{yX})_j = D(X_j, Y)3.
  • Compute (RyX)j=D(Xj,Y)(R_{yX})_j = D(X_j, Y)4.
  • Define the knockoff threshold at FDR level (RyX)j=D(Xj,Y)(R_{yX})_j = D(X_j, Y)5:

(RyX)j=D(Xj,Y)(R_{yX})_j = D(X_j, Y)6

  • Select (RyX)j=D(Xj,Y)(R_{yX})_j = D(X_j, Y)7. Theoretical guarantees (Thm 3.3) ensure (RyX)j=D(Xj,Y)(R_{yX})_j = D(X_j, Y)8.

5. Oracle Screening Properties and Theoretical Guarantees

Consider a sequence of problems with (RyX)j=D(Xj,Y)(R_{yX})_j = D(X_j, Y)9 features, true active set size RXX∈Rp×pR_{XX} \in \mathbb{R}^{p \times p}0, and oracle pseudo-truth RXX∈Rp×pR_{XX} \in \mathbb{R}^{p \times p}1 (with RXX∈Rp×pR_{XX} \in \mathbb{R}^{p \times p}2 for RXX∈Rp×pR_{XX} \in \mathbb{R}^{p \times p}3, RXX∈Rp×pR_{XX} \in \mathbb{R}^{p \times p}4 otherwise). Key assumptions include uniqueness of RXX∈Rp×pR_{XX} \in \mathbb{R}^{p \times p}5, bounded eigenvalues of RXX∈Rp×pR_{XX} \in \mathbb{R}^{p \times p}6, and regularity of the penalty:

  • RXX∈Rp×pR_{XX} \in \mathbb{R}^{p \times p}7
  • RXX∈Rp×pR_{XX} \in \mathbb{R}^{p \times p}8
  • Minimum signal: RXX∈Rp×pR_{XX} \in \mathbb{R}^{p \times p}9

Under these, Theorem 3.1 shows existence of a local minimizer (RXX)j,k=D(Xj,Xk)(R_{XX})_{j,k} = D(X_j, X_k)0 with

(RXX)j,k=D(Xj,Xk)(R_{XX})_{j,k} = D(X_j, X_k)1

Theorem 3.2 (sparsistency) asserts that, for SCAD/MCP, if (RXX)j,k=D(Xj,Xk)(R_{XX})_{j,k} = D(X_j, X_k)2 and (RXX)j,k=D(Xj,Xk)(R_{XX})_{j,k} = D(X_j, X_k)3,

(RXX)j,k=D(Xj,Xk)(R_{XX})_{j,k} = D(X_j, X_k)4

i.e., all inactive features are correctly excluded as (RXX)j,k=D(Xj,Xk)(R_{XX})_{j,k} = D(X_j, X_k)5 (Naylor et al., 26 Aug 2025). This provides formal oracle-screening guarantees in high-dimensional regimes.

6. FDR Threshold Choice and Output Guarantees

The FDR level (RXX)j,k=D(Xj,Xk)(R_{XX})_{j,k} = D(X_j, X_k)6 is a direct control lever for the model complexity. Practical choices are (RXX)j,k=D(Xj,Xk)(R_{XX})_{j,k} = D(X_j, X_k)7; lower values yield sparser selections and lower FDP but may reduce power. If no variables are selected for a conservative (RXX)j,k=D(Xj,Xk)(R_{XX})_{j,k} = D(X_j, X_k)8, a variant “SmRMR(RXX)j,k=D(Xj,Xk)(R_{XX})_{j,k} = D(X_j, X_k)9” procedure increments max⁡w≥0Q(w):=w⊤RyX−12w⊤RXXw,\max_{w \geq 0} Q(w) := w^\top R_{yX} - \frac{1}{2} w^\top R_{XX} w,0 until at least one selection is made, trading a marginal FDR increase for guaranteed nonempty sets. Cross-validation for max⁡w≥0Q(w):=w⊤RyX−12w⊤RXXw,\max_{w \geq 0} Q(w) := w^\top R_{yX} - \frac{1}{2} w^\top R_{XX} w,1 is performed only during the knockoff stage, not at initial screening.

7. Empirical Comparisons and Implementation Considerations

SmRMR with SCAD/MCP matches the predictive accuracy of HSIC-LASSO while selecting max⁡w≥0Q(w):=w⊤RyX−12w⊤RXXw,\max_{w \geq 0} Q(w) := w^\top R_{yX} - \frac{1}{2} w^\top R_{XX} w,2–max⁡w≥0Q(w):=w⊤RyX−12w⊤RXXw,\max_{w \geq 0} Q(w) := w^\top R_{yX} - \frac{1}{2} w^\top R_{XX} w,3 fewer features and maintaining FDR below max⁡w≥0Q(w):=w⊤RyX−12w⊤RXXw,\max_{w \geq 0} Q(w) := w^\top R_{yX} - \frac{1}{2} w^\top R_{XX} w,4—a more conservative selection than HSIC-LASSO, which often has higher FDP. Classical discrete mRMR is impractical for large max⁡w≥0Q(w):=w⊤RyX−12w⊤RXXw,\max_{w \geq 0} Q(w) := w^\top R_{yX} - \frac{1}{2} w^\top R_{XX} w,5 and less effective in redundancy control. On GWAS and high-dimensional biological datasets, SmRMR yields similar or better test performance than HSIC-LASSO with fewer selected SNPs/genes, simplifying interpretation.

Key implementation considerations:

  • Each mRMR fit incurs max⁡w≥0Q(w):=w⊤RyX−12w⊤RXXw,\max_{w \geq 0} Q(w) := w^\top R_{yX} - \frac{1}{2} w^\top R_{XX} w,6 complexity (kernel or max⁡w≥0Q(w):=w⊤RyX−12w⊤RXXw,\max_{w \geq 0} Q(w) := w^\top R_{yX} - \frac{1}{2} w^\top R_{XX} w,7-statistic construction); block HSIC or projection correlation can accelerate screening.
  • Knockoff requires max⁡w≥0Q(w):=w⊤RyX−12w⊤RXXw,\max_{w \geq 0} Q(w) := w^\top R_{yX} - \frac{1}{2} w^\top R_{XX} w,8, necessitating the two-stage split; data recycling may increase power.
  • Projection correlation (PCmax⁡w≥0Q(w):=w⊤RyX−12w⊤RXXw,\max_{w \geq 0} Q(w) := w^\top R_{yX} - \frac{1}{2} w^\top R_{XX} w,9) may be preferable over HSIC for heavy-tailed or non-Euclidean data.

Code is publicly available at https://github.com/PeterJackNaylor/SmRMR (Naylor et al., 26 Aug 2025).

Definition Search Book Streamline Icon: https://streamlinehq.com
References (1)

Topic to Video (Beta)

No one has generated a video about this topic yet.

Whiteboard

No one has generated a whiteboard explanation for this topic yet.

Follow Topic

Get notified by email when new papers are published related to Weighted and Penalized mRMR.