Estimating the proportion of true null hypotheses with application in microarray data (1904.13282v2)
Abstract: A new formulation for the proportion of true null hypotheses $(\pi_0)$, based on the sum of all $p$-values and the average of expected $p$-value under the false null hypotheses has been proposed in the current work. This formulation of the parameter of interest $\pi_0$ has also been used to construct a new estimator for the same. The proposed estimator removes the problem of choosing tuning parameters in the existing estimators. Though the formulation is quite general, computation of the new estimator demands use of an initial estimate of $\pi_0$. The issue of choosing an appropriate initial estimator is also discussed in this work. The current work assumes normality of each gene expression level and also assumes similar tests for all the hypotheses. Extensive simulation study shows that, the proposed estimator performs better than its closest competitor, the estimator proposed in Cheng et al., 2015 over a substantial continuous subinterval of the parameter space, under independence and weak dependence among the gene expression levels. The proposed method of estimation is applied to two real gene expression level data-sets and the results are in line with what is obtained by the competing method.