Fisher Criterion: Discriminability & Estimation
- Fisher criterion is a family of optimization principles that quantify discriminability and estimation precision via class scatter matrices or Fisher information.
- Generalized extensions address classical limitations by jointly optimizing feature selection and capturing redundancy in high-dimensional settings.
- Modern adaptations apply the criterion across supervised, unsupervised, and quantum settings to enhance model selection and parameter estimation.
Searching arXiv for recent and foundational papers on Fisher criterion across discriminant analysis, information-based estimation, and related modern extensions. to=arxiv_search.search 】【:】【“】【json {"3query3 criterion\"3 OR title:\3"Fisher criterion\"","max_results":3all:\3query3,"sort_by":"relevance"}【อ่านข้อความเต็มjson to=arxiv_search.search code 手机天天彩票: Done &&&3query3&&&)","title":"The impact of the point spread function fitting radius on photometric uncertainty based on the Fisher information matrix","authors":["J. Mendez", "F. Andrés-Trias", "A. Márquez-Lago", "B. Cuevas-Tello"],"abstract":"In point spread function (PSF) photometry, the selection of the fitting aperture radius plays a critical role in determining the precision of flux and background estimations. Traditional methods often rely on maximizing the signal-to-noise ratio (S/N) as a criterion for aperture selection. However, S/N-based approaches do not necessarily provide the optimal precision for joint estimation problems as they do not account for the statistical limits imposed by the Fisher information in the context of the Cram\u3query3query3e9r-Rao lower bound (CRLB). This study aims to establish an alternative criterion for selecting the optimal fitting radius based on Fisher information rather than S/N. Fisher information serves as a fundamental measure of estimation precision, providing theoretical guarantees on the achievable accuracy for parameter estimation. By leveraging Fisher information, we seek to define an aperture selection strategy that minimizes the loss of precision. We conducted a series of numerical experiments that analyze the behavior of Fisher information and estimator performance as a function of the PSF aperture radius. Specifically, we revisited fundamental photometric models and explored the relationship between aperture size and information content. We compared the empirical variance of classical estimators, such as maximum likelihood and stochastic weighted least squares, against the theoretical CRLB derived from the Fisher information matrix. Our results indicate that aperture selection based on the Fisher information provides a more robust framework for achieving optimal estimation precision.","categories":["astro-ph.IM", "stat.ME"]},{"arxiv_id":"(&&&3all:\3&&&)","title":"Fisher Discriminative Least Squares Regression for Image Classification","authors":["Zhe Chen", "Xing Xu", "Jiwen Lu", "Vishal M. Patel"],"abstract":"Discriminative least squares regression (DLSR) has been shown to achieve promising performance in multi-class image classification tasks. Its key idea is to force the regression labels of different classes to move in opposite directions by means of the proposed the joint use of the epsilon-draggings technique, yielding discriminative regression model exhibiting wider margins, and the Fisher criterion. The epsilon-draggings technique ignores an important problem: its non-negative relaxation matrix is dynamically updated in optimization, which means the dragging values can also cause the labels from the same class to be uncorrelated. In order to learn a more powerful discriminative projection, as well as regression labels, we propose a Fisher regularized DLSR (FDLSR) framework by constraining the relaxed labels using the Fisher criterion. On one hand, the Fisher criterion improves the intra-class compactness of the relaxed labels during relaxation learning. On the other hand, it is expected further to enhance the inter-class separability of epsilon-draggings technique. FDLSR for the first time ever attempts to integrate the Fisher discriminant criterion and epsilon-draggings technique into one unified model because they are absolutely complementary in learning discriminative projection. Extensive experiments on various datasets demonstrate that the proposed FDLSR method achieves performance that is superior to other state-of-the-art classification methods.","categories":["cs.CV", "cs.LG"]},{"arxiv_id":"(&&&3 OR title:\3&&&)","title":"An Extension of Fisher's Criterion: Theoretical Results with a Neural Network Realization","authors":["Mahadewi Tandrasasmita", "P. Sadeghi", "P. Vial", "R. Jäntti"],"abstract":"Fisher's criterion is a widely used tool in machine learning for feature selection. For large search spaces, Fisher's criterion can provide a scalable solution to select features. A challenging limitation of Fisher's criterion, however, is that it performs poorly when mean values of class-conditional distributions are close to each other. Motivated by this challenge, we propose an extension of Fisher's criterion to overcome this limitation. The proposed extension utilizes the available heteroscedasticity of class-conditional distributions to distinguish one class from another. Additionally, we describe how our theoretical results can be casted into a neural network framework, and conduct a proof-of-concept experiment to demonstrate the viability of our approach to solve classification problems.","categories":["stat.ML", "cs.LG", "cs.NE"]},{"arxiv_id":"(Gu et al., 2012)","title":"Generalized Fisher Score for Feature Selection","authors":["Feiping Nie", "Dong Xu", "Ivor Tsang", "Changshui Zhang"],"abstract":"Fisher score is one of the most widely used supervised feature selection methods. However, it selects each feature independently according to their scores under the Fisher criterion, which leads to a suboptimal subset of features. In this paper, we present a generalized Fisher score to jointly select features. It aims at finding an subset of features, which maximize the lower bound of traditional Fisher score. The resulting feature selection problem is a mixed integer programming, which can be reformulated as a quadratically constrained linear programming (QCLP). It is solved by cutting plane algorithm, in each iteration of which a multiple kernel learning problem is solved alternatively by multivariate ridge regression and projected gradient descent. Experiments on benchmark data sets indicate that the proposed method outperforms Fisher score as well as many other state-of-the-art feature selection methods.","categories":["cs.LG", "cs.CV"]},{"arxiv_id":"(Jouvin et al., 2020)","title":"A Bayesian Fisher-EM algorithm for discriminative Gaussian subspace clustering","authors":["Tatiana Ledoux", "Mikael Chavent", "Julie Saracco", "Frédéric Bertrand"],"abstract":"High-dimensional data clustering has become and remains a challenging task for modern statistics and machine learning, with a wide range of applications. We consider in this work the powerful discriminative latent mixture model, and we extend it to the Bayesian framework. Modeling data as a mixture of Gaussians in a low-dimensional discriminative subspace, a Gaussian prior distribution is introduced over the latent group means and a family of twelve submodels are derived considering different covariance structures. Model inference is done with a variational EM algorithm, while the discriminative subspace is estimated via a Fisher-step maximizing an unsupervised Fisher criterion. An empirical Bayes procedure is proposed for the estimation of the prior hyper-parameters, and an integrated classification likelihood criterion is derived for selecting both the number of clusters and the submodel. The performances of the resulting Bayesian Fisher-EM algorithm are investigated in two thorough simulated scenarios, regarding both dimensionality as well as noise and assessing its superiority with respect to state-of-the-art Gaussian subspace clustering models. In addition to standard real data benchmarks, an application to single image denoising is proposed, displaying relevant results. This work comes with a reference implementation for the R software in the FisherEM package accompanying the paper.","categories":["stat.ML", "math.ST", "stat.CO"]},{"arxiv_id":"(Zhang et al., 2020)","title":"A sufficient Entanglement Criterion Based On Quantum Fisher Information and Variance","authors":["Che-Ming Li", "Pei-Wen Liu", "Kang-Da Wu", "Teiko Heinz", "Yuan-Yuan Zhao", "Renjie-Shi Wang", "Felix Huber"],"abstract":"We derive criterion in the form of inequality based on quantum Fisher information and quantum variance to detect multipartite entanglement. It can be regarded as complementary of the well-established PPT criterion in the sense that it can also detect bound entangled states. The inequality is motivated by Y.Akbari-Kourbolagh et al.[Phys. Rev A. 99, 3query3all:\3 OR title:\33query34 (3 OR title:\3query3all:\39)] which introduced a multipartite entanglement criterion based on quantum Fisher information. Our criterion is experimentally measurable for detecting any N-qudit pure state mixed with white noisy. We take several examples to illustrate that our criterion has good performance for detecting certain entangled states.","categories":["quant-ph"]},{"arxiv_id":"(Yang et al., 2020)","title":"Quantum Fisher information-based detection of genuine tripartite entanglement","authors":["Xiao-Hui Jia", "Xian-Rong Jin", "Shao-Ming Fei"],"abstract":"Genuine multipartite entanglement plays important roles in quantum information processing. The detection of genuine multipartite entanglement has been long time a challenging problem in the theory of quantum entanglement. We propose a criterion for detecting genuine tripartite entanglement of arbitrary dimensional tripartite states based on quantum Fisher information. We show that this criterion is more effective for some states in detecting genuine tripartite entanglement by detailed example.","categories":["quant-ph"]},{"arxiv_id":"(Gohain et al., 2022)","title":"Robust Information Criterion for Model Selection in Sparse High-Dimensional Linear Regression Models","authors":["Arian Maleki","Shervin A. Aviyente"],"abstract":"Model selection in linear regression models is a major challenge when dealing with high-dimensional data where the number of available measurements (sample size) is much smaller than the dimension of the parameter space. Traditional methods for model selection such as Akaike information criterion, Bayesian information criterion (BIC) and minimum description length are heavily prone to overfitting in the high-dimensional setting. In this regard, extended BIC (EBIC), which is an extended version of the original BIC and extended Fisher information criterion (EFIC), which is a combination of EBIC and Fisher information criterion, are consistent estimators of the true model as the number of measurements grows very large. However, EBIC is not consistent in high signal-to-noise-ratio (SNR) scenarios where the sample size is fixed and EFIC is not invariant to data scaling resulting in unstable behaviour. In this paper, we propose a new form of the EBIC criterion called EBIC-Robust, which is invariant to data scaling and consistent in both large sample size and high-SNR scenarios. Analytical proofs are presented to guarantee its consistency. Simulation results indicate that the performance of EBIC-Robust is quite superior to that of both EBIC and EFIC.","categories":["math.ST", "stat.TH", "cs.IT"]}]
The Fisher criterion denotes a family of optimization principles that measure discriminability or inferential precision through a ratio, difference, or information functional derived from class scatter or Fisher information. In its classical usage in discriminant analysis, it seeks projections that maximize between-class separation relative to within-class dispersion; in later work it appears in feature selection, relaxed-label regression, unsupervised discriminative subspace learning, Fisher-information-based estimation and model selection, and quantum-information witnesses. This suggests that the term is best understood as a unifying statistical idea rather than a single fixed formula (&&&3all:\3&&&, Gohain et al., 2022, &&&3all:\3query3&&&, &&&3all:\3all:\3&&&).
3all:\3. Classical discriminant formulation
In the supervised multiclass setting, let PRESERVED_PLACEHOLDER_3query3^ have PRESERVED_PLACEHOLDER_3all:\3^ classes, with class means PRESERVED_PLACEHOLDER_3 OR title:\3, global mean , and class sizes . The classical within-class and between-class scatter matrices are
Fisher’s Linear Discriminant seeks a projection maximizing class separability. Two closely related objectives are standard: for the trace-ratio form, and
for the ratio-trace form. The latter yields the generalized eigenproblem
When PRESERVED_PLACEHOLDER_3all:\3query3^ is nonsingular, the discriminant subspace is given by the top eigenvectors of PRESERVED_PLACEHOLDER_3all:\3all:\3, and the subspace dimension satisfies PRESERVED_PLACEHOLDER_3all:\3 OR title:\3^ because PRESERVED_PLACEHOLDER_3all:\33^ (&&&3all:\3&&&).
The central interpretation is that the criterion rewards large separation among class means while penalizing large within-class variance. In the scalar-feature setting this becomes the familiar Fisher score. For feature PRESERVED_PLACEHOLDER_3all:\3max_results7^ a common multiclass form is
PRESERVED_PLACEHOLDER_3all:\35
and in the binary case it reduces to
PRESERVED_PLACEHOLDER_3all:\36
These scalar forms inherit the same logic as LDA: numerator terms encode mean separation, denominator terms encode intra-class spread (Gu et al., 2012, &&&3 OR title:\3&&&).
3 OR title:\3. Feature selection, joint optimization, and label-space regularization
A major limitation of the standard Fisher score is that it ranks features independently. The data describe two specific consequences: redundancy, because highly correlated variables can all receive high scores, and missed joint effects, because features that are weak individually can be strong jointly. The “Generalized Fisher Score” reformulates selection as a joint combinatorial problem with binary selector PRESERVED_PLACEHOLDER_3all:\3,7^ PRESERVED_PLACEHOLDER_3all:\38, and objective
PRESERVED_PLACEHOLDER_3all:\39
The paper shows that maximizing this lower bound of the traditional Fisher score can be reformulated as a mixed-integer problem, then as a QCLP, and solved by a cutting plane procedure in which a multiple kernel learning subproblem is handled by multivariate ridge regression and projected gradient descent (Gu et al., 2012).
A distinct modernization appears in Fisher Discriminative Least Squares Regression. Classical least-squares regression to one-hot labels is augmented by PRESERVED_PLACEHOLDER_3 OR title:\3query3-dragging and a Fisher regularizer in the learned label space. With relaxed labels PRESERVED_PLACEHOLDER_3 OR title:\3all:\3, sign matrix PRESERVED_PLACEHOLDER_3 OR title:\3 OR title:\3, nonnegative relaxation PRESERVED_PLACEHOLDER_3 OR title:\33, and regressor PRESERVED_PLACEHOLDER_3 OR title:\3max_results7^ the final objective is
PRESERVED_PLACEHOLDER_3 OR title:\35
The label-space Fisher term is
PRESERVED_PLACEHOLDER_3 OR title:\36
with gradient
PRESERVED_PLACEHOLDER_3 OR title:\37
Here the Fisher criterion is no longer applied to feature projections directly; it regularizes relaxed class targets so that PRESERVED_PLACEHOLDER_3 OR title:\38-dragging widens margins while Fisher regularization restores intra-class compactness and inter-class separability. The resulting block coordinate descent has closed-form updates for PRESERVED_PLACEHOLDER_3 OR title:\39, 3query3, and 3all:\3^ (&&&3all:\3&&&).
3. Unsupervised and heteroscedastic extensions
In clustering, labels are unavailable, so the Fisher criterion is adapted through soft assignments. In Bayesian Fisher-EM, hard class memberships are replaced by responsibilities 3 OR title:\3, producing soft cluster sizes 3, soft means 4, and soft scatter matrices
5
The discriminative subspace 6 is then updated by maximizing the unsupervised Fisher criterion
7
or, in the paper’s numerically stable Fisher-step,
8
with 9. The generalized eigenproblem may be solved directly or via an SVD criterion on 3query3. This embeds Fisher’s separability principle inside a variational EM procedure for Gaussian subspace clustering (Jouvin et al., 2020).
A different extension targets a known failure mode of the classical score: it can perform poorly when class means are close but class-conditional variances differ. Under binary Gaussian class-conditionals 3all:\3, the proposed extension defines the classical ratio
3 OR title:\3^
a class-specific threshold
3
and the extended score
4
The threshold 5 arises from comparing weighted self-overlap and cross-overlap integrals of the class-conditional Gaussians. In the homoscedastic equal-prior limit, 6, so the extension reduces exactly to the classical Fisher criterion. The same paper also describes a neural-network realization using random projections, class-specific KDE activations, and top-7 node selection by 8 (&&&3 OR title:\3&&&).
4. Fisher-information criteria in estimation and model selection
A separate usage of the term appears in estimation theory. Here “Fisher criterion” refers not to class scatter, but to optimization based on the Fisher information matrix and the Cramér–Rao lower bound. The distinction is explicit in the sparse-regression model-selection literature, where Fisher-information-based criteria such as FIC and EFIC are contrasted with Fisher’s linear discriminant analysis (Gohain et al., 2022).
In PSF photometry, for parameters 9 with Poisson means 3query3, the Fisher information matrix over aperture 3all:\3^ has entries
3 OR title:\3^
and the CRLB gives 3. For joint flux-background estimation, the marginal information for flux is
4
The proposed criterion is to choose the smallest fitting radius 5 such that
6
where 7 is the full pixel window and 8 is a tolerated information loss. Because signal-to-noise maximization ignores the off-diagonal coupling term 9, the S/N-optimal radius can produce substantial flux-information losses; the reported discrepancies are typically 3query3–3all:\3^ relative to larger Fisher-guided apertures (&&&3all:\3query3&&&).
In sparse high-dimensional linear regression, Fisher information enters model selection through the sample Fisher matrix
3 OR title:\3^
leading to Fisher-information-aware criteria. EFIC retains the Fisher log-determinant term, whereas EBIC-Robust modifies the normalization to obtain scale invariance: 3 The paper proves consistency both as 4 with fixed 5 and as 6 with 7, provided 8 in the latter regime (Gohain et al., 2022).
5. Quantum Fisher information criteria
In quantum information, Fisher criteria are built from quantum Fisher information (QFI) rather than classical scatter matrices. One line of work gives separability and multipartite-entanglement witnesses. For any separable bipartite state 9 and local Hermitian observables 3query3^ and 3all:\3,
3 OR title:\3^
where the papers use the SLD-based convention 3. Violation implies entanglement, and multipartite generalizations replace the right-hand side by averages or sums of pairwise difference variances. The criterion is experimentally measurable for white-noise mixtures because
4
for 5 (Zhang et al., 2020).
A related tripartite criterion bounds the sum of QFI values over local generator sets. For equal local dimension 6, any biseparable tripartite state satisfies
7
with explicit bounds 8 for one qudit and 9 for two qudits when Gell-Mann generators are used. In the qubit case this becomes the threshold 3query3; exceeding it certifies genuine tripartite entanglement (Yang et al., 2020).
Another line applies Fisher information to channel incompatibility. Given channels 3all:\3^ and output bases 3 OR title:\3, one defines
3
and solves the SDP
4
If the optimum is strictly larger than 5, the channel tuple is incompatible. For depolarizing channels with 6 mutually unbiased bases, the criterion yields incompatibility whenever 7; for noisy Schur channels it yields analytic bounds such as 8 (&&&3all:\3all:\3&&&).
6. Scope, limitations, and recurring misconceptions
A common misconception is that “Fisher criterion” always means Fisher’s linear discriminant. The sparse-regression literature explicitly distinguishes Fisher-information-based criteria from Fisher’s linear discriminant analysis, and the photometry and quantum papers use the term in the CRLB/QFI sense rather than the scatter-ratio sense (Gohain et al., 2022, &&&3all:\3query3&&&).
A second misconception is that the classical Fisher score is universally adequate for feature ranking. The generalized Fisher score paper shows that independent ranking can be suboptimal because it ignores redundancy and joint discriminative effects, while the heteroscedastic extension shows that the classical mean-separation ratio can be near zero even when variance differences still carry discriminative information (Gu et al., 2012, &&&3 OR title:\3&&&).
A third misconception is that Fisher-based optimization always produces an interior optimum. In PSF photometry, the flux information 9 is monotone and saturating in common Gaussian-PSF regimes, so the practical rule is not to maximize 3query3^ over 3all:\3^ but to choose the smallest radius achieving a target information fraction. This contrasts with the S/N curve, which can have an interior maximum (&&&3all:\3query3&&&).
Finally, most Fisher criteria are sufficient or asymptotic rather than universally exact. FDLSR is empirically robust but does not have a joint global-optimality guarantee; BFEM relies on soft responsibilities and regularization of ill-conditioned scatter matrices; EBIC-Robust requires sparse-Riesz-type identifiability conditions for its consistency theorems; QFI-based entanglement and incompatibility inequalities are witnesses, so non-violation does not imply separability or compatibility (&&&3all:\3&&&, Jouvin et al., 2020, Gohain et al., 2022, Zhang et al., 2020, &&&3all:\3all:\3&&&).
Across these formulations, the invariant theme is the same: a Fisher criterion evaluates a representation, model, aperture, or quantum object through the statistical structure that separates alternatives or bounds attainable precision. What changes across domains is the object being optimized—scatter matrices in discriminant analysis, relaxed labels in regression, soft cluster structure in unsupervised learning, Fisher matrices in estimation and model selection, and quantum Fisher information or Fisher-induced SDPs in quantum information.