IDfN: Individual Deviation from Normality
- IDfN is a domain-specific measure that quantifies how an individual unit departs from an explicitly modeled notion of normality across diverse settings.
- It unifies various methodologies—from hierarchical Bayesian modeling and anomaly detection to finite-size equilibrium analysis and spatial testing—under a comparative framework.
- Applications of IDfN include detecting anomalous behavior, informing model comparison with Bayes factors, and providing both scalar scores and directional diagnostics.
Individual Deviation from Normality (IDfN) is used in recent literature as a domain-specific measure of how a single unit departs from an explicitly modeled notion of “normal” behavior. The unit may be a subject-specific effect in a hierarchical Bayesian model, an unlabeled instance in anomaly detection, a sample drawn from a finite- equilibrium law, a laboratory measurement conditioned on patient history, an entity-time pair in streaming multivariate data, or a projection direction in a multivariate spatial process. The literature therefore suggests that IDfN is not a single canonical statistic, but a family of structurally related quantities defined relative to the normality concept adopted in each field (Faulkenberry, 2021, Soenen et al., 2023, Shim, 28 Oct 2025, Shah et al., 18 May 2026, Kor et al., 22 Sep 2025, Chen et al., 2020).
1. Conceptual scope and terminological variation
A recurrent source of ambiguity is that “normality” has different meanings across the cited works. In behavioral individual-differences modeling, normality refers to the additive-scale hierarchical specification of subject effects around a group effect. In AD-MERCS, it refers to local density-supported patterns in low-dimensional subspaces. In the finite- statistical-mechanics formulation, it refers to the Gaussian equilibrium law recovered only in the thermodynamic limit. In NORMA, it refers to a patient-specific but population-anchored forecast distribution for the next biomarker value. In online change-point detection, it refers to low reconstruction error under an autoencoder trained on normal windows. In multivariate spatial testing, it refers to Gaussianity of all one-dimensional projections under the union-intersection principle (Faulkenberry, 2021, Soenen et al., 2023, Shim, 28 Oct 2025, Shah et al., 18 May 2026, Kor et al., 22 Sep 2025, Chen et al., 2020).
| Setting | Unit of analysis | IDfN quantity |
|---|---|---|
| Behavioral tasks | Subject | Deviation structure of around |
| AD-MERCS | Instance | Final anomaly score |
| Finite- equilibrium | Sample from | Short-tailed deviation parameterized by |
| NORMA | Observation 0 | 1 |
| Quickest CPD | Entity-time pair 2 | 3 |
| Spatial MVN testing | Direction 4 | 5 |
This variation is substantive rather than merely terminological. A plausible implication is that IDfN is best treated as a comparative framework for per-unit departure scores, not as a universally fixed estimator. What is shared across the formulations is the presence of a reference model of normality and a mapping from an individual unit to a scalar or structured deviation.
2. Hierarchical Bayesian constraint models in behavioral individual differences
In the setting of Haaf and Rouder’s hierarchical Bayesian mixed-effects models, IDfN refers to the way individual-level effects are allowed to deviate around a common group effect, and how that deviation is constrained or left unconstrained (Faulkenberry, 2021). In the simplest two-condition case, with response time 6 for subject 7 on trial 8 and condition indicator 9, the within-subject model is
0
where 1 is the grand-mean intercept, 2 is the subject-specific random intercept, 3 is the subject-specific random effect, and 4 is the residual variance.
The individual-difference structure is determined by the prior on 5. Haaf and Rouder consider four models. In the unconstrained model 6,
7
with a 8-prior reparameterization in which 9, 0, 1, and 2. The positive-effects model 3 truncates that normal prior below at 4. The common-effect model 5 sets 6 for all 7. The null model 8 sets 9 for all 0. Intercepts are typically given 1 with its own weakly informative hyperprior. The only difference across 2 is the way 3 is constrained or left free on the additive scale.
A common criticism is that the observed data are assumed to be drawn from a normal distribution even though response-time distributions are well known to be non-normal. Faulkenberry examined a shifted-lognormal alternative by fixing a small shift 4, defining 5, and retaining the same hierarchical priors on 6, 7, 8, 9, and 0. Under this formulation, 1 is interpreted multiplicatively on the original response-time scale, with 2 as the corresponding factor.
Model comparison proceeds through Bayes factors,
3
For 4, 5, and 6, Haaf and Rouder use the closed-form results of Rouder et al. (2012) in the BayesFactor R package. For 7 versus 8, the encompassing-prior method compares the posterior proportion of draws satisfying 9 to the corresponding prior proportion.
Faulkenberry’s two numerical-cognition case studies show that the overall pattern of inference is essentially unchanged under the shifted-lognormal alternative (Faulkenberry, 2021). In the size-congruity dataset (0, 19 499 RTs), the normal model yielded posterior 1 ms, individual 2 shrunk into 3 ms from an observed range 4 ms, 5, and overwhelming preference for 6 over 7 (8) and 9 (0). The lognormal model gave 1, corresponding to 2 or approximately 3 ms, with the same shrinkage pattern and 4. In the unit-decade compatibility dataset (5, 11 600 RTs), the normal model gave posterior 6 ms, individual 7 ms from observed 8 ms, and 9; the lognormal model gave 0, 1 or approximately 2 ms, and 3. In both datasets, the positive-effects model best predicted the data.
The resulting controversy is interpretive rather than decisional. Under the normal response-time model, 4 is an additive difference in milliseconds; under the log model, it becomes multiplicative on the original scale. Faulkenberry therefore recommends the normal-response-time formulation as a pragmatic approach for modeling individual differences in behavioral tasks, because it preserves an additive interpretation without materially affecting inferences about IDfN (Faulkenberry, 2021).
3. AD-MERCS and instance-level deviation in unsupervised anomaly detection
AD-MERCS models both normality and abnormality in unsupervised anomaly detection by identifying low-dimensional subspaces in which patterns exist and conditions that characterize instances that deviate from those patterns (Soenen et al., 2023). The method builds on MERCS, “multi-directional ensembles of regression and classification trees.” Given an unlabeled dataset 5, one decision tree 6 is grown for each attribute 7 as target, using standard top-down splitting such as CART to maximize information gain or variance reduction. Each tree identifies a low-dimensional subspace 8, and each leaf 9 defines a rectangular cell in 0 within which the data are assumed to share the same functional pattern for 1.
Instead of residuals 2, AD-MERCS fits a local Gaussian kernel density estimate within each leaf: 3 with bandwidth 4 chosen by the method of Botev et al. (2010). Because raw densities are incomparable across leaves and unbounded above, they are squashed into a likelihood score 5 using a threshold 6 chosen so that exactly 7 of the training densities in 8 fall below 9; in practice 00. If a leaf is noisier than some ancestor node, the ancestor’s density estimate is reused so that scoring never relies on a weaker pattern than what was available higher in the tree.
The initial instance-level IDfN is formed from per-tree local anomaly probabilities
01
and combined through a noisy-OR model: 02 where 03 is an inhibition parameter. AD-MERCS then iteratively refines the score by introducing context scores 04 for anomalous contexts. Instance scores are recomputed by blending normal-pattern anomaly and context anomaly, and context scores are recomputed by a noisy-AND over the instances they contain. After convergence, the final 05 is the IDfN score of 06.
The paper states several properties of this construction (Soenen et al., 2023). By construction, each 07. The noisy-OR/AND updates guarantee monotonic non-decrease of 08 throughout iterations. The split-selection criterion ensures that each tree isolates subspaces of maximal predictive strength for its target. No asymptotic bounds are given, but empirically the method converges in few iterations, typically 09–10.
The examples clarify why density-based scoring differs from mean-residual scoring. In a bimodal leaf with target values 11, a point at 12 has zero residual to the mean 13 and would raise no alarm under residual-based scoring, but 14, hence 15, 16, and the point is maximally anomalous. In the Zoo dataset, the animal “scorpion” is assigned high scores in multiple contexts, including “(has_spine = false)” predicting “has_tail,” yielding the explanation “Among animals with no backbone, having a tail is extremely unlikely,” and another context in which “does not lay eggs” predicts “has_teeth.” The final 17 is near 18, with two concrete subspace-context explanations. This makes IDfN not only a scalar anomaly score but also an explanatory device via low-dimensional deviations from learned normality (Soenen et al., 2023).
4. Finite-19 equilibrium distributions and detectable non-Gaussianity
In the finite-20 statistical-mechanics formulation, IDfN is the systematic short-tailed deviation from Gaussian equilibrium induced by finite particle number 21 (Shim, 28 Oct 2025). For a one-dimensional ideal-gas-like system with total kinetic energy 22, maximizing Havrda–Charvát (Tsallis) entropy yields a compact-support 23-Gaussian density
24
with
25
Taking 26, the support is 27. In the thermodynamic limit 28, 29 and 30 converges to the Maxwell-Boltzmann Gaussian law. For finite 31, the support is strictly compact and the tails are shorter than Gaussian.
This formulation makes the departure from Gaussianity fully parameterized by 32. As 33 decreases toward the minimum 34, 35 falls below 36 and the support shrinks; as 37 grows, all finite-38 effects vanish smoothly. A plausible implication is that 39 serves as a physically interpretable non-Gaussianity index in settings where the system is literally or notionally finite.
Shim studies how five standard normality tests respond to this IDfN: Kolmogorov-Smirnov, Anderson-Darling, Cramér-von Mises, Jarque-Bera, and Shapiro-Wilk (Shim, 28 Oct 2025). The Monte Carlo design uses system sizes 40, sample sizes 41 for 42, and for 43 extends to 44, with 45 independent draws per 46. KS and AD calibrated under the correct custom null 47 maintain nominal size 48. Against the Gaussian null, all five statistics tend to grow as 49 becomes smaller, but the moment-based tests JB and SW detect short-tailed compact-support deviations far more efficiently than the ECDF-based tests KS, AD, and CvM.
The power map is highly uneven (Shim, 28 Oct 2025). For strongly non-Gaussian cases 50, JB and SW achieve nearly 51 power even for 52, whereas KS, AD, and CvM plateau around 53–54. For 55, ECDF tests rarely exceed 56–57 power at 58, while JB and SW still exceed 59–60. At the near-Gaussian case 61, ECDF-based tests remain below 62 power up to 63, reaching approximately 64 only at 65; JB and SW exceed 66 power by 67 and hit 68 by 69.
The practical guidance follows directly. When one suspects mild, short-tailed deviation from normality due solely to finite-size constraints, Jarque-Bera or Shapiro-Wilk are preferred. As a rule of thumb, for 70, 71 suffices with JB or SW for 72 power at the 73 level; for 74, 75–2000 suffices; for 76, 77–5000 is needed for JB or SW, whereas 78 would be required for KS, AD, or CvM (Shim, 28 Oct 2025).
5. Personalized laboratory interpretation and NORMA
In laboratory medicine, IDfN is formalized as a personalized deviation score derived from a conditional forecast distribution for the next biomarker value (Shah et al., 18 May 2026). NORMA models blood biomarkers from longitudinal histories using a conditional transformer that learns
79
where 80 is the vector of laboratory values at time 81, 82 is the patient’s history, and 83 denotes population-level “healthy” prior information such as central 84 intervals 85. The prediction head is either Gaussian, 86, or quantile-based with five quantiles 87.
The input representation combines static covariates, laboratory history, and a “normal” query token. The context token is
88
the history token is
89
with 90 and 91, and the query token is
92
These tokens are processed by transformer decoder layers with multi-head self-attention and feed-forward blocks. The training objective is either the Gaussian negative log-likelihood
93
or the quantile pinball loss
94
From the learned distribution, NORMA extracts either 95 and 96, or 97 and
98
The IDfN score is then the personalized z-score
99
and the personalized 00 reference interval is
01
or directly 02 for the quantile head.
The main methodological issue in this literature is over-personalization. The paper states that purely personalized intervals can overfit to sparse data, inflate false-positive rates, and include unrecognized or subclinical disease; purely personalized intervals routinely overfit, classifying up to 03 of measurements as abnormal, without corresponding associations with adverse clinical outcomes (Shah et al., 18 May 2026). Population intervals, by contrast, ignore stable intra-patient variability. NORMA balances these two extremes by conditioning on both 04 and 05, and by using a “normal” query token at each step. The paper presents a Bayesian-shrinkage analogy,
06
but states that NORMA learns a non-parametric, time-aware version of this anchoring.
The empirical results span nearly 07 billion longitudinal laboratory measurements from over 08 million individuals across North America, the Middle East, and East Asia (Shah et al., 18 May 2026). In CHS, abnormal flags were 09 for 10, 11 for 12, and 13 for 14; among 15-normal measurements, 16 reclassified 17 and NORMA 18; median lead time for 19 versus 20 was 21 months (IQR 22–23). In eICU, abnormal flags were 24, 25, and 26, with reclassification 27 for 28 and 29 for NORMA, and lead time 30 h (31–32). In INSPIRE, abnormal flags were 33, 34, and 35, with reclassification 36 and 37, and lead time 38 h (39–40). Among 41-normal cases in eICU, the positive predictive value per 42 reclassified for in-hospital mortality was 43 for 44 versus 45 for NORMA, for acute kidney injury 46 versus 47, and for prolonged length of stay 48 versus 49. The clinical interpretation is that IDfN in this setting is a calibrated patient-specific abnormality score anchored to a model of normal variation rather than to either raw population intervals or unconstrained personal baselines.
6. Online multi-entity change-point detection
In online change-point detection from multi-entity, multivariate time series, IDfN is defined as a reconstruction-error-based per-entity deviation under an autoencoder trained on normal behavior (Kor et al., 22 Sep 2025). Each sensor channel is first z-score normalized using training-set means 50 and standard deviations 51, and for entity 52 at time 53 a sliding window of length 54 is formed,
55
A “Simple Autoencoder” then reconstructs the flattened window. The IDfN at time 56 for entity 57 is the mean squared error
58
The architecture is a stack of fully connected dense layers with ReLU activations, a symmetric decoder, a final linear output layer, and dropout of 59 after each hidden layer. Training is done on normal data only with loss
60
using Adam with initial learning rate 61, batch size 62, up to 63 epochs with early stopping after 64 epochs without validation improvement, and learning-rate reduction on plateau by a factor of 65 down to 66. The recommended default window length is 67.
Per-entity IDfNs are aggregated into system-wide anomaly scores (SWAS) in three ways (Kor et al., 22 Sep 2025). The mean score is
68
the variance score is
69
and the distributional score uses a Gaussian-kernel KDE of the current IDfN distribution compared with a reference KDE over all training IDfNs, with SWAS
70
the Wasserstein-1 distance. These statistics are converted to deviation scores 71, 72, or 73, accumulated by CUSUM,
74
and then tested by an adaptive sequential density-based thresholding rule after the log transform 75. A change is declared at time 76 if the KDE density estimate at 77 falls below a threshold 78, taken as the 79th percentile of the training CUSUM-KDE densities, after a burn-in of 80 samples.
The validation datasets include auto-regressive series with 81 entities and 82, coupled Chen chaotic oscillators with 83, 84, and 85, and two Unity crowd-simulation settings using upper-leg acceleration streams with 86: train-station evacuation with approximately 87 agents and a change around 88 s, and bidirectional corridor collision with approximately 89 agents and a change around 90 s (Kor et al., 22 Sep 2025). The paper reports that IDfN trajectories show a clear rise at the ground-truth event time, KDE heatmaps visualize a shifting density of reconstruction errors with a sudden mode change at the incident, aggregated SWAS plots spike near the true change, and CUSUM curves cross adaptive thresholds in tight proximity to the ground-truth change point. In this formulation, IDfN is explicitly local—defined per entity and per time step—but is designed to support system-level inference after aggregation.
7. Directional IDfN in multivariate spatial normality testing
In multivariate spatial statistics, IDfN is defined as a directional departure from Gaussianity after accounting for spatial dependence (Chen et al., 2020). The overall null is
91
By the union-intersection principle, 92 holds if and only if every projection
93
is Gaussian for every unit vector 94. In practice, the continuum of directions is replaced by a finite set 95, such as the coordinate axes, the eigenvectors of an estimate of 96, or random draws from the unit sphere.
For a fixed direction 97, projected residuals are standardized as
98
and the sample skewness and kurtosis are
99
Under spatial dependence, the asymptotic variances of 00 and 01 depend on the covariance structure of the projected process. Horváth et al. (2020) propose kernel-smoothed estimators
02
03
where 04 is the sample autocovariance at lag 05, 06 is a univariate kernel such as Bartlett, and 07 is a bandwidth vector.
The directional Jarque-Bera-type statistic is then
08
Under 09, the standardized skewness and kurtosis are jointly Gaussian and asymptotically independent, so 10. At nominal level 11, the critical value is 12; for 13, 14. The global decision rule rejects multivariate normality when
15
or equivalently when any directional 16-value is at most 17, possibly with Bonferroni correction or Benjamini-Hochberg control.
In this framework, the directional score itself is interpreted as IDfN: 18 A large 19 means that, along direction 20, the data exhibit unusually large skewness or excess kurtosis after accounting for spatial dependence. This provides both a global multivariate normality test and a directional diagnostic. The method thereby makes it possible to locate which linear combinations of variables are most non-Gaussian, rank them by IDfN, and map out a “normality contour” on the unit sphere in 21 (Chen et al., 2020).
Across these formulations, IDfN functions less as a single theory than as a recurring analytic pattern: choose a reference notion of normality, define an individual unit relative to that reference, and quantify the unit’s departure in a way that is appropriate to the data-generating assumptions of the field. The main controversies concern which normality model is substantively appropriate and how much interpretability is lost when one moves from additive, physically transparent, or clinically transparent scales to transformed or highly adaptive representations. The cited literature consistently treats those choices as consequential for interpretation, but not always for the qualitative ordering of evidence or detection performance (Faulkenberry, 2021, Soenen et al., 2023, Shim, 28 Oct 2025, Shah et al., 18 May 2026, Kor et al., 22 Sep 2025, Chen et al., 2020).