Papers
Topics
Authors
Recent
Search
2000 character limit reached

IDfN: Individual Deviation from Normality

Updated 12 July 2026
  • IDfN is a domain-specific measure that quantifies how an individual unit departs from an explicitly modeled notion of normality across diverse settings.
  • It unifies various methodologies—from hierarchical Bayesian modeling and anomaly detection to finite-size equilibrium analysis and spatial testing—under a comparative framework.
  • Applications of IDfN include detecting anomalous behavior, informing model comparison with Bayes factors, and providing both scalar scores and directional diagnostics.

Individual Deviation from Normality (IDfN) is used in recent literature as a domain-specific measure of how a single unit departs from an explicitly modeled notion of “normal” behavior. The unit may be a subject-specific effect in a hierarchical Bayesian model, an unlabeled instance in anomaly detection, a sample drawn from a finite-NN equilibrium law, a laboratory measurement conditioned on patient history, an entity-time pair in streaming multivariate data, or a projection direction in a multivariate spatial process. The literature therefore suggests that IDfN is not a single canonical statistic, but a family of structurally related quantities defined relative to the normality concept adopted in each field (Faulkenberry, 2021, Soenen et al., 2023, Shim, 28 Oct 2025, Shah et al., 18 May 2026, Kor et al., 22 Sep 2025, Chen et al., 2020).

1. Conceptual scope and terminological variation

A recurrent source of ambiguity is that “normality” has different meanings across the cited works. In behavioral individual-differences modeling, normality refers to the additive-scale hierarchical specification of subject effects around a group effect. In AD-MERCS, it refers to local density-supported patterns in low-dimensional subspaces. In the finite-NN statistical-mechanics formulation, it refers to the Gaussian equilibrium law recovered only in the thermodynamic limit. In NORMA, it refers to a patient-specific but population-anchored forecast distribution for the next biomarker value. In online change-point detection, it refers to low reconstruction error under an autoencoder trained on normal windows. In multivariate spatial testing, it refers to Gaussianity of all one-dimensional projections under the union-intersection principle (Faulkenberry, 2021, Soenen et al., 2023, Shim, 28 Oct 2025, Shah et al., 18 May 2026, Kor et al., 22 Sep 2025, Chen et al., 2020).

Setting Unit of analysis IDfN quantity
Behavioral tasks Subject ii Deviation structure of δi\delta_i around ν\nu
AD-MERCS Instance xix_i Final anomaly score δi[0,1]\delta_i \in [0,1]
Finite-NN equilibrium Sample from pNp_N Short-tailed deviation parameterized by qNq_N
NORMA Observation NN0 NN1
Quickest CPD Entity-time pair NN2 NN3
Spatial MVN testing Direction NN4 NN5

This variation is substantive rather than merely terminological. A plausible implication is that IDfN is best treated as a comparative framework for per-unit departure scores, not as a universally fixed estimator. What is shared across the formulations is the presence of a reference model of normality and a mapping from an individual unit to a scalar or structured deviation.

2. Hierarchical Bayesian constraint models in behavioral individual differences

In the setting of Haaf and Rouder’s hierarchical Bayesian mixed-effects models, IDfN refers to the way individual-level effects are allowed to deviate around a common group effect, and how that deviation is constrained or left unconstrained (Faulkenberry, 2021). In the simplest two-condition case, with response time NN6 for subject NN7 on trial NN8 and condition indicator NN9, the within-subject model is

ii0

where ii1 is the grand-mean intercept, ii2 is the subject-specific random intercept, ii3 is the subject-specific random effect, and ii4 is the residual variance.

The individual-difference structure is determined by the prior on ii5. Haaf and Rouder consider four models. In the unconstrained model ii6,

ii7

with a ii8-prior reparameterization in which ii9, δi\delta_i0, δi\delta_i1, and δi\delta_i2. The positive-effects model δi\delta_i3 truncates that normal prior below at δi\delta_i4. The common-effect model δi\delta_i5 sets δi\delta_i6 for all δi\delta_i7. The null model δi\delta_i8 sets δi\delta_i9 for all ν\nu0. Intercepts are typically given ν\nu1 with its own weakly informative hyperprior. The only difference across ν\nu2 is the way ν\nu3 is constrained or left free on the additive scale.

A common criticism is that the observed data are assumed to be drawn from a normal distribution even though response-time distributions are well known to be non-normal. Faulkenberry examined a shifted-lognormal alternative by fixing a small shift ν\nu4, defining ν\nu5, and retaining the same hierarchical priors on ν\nu6, ν\nu7, ν\nu8, ν\nu9, and xix_i0. Under this formulation, xix_i1 is interpreted multiplicatively on the original response-time scale, with xix_i2 as the corresponding factor.

Model comparison proceeds through Bayes factors,

xix_i3

For xix_i4, xix_i5, and xix_i6, Haaf and Rouder use the closed-form results of Rouder et al. (2012) in the BayesFactor R package. For xix_i7 versus xix_i8, the encompassing-prior method compares the posterior proportion of draws satisfying xix_i9 to the corresponding prior proportion.

Faulkenberry’s two numerical-cognition case studies show that the overall pattern of inference is essentially unchanged under the shifted-lognormal alternative (Faulkenberry, 2021). In the size-congruity dataset (δi[0,1]\delta_i \in [0,1]0, 19 499 RTs), the normal model yielded posterior δi[0,1]\delta_i \in [0,1]1 ms, individual δi[0,1]\delta_i \in [0,1]2 shrunk into δi[0,1]\delta_i \in [0,1]3 ms from an observed range δi[0,1]\delta_i \in [0,1]4 ms, δi[0,1]\delta_i \in [0,1]5, and overwhelming preference for δi[0,1]\delta_i \in [0,1]6 over δi[0,1]\delta_i \in [0,1]7 (δi[0,1]\delta_i \in [0,1]8) and δi[0,1]\delta_i \in [0,1]9 (NN0). The lognormal model gave NN1, corresponding to NN2 or approximately NN3 ms, with the same shrinkage pattern and NN4. In the unit-decade compatibility dataset (NN5, 11 600 RTs), the normal model gave posterior NN6 ms, individual NN7 ms from observed NN8 ms, and NN9; the lognormal model gave pNp_N0, pNp_N1 or approximately pNp_N2 ms, and pNp_N3. In both datasets, the positive-effects model best predicted the data.

The resulting controversy is interpretive rather than decisional. Under the normal response-time model, pNp_N4 is an additive difference in milliseconds; under the log model, it becomes multiplicative on the original scale. Faulkenberry therefore recommends the normal-response-time formulation as a pragmatic approach for modeling individual differences in behavioral tasks, because it preserves an additive interpretation without materially affecting inferences about IDfN (Faulkenberry, 2021).

3. AD-MERCS and instance-level deviation in unsupervised anomaly detection

AD-MERCS models both normality and abnormality in unsupervised anomaly detection by identifying low-dimensional subspaces in which patterns exist and conditions that characterize instances that deviate from those patterns (Soenen et al., 2023). The method builds on MERCS, “multi-directional ensembles of regression and classification trees.” Given an unlabeled dataset pNp_N5, one decision tree pNp_N6 is grown for each attribute pNp_N7 as target, using standard top-down splitting such as CART to maximize information gain or variance reduction. Each tree identifies a low-dimensional subspace pNp_N8, and each leaf pNp_N9 defines a rectangular cell in qNq_N0 within which the data are assumed to share the same functional pattern for qNq_N1.

Instead of residuals qNq_N2, AD-MERCS fits a local Gaussian kernel density estimate within each leaf: qNq_N3 with bandwidth qNq_N4 chosen by the method of Botev et al. (2010). Because raw densities are incomparable across leaves and unbounded above, they are squashed into a likelihood score qNq_N5 using a threshold qNq_N6 chosen so that exactly qNq_N7 of the training densities in qNq_N8 fall below qNq_N9; in practice NN00. If a leaf is noisier than some ancestor node, the ancestor’s density estimate is reused so that scoring never relies on a weaker pattern than what was available higher in the tree.

The initial instance-level IDfN is formed from per-tree local anomaly probabilities

NN01

and combined through a noisy-OR model: NN02 where NN03 is an inhibition parameter. AD-MERCS then iteratively refines the score by introducing context scores NN04 for anomalous contexts. Instance scores are recomputed by blending normal-pattern anomaly and context anomaly, and context scores are recomputed by a noisy-AND over the instances they contain. After convergence, the final NN05 is the IDfN score of NN06.

The paper states several properties of this construction (Soenen et al., 2023). By construction, each NN07. The noisy-OR/AND updates guarantee monotonic non-decrease of NN08 throughout iterations. The split-selection criterion ensures that each tree isolates subspaces of maximal predictive strength for its target. No asymptotic bounds are given, but empirically the method converges in few iterations, typically NN09–NN10.

The examples clarify why density-based scoring differs from mean-residual scoring. In a bimodal leaf with target values NN11, a point at NN12 has zero residual to the mean NN13 and would raise no alarm under residual-based scoring, but NN14, hence NN15, NN16, and the point is maximally anomalous. In the Zoo dataset, the animal “scorpion” is assigned high scores in multiple contexts, including “(has_spine = false)” predicting “has_tail,” yielding the explanation “Among animals with no backbone, having a tail is extremely unlikely,” and another context in which “does not lay eggs” predicts “has_teeth.” The final NN17 is near NN18, with two concrete subspace-context explanations. This makes IDfN not only a scalar anomaly score but also an explanatory device via low-dimensional deviations from learned normality (Soenen et al., 2023).

4. Finite-NN19 equilibrium distributions and detectable non-Gaussianity

In the finite-NN20 statistical-mechanics formulation, IDfN is the systematic short-tailed deviation from Gaussian equilibrium induced by finite particle number NN21 (Shim, 28 Oct 2025). For a one-dimensional ideal-gas-like system with total kinetic energy NN22, maximizing Havrda–Charvát (Tsallis) entropy yields a compact-support NN23-Gaussian density

NN24

with

NN25

Taking NN26, the support is NN27. In the thermodynamic limit NN28, NN29 and NN30 converges to the Maxwell-Boltzmann Gaussian law. For finite NN31, the support is strictly compact and the tails are shorter than Gaussian.

This formulation makes the departure from Gaussianity fully parameterized by NN32. As NN33 decreases toward the minimum NN34, NN35 falls below NN36 and the support shrinks; as NN37 grows, all finite-NN38 effects vanish smoothly. A plausible implication is that NN39 serves as a physically interpretable non-Gaussianity index in settings where the system is literally or notionally finite.

Shim studies how five standard normality tests respond to this IDfN: Kolmogorov-Smirnov, Anderson-Darling, Cramér-von Mises, Jarque-Bera, and Shapiro-Wilk (Shim, 28 Oct 2025). The Monte Carlo design uses system sizes NN40, sample sizes NN41 for NN42, and for NN43 extends to NN44, with NN45 independent draws per NN46. KS and AD calibrated under the correct custom null NN47 maintain nominal size NN48. Against the Gaussian null, all five statistics tend to grow as NN49 becomes smaller, but the moment-based tests JB and SW detect short-tailed compact-support deviations far more efficiently than the ECDF-based tests KS, AD, and CvM.

The power map is highly uneven (Shim, 28 Oct 2025). For strongly non-Gaussian cases NN50, JB and SW achieve nearly NN51 power even for NN52, whereas KS, AD, and CvM plateau around NN53–NN54. For NN55, ECDF tests rarely exceed NN56–NN57 power at NN58, while JB and SW still exceed NN59–NN60. At the near-Gaussian case NN61, ECDF-based tests remain below NN62 power up to NN63, reaching approximately NN64 only at NN65; JB and SW exceed NN66 power by NN67 and hit NN68 by NN69.

The practical guidance follows directly. When one suspects mild, short-tailed deviation from normality due solely to finite-size constraints, Jarque-Bera or Shapiro-Wilk are preferred. As a rule of thumb, for NN70, NN71 suffices with JB or SW for NN72 power at the NN73 level; for NN74, NN75–2000 suffices; for NN76, NN77–5000 is needed for JB or SW, whereas NN78 would be required for KS, AD, or CvM (Shim, 28 Oct 2025).

5. Personalized laboratory interpretation and NORMA

In laboratory medicine, IDfN is formalized as a personalized deviation score derived from a conditional forecast distribution for the next biomarker value (Shah et al., 18 May 2026). NORMA models blood biomarkers from longitudinal histories using a conditional transformer that learns

NN79

where NN80 is the vector of laboratory values at time NN81, NN82 is the patient’s history, and NN83 denotes population-level “healthy” prior information such as central NN84 intervals NN85. The prediction head is either Gaussian, NN86, or quantile-based with five quantiles NN87.

The input representation combines static covariates, laboratory history, and a “normal” query token. The context token is

NN88

the history token is

NN89

with NN90 and NN91, and the query token is

NN92

These tokens are processed by transformer decoder layers with multi-head self-attention and feed-forward blocks. The training objective is either the Gaussian negative log-likelihood

NN93

or the quantile pinball loss

NN94

From the learned distribution, NORMA extracts either NN95 and NN96, or NN97 and

NN98

The IDfN score is then the personalized z-score

NN99

and the personalized ii00 reference interval is

ii01

or directly ii02 for the quantile head.

The main methodological issue in this literature is over-personalization. The paper states that purely personalized intervals can overfit to sparse data, inflate false-positive rates, and include unrecognized or subclinical disease; purely personalized intervals routinely overfit, classifying up to ii03 of measurements as abnormal, without corresponding associations with adverse clinical outcomes (Shah et al., 18 May 2026). Population intervals, by contrast, ignore stable intra-patient variability. NORMA balances these two extremes by conditioning on both ii04 and ii05, and by using a “normal” query token at each step. The paper presents a Bayesian-shrinkage analogy,

ii06

but states that NORMA learns a non-parametric, time-aware version of this anchoring.

The empirical results span nearly ii07 billion longitudinal laboratory measurements from over ii08 million individuals across North America, the Middle East, and East Asia (Shah et al., 18 May 2026). In CHS, abnormal flags were ii09 for ii10, ii11 for ii12, and ii13 for ii14; among ii15-normal measurements, ii16 reclassified ii17 and NORMA ii18; median lead time for ii19 versus ii20 was ii21 months (IQR ii22–ii23). In eICU, abnormal flags were ii24, ii25, and ii26, with reclassification ii27 for ii28 and ii29 for NORMA, and lead time ii30 h (ii31–ii32). In INSPIRE, abnormal flags were ii33, ii34, and ii35, with reclassification ii36 and ii37, and lead time ii38 h (ii39–ii40). Among ii41-normal cases in eICU, the positive predictive value per ii42 reclassified for in-hospital mortality was ii43 for ii44 versus ii45 for NORMA, for acute kidney injury ii46 versus ii47, and for prolonged length of stay ii48 versus ii49. The clinical interpretation is that IDfN in this setting is a calibrated patient-specific abnormality score anchored to a model of normal variation rather than to either raw population intervals or unconstrained personal baselines.

6. Online multi-entity change-point detection

In online change-point detection from multi-entity, multivariate time series, IDfN is defined as a reconstruction-error-based per-entity deviation under an autoencoder trained on normal behavior (Kor et al., 22 Sep 2025). Each sensor channel is first z-score normalized using training-set means ii50 and standard deviations ii51, and for entity ii52 at time ii53 a sliding window of length ii54 is formed,

ii55

A “Simple Autoencoder” then reconstructs the flattened window. The IDfN at time ii56 for entity ii57 is the mean squared error

ii58

The architecture is a stack of fully connected dense layers with ReLU activations, a symmetric decoder, a final linear output layer, and dropout of ii59 after each hidden layer. Training is done on normal data only with loss

ii60

using Adam with initial learning rate ii61, batch size ii62, up to ii63 epochs with early stopping after ii64 epochs without validation improvement, and learning-rate reduction on plateau by a factor of ii65 down to ii66. The recommended default window length is ii67.

Per-entity IDfNs are aggregated into system-wide anomaly scores (SWAS) in three ways (Kor et al., 22 Sep 2025). The mean score is

ii68

the variance score is

ii69

and the distributional score uses a Gaussian-kernel KDE of the current IDfN distribution compared with a reference KDE over all training IDfNs, with SWAS

ii70

the Wasserstein-1 distance. These statistics are converted to deviation scores ii71, ii72, or ii73, accumulated by CUSUM,

ii74

and then tested by an adaptive sequential density-based thresholding rule after the log transform ii75. A change is declared at time ii76 if the KDE density estimate at ii77 falls below a threshold ii78, taken as the ii79th percentile of the training CUSUM-KDE densities, after a burn-in of ii80 samples.

The validation datasets include auto-regressive series with ii81 entities and ii82, coupled Chen chaotic oscillators with ii83, ii84, and ii85, and two Unity crowd-simulation settings using upper-leg acceleration streams with ii86: train-station evacuation with approximately ii87 agents and a change around ii88 s, and bidirectional corridor collision with approximately ii89 agents and a change around ii90 s (Kor et al., 22 Sep 2025). The paper reports that IDfN trajectories show a clear rise at the ground-truth event time, KDE heatmaps visualize a shifting density of reconstruction errors with a sudden mode change at the incident, aggregated SWAS plots spike near the true change, and CUSUM curves cross adaptive thresholds in tight proximity to the ground-truth change point. In this formulation, IDfN is explicitly local—defined per entity and per time step—but is designed to support system-level inference after aggregation.

7. Directional IDfN in multivariate spatial normality testing

In multivariate spatial statistics, IDfN is defined as a directional departure from Gaussianity after accounting for spatial dependence (Chen et al., 2020). The overall null is

ii91

By the union-intersection principle, ii92 holds if and only if every projection

ii93

is Gaussian for every unit vector ii94. In practice, the continuum of directions is replaced by a finite set ii95, such as the coordinate axes, the eigenvectors of an estimate of ii96, or random draws from the unit sphere.

For a fixed direction ii97, projected residuals are standardized as

ii98

and the sample skewness and kurtosis are

ii99

Under spatial dependence, the asymptotic variances of δi\delta_i00 and δi\delta_i01 depend on the covariance structure of the projected process. Horváth et al. (2020) propose kernel-smoothed estimators

δi\delta_i02

δi\delta_i03

where δi\delta_i04 is the sample autocovariance at lag δi\delta_i05, δi\delta_i06 is a univariate kernel such as Bartlett, and δi\delta_i07 is a bandwidth vector.

The directional Jarque-Bera-type statistic is then

δi\delta_i08

Under δi\delta_i09, the standardized skewness and kurtosis are jointly Gaussian and asymptotically independent, so δi\delta_i10. At nominal level δi\delta_i11, the critical value is δi\delta_i12; for δi\delta_i13, δi\delta_i14. The global decision rule rejects multivariate normality when

δi\delta_i15

or equivalently when any directional δi\delta_i16-value is at most δi\delta_i17, possibly with Bonferroni correction or Benjamini-Hochberg control.

In this framework, the directional score itself is interpreted as IDfN: δi\delta_i18 A large δi\delta_i19 means that, along direction δi\delta_i20, the data exhibit unusually large skewness or excess kurtosis after accounting for spatial dependence. This provides both a global multivariate normality test and a directional diagnostic. The method thereby makes it possible to locate which linear combinations of variables are most non-Gaussian, rank them by IDfN, and map out a “normality contour” on the unit sphere in δi\delta_i21 (Chen et al., 2020).

Across these formulations, IDfN functions less as a single theory than as a recurring analytic pattern: choose a reference notion of normality, define an individual unit relative to that reference, and quantify the unit’s departure in a way that is appropriate to the data-generating assumptions of the field. The main controversies concern which normality model is substantively appropriate and how much interpretability is lost when one moves from additive, physically transparent, or clinically transparent scales to transformed or highly adaptive representations. The cited literature consistently treats those choices as consequential for interpretation, but not always for the qualitative ordering of evidence or detection performance (Faulkenberry, 2021, Soenen et al., 2023, Shim, 28 Oct 2025, Shah et al., 18 May 2026, Kor et al., 22 Sep 2025, Chen et al., 2020).

Topic to Video (Beta)

No one has generated a video about this topic yet.

Whiteboard

No one has generated a whiteboard explanation for this topic yet.

Follow Topic

Get notified by email when new papers are published related to Individual Deviation from Normality (IDfN).