Papers
Topics
Authors
Recent
Search
2000 character limit reached

Identifiability Defect in Statistical Models

Updated 16 July 2026
  • Identifiability defect is the failure of the parameter-to-observation map to be injective, resulting in different parameters producing the same model outputs.
  • It often arises from design-induced confounding, hidden symmetry, and null-space degeneracy, affecting areas like psychometrics, causal learning, and neural architectures.
  • This defect complicates unique estimation and interpretation, necessitating additional structural constraints or reparameterization to achieve identifiable solutions.

An identifiability defect is a failure of injectivity in the map from latent or structural objects to observables: distinct parameter values, latent structures, decompositions, or internal representations induce the same distribution of observed data, the same model output, or the same input–output behavior. Across the arXiv literature, this defect appears in psychometric latent variable models, causal and unsupervised representation learning, neural attention mechanisms, inverse problems, algebraic statistics, functional regression, and privacy-constrained econometrics. Its common consequence is that estimation, interpretation, or decomposition can cease to be unique even with arbitrarily large samples, because the ambiguity is structural rather than merely finite-sample (Gu et al., 2017, Silbert et al., 2016, Ghosh et al., 2022, Brunner et al., 2019, Chan et al., 2021, Choudhary et al., 2014, Scheipl et al., 2015, Komarova et al., 2020).

1. Formal characterizations

In the standard statistical formulation, identifiability means that Pθ=Pθ′P_\theta = P_{\theta'} for all observables implies θ=θ′\theta=\theta'; an identifiability defect is the existence of θ≠θ′\theta\neq\theta' with the same observable law. The same idea is restated in several equivalent ways across fields: injectivity of a parameter-to-distribution map in the DINA model, local one-to-one-ness of a coefficient map in linear compartmental models, birationality of a secant map in algebraic geometry, and uniqueness of an internal attention matrix given a head output in transformers (Gu et al., 2017, Chan et al., 2021, Casarotti et al., 2019, Brunner et al., 2019).

The defect can also be expressed geometrically. In linear compartmental models it is the locus where the Jacobian of the coefficient map drops rank, equivalently where fibers of the coefficient map have positive dimension; in projective geometry it is the failure of the hh-secant map to be birational, often linked to secant defect; in bilinear inverse problems it is the presence of nonzero rank-2 matrices in the null space of the lifted operator; and under differential privacy it appears as a weak limit that is a non-degenerate random set rather than a singleton parameter value (Chan et al., 2021, Casarotti et al., 2019, Choudhary et al., 2014, Komarova et al., 2020).

Domain Carrier of the defect Identifiable object when possible
DINA and restricted latent class models Multiple parameter sets or latent classes yield the same response distribution Full parameters, or grouped proportions ν\boldsymbol{\nu}
GRT and GRTwIND Rotated or sheared perceptual/decisional parameterizations are empirically equivalent Parameters only after fixing additional structure
Transformers Nontrivial left null space of TT makes A↦ATA \mapsto AT non-injective Effective attention, not raw attention
Bilinear inverse problems Rank-2 null space of the lifted map Factor pair up to scaling
Linear compartmental models Singular locus of the coefficient map Parameters off the singular locus
Differential privacy Weak limit of estimators is a random set Singleton or deterministically locatable point in that set

This suggests that an identifiability defect is best understood as a model-specific equivalence relation on hidden structure, with observables unable to distinguish members of the same equivalence class.

2. Structural sources

A first source is design-induced confounding. In cognitive diagnosis, incomplete or poorly separated QQ-matrices create indistinguishable latent classes; in bifactor models and two-tier models, non-identifiability can be checked by inspecting the factor loading structure; in function-on-function regression, the defect persists when the functional covariate’s empirical covariance has a kernel that overlaps that of the roughness penalty; and in GRT, rotations and shears can move interaction structure between perceptual distributions and decision bounds without changing response probabilities (Gu et al., 2017, Fang et al., 2020, Scheipl et al., 2015, Silbert et al., 2016).

A second source is hidden symmetry. In nonlinear representation learning, measure-preserving automorphisms can produce observationally equivalent models whose latent spaces differ by nonlinear transformations that are not permutations plus coordinate-wise bijections. In tensor decomposition, generic points can admit multiple decompositions over C\mathbb{C}, and in some cases only one of these is real. In bilinear inverse problems, scaling ambiguity is unavoidable, and further non-uniqueness is controlled by the geometry of the rank-2 null space after lifting (Ghosh et al., 2022, Angelini et al., 2016, Choudhary et al., 2014).

A third source is null-space degeneracy inside the model architecture itself. In transformers, attention weights are not identifiable when sequence length exceeds the attention head dimension, because a nontrivial left null space allows infinitely many distinct attention matrices to produce the same head output. In linear compartmental models, factors of the singular-locus polynomial identify parameter hyperplanes on which rank drops and identifiability fails (Brunner et al., 2019, Chan et al., 2021).

3. Latent-variable and psychometric models

In the DINA model, an identifiability defect occurs precisely when different (s,g,p)(\mathbf{s},\mathbf{g},\mathbf{p}) induce the same response distribution. The complete characterization is structural: all DINA parameters are identifiable if and only if the θ=θ′\theta=\theta'0-matrix satisfies Condition 1 and Condition 2. Condition 1 requires completeness and at least three items per attribute; Condition 2 requires that any two different columns of the submatrix θ=θ′\theta=\theta'1 are distinct. When these conditions fail, the paper constructs explicit alternative parameter sets, including cases with infinitely many solutions that attain the same likelihood, so adding more examinees does not fix the problem (Gu et al., 2017).

A related line of work treats non-identifiability through equivalence classes of latent profiles. For incomplete θ=θ′\theta=\theta'2-matrices, profiles with identical ideal response vectors are inseparable and form equivalence classes θ=θ′\theta=\theta'3. The identifiable object is then not the full vector of class proportions θ=θ′\theta=\theta'4, but the grouped proportions θ=θ′\theta=\theta'5. This yields the notions of marginal identifiability for individual attributes, the indicator θ=θ′\theta=\theta'6, and the marginal identifiability rate θ=θ′\theta=\theta'7, which quantifies the proportion of the population for which attribute θ=θ′\theta=\theta'8 is identifiable. In restricted latent class models more generally, the literature distinguishes strict identifiability, θ=θ′\theta=\theta'9-partial identifiability for two-parameter models, and generic identifiability for multi-parameter models, all driven by the design matrix through the θ≠θ′\theta\neq\theta'0-structure (Zhang et al., 2013, Gu et al., 2018).

The bifactor literature frames the defect in similar terms but at the level of loading patterns. The bifactor model and its extensions may suffer from non-identifiability, which can further lead to inconsistent parameter estimation and invalid inference; the 2020 characterization covers the linear and dichotomous bifactor models, the linear extended bifactor model with correlated subdimensions, and analogous results for two-tier models, with identifiability checked through the factor loading structure (Fang et al., 2020).

4. Perception, representation learning, and neural architectures

In General Recognition Theory, the defect is a reparameterization equivalence between perceptual and decisional structure. In fully parameterized θ≠θ′\theta\neq\theta'1 Gaussian GRT models, failures of decisional separability are not, in general, identifiable, because invertible linear transformations can align oblique decision bounds with the coordinate axes while preserving all response probabilities. In GRTwIND, the intended remedy is universal perception, but the paper shows that a model with universal perception and subject-specific failures of decisional separability is mathematically, and thereby empirically, equivalent to a model with decisional separability and failure of universal perception. It also proves that means and marginal variances are not, in general, simultaneously identifiable in θ≠θ′\theta\neq\theta'2 Gaussian GRT models, including GRTwIND (Silbert et al., 2016).

In unsupervised representation learning, the defect appears as observational equivalence under nonlinear latent reparameterization. The note on Wang and Jordan’s Theorem 11 constructs two representations θ≠θ′\theta\neq\theta'3 and θ≠θ′\theta\neq\theta'4 with the same θ≠θ′\theta\neq\theta'5-algebra, compact rectangular support, and identical observational distribution, yet related by a nonlinear measure-preserving automorphism that is not a permutation plus coordinate-wise bijections. This shows that support factorization and shared information content do not by themselves rule out nontrivial latent mixing. A separate line partly repairs this impossibility by weakening the target: under a generic nonlinear diffeomorphism, continuous factors remain unidentifiable, but quantized factors become identifiable when the latent density has independent discontinuities that form an axis-aligned grid with a backbone (Ghosh et al., 2022, Barin-Pacela et al., 2023).

In transformers, the defect concerns internal explanatory variables rather than model parameters. For a fixed input, attention weights are identifiable only if they are uniquely determined by the head output θ≠θ′\theta\neq\theta'6. When sequence length exceeds head dimension, the left null space of θ≠θ′\theta\neq\theta'7 is nontrivial, so there are infinitely many valid attention matrices producing the same output; raw attention is therefore not directly interpretable. The paper proposes effective attention as the projection of attention onto the complement of the null space component, and separately shows that input tokens retain to a large degree their identity across the model, with evidence that identity information is mainly encoded in the angle of embeddings while contextual embeddings are generated by strong mixing of input information (Brunner et al., 2019).

5. Inverse problems, geometry, and privacy

In linear compartmental models, an identifiability defect is the singular locus of the coefficient map: the set of parameter values where the Jacobian loses rank and the parameter–data map ceases to be locally one-to-one. In square-Jacobian cases this locus is cut out by the singular-locus polynomial, whose factors reveal whether setting a leak or edge parameter to zero forces loss of identifiability. The paper proves a special case of the conjecture that removing a leak from an identifiable model preserves identifiability, proves equivalence between this conjecture and the claim that leak terms do not divide the singular-locus equation, and proves a case in which removing a dividing edge makes the resulting model unidentifiable (Chan et al., 2021).

In bilinear inverse problems, the observation is bilinear in two unknown factors, so identifiability is only meaningful up to the scaling transformation θ≠θ′\theta\neq\theta'8. After lifting to rank-one matrix recovery, non-uniqueness is controlled by the rank-2 null space of the lifted linear operator. The paper develops deterministic identifiability conditions and scaling laws that trade off probability of robust identifiability with the complexity of that rank-2 null space, with blind deconvolution as the canonical example (Choudhary et al., 2014).

In algebraic geometry and tensor decomposition, identifiability is tied to secant geometry. A projective variety is θ≠θ′\theta\neq\theta'9-identifiable when the generic element of its hh0-secant variety uniquely determines hh1 points on the variety; positive secant defect forces failure of generic hh2-identifiability. The secant-defect approach improves known bounds for Segre, Segre–Veronese, and Grassmann varieties. Over hh3, the picture can be subtler: there are nonempty Euclidean open subsets of tensor spaces whose elements have several decompositions over hh4, but only one formed by real summands, so they are identifiable over hh5 but not over hh6 (Casarotti et al., 2019, Angelini et al., 2016).

Under differential privacy, the defect takes a different form. Privacy mechanisms are finite-sample randomization procedures, so the paper defines identification as a property of the limit of experiments and shows that the set of differentially private estimators converges weakly to a random set. In particular instances of regression discontinuity design, parameters turn out to be neither point nor partially identified. Identification becomes possible only if the target parameter can be deterministically located within the random set; in that case, full exploration of the random set of weak limits can allow the data curator to select a sequence of differentially private estimators converging to the target parameter in probability (Komarova et al., 2020).

6. Consequences, diagnostics, and resolution strategies

The immediate consequences of an identifiability defect are domain-specific but structurally parallel. In DINA and related cognitive diagnosis models, non-identifiability prevents unique recovery of slip, guess, and latent class proportions; in bifactor models it can lead to inconsistent parameter estimation and invalid inference; in function-on-function regression it can produce arbitrarily large errors for coefficient surface estimates despite accurate predictions of the responses; in GRT it collapses distinct theoretical interpretations into observational equivalence classes; and under differential privacy it yields asymptotic random sets rather than point limits (Gu et al., 2017, Fang et al., 2020, Scheipl et al., 2015, Silbert et al., 2016, Komarova et al., 2020).

Diagnostics therefore focus on the structural carrier of the ambiguity. In cognitive diagnosis, practical checks are completeness of hh7, at least three items per attribute, and distinct columns in hh8; for incomplete hh9-matrices one can compute equivalence classes and the marginal identifiability rates ν\boldsymbol{\nu}0. In bifactor models, identifiability is checked by inspecting factor loading structure. In function-on-function regression, the key diagnostic is whether ν\boldsymbol{\nu}1 is nontrivial, operationalized through the condition number of ν\boldsymbol{\nu}2 and an overlap measure between the empirical covariance kernel and the penalty null space. In transformers, comparison of raw attention with effective attention exposes null-space components that do not affect the head output. In linear compartmental models, the singular-locus equation plays the analogous role (Gu et al., 2017, Zhang et al., 2013, Fang et al., 2020, Scheipl et al., 2015, Brunner et al., 2019, Chan et al., 2021).

Resolution strategies are correspondingly structural. In DINA, the remedy is redesign of the ν\boldsymbol{\nu}3-matrix so that attributes are structurally distinguishable; in restricted latent class models, one may settle for ν\boldsymbol{\nu}4-partial identifiability when full identification is impossible. In GRT, one must fix orthogonality of perceptual dimensions, often by assuming decisional separability in the chosen coordinate system. In transformers, effective attention is a more behaviorally grounded object than raw attention. In function-on-function regression, the paper recommends avoiding aggressive rank-reducing preprocessing and curve-wise centering, preferring penalties with smaller null spaces, and using full-rank or constraint-based penalties when overlap is detected. In representation learning, a plausible implication is that replacing exact continuous recovery by quantized factor identifiability changes the target from an impossible one to a provably attainable one under explicit discontinuity assumptions. Under differential privacy, identification can be recovered only when the data curator can select from the limiting random set by a deterministic rule that locates the target parameter (Gu et al., 2018, Silbert et al., 2016, Brunner et al., 2019, Scheipl et al., 2015, Barin-Pacela et al., 2023, Komarova et al., 2020).

Taken together, these results show that an identifiability defect is not a single pathology but a family of structural ambiguities: duplicated design columns, hidden symmetries, rank-deficient operators, singular loci, secant defects, null spaces, and privacy-induced random-set limits. What unifies them is that more data alone do not remove the ambiguity. Only additional structure—design constraints, side information, stronger model restrictions, or a redefinition of the identifiable target—can do so.

Topic to Video (Beta)

No one has generated a video about this topic yet.

Whiteboard

No one has generated a whiteboard explanation for this topic yet.

Follow Topic

Get notified by email when new papers are published related to Identifiability Defect.