---
title: Identifiability Defect in Statistical Models
url: https://www.emergentmind.com/topics/identifiability-defect
type: topic
---

# Identifiability Defect in Statistical Models

An identifiability defect is a failure of injectivity in the map from latent or structural objects to observables: distinct parameter values, latent structures, decompositions, or internal representations induce the same distribution of observed data, the same model output, or the same input–output behavior. Across the arXiv literature, this defect appears in psychometric latent variable models, causal and unsupervised representation learning, neural attention mechanisms, inverse problems, algebraic statistics, functional regression, and privacy-constrained econometrics. Its common consequence is that estimation, interpretation, or decomposition can cease to be unique even with arbitrarily large samples, because the ambiguity is structural rather than merely finite-sample [1711.03174] [1606.05598] [2202.06844] [1908.04211] [2102.04417] [1402.2637] [1506.03627] [2006.14732].

## 1. Formal characterizations

In the standard statistical formulation, identifiability means that \(P_\theta = P_{\theta'}\) for all observables implies \(\theta=\theta'\); an identifiability defect is the existence of \(\theta\neq\theta'\) with the same observable law. The same idea is restated in several equivalent ways across fields: injectivity of a parameter-to-distribution map in the DINA model, local one-to-one-ness of a coefficient map in linear compartmental models, birationality of a secant map in algebraic geometry, and uniqueness of an internal attention matrix given a head output in transformers [1711.03174] [2102.04417] [1911.00780] [1908.04211].

The defect can also be expressed geometrically. In linear compartmental models it is the locus where the Jacobian of the coefficient map drops rank, equivalently where fibers of the coefficient map have positive dimension; in projective geometry it is the failure of the \(h\)-secant map to be birational, often linked to secant defect; in bilinear inverse problems it is the presence of nonzero rank-2 matrices in the null space of the lifted operator; and under differential privacy it appears as a weak limit that is a non-degenerate random set rather than a singleton parameter value [2102.04417] [1911.00780] [1402.2637] [2006.14732].

| Domain | Carrier of the defect | Identifiable object when possible |
|---|---|---|
| DINA and restricted latent class models | Multiple parameter sets or latent classes yield the same response distribution | Full parameters, or grouped proportions \(\boldsymbol{\nu}\) |
| GRT and GRTwIND | Rotated or sheared perceptual/decisional parameterizations are empirically equivalent | Parameters only after fixing additional structure |
| Transformers | Nontrivial left null space of \(T\) makes \(A \mapsto AT\) non-injective | Effective attention, not raw attention |
| Bilinear inverse problems | Rank-2 null space of the lifted map | Factor pair up to scaling |
| Linear compartmental models | Singular locus of the coefficient map | Parameters off the singular locus |
| Differential privacy | Weak limit of estimators is a random set | Singleton or deterministically locatable point in that set |

This suggests that an identifiability defect is best understood as a model-specific equivalence relation on hidden structure, with observables unable to distinguish members of the same equivalence class.

## 2. Structural sources

A first source is **design-induced confounding**. In cognitive diagnosis, incomplete or poorly separated \(Q\)-matrices create indistinguishable latent classes; in bifactor models and two-tier models, non-identifiability can be checked by inspecting the factor loading structure; in function-on-function regression, the defect persists when the functional covariate’s empirical covariance has a kernel that overlaps that of the roughness penalty; and in GRT, rotations and shears can move interaction structure between perceptual distributions and decision bounds without changing response probabilities [1711.03174] [2012.12196] [1506.03627] [1606.05598].

A second source is **hidden symmetry**. In nonlinear representation learning, measure-preserving automorphisms can produce observationally equivalent models whose latent spaces differ by nonlinear transformations that are not permutations plus coordinate-wise bijections. In tensor decomposition, generic points can admit multiple decompositions over \(\mathbb{C}\), and in some cases only one of these is real. In bilinear inverse problems, scaling ambiguity is unavoidable, and further non-uniqueness is controlled by the geometry of the rank-2 null space after lifting [2202.06844] [1608.07197] [1402.2637].

A third source is **null-space degeneracy inside the model architecture itself**. In transformers, attention weights are not identifiable when sequence length exceeds the attention head dimension, because a nontrivial left null space allows infinitely many distinct attention matrices to produce the same head output. In linear compartmental models, factors of the singular-locus polynomial identify parameter hyperplanes on which rank drops and identifiability fails [1908.04211] [2102.04417].

## 3. Latent-variable and psychometric models

In the DINA model, an identifiability defect occurs precisely when different \((\mathbf{s},\mathbf{g},\mathbf{p})\) induce the same response distribution. The complete characterization is structural: all DINA parameters are identifiable if and only if the \(Q\)-matrix satisfies Condition 1 and Condition 2. Condition 1 requires completeness and at least three items per attribute; Condition 2 requires that any two different columns of the submatrix \(Q^*\) are distinct. When these conditions fail, the paper constructs explicit alternative parameter sets, including cases with infinitely many solutions that attain the same likelihood, so adding more examinees does not fix the problem [1711.03174].

A related line of work treats non-identifiability through **equivalence classes of latent profiles**. For incomplete \(Q\)-matrices, profiles with identical ideal response vectors are inseparable and form equivalence classes \([\alpha]\). The identifiable object is then not the full vector of class proportions \(\nu_\alpha\), but the grouped proportions \(\nu_{[\alpha]}\). This yields the notions of marginal identifiability for individual attributes, the indicator \(\delta_{[\alpha],k}\), and the marginal identifiability rate \(\zeta_k\), which quantifies the proportion of the population for which attribute \(k\) is identifiable. In restricted latent class models more generally, the literature distinguishes strict identifiability, \(\nu\)-partial identifiability for two-parameter models, and generic identifiability for multi-parameter models, all driven by the design matrix through the \(\Gamma\)-structure [1303.0426] [1803.04353].

The bifactor literature frames the defect in similar terms but at the level of loading patterns. The bifactor model and its extensions may suffer from non-identifiability, which can further lead to inconsistent parameter estimation and invalid inference; the 2020 characterization covers the linear and dichotomous bifactor models, the linear extended bifactor model with correlated subdimensions, and analogous results for two-tier models, with identifiability checked through the factor loading structure [2012.12196].

## 4. Perception, representation learning, and neural architectures

In General Recognition Theory, the defect is a reparameterization equivalence between perceptual and decisional structure. In fully parameterized \(2\times 2\) Gaussian GRT models, failures of decisional separability are not, in general, identifiable, because invertible linear transformations can align oblique decision bounds with the coordinate axes while preserving all response probabilities. In GRTwIND, the intended remedy is universal perception, but the paper shows that a model with universal perception and subject-specific failures of decisional separability is mathematically, and thereby empirically, equivalent to a model with decisional separability and failure of universal perception. It also proves that means and marginal variances are not, in general, simultaneously identifiable in \(2\times 2\) Gaussian GRT models, including GRTwIND [1606.05598].

In unsupervised representation learning, the defect appears as **observational equivalence under nonlinear latent reparameterization**. The note on Wang and Jordan’s Theorem 11 constructs two representations \(Z\) and \(Z'\) with the same \(\sigma\)-algebra, compact rectangular support, and identical observational distribution, yet related by a nonlinear measure-preserving automorphism that is not a permutation plus coordinate-wise bijections. This shows that support factorization and shared information content do not by themselves rule out nontrivial latent mixing. A separate line partly repairs this impossibility by weakening the target: under a generic nonlinear diffeomorphism, continuous factors remain unidentifiable, but quantized factors become identifiable when the latent density has independent discontinuities that form an axis-aligned grid with a backbone [2202.06844] [2306.16334].

In transformers, the defect concerns internal explanatory variables rather than model parameters. For a fixed input, attention weights are identifiable only if they are uniquely determined by the head output \(AT\). When sequence length exceeds head dimension, the left null space of \(T\) is nontrivial, so there are infinitely many valid attention matrices producing the same output; raw attention is therefore not directly interpretable. The paper proposes effective attention as the projection of attention onto the complement of the null space component, and separately shows that input tokens retain to a large degree their identity across the model, with evidence that identity information is mainly encoded in the angle of embeddings while contextual embeddings are generated by strong mixing of input information [1908.04211].

## 5. Inverse problems, geometry, and privacy

In linear compartmental models, an identifiability defect is the singular locus of the coefficient map: the set of parameter values where the Jacobian loses rank and the parameter–data map ceases to be locally one-to-one. In square-Jacobian cases this locus is cut out by the singular-locus polynomial, whose factors reveal whether setting a leak or edge parameter to zero forces loss of identifiability. The paper proves a special case of the conjecture that removing a leak from an identifiable model preserves identifiability, proves equivalence between this conjecture and the claim that leak terms do not divide the singular-locus equation, and proves a case in which removing a dividing edge makes the resulting model unidentifiable [2102.04417].

In bilinear inverse problems, the observation is bilinear in two unknown factors, so identifiability is only meaningful up to the scaling transformation \((\vec{x},\vec{y})\mapsto(\alpha\vec{x},\alpha^{-1}\vec{y})\). After lifting to rank-one matrix recovery, non-uniqueness is controlled by the rank-2 null space of the lifted linear operator. The paper develops deterministic identifiability conditions and scaling laws that trade off probability of robust identifiability with the complexity of that rank-2 null space, with blind deconvolution as the canonical example [1402.2637].

In algebraic geometry and tensor decomposition, identifiability is tied to secant geometry. A projective variety is \(h\)-identifiable when the generic element of its \(h\)-secant variety uniquely determines \(h\) points on the variety; positive secant defect forces failure of generic \(h\)-identifiability. The secant-defect approach improves known bounds for Segre, Segre–Veronese, and Grassmann varieties. Over \(\mathbb{R}\), the picture can be subtler: there are nonempty Euclidean open subsets of tensor spaces whose elements have several decompositions over \(\mathbb{C}\), but only one formed by real summands, so they are identifiable over \(\mathbb{R}\) but not over \(\mathbb{C}\) [1911.00780] [1608.07197].

Under differential privacy, the defect takes a different form. Privacy mechanisms are finite-sample randomization procedures, so the paper defines identification as a property of the limit of experiments and shows that the set of differentially private estimators converges weakly to a random set. In particular instances of regression discontinuity design, parameters turn out to be neither point nor partially identified. Identification becomes possible only if the target parameter can be deterministically located within the random set; in that case, full exploration of the random set of weak limits can allow the data curator to select a sequence of differentially private estimators converging to the target parameter in probability [2006.14732].

## 6. Consequences, diagnostics, and resolution strategies

The immediate consequences of an identifiability defect are domain-specific but structurally parallel. In DINA and related cognitive diagnosis models, non-identifiability prevents unique recovery of slip, guess, and latent class proportions; in bifactor models it can lead to inconsistent parameter estimation and invalid inference; in function-on-function regression it can produce arbitrarily large errors for coefficient surface estimates despite accurate predictions of the responses; in GRT it collapses distinct theoretical interpretations into observational equivalence classes; and under differential privacy it yields asymptotic random sets rather than point limits [1711.03174] [2012.12196] [1506.03627] [1606.05598] [2006.14732].

Diagnostics therefore focus on the structural carrier of the ambiguity. In cognitive diagnosis, practical checks are completeness of \(Q\), at least three items per attribute, and distinct columns in \(Q^*\); for incomplete \(Q\)-matrices one can compute equivalence classes and the marginal identifiability rates \(\zeta_k\). In bifactor models, identifiability is checked by inspecting factor loading structure. In function-on-function regression, the key diagnostic is whether \(\ker(D_s^\top D_s)\cap\ker(P_s)\) is nontrivial, operationalized through the condition number of \(D_s^\top D_s\) and an overlap measure between the empirical covariance kernel and the penalty null space. In transformers, comparison of raw attention with effective attention exposes null-space components that do not affect the head output. In linear compartmental models, the singular-locus equation plays the analogous role [1711.03174] [1303.0426] [2012.12196] [1506.03627] [1908.04211] [2102.04417].

Resolution strategies are correspondingly structural. In DINA, the remedy is redesign of the \(Q\)-matrix so that attributes are structurally distinguishable; in restricted latent class models, one may settle for \(\nu\)-partial identifiability when full identification is impossible. In GRT, one must fix orthogonality of perceptual dimensions, often by assuming decisional separability in the chosen coordinate system. In transformers, effective attention is a more behaviorally grounded object than raw attention. In function-on-function regression, the paper recommends avoiding aggressive rank-reducing preprocessing and curve-wise centering, preferring penalties with smaller null spaces, and using full-rank or constraint-based penalties when overlap is detected. In representation learning, a plausible implication is that replacing exact continuous recovery by quantized factor identifiability changes the target from an impossible one to a provably attainable one under explicit discontinuity assumptions. Under differential privacy, identification can be recovered only when the data curator can select from the limiting random set by a deterministic rule that locates the target parameter [1803.04353] [1606.05598] [1908.04211] [1506.03627] [2306.16334] [2006.14732].

Taken together, these results show that an identifiability defect is not a single pathology but a family of structural ambiguities: duplicated design columns, hidden symmetries, rank-deficient operators, singular loci, secant defects, null spaces, and privacy-induced random-set limits. What unifies them is that more data alone do not remove the ambiguity. Only additional structure—design constraints, side information, stronger model restrictions, or a redefinition of the identifiable target—can do so.

Source: https://www.emergentmind.com/topics/identifiability-defect