Papers
Topics
Authors
Recent
Search
2000 character limit reached

JE-IRT: Joint Embedding in Item Response Theory

Updated 13 July 2026
  • JE-IRT is a framework that jointly learns latent representations and item parameters to capture both content structure and probabilistic response behavior.
  • It employs content-derived embeddings for adaptive test item calibration and geometric embeddings for multidimensional LLM evaluation using low-dimensional neural mappings.
  • JE-IRT facilitates transfer to unseen items or models by merging IRT’s probabilistic semantics with deep representation learning, ensuring robust predictive performance.

Searching arXiv for recent JE-IRT papers and related IRT embedding work. JE-IRT, usually expanded as Joint Embedding Item Response Theory, denotes a family of IRT extensions in which latent representations are learned jointly with response models rather than treating all item parameters as independent free scalars. In recent arXiv usage, the label has been applied in two closely related but non-identical ways: a content-derived psychometric calibration framework in which item features are mapped to low-dimensional embeddings that determine 3PL parameters for adaptive testing (Sharpnack et al., 8 Jul 2026), and a geometric evaluation framework in which both LLMs and questions are embedded in a shared space whose geometry controls correctness probabilities (Yao et al., 26 Sep 2025). Across both usages, the unifying idea is that embeddings mediate transfer to unseen items or models while preserving an explicit IRT-style probabilistic semantics.

1. Scope of the term

The recent literature uses JE-IRT to connect IRT with learned representations rather than with fixed, per-item parameter tables. A useful distinction is between one-sided embedding and two-sided embedding formulations.

Usage Embedded objects Main objective
Content-derived JE-IRT Items Feature-only calibration of new test items
Geometric JE-IRT Models and questions Multidimensional evaluation of LLM abilities

Earlier Bayesian IRT work had already emphasized joint estimation of learner and item parameters, including hierarchical and temporal extensions, as a core design principle for interpretability and regularization (Wilson et al., 2016). JE-IRT extends that jointness from parameter estimation alone to representation learning: the latent structure is no longer only a collection of abilities and difficulties, but also an embedding space that can encode content, topic, and transfer structure.

This suggests a broader conceptual reading of JE-IRT: IRT is retained as the response model, while embeddings absorb structure that standard scalar ability–difficulty parameterizations cannot express directly.

2. Content-derived JE-IRT for item calibration

In the adaptive-testing formulation, JE-IRT begins from the 3-Parameter Logistic model

P(Yij=1∣θi)=cj+(1−cj) σ(aj(θi−bj)),P(Y_{ij}=1 \mid \theta_i) = c_j + (1-c_j)\,\sigma(a_j(\theta_i-b_j)),

but does not estimate (aj,bj)(a_j,b_j) as unrestricted item-specific parameters. Instead, each item jj has content features xj∈RFx_j \in \mathbb{R}^F, and a feature network z(⋅)z(\cdot) produces a low-dimensional embedding

hj=z(xj)∈Rd.h_j = z(x_j) \in \mathbb{R}^d.

Discrimination and difficulty are then derived by generalized linear mappings,

(λaj,λbj)=Whj+u,aj=exp⁡(λaj),bj=λbj,(\lambda_{a_j},\lambda_{b_j}) = W h_j + u,\qquad a_j=\exp(\lambda_{a_j}),\qquad b_j=\lambda_{b_j},

with the exponential link enforcing aj>0a_j>0 (Sharpnack et al., 8 Jul 2026).

A central modeling decision is the treatment of the guessing parameter. Rather than estimating cjc_j per item, the approach fixes cc globally, or by heuristic group, because simultaneous estimation of (aj,bj)(a_j,b_j)0 and (aj,bj)(a_j,b_j)1 is described as problematic for identifiability and convergence (Sharpnack et al., 8 Jul 2026). In operational terms, this choice keeps the production model simple while moving model-capacity and hyperparameter search into a pre-launch stage.

The intended use case is high-stakes computerized adaptive testing, where new items must be calibrated continuously but on-the-fly neural architecture tuning would be impractical and could threaten validity. The learned embedding therefore functions as a compact content-derived summary that can feed a simple explanatory IRT layer in production. In the reported application, the framework is positioned as a first step toward a compact item embedding for the Scalable Parametric Item Calibration Engine (SPICE), the fully Bayesian engine at the core of the S2A3 adaptive-testing system (Sharpnack et al., 8 Jul 2026).

3. Estimation, model selection, and empirical behavior in adaptive testing

The calibration paper fits the feature network and latent abilities jointly with Monte Carlo Expectation-Maximization (MCEM), explicitly eliminating any separate ability-estimation or pre-calibration stage (Sharpnack et al., 8 Jul 2026). In the E-step, for each session (aj,bj)(a_j,b_j)2, the method samples abilities from

(aj,bj)(a_j,b_j)3

using a standard normal prior and a discrete posterior approximation on a grid such as 60 points over (aj,bj)(a_j,b_j)4. In the M-step, the network parameters (aj,bj)(a_j,b_j)5 are updated with Adam by maximizing the expected complete-data log-likelihood

(aj,bj)(a_j,b_j)6

alternating E- and M-steps for (aj,bj)(a_j,b_j)7 iterations such as (aj,bj)(a_j,b_j)8. Initialization uses external ability proxies such as grade rates to accelerate convergence (Sharpnack et al., 8 Jul 2026).

Evaluation is deliberately item-split rather than response-split. Entire items are held out, typically in an 80/20 train/test partition, so that held-out items are evaluated in the feature-only regime: only their content is provided during inference, while all responses to those items are reserved for assessment. The primary metric is binary cross-entropy on held-out responses, with bootstrap confidence intervals from 200 resamplings (Sharpnack et al., 8 Jul 2026).

The experiments are conducted on two Duolingo English Test practice task types: yes/no vocabulary and vocabulary-in-context. The main empirical result is that a shallow two-layer ReLU network with (aj,bj)(a_j,b_j)9 and hand-engineered scalar features matches or beats larger architectures on held-out items for both task types (Sharpnack et al., 8 Jul 2026). The same study reports several architecture-level regularities: scalar linguistic features outperform high-dimensional contextual BERT embeddings at moderate item-bank size; exponential and softplus links for jj0 are both viable, with exponential used as default; ReLU activations yield a wider and more structured discrimination distribution; and deeper stacks of sigmoid or softplus nonlinearities can collapse to degenerate solutions in which items receive similar parameters (Sharpnack et al., 8 Jul 2026).

Operationally, these findings matter because the calibration target is not merely in-sample fit but reliable transfer to never-piloted items. The reported item-split design therefore tests exactly the regime that adaptive item-bank growth requires.

4. Geometric JE-IRT for LLM evaluation

A second formulation uses JE-IRT as a geometric model of LLM evaluation. Here, every model jj1 is assigned a learnable embedding jj2, and every question jj3 is mapped to an embedding

jj4

typically by a content-aware neural mapping (Yao et al., 26 Sep 2025). The geometry is interpretable by construction: the direction of a question embedding encodes semantics or topical specialization, while its norm encodes difficulty.

Correctness is controlled by a projected, question-specific ability

jj5

and the response probability is

jj6

This replaces a scalar ability parameter with a directional notion of competence. A model does not possess a single globally valid rank; rather, its effective ability depends on alignment with the question direction (Yao et al., 26 Sep 2025).

The formulation is explicitly contrasted with traditional 2PL IRT,

jj7

which imposes a global total order over models once jj8 is scalar. The geometric JE-IRT framework proves that such an order need not exist: there can be two models jj9 and two questions xj∈RFx_j \in \mathbb{R}^F0 such that xj∈RFx_j \in \mathbb{R}^F1 has higher projected ability on xj∈RFx_j \in \mathbb{R}^F2 but lower projected ability on xj∈RFx_j \in \mathbb{R}^F3 (Yao et al., 26 Sep 2025). It also establishes a smoothness result stating that geometrically similar questions induce similar predicted probabilities, up to a bound involving model norm, angular discrepancy, and the difference between question norms (Yao et al., 26 Sep 2025).

The interpretive consequence is that JE-IRT treats LLM evaluation as intrinsically multidimensional. Difficulty is separated from topic through norm-versus-direction geometry, and specialization replaces universal ordering.

5. Empirical findings in the geometric formulation

The empirical analysis supporting the geometric JE-IRT model reports several failures of scalar IRT on LLM benchmarks. When 2PL is fit in this setting, many items have zero or negative discrimination, and a large fraction exhibit flat or non-monotonic item characteristic curves. The observed response patterns do not support a simple “stronger model answers every question a weaker model answers” ordering (Yao et al., 26 Sep 2025).

Against that background, JE-IRT is reported to obtain accurate predictions with relatively low-dimensional embeddings, including dimensions in the 16–64 range, while outperforming baselines built on much higher-dimensional per-model vectors of roughly 252 dimensions (Yao et al., 26 Sep 2025). The learned question geometry also yields interpretable out-of-distribution behavior: performance degradation under leave-one-benchmark-out evaluation is larger when the mean direction of the held-out benchmark is less aligned with the remainder of the training data. In the reported examples, MathQA and LogiQA show small drops under high alignment, whereas PIQA and GSM8K drop more sharply under lower alignment (Yao et al., 26 Sep 2025).

Difficulty behaves geometrically as well. Larger question-embedding norms are associated with lower accuracy, and using norm alone to predict whether models answer a question correctly yields ROC AUC values around xj∈RFx_j \in \mathbb{R}^F4–xj∈RFx_j \in \mathbb{R}^F5 (Yao et al., 26 Sep 2025). Binned analyses show monotonic accuracy decline as xj∈RFx_j \in \mathbb{R}^F6 increases.

A further practical property is incremental extensibility. Once question embeddings are learned from content, a new LLM can be inserted by fitting only a single embedding vector xj∈RFx_j \in \mathbb{R}^F7. The reported experiments state that with just 10% of a held-out model’s data, test accuracy is already within 0.5% of joint training (Yao et al., 26 Sep 2025). The same learned space supports unsupervised structural analysis: k-means clusters of question embeddings are only partially aligned with human-defined subject labels, with homogeneity exceeding completeness, suggesting that the internal taxonomy induced by model behavior fragments some human subject categories (Yao et al., 26 Sep 2025).

6. Relation to adjacent IRT and neural-IRT literatures

JE-IRT sits within a broader movement that augments IRT rather than abandoning it. One nearby line is Deep-IRT, which combines the dynamic key-value memory network with IRT so that a neural system estimates student ability and item difficulty over time, and the final correctness probability is computed directly through an IRT equation (Yeung, 2019). Deep-IRT is therefore interpretable in the sense that both student and item states are exposed in psychologically meaningful terms, but it is not organized around a shared joint embedding geometry of the JE-IRT type.

Another important background is the Bayesian literature on joint estimation of proficiency and item parameters. Hierarchical IRT and temporal IRT infer student and item parameters together, regularized by priors and optionally by group structure or temporal smoothness, and were reported to match or outperform Deep Knowledge Tracing while retaining interpretability (Wilson et al., 2016). JE-IRT can be read as extending this jointness from estimation to learned latent representations.

The contrast with traditional 2PL and 3PL scoring is also instructive. A formal analysis of high-stakes IRT scoring shows that the principle of consistent order can fail under 2PL and 3PL: a student who answers more items, including harder ones, may still receive a lower score because estimated ability is an increasing function of the accumulated discriminations of correct items rather than of their difficulties (Lacourly et al., 2018). This does not invalidate IRT, but it clarifies why scalar orderings can become difficult to interpret in heterogeneous settings. A plausible implication is that embedding-based formulations become attractive when specialization, content structure, or multidimensionality are central.

A separate neighboring development, xj∈RFx_j \in \mathbb{R}^F8-IRT, addresses discrimination estimation directly by factorizing discrimination into sign and magnitude, xj∈RFx_j \in \mathbb{R}^F9, thereby improving sign recovery and difficulty estimation relative to z(⋅)z(\cdot)0-IRT (Ferreira-Junior et al., 2023). This line tackles a different technical problem from JE-IRT, but it underscores a common theme: modern IRT research increasingly modifies the parameterization itself rather than merely changing the optimizer.

7. Limitations, interpretations, and likely development paths

The current JE-IRT literature makes clear that embedding-based IRT does not eliminate classical psychometric constraints; it relocates them. In the calibration setting, the guessing parameter is fixed globally because learning z(⋅)z(\cdot)1 jointly with z(⋅)z(\cdot)2 raises identifiability and convergence problems (Sharpnack et al., 8 Jul 2026). In the same setting, BERT-based item features are reported to overfit at moderate item-bank size, and deeper stacks of squashing nonlinearities can collapse to nearly indistinguishable item parameters (Sharpnack et al., 8 Jul 2026). In the geometric LLM setting, the learned taxonomy only partially matches human subject labels, and the model rejects any single universal ranking of systems across all question directions (Yao et al., 26 Sep 2025).

These are limitations in one sense, but they are also interpretive statements about the domains being modeled. The adaptive-testing results suggest that content-derived low-dimensional embeddings may be sufficient for robust feature-only calibration when the operational objective is immediate deployment in a production IRT pipeline (Sharpnack et al., 8 Jul 2026). The LLM results suggest that benchmark behavior is better described by directional specialization and difficulty norms than by a one-dimensional ability scale (Yao et al., 26 Sep 2025).

Taken together, the two usages indicate a common research trajectory: use embeddings to expose structure that classical scalar IRT either compresses or cannot transfer across. In psychometric calibration, that structure is item content; in LLM evaluation, it is the geometry relating model competencies to question semantics. The shared premise is that IRT remains valuable as the response model, but its latent variables increasingly take the form of learned embeddings rather than isolated scalar parameters.

Topic to Video (Beta)

No one has generated a video about this topic yet.

Whiteboard

No one has generated a whiteboard explanation for this topic yet.

Follow Topic

Get notified by email when new papers are published related to JE-IRT.