Papers
Topics
Authors
Recent
Search
2000 character limit reached

Learning Item Embeddings and Hyperparameters for IRT Calibration via Monte Carlo EM

Published 8 Jul 2026 in stat.AP | (2607.06905v1)

Abstract: High-stakes computerized adaptive tests (CATs) must continually calibrate new items in their item bank. When an item is new, few responses are available, so item parameter estimates -- and thus test scores -- are poor. Item features and explanatory item response theory (IRT) models mitigate this by folding item content into calibration. Neural IRT models, whose item parameters are neural-net outputs, are powerful, but tuning hyperparameters and architectures in real time while a CAT is scoring is impractical and threatens validity. We propose a pre-launch step that fits a neural net to produce low-dimensional item embeddings, so the production system can use a simple linear explanatory IRT model on top of them. We use a neural parameterization of the 3-parameter logistic (3PL) model in which a feature network maps each item's content features to a representation hj=z(xj)∈R<sup>dh_j = z(x_j) \in \mathbb{R}<sup>d, from which the discrimination and difficulty (a,b)(a,b) follow generalized linear forms; the guessing parameter cc is fixed to a global constant to avoid identifiability issues. The feature network and latent abilities θθ are fit jointly via Monte Carlo Expectation-Maximization (MCEM), with no separate ability-estimation or pre-calibration stage. Using an item-split protocol that holds out entire items to simulate feature-only evaluation, we apply this to two Duolingo English Test practice task types -- yes/no vocabulary and vocabulary-in-context -- searching over feature sets, architectures, and dimensions dd. A shallow two-layer ReLU network with d=6d=6 and hand-engineered scalar features matches or beats larger architectures on held-out items for both. This is a first step toward a compact, content-derived item embedding for the Scalable Parametric Item Calibration Engine (SPICE), the fully Bayesian engine at the core of the S2A3 adaptive-testing system.

Authors (2)

Summary

No one has generated a summary of this paper yet.

Paper to Video (Beta)

No one has generated a video about this paper yet.

Whiteboard

No one has generated a whiteboard explanation for this paper yet.

Open Problems

We haven't generated a list of open problems mentioned in this paper yet.

Continue Learning

We haven't generated follow-up questions for this paper yet.