Effectiveness of Intelligent Tutoring Systems (ITSs) on Teaching and Learning

Determine whether Intelligent Tutoring Systems (ITSs) meaningfully impact teaching and learning outcomes by conducting rigorous, transparent evaluations that can resolve mixed evidence and address criticisms of existing evaluation protocols.

Background

The paper reviews decades of work on Intelligent Tutoring Systems (ITSs), noting that despite broad adoption and initial optimism, empirical evidence on their effectiveness is mixed and evaluation methods have been criticized. This persistent uncertainty about impact, coupled with the absence of standardized evaluation practices, poses a foundational question for educational technology research and deployment.

Clarifying the actual effect of ITSs on teaching and learning is crucial for guiding investment, informing policy, and benchmarking newer generative AI tutoring approaches against established systems.

References

Despite initial excitement about the potential of ITSs to revolutionise education~\citep{davies2021mobilisation, seldon2020fourth}, and their broad adoption~\citep{becker2017artificial, miao2021ai}, it remains unclear if they can impact teaching and learning in a meaningful way~\citep{holmes2022artificial,zawacki2019systematic}: evidence of their effectiveness is mixed~\citep{holmes2022artificial,ilkka2018impact,foster2023edtech,kulik2016effectiveness}, and the underlying evaluation protocols have come under criticism~\citep{wollny2021we,okonkwo2021chatbots} (see Section~\ref{sec:evaluation_its} for more details).

Towards Responsible Development of Generative AI for Education: An Evaluation-Driven Approach  (2407.12687 - Jurenka et al., 2024) in Section 3.2 (Lack of transparency and common evaluation practices: lessons from EdTech)

Beyond this single deployment, generalizing the Living Library model to other institutions raises questions this paper does not settle: how to standardize metadata across institutions with different cataloging histories; how to ensure long-term model accuracy and governance as underlying models change; what best practices should govern representing historical figures responsibly across different subjects and sensitivities; how to measure educational and engagement impact rather than infer it from anecdote; and what interoperability standards would let Layer 3 corpora from different institutions be queried together.

The Living Library: Transforming Archival Collections into Conversational Knowledge Systems -- Lessons from the Theodore Roosevelt Presidential Library  (2609.09368 - Wang et al., 8 Sep 2026) in Section 10, “Limitations and Future Work,” paragraph “Open questions for the framework more broadly”

This study leaves several practical questions untested: whether instructors can interpret LO-level feedback, whether that feedback changes teaching decisions, and whether students who receive it learn more than students who do not.

Mechanics Cognitive Diagnostic: Testing Fine-Grained Learning Objectives in Introductory Physics  (2609.09584 - Le et al., 9 Sep 2026) in Section 7, Limitations