Effectiveness of Intelligent Tutoring Systems (ITSs) on Teaching and Learning
Determine whether Intelligent Tutoring Systems (ITSs) meaningfully impact teaching and learning outcomes by conducting rigorous, transparent evaluations that can resolve mixed evidence and address criticisms of existing evaluation protocols.
References
Despite initial excitement about the potential of ITSs to revolutionise education~\citep{davies2021mobilisation, seldon2020fourth}, and their broad adoption~\citep{becker2017artificial, miao2021ai}, it remains unclear if they can impact teaching and learning in a meaningful way~\citep{holmes2022artificial,zawacki2019systematic}: evidence of their effectiveness is mixed~\citep{holmes2022artificial,ilkka2018impact,foster2023edtech,kulik2016effectiveness}, and the underlying evaluation protocols have come under criticism~\citep{wollny2021we,okonkwo2021chatbots} (see Section~\ref{sec:evaluation_its} for more details).
Beyond this single deployment, generalizing the Living Library model to other institutions raises questions this paper does not settle: how to standardize metadata across institutions with different cataloging histories; how to ensure long-term model accuracy and governance as underlying models change; what best practices should govern representing historical figures responsibly across different subjects and sensitivities; how to measure educational and engagement impact rather than infer it from anecdote; and what interoperability standards would let Layer 3 corpora from different institutions be queried together.
This study leaves several practical questions untested: whether instructors can interpret LO-level feedback, whether that feedback changes teaching decisions, and whether students who receive it learn more than students who do not.