- The paper demonstrates that teaching resists automation due to its reliance on contextual human judgment and integrated, non-modular approaches.
- It critiques assumptions that discrete, modular tasks in education can be automated, showing the importance of adaptive, relational teaching methods.
- The study highlights that AI systems improve factual recall but fall short in addressing interpretive and contextual evaluation critical for effective teaching.
Teaching Resists Automation: Human Judgment, Non-Modular Work, and the Limits of Delegation
Problematic Assumptions in Educational Automation
The paper critiques widespread assumptions underpinning claims about the automability of teaching, particularly those that characterize instructional work as modular and procedural. Automation narratives, which extrapolate from the compartmentalization of tasks seen in complex systems design, misrepresent teaching as a series of discrete, independently optimizable components. Empirical and conceptual work demonstrates that while tasks such as lesson planning, feedback generation, and assessment may seem modular in abstraction, their effective execution is deeply contextual, interpretive, and relational. Instructional efficacy derives not from the mechanical completion of isolated tasks, but from ongoing professional judgment exercised across time, learner profiles, and evolving classroom dynamics.
Empirical Limits of LLMs and Retrieval-Augmented Generation
Recent research into LLMs and RAG systems illuminates both the utility and intrinsic limitations of these technologies in educational contexts. LLMs, trained on large textual corpora, can generate fluent explanations, draft feedback, and simulate conversational support. Their statistical pattern-matching delivers strong performance on well-specified, correctness-evaluable tasks. However, systematic weaknesses persist in tasks requiring interpretive inference, nuanced judgment, and authentic contextual reasoning. Notably, LLM-assessment models systematically undervalue high-quality student responses reliant on inference or context, while rewarding structurally coherent but substantively shallow work.
RAG architectures mitigate certain issues—particularly those related to factual grounding—by conditioning outputs on retrieved content from authoritative sources. Empirical studies confirm improvements in factual accuracy and reference fidelity with RAG integration, but instructional value remains contingent on the teacher's interpretive work. Improved outcomes correlate not with AI-driven conversational fluency, but with the learner’s meaning-making processes and the instructor’s pedagogical interpretation.
Human-Factor Driven Complexity and the Epistemic Ceiling
The paper contends that the core irreducibility of teaching lies in its human-factor-driven complexity. Instructional success is predicated on interpreting partially observable internal states, including motivation, affect, social belonging, and perceived competence. Teachers operate under uncertainty, continually making high-stakes, context-sensitive decisions that cannot be reliably formalized or reduced to deterministic logic. The Situated Expectancy-Value Theory (SEVT) conceptualizes learning as shaped by dynamically constructed task values and fluctuating expectations—features not amenable to exhaustive modeling or algorithmic control.
Learners frequently misalign stated intentions, actual understanding, and strategic behavior. Addressing these gaps requires teacher judgment and relational scaffolding, not procedural automation. Even advanced AI systems lack the situational awareness and context sensitivity essential for responding effectively to these emergent issues. Instruction is fundamentally an interpretive process that manages uncertainty, not one that eliminates it via comprehensive specification.
Implications for Learning Design and Policy
The non-modular nature of teaching has direct implications for educational practice and policy. AI systems are productively deployed in support roles where task outputs are literal and structurally consistent—drafting materials, retrieving resources, generating summaries. However, tasks involving interpretive evaluation, motivational intervention, or conflict mediation necessitate retention of teacher agency. The instructional value of AI support is always filtered through human judgment.
Policy frameworks should prioritize professional development to equip educators with skills to discern appropriate AI integration points, emphasizing contexts where human judgment is central. Governance structures must safeguard teacher autonomy, ensuring AI is framed as a supplementary infrastructure rather than an evaluative authority, especially in assessment and feedback contexts. Transparency, accountability, and reference grounding are essential in AI-assisted decision-making.
Conclusion
The paper robustly demonstrates that teaching resists automation due to its inherently non-modular, judgment-centered, and human-factor driven complexity. While AI and RAG systems can support bounded instructional functions and information access, they fail to replicate the interpretive and relational labor integral to effective teaching. This limitation is rooted not in current technical constraints but in the epistemic structure of educational practice itself. Future developments in AI may augment support functions, but the ceiling of automation in teaching will persist as long as instruction depends on emergent human cognition, motivation, and social interaction. Preserving teacher agency and instructional quality requires a nuanced understanding of these principled limits.