Time Machine Experiments: Using AI's Temporal Knowledge Boundaries to Study the Human Mind

This presentation introduces a new experimental paradigm in which a language model's training cutoff date becomes a manipulable variable in human-computer interaction research. The authors demonstrate that when contemporary participants interact with an AI trained exclusively on pre-1930 material, their perception of historical morality shifts significantly compared to those interacting with a modern model. The study offers a framework for using historically bounded AI systems as experimental instruments to probe how people judge the past, while carefully defining what such experiments can and cannot tell us about history itself.
Script
Can you change how someone views the moral past by letting them talk to an AI that knows nothing after 1930? The authors of this paper built exactly that system and discovered it dramatically reduced the illusion that people were more moral 100 years ago.
Traditional archives preserve historical material but cannot answer new questions. Living witnesses can respond to questions but carry knowledge of everything that happened since. A language model trained only on material published before 1930 offers something different: an interactive system with a testable information boundary.
240 participants were randomly assigned to interact with either Talkie, a model trained on 260 billion tokens published before 1930, or a contemporary comparison model. They predicted how their assigned model would complete morally contested sentence stems, receiving feedback across multiple attempts on topics like women's employment, capital punishment, and interracial marriage.
Before the interaction, both groups believed people 100 years ago were more moral than people today. After talking with the 1930-bounded model, that perceived decline nearly disappeared, dropping by 0.8 scale points. Interacting with the modern model produced almost no change. The effect size was 0.56, and the shift came from revising judgments about the past, not about the present.
The paradigm opens a design space. You could deploy multiple period agents instead of one, create persistent interactions across weeks instead of single sessions, or build immersive environments instead of text interfaces. Each expansion increases the system's experiential realism but also the difficulty of maintaining temporal integrity and specifying what participants actually encountered.
This framework makes it possible to treat a system's knowledge cutoff as an experimentally assignable property rather than an aesthetic feature of a chatbot persona. The study demonstrates feasibility while defining clear limits: participants encountered one model's historically bounded responses, not actual contact with people from 1930. If you want to explore how temporal perspectives shape human judgment, you can learn more about this approach and create your own research videos at emergentmind.com.