Papers
Topics
Authors
Recent
Search
2000 character limit reached

Time Machine Experiments: Using Historically-Bounded AI for Inquiry into the Human Mind

Published 14 Sep 2026 in cs.HC and cs.CY | (2609.15468v1)

Abstract: Can interacting with someone from 1930, with no knowledge of what happened after, influence a person's perception of the past? People reason about the present against a picture of the past without observing it. The past is reconstructed from memory and testimony, but this reconstruction has been filtered through everything that happened since. Historically-bounded LLMs make that past available for interaction. As a proof-of-concept for the impact of interacting with historical minds, we ran a preregistered randomized experiment (N=240N=240), where participants interacted with an LLM trained on pre-1930 text. The interaction reduced the illusion of moral decline, the tendency to view the past as more moral than the present, compared to the contemporary-model control. This Time Machine Experiment paradigm informs new forms of interactive experiments, where temporal knowledge boundaries become experimental variables, and expands the realm of science fiction science, which turns thought experiments into actual experiments.

Summary

  • Exposing participants to a historically bounded AI modifies the human participants' perception of moral change by lowering their evaluation of the historical past.
  • The AI model, trained on material pre-1930, facilitates a unique experiment allowing participants to probe a temporally restricted viewpoint.
  • The differences in the approach between the two AI groups indicate the role of temporal content in perception.

Methodological premise

“Time Machine Experiments: Using Historically-Bounded AI for Inquiry into the Human Mind” proposes a research paradigm in which an AI system’s temporal knowledge boundary becomes an experimentally manipulated property of human–computer interaction (2609.15468). The central methodological object is not a simulation of a historical person, nor a model intended to estimate what an entire historical population believed. It is an interactive system whose training information is bounded at a specified date and whose effects on contemporary participants can be measured relative to a comparison condition.

The paper addresses a methodological problem in historical and behavioral inquiry. Archives preserve historically situated material but cannot answer unanticipated questions. Living witnesses can answer questions but possess retrospective knowledge shaped by subsequent events. Contemporary LLMs can produce historically styled responses, but prompting them to adopt an earlier persona does not reliably remove post-boundary information. A model trained on a temporally restricted corpus therefore offers a distinct experimental instrument: it can respond to participant-generated questions while maintaining a documented, testable information boundary.

The authors situate this proposal within HCI research on experience prototyping, speculative enactment, interactive probes, and systems that function simultaneously as objects of study and behavioral interventions. The paper also connects the paradigm to the “science fiction science” method, in which counterfactual or technologically inaccessible scenarios are converted into empirical designs. Here, the thought experiment is interaction with a standpoint from another time, while the historically bounded AI provides the experimental surrogate.

The proof-of-concept study tests whether interacting with an AI trained on pre-1930 material changes contemporary participants’ perception of moral change. The authors target the illusion of moral decline: the recurrent tendency to judge people in the past as more kind, honest, nice, and morally good than people today. The primary claim is not that the model reproduces historical morality, but that exposure to its historically bounded responses can alter how contemporary participants evaluate the moral past.

Conceptual and validity framework

The proposed paradigm is defined by the informational boundary of the deployed interactive system rather than by a specific model architecture or modality. A Time Machine Experiment could use text, speech, visual generation, embodied agents, or immersive environments, provided that the system’s temporal boundary can be specified and audited. Participants may question the system, predict its responses, compare it with contemporary norms, or explore an environment generated from period-bounded data.

The paper identifies four validity requirements that delimit what can be inferred.

Historical grounding concerns what material supports the system’s representation of a period. A corpus of published and preserved text is necessarily selective. It overrepresents literate, institutionally connected, and historically preserved populations, while underrepresenting everyday speech and marginalized groups. Consequently, the model should be interpreted as a generative extension of a historical record, not as a direct sample from a historical population.

Temporal integrity concerns whether post-boundary information has entered the complete deployed system through pretraining, prompting, retrieval, moderation, safety components, or other layers. The boundary must be validated for the assembled system, not merely asserted for the base model. This distinction is important because role prompting alone does not reliably suppress later knowledge.

Interaction specification concerns what participants actually encounter. Live model outputs, prompt construction, response selection, and task composition can all manufacture differences that are subsequently attributed to historical perspective. The study therefore requires disclosure of which stimuli were fixed before data collection, which were generated live, and how unscripted interaction was logged.

Comparative attribution concerns the control condition. A historically bounded model will typically differ from a contemporary model in parameter count, training regime, coherence, style, refusal behavior, capability, and predictability, in addition to temporal scope. Thus, the causal estimand is the effect of the bounded interaction condition as a package, not the isolated effect of temporal cutoff.

This framework is one of the paper’s strongest contributions because it prevents an overinterpretation that the empirical design cannot support. The study does not demonstrate what people in 1930 believed, nor does it reproduce an encounter with an actual historical individual. It estimates how contemporary participants respond to one historically bounded model relative to one contemporary model under a specified interaction protocol.

Experimental design

The preregistered experiment included 240 English-speaking participants recruited from the United States and the United Kingdom through Prolific. Participants were randomly assigned to either the 1930-AI condition (N=121N=121) or the Modern-AI condition (N=119N=119). The sample was balanced by gender, ranged from 19 to 83 years of age, and had a median session duration of approximately 22 minutes.

The treatment system was Talkie, a 13-billion-parameter LLM pretrained on 260 billion tokens of English-language material published before 1930. Its corpus included books, newspapers, periodicals, scientific journals, patents, and case law. The control system was GPT-5.5, described to participants as a contemporary model. The two conditions used the same interface, task structure, sentence stems, measurement instruments, and general procedure.

Participants first completed a baseline assessment of perceived morality across three temporal horizons: approximately 100 years ago, approximately 20 years ago, and today. They were then informed about their assigned model and completed a comprehension check. The interaction involved a familiarization conversation, a practice item, four prediction items, and a final free-form conversation.

The core intervention was a prediction task rather than unconstrained dialogue. Participants saw sentence stems describing morally or culturally contested situations and predicted whether the assigned model would complete each stem with a positive, negative, or neutral evaluation, together with the reason it would provide. They could make up to three attempts per item and received feedback after each attempt.

Figure 1

Figure 1: The prediction interface required participants to infer both the model’s evaluation and its stated grounds, with feedback across up to three attempts.

This structure served two methodological purposes. First, it standardized exposure to morally relevant content while preserving participant agency in probing the model. Second, it reduced the extent to which differences in long-context coherence or conversational capability could determine treatment exposure. The design nevertheless does not eliminate all capability confounds: the historically bounded model was smaller and less predictable than the contemporary control.

The 12-item pool was derived from repeated public-opinion questions from the Gallup World Poll, the General Social Survey, and related historical survey materials. The items concerned issues including women’s paid employment, contraception, capital punishment, physical discipline, euthanasia, abortion, interracial marriage, premarital relations, same-sex relations, women in high political office, gun-purchase permits, and the employment of openly homosexual teachers. Four items were sampled for each participant.

To evaluate predictions, the authors generated one reference completion per item and model. The references were selected from repeated model samples using criteria for coherence, non-circularity, intelligibility, brevity, and guessability. Participant responses received scores from 0 to 5 according to stance and similarity of reasoning. The scoring rubric was calibrated against human ratings, achieving Krippendorff’s α=.87\alpha = .87 in the reported validation procedure.

Figure 2

Figure 2: The study measured perceived moral decline before model disclosure, exposed participants to the assigned interaction, and repeated the measures afterward.

The main outcome was the change in perceived 100-year moral decline. For each participant, the decline score was the rating assigned to people today minus the rating assigned to people 100 years ago. Negative values indicate that the past was judged more moral than the present. A positive pre-to-post change therefore indicates attenuation of the illusion.

The authors distinguished between two components of perceived morality. Moral-value endorsement measured how much people were believed to value kindness, honesty, niceness, and goodness. Moral compliance measured how well people were believed to live up to those values in actual behavior. This distinction is analytically useful because a lower evaluation of the past could reflect changed beliefs about historical values, changed beliefs about historical conduct, or both.

Effects on perceived moral decline

At baseline, participants in both conditions expressed the illusion of moral decline. For moral-value endorsement, the 1930-AI group’s mean 100-year decline score was 0.80-0.80 with a 95% confidence interval of [1.15,0.45][-1.15,-0.45]; the Modern-AI group’s score was 0.61-0.61, with a 95% confidence interval of [0.94,0.28][-0.94,-0.28]. The conditions did not differ significantly at baseline, p=.438p=.438.

The post-interaction changes were markedly different. In the 1930-AI condition, the perceived moral value of people 100 years ago fell from 5.49 to 4.66 on the seven-point scale, whereas the rating of people today did not significantly change. In the Modern-AI condition, the 100-year rating fell only from 5.54 to 5.39. The resulting difference in pre-to-post change in perceived 100-year decline was $0.79$ scale points, with a 95% confidence interval of [0.44,1.15][0.44,1.15], Welch’s N=119N=1190, N=119N=1191, and Cohen’s N=119N=1192.

Figure 3

Figure 3: Interaction with the 1930-bounded model substantially reduced perceived moral decline by lowering evaluations of the historical period rather than increasing evaluations of the present.

Within the 1930-AI condition, the decline score increased from N=119N=1193 to N=119N=1194, corresponding to a paired effect of N=119N=1195. The post-interaction mean therefore indicated that the perceived decline had disappeared and slightly reversed. The authors emphasize that this change was driven primarily by a less favorable evaluation of the past, not by a more favorable evaluation of contemporary people.

The effect generalized to moral compliance. The change in perceived 100-year decline was N=119N=1196 points in the 1930-AI condition and N=119N=1197 points in the Modern-AI condition. The between-condition difference was N=119N=1198 points, 95% CI N=119N=1199, α=.87\alpha = .870, and α=.87\alpha = .871. Both primary between-condition effects survived Holm correction.

The magnitude was smaller but still detectable at the 20-year horizon. For both endorsement and compliance, the between-condition difference was approximately α=.87\alpha = .872 points, with Holm-adjusted α=.87\alpha = .873. This pattern is consistent with the intended temporal specificity of the treatment: the historical model supplied information most directly relevant to the approximately 100-year comparison, while the 20-year period was not represented in its corpus. The residual 20-year effect may reflect participants recalibrating an intermediate historical point after revising their evaluation of the more distant past.

Direct comparative judgments supported the scale-based results. The proportion of participants judging current moral endorsement as lower than 100 years ago decreased from 61.2% to 40.5% in the 1930-AI condition, compared with a smaller change from 53.8% to 50.4% in the Modern-AI condition. For behavioral compliance, the corresponding proportions changed from 57.9% to 41.3% and from 58.8% to 54.6%, respectively.

The authors’ regression adjustment produced nearly identical estimates. Controlling for age, gender, education, country, political self-placement, and AI use, the 1930-AI coefficient was α=.87\alpha = .874 for endorsement, with HC3 α=.87\alpha = .875, 95% CI α=.87\alpha = .876, α=.87\alpha = .877, and α=.87\alpha = .878 for compliance, with HC3 α=.87\alpha = .879, 95% CI 0.80-0.800, 0.80-0.801. This stability is expected under random assignment and provides a robustness check against chance imbalance rather than a replacement for the randomized comparison.

Reflective insight and model predictability

Participants also rated whether the interaction changed their perspective, altered how they viewed the past and present, and exposed them to a previously unconsidered worldview. The three-item scale had high internal consistency, 0.80-0.802.

Reflective insight was higher in the 1930-AI condition, with means of 4.15 versus 3.43 on a seven-point scale. The difference was 0.80-0.803 points, 95% CI 0.80-0.804, 0.80-0.805, and 0.80-0.806. The largest item-level difference concerned changed views of the past and present, 0.80-0.807. The difference for encountering a previously unconsidered perspective was 0.80-0.808. By contrast, the difference for reconsidering participants’ own values was small and nonsignificant, 0.80-0.809.

Figure 4

Figure 4: The historically bounded interaction increased reported perspective change and historical reflection, but produced little evidence of changes in participants’ core values.

This distinction is important for interpreting the principal outcome. The treatment appears to have changed the historical reference frame against which participants assessed contemporary society rather than substantially altering their own moral commitments. The qualitative responses support this interpretation: participants frequently reported that their values had remained stable while describing greater awareness of how moral judgments depend on historical and social context.

Participants were less accurate at predicting the 1930-AI than the Modern-AI. Mean prediction performance was 4.27 versus 4.43 on the five-point scale, a difference of [1.15,0.45][-1.15,-0.45]0, 95% CI [1.15,0.45][-1.15,-0.45]1, [1.15,0.45][-1.15,-0.45]2, and [1.15,0.45][-1.15,-0.45]3. Continuous similarity analyses converged on the same pattern: participant completions were closer to the contemporary model’s canonical responses in embedding space, while responses in the 1930-AI condition had lower negative log-likelihood under the historical Talkie base model.

Figure 5

Figure 5: Participants were closer to the Modern-AI reference in semantic embedding space, whereas their responses were more compatible with the historical base model under the common NLL measure.

Prediction accuracy was not associated with reflective insight, [1.15,0.45][-1.15,-0.45]4, [1.15,0.45][-1.15,-0.45]5, or with change in perceived moral decline, [1.15,0.45][-1.15,-0.45]6, [1.15,0.45][-1.15,-0.45]7. The absence of these associations weakens an explanation based on successful acquisition of a predictive model of the assigned AI. It is consistent with the alternative interpretation that the effects arose from encountering unfamiliar evaluative positions, although the correlational null results do not identify the mechanism.

Retry behavior provides limited evidence that participants progressively approximated model responses. Completion cosine increased modestly between the first and second attempts in both conditions, but all confidence intervals for changes in historical-model NLL included zero across retry transitions.

Figure 6

Figure 6: Retried responses moved modestly toward canonical completions in embedding space, while changes in historical-model likelihood were uncertain.

The qualitative interaction data reveal a further treatment difference. Participants assigned to the 1930-AI frequently asked about women’s rights, racial equality, interracial relationships, homosexuality, abortion, religion, and immigration. They often challenged apparent inconsistencies between abstract commitments to equality and judgments concerning particular groups. Modern-AI participants more often asked about general principles such as autonomy, fairness, harm, honesty, and individual rights, or about the system’s political orientation.

These patterns imply that the temporal boundary affected not only the responses participants received but also the questions they considered worth asking. Participants used the historically bounded system to locate discontinuities between period norms and contemporary values. That feature strengthens the ecological validity of interactive inquiry but complicates causal interpretation: the exposure was partly participant-selected, and the selected domains were precisely those most likely to display historical moral divergence.

Design space and methodological scope

The paper places the experiment at the origin of a three-dimensional design space: social complexity, temporal structure, and experiential fidelity.

Figure 7

Figure 7: Time Machine Experiments can vary the number of interacting agents, the structure and duration of temporal exposure, and the fidelity of the experiential interface.

The present design is a dyadic, single-session, text-based encounter with one historically bounded model. More complex designs could introduce multiple period agents, persistent interactions, staggered model cutoffs, voice and persona, navigable environments, or immersive VR. Each expansion increases the burden of validity.

Higher social complexity makes interaction less specifiable because emergent behavior among multiple agents cannot be fully fixed in advance. Persistent contact creates more opportunities for post-boundary information to enter the interaction. Earlier or sparser historical corpora make historical grounding more difficult. Conversely, a ladder of models with staggered cutoffs could help separate temporal boundary from model-family and capability effects.

Figure 8

Figure 8: Across moral endorsement and behavioral compliance, the largest post-interaction shift occurred for the 100-year rating in the 1930-AI condition.

The framework has potential applications in historical education, heritage interpretation, interactive archival research, and experimental studies of nostalgia and temporal judgment. The paper’s methodological contribution is strongest where it treats the historical boundary as an assignable interaction variable rather than as an aesthetic feature of a simulated persona. Participants can actively interrogate the boundary, discover what the system does not know, and compare its responses with present-day expectations.

However, the proposed applications must preserve the paper’s distinction between historical evidence and generative reconstruction. A period model can make an archive more interrogable, but it cannot by itself solve archival selection bias or establish the prevalence of a belief in the historical population. Its outputs remain model-mediated extensions of an uneven textual record.

Limitations and open questions

The most important limitation is comparative attribution. Talkie and GPT-5.5 differed in temporal knowledge, scale, training data, instruction tuning, linguistic style, capability, coherence, and predictability. Participants also knew whether they were interacting with a historical or contemporary model. The observed effect could therefore reflect the bounded historical content, the novelty of the interaction, the model’s lower predictability, the framing of the treatment, or their combination. The study estimates the effect of the complete interaction package rather than the temporal cutoff in isolation.

A stronger design would compare the historical model with a contemporary model matched more closely in architecture, training budget, interface behavior, and post-training, while varying the corpus cutoff. Staggered cutoff models could further distinguish effects specific to 1930 from generic effects of historical distance. Such comparisons would still require auditing the entire deployed system for post-boundary leakage.

Historical grounding is also limited. The pre-1930 corpus is dominated by written and preserved material, and the authors acknowledge that it cannot represent the full distribution of historical beliefs or conduct. The fact that the model’s responses sometimes align with surviving survey marginals is evidence about the model’s historical voice, not a population-level validation of 1930 public opinion. The prediction items themselves were a purposive sample of contested moral topics, and the qualitative analysis was descriptive, based on one author’s reading without inter-rater reliability.

The outcome measurement is immediate and self-reported. Participants were assessed minutes after a single interaction, so the persistence of the change is unknown. The study also does not establish whether the effect transfers to behavior, political judgment, memory, or beliefs outside moral decline. The modest 20-year effects are compatible with several explanations, including anchoring changes induced by the revised 100-year judgment.

Finally, the mechanism remains unresolved. Participants did not need to predict the historical model accurately for belief change or reflective insight to occur, but this does not establish that unfamiliarity itself caused the effect. The treatment included historical framing, novel model behavior, morally contested content, feedback, and participant-selected questioning. The paper therefore leaves open whether the primary causal ingredient is historical information, exposure to norm divergence, interactive perspective-taking, model unpredictability, or the combination of these factors.

Conclusion

The paper introduces historically bounded AI as an experimental instrument for studying how contemporary people respond to temporally situated perspectives. In a preregistered randomized study with 240 participants, interaction with a model trained on pre-1930 text reduced perceived moral decline relative to interaction with a contemporary model, with a between-condition effect of [1.15,0.45][-1.15,-0.45]8 for moral-value endorsement and [1.15,0.45][-1.15,-0.45]9 for moral compliance. The same interaction increased reported reflective insight, particularly regarding participants’ views of the past and present.

The findings support the feasibility of the paradigm while also defining its evidential limits. The study demonstrates an effect of a bounded interaction condition, not contact with a historical person and not recovery of historical population beliefs. Its principal methodological value lies in making temporal knowledge boundaries experimentally assignable, interactively probeable, and subject to explicit validity requirements (2609.15468).

Whiteboard

Explain it Like I'm 14

论文主题

这篇论文研究了一种像“时间机器”一样的人工智能实验。研究人员想知道:如果今天的人能和一个只知道 1930 年以前信息的 AI 聊天,他们对过去的看法会不会改变?

这个 AI 不是假装自己生活在过去,而是使用主要来自 1930 年以前的书籍、报纸和其他文字训练。因此,它不会像现代 AI 那样知道后来发生的所有事情。

研究人员特别关注一个常见想法:人们常常认为过去的人比今天的人更善良、更诚实、更有道德。 这被称为“道德衰退错觉”。

研究人员想回答什么问题?

论文主要想回答以下问题:

  • 和一个只了解过去的 AI 交谈,是否会改变人们对过去的看法?
  • 人们是否会因此减少“过去比现在更有道德”的想法?
  • 这种 AI 能不能成为一种新的科学研究工具,帮助我们研究人类如何理解历史和时间?

研究人员并不是想证明这个 AI 就代表了真实的历史人物或所有过去的人。它更像是一个根据历史文字制作出来的“实验伙伴”。

研究方法:他们是怎样做实验的?

研究人员招募了 240 名来自美国和英国的英语使用者。参与者被随机分成两组:

  • 一组和只使用 1930 年以前资料的 AI 交谈,称为“1930-AI”。
  • 另一组和现代 AI 交谈,作为比较对象。

随机分组很重要,因为这样两组人在开始时大致是公平的。最后出现的不同结果,就更可能是由 AI 的不同造成的,而不是由参与者本身造成的。

实验过程

参与者先回答问题,例如:

你认为 100 年前的人有多重视善良、诚实和做好人?

他们要分别评价今天的人、约 20 年前的人和约 100 年前的人。

接着,他们和自己的 AI 进行互动,包括:

  1. 先进行一小段自由聊天;
  2. 阅读一些与道德和社会观念有关的句子;
  3. 猜测自己的 AI 会如何评价这些句子;
  4. 猜测 AI 会给出什么理由;
  5. 根据 AI 的回答修改自己的猜测;
  6. 最后再次和 AI 自由聊天。

例如,参与者可能需要完成这样的句子:

“如果丈夫已经能够养家,已婚女性去工作是_,因为_。”

参与者要预测 AI 会认为这件事是好、坏,还是不好说,并猜测 AI 会怎样解释。

这项任务有点像一个“猜 AI 想法的游戏”。通过这个过程,参与者会更认真地了解 AI 的观点,而不只是随便读几句话。

互动结束后,参与者再次回答关于过去和现在道德水平的问题。研究人员比较他们前后的答案。

研究结果

人们原本确实认为过去更有道德

在实验开始时,两组参与者都认为:

  • 100 年前的人比今天的人更重视善良、诚实和做好人;
  • 过去的人在行为上也更符合这些道德标准。

这说明“过去更有道德”的想法确实存在于参与者中。

和历史 AI 聊天后,这种看法明显减弱

与 1930-AI 互动后,参与者对 100 年前人的道德评价明显下降。

简单来说,他们不再那么相信:

“过去的人一定比现在的人更善良、更诚实。”

在 1930-AI 组中,参与者认为过去更有道德的差距几乎消失了,甚至出现了轻微反转。

相比之下,和现代 AI 聊天的人,想法变化很小。

重要的是,参与者对今天人的评价并没有明显提高。变化主要来自于他们降低了对过去的理想化评价

历史 AI 也让人们更多反思

和 1930-AI 交谈的人更常表示:

  • 自己重新思考了原来的观点;
  • 自己改变了对过去和现在的看法;
  • 自己接触到了以前没有考虑过的角度。

这说明历史 AI 不只是提供答案,也可能促使人们重新检查自己的假设。

为什么这些结果重要?

这些结果说明,人们对过去的印象并不一定完全来自真实历史。我们常常会把过去想象得更简单、更善良,忘记过去也存在不公平、歧视和其他问题。

与一个“不了解后来历史”的 AI 对话,可能帮助人们发现:

  • 过去并不一定像记忆中那么美好;
  • 人们对历史的判断可能受到今天生活经验的影响;
  • 所谓“社会越来越糟”有时可能只是我们的错觉。

这项研究还展示了一种新的实验方法。过去,研究人员只能阅读历史资料,或者询问今天仍然活着的人。但历史资料不能回答新的问题,而现代人又已经知道后来发生的事情,可能会受到“事后之明”的影响。

历史 AI 提供了第三种可能:它可以对问题作出回应,同时尽量保持在某个历史时期的信息范围内。

这项研究的局限

研究人员也强调,这种 AI 并不等于真正的历史人物,也不能代表 1930 年以前所有人的思想。它学习的资料主要是保存下来的文字,而这些文字可能:

  • 只代表一部分人;
  • 更常来自有教育或有权力的人;
  • 忽略了许多普通人和弱势群体的声音。

此外,两个 AI 不只在“知道哪些年代的信息”上不同。它们还可能在语言风格、回答能力和回答长度上不同。因此,研究结果不能完全证明“时间限制本身”造成了变化,只能说明:与这种历史 AI 进行整体互动,会影响人们的看法。

研究规模也比较小,参与者只来自美国和英国,实验时间大约 22 分钟。因此,还需要在更多国家、更多年龄群体和更长时间内重复研究。

研究的影响和意义

这篇论文提出了“时间机器实验”这一新方法。它的核心想法是:

让 AI 的知识停留在某个历史时间点,再观察今天的人如何与它互动,以及这种互动如何改变人的想法。

未来,这种方法可能被用于研究:

  • 人们如何理解不同历史时期;
  • 我们为什么会美化过去;
  • 和历史观点互动是否能减少偏见;
  • 人们如何看待未来或不同文化;
  • AI 是否能帮助人们进行更深入的反思。

总的来说,这项研究表明,AI 不只是一个回答问题的工具,也可以成为研究人类思想的“实验仪器”。不过,我们必须小心使用它,清楚说明 AI 代表什么、不代表什么,并检查它是否真的遵守了设定的历史时间范围。

Knowledge Gaps

Knowledge gaps, limitations, and open questions

  • Temporal integrity of the deployed system remains incompletely established. The paper reports that Talkie was trained on pre-1930 text, but it does not fully demonstrate that the deployed interface, moderation layer, system prompts, post-training, filtering, or other components introduced no post-1930 information.
  • The historical model cannot be treated as a representative historical person or population. The study does not determine whose perspectives are overrepresented or absent in the corpus by gender, class, race, geography, literacy, occupation, or political position.
  • Corpus-selection bias is not quantified. The effects of OCR errors, preservation practices, publication inequalities, English-language coverage, and the overrepresentation of elite or institutional writing on model outputs remain unclear.
  • The causal effect of temporal boundedness is not isolated from model differences. Talkie and the contemporary model likely differ in parameter count, fluency, coherence, refusal behavior, verbosity, training objectives, and general capability, so the observed effect may reflect the historically bounded condition as a package rather than the cutoff itself.
  • The contemporary control is not demonstrably matched to the historical model on interaction quality. Differences in perceived credibility, novelty, conversational naturalness, difficulty, or engagement could independently explain the treatment effect.
  • The study does not test alternative control conditions that would separate historical content from presentation effects. For example, it does not compare Talkie with a contemporary model stylistically matched to Talkie, a scripted historical text, a prompted historical persona, or a modern model given equivalent information about historical moral norms.
  • The historical model’s actual knowledge boundary is tested only to a limited extent. A systematic, adversarial audit across direct and indirect post-1930 queries is needed to establish whether the model can reveal later information through associations, extrapolation, or memorized contamination.
  • The intervention exposes participants to only a small and curated set of moral prompts. Four main prediction items cannot establish whether the effect generalizes across moral foundations, political issues, cultures, historical periods, or less explicitly evaluative topics.
  • The item-construction process may have shaped the findings. Reformulating historical survey questions into sentence completions with required justifications may alter their meaning, introduce present-day assumptions, or make participants respond to linguistic artifacts rather than historical perspectives.
  • Reference completions may not represent the models’ natural behavior. Selecting one pre-generated response per item after quality screening and prioritizing the modal stance can suppress output variability and create fixed targets that participants learn rather than encounter naturally.
  • The LLM-based scoring procedure remains a potential source of measurement error. Although calibrated against three human raters, the paper does not establish whether scoring is unbiased across conditions, moral positions, linguistic styles, or reasoning quality.
  • The prediction task may confound learning, compliance, and belief change. Repeated attempts, feedback, performance bonuses, and the requirement to infer the model’s answers could cause participants to optimize for task performance rather than genuinely revise their beliefs about moral change.
  • Demand characteristics and framing effects are not fully ruled out. Participants learned that their model was historically bounded and were explicitly asked about changes in their views, which may have encouraged hypothesis-consistent responses.
  • The study cannot determine whether participants changed beliefs about historical morality or merely lowered ratings of the past. The primary effect arose mainly from reduced ratings for people 100 years ago, and the design does not establish whether this reflects improved historical accuracy, generalized skepticism, negative reactions to the model, or assimilation to its responses.
  • The accuracy of participants’ revised beliefs is not validated. The study does not compare post-intervention judgments with independent historical evidence or assess whether the intervention reduces bias rather than introducing a different distortion.
  • The measures rely heavily on self-report and immediate post-interaction judgments. It remains unknown whether the observed changes persist over time or affect political attitudes, historical reasoning, behavior, or real-world decisions.
  • The duration and intensity of exposure are limited. A median session of approximately 22 minutes does not show whether repeated, longer-term, or self-directed interactions produce durable, cumulative, or counterproductive effects.
  • Potential harms of historical interaction are not examined. Future work should test whether historically bounded models normalize discriminatory beliefs, reproduce archival prejudice, increase nostalgia, or influence vulnerable users differently.
  • The sample limits generalizability. Participants were English-speaking residents of the United States or United Kingdom recruited through Prolific; effects may differ by age, education, political ideology, cultural background, historical knowledge, and familiarity with AI.
  • Cross-cultural and multilingual validity is unresolved. The paradigm is based on an English-language, largely Western corpus and does not establish whether similar effects occur for other languages, regions, historical traditions, or culturally distinct moral frameworks.
  • Individual differences in susceptibility are not fully explained. The paper does not identify whether effects vary according to participants’ baseline nostalgia, trust in AI, perceived model competence, prior beliefs about moral decline, religiosity, political orientation, or historical literacy.
  • The contribution of free-form conversation is unclear. Because participants could ask arbitrary questions before and after the structured task, the study cannot determine which interaction content, conversational strategy, or participant-generated question caused the observed change.
  • The mechanism linking interaction to reduced moral decline is not established. Reflective insight is measured, but the study does not test whether reflection mediates the effect, whether prediction accuracy matters, or whether exposure to disagreement, uncertainty, or model inconsistency drives belief revision.
  • The qualitative analysis is exploratory and potentially subjective. One author coded the responses without a formal coding scheme or inter-rater reliability, limiting confidence in claims about recurring conversational themes.
  • The study does not compare different historical cutoff dates. It remains unknown whether effects vary smoothly with the cutoff, depend on specific historical periods, or result specifically from the contrast between pre-1930 and contemporary knowledge.
  • The paradigm’s applicability beyond text dialogue is untested. The paper proposes speech, images, audiovisual environments, and embodied agents, but provides no evidence that the validity requirements or psychological effects transfer to these modalities.
  • The reproducibility of the findings may be sensitive to model and interface updates. Hosted models, moderation systems, prompts, and APIs can change over time, yet the study does not establish whether the effect survives exact replication with version-locked systems.
  • The boundaries of permissible inference from Time Machine Experiments require further empirical validation. The paper proposes four validity requirements, but does not show how failures in historical grounding, temporal integrity, interaction specification, or comparative attribution quantitatively alter conclusions.

Practical Applications

Immediate Applications

The paper’s findings support applications that use historically bounded AI as an interactive research instrument or reflective intervention, rather than as an authoritative reconstruction of historical people or beliefs.

  • Behavioral-science experiments on historical perspective and biasAcademia / psychology / HCI Researchers can deploy historically bounded LLMs in randomized studies to examine how people form beliefs about the past, respond to historical viewpoints, and revise judgments after interactive exposure. The paper provides a reusable workflow: establish a temporal cutoff, assign participants to historical and contemporary model conditions, standardize core prompts, log interactions, and measure pre/post changes in beliefs or reflection. Dependencies: The model’s corpus, cutoff, prompts, interface, and moderation components must be documented. Effects should be attributed to the complete model package—not automatically to temporal boundedness alone.
  • Interventions addressing the “illusion of moral decline”Education / civic organizations / public communication Museums, schools, libraries, and civic-education programs could offer guided dialogues with a pre-1930 model to help users question idealized narratives about the past. The experiment found that interaction with the historical model reduced participants’ perception that people 100 years ago were morally superior to people today. Potential product: A museum or classroom module in which students predict how a historical AI will evaluate social situations, receive feedback, and discuss why their expectations differed. Dependencies: The intervention should be framed as perspective-taking and critical reflection, not as proof of what historical populations believed. It requires age-appropriate content controls, historical contextualization, and human facilitation for sensitive topics.
  • Interactive historical-literacy toolsEducation / museums / libraries Institutions can supplement archival materials with conversational interfaces that answer user-generated questions from a bounded historical corpus. Such systems could help learners explore how concepts, arguments, and social norms were represented in documents before a specified date. Potential workflow:
  1. User asks a question about a historical period.
  2. The system responds using the bounded corpus.
  3. The interface displays source provenance, uncertainty, and links to original documents.
  4. The learner compares the response with contemporary evidence and scholarly interpretation. Dependencies: Retrieval or citation mechanisms, corpus provenance, OCR quality, and clear disclosure that generated responses are model-mediated interpretations.
  • Research on presentism and hindsight biasHistory / social science / law Scholars can use bounded models to study how contemporary users project current knowledge and moral standards onto earlier periods. Participants might be asked to predict a historical model’s responses, challenge it with present-day concepts, or compare it with a modern model. Potential application: Experiments on hindsight bias, historical empathy, political nostalgia, attitudes toward social change, or judgments of past legal decisions. Dependencies: Carefully matched control conditions are necessary because historical and contemporary models may differ in capability, tone, verbosity, refusal behavior, and cultural representation.
  • Design and evaluation of temporal-perspective interfacesSoftware / HCI / product design Product teams can use the paper’s structured prediction task to evaluate interfaces intended to promote reflection. Rather than relying only on open-ended conversation, a system can ask users to predict an AI’s response, provide feedback, and measure changes in understanding. Potential tools: Reflection dashboards, interactive timelines, “compare perspectives across eras” interfaces, and conversation logs for qualitative analysis. Dependencies: Predefined tasks reduce uncontrolled variation, but they may also constrain natural interaction. Designers should separately test structured and free-form modes.
  • Museum and archival exploration workflowsCultural heritage / libraries / public history Archives can use bounded AI as a conversational layer over digitized newspapers, books, patents, scientific journals, or case law. Users could ask questions that are not answered verbatim in any single document while the system identifies the relevant source period and uncertainty. Dependencies: The corpus will reflect preservation and digitization biases. The system should not be presented as representing an entire society, demographic group, or historical “mind.”
  • Policy communication about social progress and declineGovernment / public policy / civic communication Policymakers and communicators can use historically grounded interactive exhibits to test whether public beliefs about social decline are based on evidence or nostalgia. This may improve public interpretation of long-term survey data and reduce support for policies justified solely by an idealized past. Dependencies: The tool should not be used to advocate a predetermined political position. Claims should be triangulated with historical surveys, administrative data, and expert review.
  • Internal training for researchers using historical AIAcademia / research governance Universities and research teams can immediately adopt the paper’s four-part validity checklist when developing historical-model studies: historical grounding, temporal integrity, interaction specification, and comparative attribution. This can become a protocol for preregistration, model cards, ethics review, and replication packages. Dependencies: Teams need access to model-training information, deployment logs, prompt specifications, and tests for post-cutoff leakage.
  • Public-facing critical-thinking exercisesDaily life / media literacy Individuals could use a historically bounded system to challenge assumptions such as “people were more respectful in the past” or “society has continuously declined.” A guided comparison between archival evidence, the bounded model, and a contemporary model could encourage users to distinguish memory, nostalgia, and evidence. Dependencies: Unguided use may reinforce stereotypes or generate fabricated historical claims. Source links and warnings about model limitations are essential.

Long-Term Applications

The following possibilities require additional validation, larger-scale deployment, improved model control, or research beyond the paper’s proof of concept.

  • Cross-temporal experiments on political attitudes and polarizationPolitical science / policy Future studies could test whether interacting with models bounded before major political transformations changes attitudes toward nationalism, democracy, gender roles, immigration, reproductive rights, colonialism, or economic policy. Multiple cutoff dates could distinguish the effects of different information environments. Dependencies: Such studies require representative samples, culturally diverse corpora, carefully neutral framing, and safeguards against normalizing discriminatory historical views. A single historical model cannot establish how a population actually thought.
  • Longitudinal interventions for nostalgia and political decision-makingMental health / civic policy / communications Repeated, evidence-linked interactions might help users evaluate “return to the past” narratives before voting or making policy judgments. Researchers could measure whether effects persist beyond the immediate post-interaction survey and whether they influence actual information-seeking or political choices. Dependencies: The current evidence measures short-term self-reported belief change. Persistence, behavioral transfer, and possible reactance remain untested.
  • Historically bounded AI tutors for educationEducation Future systems could provide period-specific tutors for history, literature, science, law, and economics. Students might debate with an AI limited to the knowledge available in 1789, 1850, or 1929, then compare its reasoning with later developments. This could teach scientific paradigm shifts, historical contingency, and the difference between contemporary knowledge and period knowledge. Dependencies: Models need reliable uncertainty expression, curriculum alignment, multilingual support, citation, teacher oversight, and resistance to anachronistic leakage.
  • Simulation of historical decision environmentsBusiness / policy / military history / organizational research Bounded models could support counterfactual experiments in which participants make decisions using only information available at a historical point—for example, evaluating an investment, scientific hypothesis, public-health measure, or political crisis. Researchers could study how hindsight and outcome knowledge affect judgment. Dependencies: A LLM’s corpus is not equivalent to the information available to a specific decision-maker. The system would require carefully reconstructed information sets, role-specific access controls, and validation against primary documents.
  • Historical legal and regulatory analysisLaw / compliance / public administration Models bounded at particular dates could help researchers examine how legal reasoning and regulatory language changed over time, or simulate how a historical legal corpus would frame a contemporary issue. This could support comparative legal scholarship and training exercises. Dependencies: Legal applications require authoritative source retrieval, jurisdictional controls, citation-level verification, and strong disclaimers. Generated output must not be treated as legal advice or as evidence of historical legal consensus.
  • Embodied and multimodal “time machine” environmentsVR / robotics / cultural heritage The paradigm could extend from text chat to speech, visual scenes, avatars, virtual neighborhoods, or embodied agents whose available information is bounded by a historical date. Users might explore how a period-specific agent interprets objects, social roles, or emerging technologies. Dependencies: Every component—including image models, speech systems, retrieval tools, safety filters, and environment assets—must be checked for post-cutoff information. Multimodal systems also increase risks of fabricated details and stereotypical representation.
  • Historical perspective-taking for conflict resolutionPeacebuilding / diplomacy / social psychology Carefully designed systems could expose users to competing historical narratives while preserving temporal boundaries, helping them identify how later events reshape interpretations of earlier conflicts. Researchers could test effects on empathy, dehumanization, and willingness to engage with opposing groups. Dependencies: This is high-risk: a model may reproduce propaganda, archival silences, or dominant-group perspectives. Deployment would require local experts, plural corpora, transparent provenance, and facilitated dialogue rather than autonomous persuasion.
  • AI-assisted discovery of historically plausible hypothesesHumanities / social science / science and technology studies Researchers could compare outputs from models trained to different cutoff dates to identify concepts, inventions, or arguments that were plausible before they emerged historically. This could generate hypotheses about technological forecasting, cultural diffusion, and the development of social norms. Dependencies: Model outputs are hypothesis-generating, not historical evidence. Results require archival verification, controls for corpus size and genre imbalance, and methods for distinguishing genuine historical patterns from training artifacts.
  • Personalized reflection and anti-nostalgia applicationsMental health / coaching / daily life A future consumer product could help users examine autobiographical or intergenerational nostalgia by comparing personal memories with period-specific documents and survey data. It might prompt questions such as whether a remembered past was broadly experienced or selectively remembered. Dependencies: Such systems would handle sensitive personal and political beliefs. Privacy, data minimization, non-manipulative design, and clinical validation would be required before use in mental-health contexts.
  • Benchmarking temporal integrity and historical AI safetyAI engineering / governance The paper’s validity requirements could become a standard evaluation suite for “historically bounded” models. Benchmarks could probe whether systems reveal post-cutoff facts through direct questions, indirect causal reasoning, prompts, retrieval, safety layers, or multimodal inputs. Potential outputs: Temporal-integrity certificates, model cards specifying cutoff boundaries, leakage test suites, audit logs, and deployment standards for research and education. Dependencies: No finite probe can prove the complete absence of leakage. Certification would need layered audits, reproducible training data documentation, adversarial testing, and explicit limits on the claims made.
  • Large-scale studies of cultural change using multiple historical modelsSociology / anthropology / economics Researchers could compare participant responses across models bounded at 1800, 1900, 1930, 1950, and later dates to map how perceived norms and judgments shift as information environments change. This could produce richer models of cultural evolution and collective memory. Dependencies: Differences between models may reflect architecture, training scale, language, corpus availability, or alignment—not only historical time. Large preregistered studies and matched model families would be necessary.
  • Policy evaluation of nostalgia-driven interventionsGovernment / political communication Governments or civil-society organizations could experimentally test whether evidence-based historical interactions reduce support for policies justified by claims of moral or social decline. Such work might inform public-information campaigns and deliberative-democracy tools. Dependencies: Manipulating political beliefs raises ethical and democratic concerns. Studies should prioritize informed consent, neutrality, transparency, independent oversight, and measurement of unintended effects such as distrust or polarization.

Glossary

  • Archival materials: Historical records preserved for later study, such as documents, newspapers, or other primary sources. “Archival materials provide historically situated evidence, but they are fixed and reflect selective processes of production and preservation”
  • Causal attribution: The process of determining whether an observed effect was caused by a particular intervention or experimental condition. “the comparison condition must support the causal attribution drawn from it”
  • Contemporary frontier model: A highly capable, state-of-the-art AI model available at the present time. “interacting with a current frontier model: gpt-5.5”
  • Corpus: A structured collection of texts used for linguistic analysis or model training. “13 billion parameters pretrained on 260 billion tokens of English-language text published before 1930”
  • Counterfactual: Describing a hypothetical situation that differs from what actually occurred. “design fiction and speculative design, which construct alternative or counterfactual technological scenarios”
  • Degrees of freedom: The number of independent values available to vary when estimating a statistical quantity. “a tt statistic with degrees of freedom in parentheses”
  • Embodied agent: An artificial agent represented as having a physical or virtual body through which it interacts with users or environments. “text, speech, images, audiovisual environments, embodied agents, or combinations of these”
  • Empirical studies in HCI: Research based on observed or measured human behavior in human–computer interaction settings. “Human-centered computing~Empirical studies in HCI”
  • Experimental vignette: A standardized, hypothetical scenario presented to study participants as part of an experiment. “Experimental vignettes provide standardized and manipulable scenarios”
  • External validity: The extent to which findings generalize beyond the particular participants, setting, or conditions of a study. “Although these are not new kinds of validity”
  • Hindsight bias: The tendency to see past events as having been more predictable after their outcomes are known. “their recollections and retrospective judgments may be shaped by subsequent experiences and outcome knowledge”
  • Historically-bounded LLM: A LLM trained on data available only before a specified historical cutoff. “There now exist LLMs trained exclusively on historical text corpora”
  • Human–computer interaction (HCI): The interdisciplinary study and design of interactions between people and computing systems. “Human-computer interaction (HCI) research has contributed to this advancement”
  • Illusion of moral decline: The belief that people in the past were more moral than people in the present, despite evidence that does not support this conclusion. “the `Illusion of Moral Decline', which describes people’s tendency to perceive past societies as more moral than contemporary ones”
  • Information environment: The body of information available to a system and the constraints governing that availability. “an information environment intentionally bounded in time”
  • Inter-rater reliability: The degree to which multiple evaluators produce consistent judgments when assessing the same material. “No inter-rater reliability was computed”
  • LLM: A machine-learning model trained on large amounts of text to generate and interpret language. “LLMs offer HCI another such methodological opportunity”
  • Manipulation: An experimentally controlled change introduced to test its effect on participants’ responses or behavior. “an ordinary manipulation has no informational boundary that can leak”
  • Methodological apparatus: A tool, system, or procedure used to conduct research or investigate a phenomenon. “Interactive systems have long served in HCI and adjacent behavioral sciences both as objects of evaluation and as methodological apparatuses”
  • Moral compliance: The extent to which people’s behavior conforms to their own stated moral values. “The pattern held for moral compliance, the second measure”
  • Ordinary least squares regression: A statistical method that estimates relationships by minimizing the squared differences between observed and predicted values. “we fit an ordinary least squares regression with the change in perceived moral decline as the outcome”
  • Optical character recognition (OCR): Technology that converts text in images or scanned documents into machine-readable characters. “instead involved conventional OCR”
  • Post-training: Training performed after a model’s initial pretraining to adapt its behavior, capabilities, or interaction style. “Talkie is also notable for combining relatively large-scale historical pretraining with post-training specifically designed to support conversational interaction”
  • Preregistered randomized experiment: An experiment whose design and analysis plan are specified in advance and whose participants are assigned randomly to conditions. “we conducted a preregistered randomized experiment (N=240N=240)”
  • Proof of concept: An initial demonstration that a proposed method or system is feasible and potentially effective. “To demonstrate the potential of the Time Machine Experiment as an instrument for scientific inquiry, we conducted a proof-of-concept experiment”
  • Qualitative data: Non-numerical information, such as written responses or interview transcripts, analyzed for meaning or recurring themes. “We used these qualitative data to contextualize the quantitative findings”
  • Randomized controlled comparison: An experimental comparison in which participants are assigned by chance to different conditions to estimate causal effects. “Participants were randomly assigned to either of the two conditions, which identifies the causal effect of the model on changes in participant evaluations”
  • Reflective insight: A measured change in perspective or understanding resulting from reconsidering one’s assumptions, values, or experiences. “We measured reflective insight using three items adapted from the Insight subscale”
  • Speculative design: A design practice that creates hypothetical technologies or situations to provoke reflection about social assumptions and possible futures. “Our conceptualization of the Time Machine Experiment is situated across several HCI traditions, such as design fiction and speculative design”
  • Temporal boundary: A cutoff that limits the time period from which a system may obtain information. “The model was accessed through a moderated interface and presented to participants under the name 1930-AI”
  • Temporal generalization: The ability of a model or method to perform across different time periods or historical conditions. “This development has attracted interest not only for the evaluation of temporal generalization and forecasting”
  • Temporal integrity: The extent to which a system actually preserves its stated historical information cutoff after deployment. “Temporal integrity. Does the boundary hold in the system as deployed, and not only in the base model?”
  • Temporal knowledge cutoff: The latest date represented in the information available to a model. “where temporal knowledge boundaries become experimental variables”
  • Thought experiment: A hypothetical scenario used to reason about an idea without directly carrying it out. “contact with another time is the thought experiment”
  • VLM-based recognition model: A vision-LLM used to interpret visual information and associate it with language. “they can hallucinate contemporary facts into historical text”
  • Welch’s tt-test: A version of the independent-samples tt-test that does not assume equal variances between groups. “Our primary outcome was ΔD100\Delta D_{100} for endorsement, compared between conditions with a Welch tt-test”
  • Within-participant counterpart: A comparison involving measurements taken from the same participant under different time points or conditions. “dzd_z its within-participant counterpart”

Open Problems

We haven't generated a list of open problems mentioned in this paper yet.

Tweets

Sign up for free to view the 2 tweets with 1663 likes about this paper.