Time Machine Experiments: Using Historically-Bounded AI for Inquiry into the Human Mind
Abstract: Can interacting with someone from 1930, with no knowledge of what happened after, influence a person's perception of the past? People reason about the present against a picture of the past without observing it. The past is reconstructed from memory and testimony, but this reconstruction has been filtered through everything that happened since. Historically-bounded LLMs make that past available for interaction. As a proof-of-concept for the impact of interacting with historical minds, we ran a preregistered randomized experiment (), where participants interacted with an LLM trained on pre-1930 text. The interaction reduced the illusion of moral decline, the tendency to view the past as more moral than the present, compared to the contemporary-model control. This Time Machine Experiment paradigm informs new forms of interactive experiments, where temporal knowledge boundaries become experimental variables, and expands the realm of science fiction science, which turns thought experiments into actual experiments.
Paper Prompts
Sign up for free to create and run prompts on this paper.
Top Community Prompts
Explain it Like I'm 14
论文主题
这篇论文研究了一种像“时间机器”一样的人工智能实验。研究人员想知道:如果今天的人能和一个只知道 1930 年以前信息的 AI 聊天,他们对过去的看法会不会改变?
这个 AI 不是假装自己生活在过去,而是使用主要来自 1930 年以前的书籍、报纸和其他文字训练。因此,它不会像现代 AI 那样知道后来发生的所有事情。
研究人员特别关注一个常见想法:人们常常认为过去的人比今天的人更善良、更诚实、更有道德。 这被称为“道德衰退错觉”。
研究人员想回答什么问题?
论文主要想回答以下问题:
- 和一个只了解过去的 AI 交谈,是否会改变人们对过去的看法?
- 人们是否会因此减少“过去比现在更有道德”的想法?
- 这种 AI 能不能成为一种新的科学研究工具,帮助我们研究人类如何理解历史和时间?
研究人员并不是想证明这个 AI 就代表了真实的历史人物或所有过去的人。它更像是一个根据历史文字制作出来的“实验伙伴”。
研究方法:他们是怎样做实验的?
研究人员招募了 240 名来自美国和英国的英语使用者。参与者被随机分成两组:
- 一组和只使用 1930 年以前资料的 AI 交谈,称为“1930-AI”。
- 另一组和现代 AI 交谈,作为比较对象。
随机分组很重要,因为这样两组人在开始时大致是公平的。最后出现的不同结果,就更可能是由 AI 的不同造成的,而不是由参与者本身造成的。
实验过程
参与者先回答问题,例如:
你认为 100 年前的人有多重视善良、诚实和做好人?
他们要分别评价今天的人、约 20 年前的人和约 100 年前的人。
接着,他们和自己的 AI 进行互动,包括:
- 先进行一小段自由聊天;
- 阅读一些与道德和社会观念有关的句子;
- 猜测自己的 AI 会如何评价这些句子;
- 猜测 AI 会给出什么理由;
- 根据 AI 的回答修改自己的猜测;
- 最后再次和 AI 自由聊天。
例如,参与者可能需要完成这样的句子:
“如果丈夫已经能够养家,已婚女性去工作是_,因为_。”
参与者要预测 AI 会认为这件事是好、坏,还是不好说,并猜测 AI 会怎样解释。
这项任务有点像一个“猜 AI 想法的游戏”。通过这个过程,参与者会更认真地了解 AI 的观点,而不只是随便读几句话。
互动结束后,参与者再次回答关于过去和现在道德水平的问题。研究人员比较他们前后的答案。
研究结果
人们原本确实认为过去更有道德
在实验开始时,两组参与者都认为:
- 100 年前的人比今天的人更重视善良、诚实和做好人;
- 过去的人在行为上也更符合这些道德标准。
这说明“过去更有道德”的想法确实存在于参与者中。
和历史 AI 聊天后,这种看法明显减弱
与 1930-AI 互动后,参与者对 100 年前人的道德评价明显下降。
简单来说,他们不再那么相信:
“过去的人一定比现在的人更善良、更诚实。”
在 1930-AI 组中,参与者认为过去更有道德的差距几乎消失了,甚至出现了轻微反转。
相比之下,和现代 AI 聊天的人,想法变化很小。
重要的是,参与者对今天人的评价并没有明显提高。变化主要来自于他们降低了对过去的理想化评价。
历史 AI 也让人们更多反思
和 1930-AI 交谈的人更常表示:
- 自己重新思考了原来的观点;
- 自己改变了对过去和现在的看法;
- 自己接触到了以前没有考虑过的角度。
这说明历史 AI 不只是提供答案,也可能促使人们重新检查自己的假设。
为什么这些结果重要?
这些结果说明,人们对过去的印象并不一定完全来自真实历史。我们常常会把过去想象得更简单、更善良,忘记过去也存在不公平、歧视和其他问题。
与一个“不了解后来历史”的 AI 对话,可能帮助人们发现:
- 过去并不一定像记忆中那么美好;
- 人们对历史的判断可能受到今天生活经验的影响;
- 所谓“社会越来越糟”有时可能只是我们的错觉。
这项研究还展示了一种新的实验方法。过去,研究人员只能阅读历史资料,或者询问今天仍然活着的人。但历史资料不能回答新的问题,而现代人又已经知道后来发生的事情,可能会受到“事后之明”的影响。
历史 AI 提供了第三种可能:它可以对问题作出回应,同时尽量保持在某个历史时期的信息范围内。
这项研究的局限
研究人员也强调,这种 AI 并不等于真正的历史人物,也不能代表 1930 年以前所有人的思想。它学习的资料主要是保存下来的文字,而这些文字可能:
- 只代表一部分人;
- 更常来自有教育或有权力的人;
- 忽略了许多普通人和弱势群体的声音。
此外,两个 AI 不只在“知道哪些年代的信息”上不同。它们还可能在语言风格、回答能力和回答长度上不同。因此,研究结果不能完全证明“时间限制本身”造成了变化,只能说明:与这种历史 AI 进行整体互动,会影响人们的看法。
研究规模也比较小,参与者只来自美国和英国,实验时间大约 22 分钟。因此,还需要在更多国家、更多年龄群体和更长时间内重复研究。
研究的影响和意义
这篇论文提出了“时间机器实验”这一新方法。它的核心想法是:
让 AI 的知识停留在某个历史时间点,再观察今天的人如何与它互动,以及这种互动如何改变人的想法。
未来,这种方法可能被用于研究:
- 人们如何理解不同历史时期;
- 我们为什么会美化过去;
- 和历史观点互动是否能减少偏见;
- 人们如何看待未来或不同文化;
- AI 是否能帮助人们进行更深入的反思。
总的来说,这项研究表明,AI 不只是一个回答问题的工具,也可以成为研究人类思想的“实验仪器”。不过,我们必须小心使用它,清楚说明 AI 代表什么、不代表什么,并检查它是否真的遵守了设定的历史时间范围。
Knowledge Gaps
Knowledge gaps, limitations, and open questions
- Temporal integrity of the deployed system remains incompletely established. The paper reports that Talkie was trained on pre-1930 text, but it does not fully demonstrate that the deployed interface, moderation layer, system prompts, post-training, filtering, or other components introduced no post-1930 information.
- The historical model cannot be treated as a representative historical person or population. The study does not determine whose perspectives are overrepresented or absent in the corpus by gender, class, race, geography, literacy, occupation, or political position.
- Corpus-selection bias is not quantified. The effects of OCR errors, preservation practices, publication inequalities, English-language coverage, and the overrepresentation of elite or institutional writing on model outputs remain unclear.
- The causal effect of temporal boundedness is not isolated from model differences. Talkie and the contemporary model likely differ in parameter count, fluency, coherence, refusal behavior, verbosity, training objectives, and general capability, so the observed effect may reflect the historically bounded condition as a package rather than the cutoff itself.
- The contemporary control is not demonstrably matched to the historical model on interaction quality. Differences in perceived credibility, novelty, conversational naturalness, difficulty, or engagement could independently explain the treatment effect.
- The study does not test alternative control conditions that would separate historical content from presentation effects. For example, it does not compare Talkie with a contemporary model stylistically matched to Talkie, a scripted historical text, a prompted historical persona, or a modern model given equivalent information about historical moral norms.
- The historical model’s actual knowledge boundary is tested only to a limited extent. A systematic, adversarial audit across direct and indirect post-1930 queries is needed to establish whether the model can reveal later information through associations, extrapolation, or memorized contamination.
- The intervention exposes participants to only a small and curated set of moral prompts. Four main prediction items cannot establish whether the effect generalizes across moral foundations, political issues, cultures, historical periods, or less explicitly evaluative topics.
- The item-construction process may have shaped the findings. Reformulating historical survey questions into sentence completions with required justifications may alter their meaning, introduce present-day assumptions, or make participants respond to linguistic artifacts rather than historical perspectives.
- Reference completions may not represent the models’ natural behavior. Selecting one pre-generated response per item after quality screening and prioritizing the modal stance can suppress output variability and create fixed targets that participants learn rather than encounter naturally.
- The LLM-based scoring procedure remains a potential source of measurement error. Although calibrated against three human raters, the paper does not establish whether scoring is unbiased across conditions, moral positions, linguistic styles, or reasoning quality.
- The prediction task may confound learning, compliance, and belief change. Repeated attempts, feedback, performance bonuses, and the requirement to infer the model’s answers could cause participants to optimize for task performance rather than genuinely revise their beliefs about moral change.
- Demand characteristics and framing effects are not fully ruled out. Participants learned that their model was historically bounded and were explicitly asked about changes in their views, which may have encouraged hypothesis-consistent responses.
- The study cannot determine whether participants changed beliefs about historical morality or merely lowered ratings of the past. The primary effect arose mainly from reduced ratings for people 100 years ago, and the design does not establish whether this reflects improved historical accuracy, generalized skepticism, negative reactions to the model, or assimilation to its responses.
- The accuracy of participants’ revised beliefs is not validated. The study does not compare post-intervention judgments with independent historical evidence or assess whether the intervention reduces bias rather than introducing a different distortion.
- The measures rely heavily on self-report and immediate post-interaction judgments. It remains unknown whether the observed changes persist over time or affect political attitudes, historical reasoning, behavior, or real-world decisions.
- The duration and intensity of exposure are limited. A median session of approximately 22 minutes does not show whether repeated, longer-term, or self-directed interactions produce durable, cumulative, or counterproductive effects.
- Potential harms of historical interaction are not examined. Future work should test whether historically bounded models normalize discriminatory beliefs, reproduce archival prejudice, increase nostalgia, or influence vulnerable users differently.
- The sample limits generalizability. Participants were English-speaking residents of the United States or United Kingdom recruited through Prolific; effects may differ by age, education, political ideology, cultural background, historical knowledge, and familiarity with AI.
- Cross-cultural and multilingual validity is unresolved. The paradigm is based on an English-language, largely Western corpus and does not establish whether similar effects occur for other languages, regions, historical traditions, or culturally distinct moral frameworks.
- Individual differences in susceptibility are not fully explained. The paper does not identify whether effects vary according to participants’ baseline nostalgia, trust in AI, perceived model competence, prior beliefs about moral decline, religiosity, political orientation, or historical literacy.
- The contribution of free-form conversation is unclear. Because participants could ask arbitrary questions before and after the structured task, the study cannot determine which interaction content, conversational strategy, or participant-generated question caused the observed change.
- The mechanism linking interaction to reduced moral decline is not established. Reflective insight is measured, but the study does not test whether reflection mediates the effect, whether prediction accuracy matters, or whether exposure to disagreement, uncertainty, or model inconsistency drives belief revision.
- The qualitative analysis is exploratory and potentially subjective. One author coded the responses without a formal coding scheme or inter-rater reliability, limiting confidence in claims about recurring conversational themes.
- The study does not compare different historical cutoff dates. It remains unknown whether effects vary smoothly with the cutoff, depend on specific historical periods, or result specifically from the contrast between pre-1930 and contemporary knowledge.
- The paradigm’s applicability beyond text dialogue is untested. The paper proposes speech, images, audiovisual environments, and embodied agents, but provides no evidence that the validity requirements or psychological effects transfer to these modalities.
- The reproducibility of the findings may be sensitive to model and interface updates. Hosted models, moderation systems, prompts, and APIs can change over time, yet the study does not establish whether the effect survives exact replication with version-locked systems.
- The boundaries of permissible inference from Time Machine Experiments require further empirical validation. The paper proposes four validity requirements, but does not show how failures in historical grounding, temporal integrity, interaction specification, or comparative attribution quantitatively alter conclusions.
Practical Applications
Immediate Applications
The paper’s findings support applications that use historically bounded AI as an interactive research instrument or reflective intervention, rather than as an authoritative reconstruction of historical people or beliefs.
- Behavioral-science experiments on historical perspective and bias — Academia / psychology / HCI Researchers can deploy historically bounded LLMs in randomized studies to examine how people form beliefs about the past, respond to historical viewpoints, and revise judgments after interactive exposure. The paper provides a reusable workflow: establish a temporal cutoff, assign participants to historical and contemporary model conditions, standardize core prompts, log interactions, and measure pre/post changes in beliefs or reflection. Dependencies: The model’s corpus, cutoff, prompts, interface, and moderation components must be documented. Effects should be attributed to the complete model package—not automatically to temporal boundedness alone.
- Interventions addressing the “illusion of moral decline” — Education / civic organizations / public communication Museums, schools, libraries, and civic-education programs could offer guided dialogues with a pre-1930 model to help users question idealized narratives about the past. The experiment found that interaction with the historical model reduced participants’ perception that people 100 years ago were morally superior to people today. Potential product: A museum or classroom module in which students predict how a historical AI will evaluate social situations, receive feedback, and discuss why their expectations differed. Dependencies: The intervention should be framed as perspective-taking and critical reflection, not as proof of what historical populations believed. It requires age-appropriate content controls, historical contextualization, and human facilitation for sensitive topics.
- Interactive historical-literacy tools — Education / museums / libraries Institutions can supplement archival materials with conversational interfaces that answer user-generated questions from a bounded historical corpus. Such systems could help learners explore how concepts, arguments, and social norms were represented in documents before a specified date. Potential workflow:
- User asks a question about a historical period.
- The system responds using the bounded corpus.
- The interface displays source provenance, uncertainty, and links to original documents.
- The learner compares the response with contemporary evidence and scholarly interpretation. Dependencies: Retrieval or citation mechanisms, corpus provenance, OCR quality, and clear disclosure that generated responses are model-mediated interpretations.
- Research on presentism and hindsight bias — History / social science / law Scholars can use bounded models to study how contemporary users project current knowledge and moral standards onto earlier periods. Participants might be asked to predict a historical model’s responses, challenge it with present-day concepts, or compare it with a modern model. Potential application: Experiments on hindsight bias, historical empathy, political nostalgia, attitudes toward social change, or judgments of past legal decisions. Dependencies: Carefully matched control conditions are necessary because historical and contemporary models may differ in capability, tone, verbosity, refusal behavior, and cultural representation.
- Design and evaluation of temporal-perspective interfaces — Software / HCI / product design Product teams can use the paper’s structured prediction task to evaluate interfaces intended to promote reflection. Rather than relying only on open-ended conversation, a system can ask users to predict an AI’s response, provide feedback, and measure changes in understanding. Potential tools: Reflection dashboards, interactive timelines, “compare perspectives across eras” interfaces, and conversation logs for qualitative analysis. Dependencies: Predefined tasks reduce uncontrolled variation, but they may also constrain natural interaction. Designers should separately test structured and free-form modes.
- Museum and archival exploration workflows — Cultural heritage / libraries / public history Archives can use bounded AI as a conversational layer over digitized newspapers, books, patents, scientific journals, or case law. Users could ask questions that are not answered verbatim in any single document while the system identifies the relevant source period and uncertainty. Dependencies: The corpus will reflect preservation and digitization biases. The system should not be presented as representing an entire society, demographic group, or historical “mind.”
- Policy communication about social progress and decline — Government / public policy / civic communication Policymakers and communicators can use historically grounded interactive exhibits to test whether public beliefs about social decline are based on evidence or nostalgia. This may improve public interpretation of long-term survey data and reduce support for policies justified solely by an idealized past. Dependencies: The tool should not be used to advocate a predetermined political position. Claims should be triangulated with historical surveys, administrative data, and expert review.
- Internal training for researchers using historical AI — Academia / research governance Universities and research teams can immediately adopt the paper’s four-part validity checklist when developing historical-model studies: historical grounding, temporal integrity, interaction specification, and comparative attribution. This can become a protocol for preregistration, model cards, ethics review, and replication packages. Dependencies: Teams need access to model-training information, deployment logs, prompt specifications, and tests for post-cutoff leakage.
- Public-facing critical-thinking exercises — Daily life / media literacy Individuals could use a historically bounded system to challenge assumptions such as “people were more respectful in the past” or “society has continuously declined.” A guided comparison between archival evidence, the bounded model, and a contemporary model could encourage users to distinguish memory, nostalgia, and evidence. Dependencies: Unguided use may reinforce stereotypes or generate fabricated historical claims. Source links and warnings about model limitations are essential.
Long-Term Applications
The following possibilities require additional validation, larger-scale deployment, improved model control, or research beyond the paper’s proof of concept.
- Cross-temporal experiments on political attitudes and polarization — Political science / policy Future studies could test whether interacting with models bounded before major political transformations changes attitudes toward nationalism, democracy, gender roles, immigration, reproductive rights, colonialism, or economic policy. Multiple cutoff dates could distinguish the effects of different information environments. Dependencies: Such studies require representative samples, culturally diverse corpora, carefully neutral framing, and safeguards against normalizing discriminatory historical views. A single historical model cannot establish how a population actually thought.
- Longitudinal interventions for nostalgia and political decision-making — Mental health / civic policy / communications Repeated, evidence-linked interactions might help users evaluate “return to the past” narratives before voting or making policy judgments. Researchers could measure whether effects persist beyond the immediate post-interaction survey and whether they influence actual information-seeking or political choices. Dependencies: The current evidence measures short-term self-reported belief change. Persistence, behavioral transfer, and possible reactance remain untested.
- Historically bounded AI tutors for education — Education Future systems could provide period-specific tutors for history, literature, science, law, and economics. Students might debate with an AI limited to the knowledge available in 1789, 1850, or 1929, then compare its reasoning with later developments. This could teach scientific paradigm shifts, historical contingency, and the difference between contemporary knowledge and period knowledge. Dependencies: Models need reliable uncertainty expression, curriculum alignment, multilingual support, citation, teacher oversight, and resistance to anachronistic leakage.
- Simulation of historical decision environments — Business / policy / military history / organizational research Bounded models could support counterfactual experiments in which participants make decisions using only information available at a historical point—for example, evaluating an investment, scientific hypothesis, public-health measure, or political crisis. Researchers could study how hindsight and outcome knowledge affect judgment. Dependencies: A LLM’s corpus is not equivalent to the information available to a specific decision-maker. The system would require carefully reconstructed information sets, role-specific access controls, and validation against primary documents.
- Historical legal and regulatory analysis — Law / compliance / public administration Models bounded at particular dates could help researchers examine how legal reasoning and regulatory language changed over time, or simulate how a historical legal corpus would frame a contemporary issue. This could support comparative legal scholarship and training exercises. Dependencies: Legal applications require authoritative source retrieval, jurisdictional controls, citation-level verification, and strong disclaimers. Generated output must not be treated as legal advice or as evidence of historical legal consensus.
- Embodied and multimodal “time machine” environments — VR / robotics / cultural heritage The paradigm could extend from text chat to speech, visual scenes, avatars, virtual neighborhoods, or embodied agents whose available information is bounded by a historical date. Users might explore how a period-specific agent interprets objects, social roles, or emerging technologies. Dependencies: Every component—including image models, speech systems, retrieval tools, safety filters, and environment assets—must be checked for post-cutoff information. Multimodal systems also increase risks of fabricated details and stereotypical representation.
- Historical perspective-taking for conflict resolution — Peacebuilding / diplomacy / social psychology Carefully designed systems could expose users to competing historical narratives while preserving temporal boundaries, helping them identify how later events reshape interpretations of earlier conflicts. Researchers could test effects on empathy, dehumanization, and willingness to engage with opposing groups. Dependencies: This is high-risk: a model may reproduce propaganda, archival silences, or dominant-group perspectives. Deployment would require local experts, plural corpora, transparent provenance, and facilitated dialogue rather than autonomous persuasion.
- AI-assisted discovery of historically plausible hypotheses — Humanities / social science / science and technology studies Researchers could compare outputs from models trained to different cutoff dates to identify concepts, inventions, or arguments that were plausible before they emerged historically. This could generate hypotheses about technological forecasting, cultural diffusion, and the development of social norms. Dependencies: Model outputs are hypothesis-generating, not historical evidence. Results require archival verification, controls for corpus size and genre imbalance, and methods for distinguishing genuine historical patterns from training artifacts.
- Personalized reflection and anti-nostalgia applications — Mental health / coaching / daily life A future consumer product could help users examine autobiographical or intergenerational nostalgia by comparing personal memories with period-specific documents and survey data. It might prompt questions such as whether a remembered past was broadly experienced or selectively remembered. Dependencies: Such systems would handle sensitive personal and political beliefs. Privacy, data minimization, non-manipulative design, and clinical validation would be required before use in mental-health contexts.
- Benchmarking temporal integrity and historical AI safety — AI engineering / governance The paper’s validity requirements could become a standard evaluation suite for “historically bounded” models. Benchmarks could probe whether systems reveal post-cutoff facts through direct questions, indirect causal reasoning, prompts, retrieval, safety layers, or multimodal inputs. Potential outputs: Temporal-integrity certificates, model cards specifying cutoff boundaries, leakage test suites, audit logs, and deployment standards for research and education. Dependencies: No finite probe can prove the complete absence of leakage. Certification would need layered audits, reproducible training data documentation, adversarial testing, and explicit limits on the claims made.
- Large-scale studies of cultural change using multiple historical models — Sociology / anthropology / economics Researchers could compare participant responses across models bounded at 1800, 1900, 1930, 1950, and later dates to map how perceived norms and judgments shift as information environments change. This could produce richer models of cultural evolution and collective memory. Dependencies: Differences between models may reflect architecture, training scale, language, corpus availability, or alignment—not only historical time. Large preregistered studies and matched model families would be necessary.
- Policy evaluation of nostalgia-driven interventions — Government / political communication Governments or civil-society organizations could experimentally test whether evidence-based historical interactions reduce support for policies justified by claims of moral or social decline. Such work might inform public-information campaigns and deliberative-democracy tools. Dependencies: Manipulating political beliefs raises ethical and democratic concerns. Studies should prioritize informed consent, neutrality, transparency, independent oversight, and measurement of unintended effects such as distrust or polarization.
Glossary
- Archival materials: Historical records preserved for later study, such as documents, newspapers, or other primary sources. “Archival materials provide historically situated evidence, but they are fixed and reflect selective processes of production and preservation”
- Causal attribution: The process of determining whether an observed effect was caused by a particular intervention or experimental condition. “the comparison condition must support the causal attribution drawn from it”
- Contemporary frontier model: A highly capable, state-of-the-art AI model available at the present time. “interacting with a current frontier model: gpt-5.5”
- Corpus: A structured collection of texts used for linguistic analysis or model training. “13 billion parameters pretrained on 260 billion tokens of English-language text published before 1930”
- Counterfactual: Describing a hypothetical situation that differs from what actually occurred. “design fiction and speculative design, which construct alternative or counterfactual technological scenarios”
- Degrees of freedom: The number of independent values available to vary when estimating a statistical quantity. “a statistic with degrees of freedom in parentheses”
- Embodied agent: An artificial agent represented as having a physical or virtual body through which it interacts with users or environments. “text, speech, images, audiovisual environments, embodied agents, or combinations of these”
- Empirical studies in HCI: Research based on observed or measured human behavior in human–computer interaction settings. “Human-centered computing~Empirical studies in HCI”
- Experimental vignette: A standardized, hypothetical scenario presented to study participants as part of an experiment. “Experimental vignettes provide standardized and manipulable scenarios”
- External validity: The extent to which findings generalize beyond the particular participants, setting, or conditions of a study. “Although these are not new kinds of validity”
- Hindsight bias: The tendency to see past events as having been more predictable after their outcomes are known. “their recollections and retrospective judgments may be shaped by subsequent experiences and outcome knowledge”
- Historically-bounded LLM: A LLM trained on data available only before a specified historical cutoff. “There now exist LLMs trained exclusively on historical text corpora”
- Human–computer interaction (HCI): The interdisciplinary study and design of interactions between people and computing systems. “Human-computer interaction (HCI) research has contributed to this advancement”
- Illusion of moral decline: The belief that people in the past were more moral than people in the present, despite evidence that does not support this conclusion. “the `Illusion of Moral Decline', which describes peopleâs tendency to perceive past societies as more moral than contemporary ones”
- Information environment: The body of information available to a system and the constraints governing that availability. “an information environment intentionally bounded in time”
- Inter-rater reliability: The degree to which multiple evaluators produce consistent judgments when assessing the same material. “No inter-rater reliability was computed”
- LLM: A machine-learning model trained on large amounts of text to generate and interpret language. “LLMs offer HCI another such methodological opportunity”
- Manipulation: An experimentally controlled change introduced to test its effect on participants’ responses or behavior. “an ordinary manipulation has no informational boundary that can leak”
- Methodological apparatus: A tool, system, or procedure used to conduct research or investigate a phenomenon. “Interactive systems have long served in HCI and adjacent behavioral sciences both as objects of evaluation and as methodological apparatuses”
- Moral compliance: The extent to which people’s behavior conforms to their own stated moral values. “The pattern held for moral compliance, the second measure”
- Ordinary least squares regression: A statistical method that estimates relationships by minimizing the squared differences between observed and predicted values. “we fit an ordinary least squares regression with the change in perceived moral decline as the outcome”
- Optical character recognition (OCR): Technology that converts text in images or scanned documents into machine-readable characters. “instead involved conventional OCR”
- Post-training: Training performed after a model’s initial pretraining to adapt its behavior, capabilities, or interaction style. “Talkie is also notable for combining relatively large-scale historical pretraining with post-training specifically designed to support conversational interaction”
- Preregistered randomized experiment: An experiment whose design and analysis plan are specified in advance and whose participants are assigned randomly to conditions. “we conducted a preregistered randomized experiment ()”
- Proof of concept: An initial demonstration that a proposed method or system is feasible and potentially effective. “To demonstrate the potential of the Time Machine Experiment as an instrument for scientific inquiry, we conducted a proof-of-concept experiment”
- Qualitative data: Non-numerical information, such as written responses or interview transcripts, analyzed for meaning or recurring themes. “We used these qualitative data to contextualize the quantitative findings”
- Randomized controlled comparison: An experimental comparison in which participants are assigned by chance to different conditions to estimate causal effects. “Participants were randomly assigned to either of the two conditions, which identifies the causal effect of the model on changes in participant evaluations”
- Reflective insight: A measured change in perspective or understanding resulting from reconsidering one’s assumptions, values, or experiences. “We measured reflective insight using three items adapted from the Insight subscale”
- Speculative design: A design practice that creates hypothetical technologies or situations to provoke reflection about social assumptions and possible futures. “Our conceptualization of the Time Machine Experiment is situated across several HCI traditions, such as design fiction and speculative design”
- Temporal boundary: A cutoff that limits the time period from which a system may obtain information. “The model was accessed through a moderated interface and presented to participants under the name 1930-AI”
- Temporal generalization: The ability of a model or method to perform across different time periods or historical conditions. “This development has attracted interest not only for the evaluation of temporal generalization and forecasting”
- Temporal integrity: The extent to which a system actually preserves its stated historical information cutoff after deployment. “Temporal integrity. Does the boundary hold in the system as deployed, and not only in the base model?”
- Temporal knowledge cutoff: The latest date represented in the information available to a model. “where temporal knowledge boundaries become experimental variables”
- Thought experiment: A hypothetical scenario used to reason about an idea without directly carrying it out. “contact with another time is the thought experiment”
- VLM-based recognition model: A vision-LLM used to interpret visual information and associate it with language. “they can hallucinate contemporary facts into historical text”
- Welch’s -test: A version of the independent-samples -test that does not assume equal variances between groups. “Our primary outcome was for endorsement, compared between conditions with a Welch -test”
- Within-participant counterpart: A comparison involving measurements taken from the same participant under different time points or conditions. “ its within-participant counterpart”







