Papers
Topics
Authors
Recent
Search
2000 character limit reached

Investigating the Effects of LLM Use on Critical Thinking Under Time Constraints: Access Timing and Time Availability

Published 9 Mar 2026 in cs.HC | (2603.08849v1)

Abstract: The impact of LLMs on critical thinking has provoked growing attention, yet this impact on actual performance may not be uniformly negative or positive. Particularly, the role of time -- the temporal context under which an LLM is provided -- remains overlooked. In a between-subjects experiment (n=393), we examined two types of time constraints for a critical thinking task requiring participants to make a reasoned decision for a real-world scenario based on diverse documents: (1) LLM access timing -- an LLM available only at the beginning (early), throughout (continuous), near the end (late), or not at all (no LLM), and (2) time availability -- insufficient or sufficient time for the task. We found a temporal reversal: LLM access from the start (early, continuous) improved performance under time pressure but impaired it with sufficient time, whereas beginning the task independently (late, no LLM) showed the opposite pattern. These findings demonstrate that time constraints fundamentally shape whether an LLM augments or undermines critical thinking, making time a central consideration when designing LLM support and evaluating human-AI collaboration in cognitive tasks.

Authors (3)

Summary

  • The paper finds a temporal reversal in a preregistered experiment with 393 participants: early or continuous LLM access improved essay performance under 10-minute constraints, while late or no access performed better with 30 minutes.
  • The paper shows that independent work before LLM access produced more unique arguments, stronger recall, and lower myside bias, whereas early access anchored deliberation and reduced engagement with source documents.
  • The paper demonstrates that self-reported critical-thinking measures barely detected these differences, highlighting the need for objective performance assessments and time-sensitive LLM designs that adapt assistance to task conditions.

Overview

This paper reports a preregistered 4×2 between-subjects experiment (n=393) examining how two temporal factors—LLM access timing (early, continuous, late, or none) and time availability (insufficient: 10 minutes; sufficient: 30 minutes)—shape critical thinking performance on an authentic performance-assessment task. The central result is a temporal reversal: LLM access from the start (early or continuous) improved essay-based critical thinking under time pressure but impaired it when time was sufficient, while beginning the task independently (late or no LLM access) showed the opposite pattern (2603.08849). The study moves beyond the correlational self-report literature on LLMs and critical thinking by measuring objective task performance in a controlled information environment.

Method

Participants completed a civic decision-making task from the iPAL (International Performance Assessment of Learning) framework: acting as city council members deciding whether to accept a company's water-contamination remediation proposal, based on seven documents varying in relevance, trustworthiness, and stance. The primary outcome was the Essay score—an arithmetic count of valid arguments, document references, and trustworthiness evaluations, minus penalties for fabrications and errors—supplemented by Recall (free-recall of document main ideas), Evaluation (rating document relevance/trustworthiness/stance), Comprehension (factual/counterfactual judgments), and Myside Bias (absolute difference between pro and con arguments). Essays were graded with an LLM grader validated against human scoring (ICC(2,1) = 0.70–0.82 for score components; 94.7% agreement, mean Cohen's κ = 0.76 for binary argument identification). Analyses used ANCOVA with Tukey HSD post-hoc tests, adjusting for LLM use frequency, attitudes, self-efficacy, and confidence in LLM capabilities.

Main results: a temporal reversal

For Essay score, ANCOVA revealed a significant interaction between access timing and time availability (F(3,381)=6.39F(3,381)=6.39, p<0.001p<0.001) and a main effect of time availability (F(1,381)=54.78F(1,381)=54.78, p<0.001p<0.001). Under insufficient time, early access yielded the highest Essay score (M=3.80M=3.80) and no-LLM access the lowest (M=1.86M=1.86), with significant post-hoc advantages for early over none (p<0.01p<0.01) and continuous over none (p<0.05p<0.05). Under sufficient time this reversed: late (M=5.77M=5.77) and no-LLM (M=5.76M=5.76) conditions outperformed continuous (p<0.001p<0.0010) and early (p<0.001p<0.0011), though these pairwise differences were reported as trends rather than individually significant contrasts.

The time-availability effect was equally striking and condition-dependent. Extra time produced large gains for participants who worked independently first—+3.91 points for late access and +3.01 points for no access (both p<0.001p<0.0012)—but only minimal gains for those with LLM access from the start (+1.22 points for continuous, p<0.001p<0.0013; non-significant for early). This implies that early LLM access prematurely constrains deliberation: the benefit typically assumed from additional time does not materialize when the model is present from the outset.

Recall showed a complementary pattern. Under sufficient time, early (p<0.001p<0.0014) and continuous (p<0.001p<0.0015) access impaired Recall relative to no access (p<0.001p<0.0016; p<0.001p<0.0017 and p<0.001p<0.0018 respectively), while moving from insufficient to sufficient time improved Recall substantially for late (+42%) and no-access (+32%) conditions but barely at all for early or continuous access. This suggests that having the LLM from the start prevents internalization of source material even when time permits deep engagement—a concrete cost beyond immediate task output.

Evaluation and Comprehension were largely insensitive to access timing, with sufficient time improving both. One exception: under sufficient time, continuous access impaired Evaluation of trustworthiness (p<0.001p<0.0019) versus no access (F(1,381)=54.78F(1,381)=54.780, F(1,381)=54.78F(1,381)=54.781), possibly reflecting reduced attention to source credibility when relying on the chatbot.

Myside bias and the value of late access

Under sufficient time, late access significantly reduced Myside Bias compared to no access (F(1,381)=54.78F(1,381)=54.782 vs. F(1,381)=54.78F(1,381)=54.783, F(1,381)=54.78F(1,381)=54.784) while maintaining comparable argument quantity (F(1,381)=54.78F(1,381)=54.785 vs. F(1,381)=54.78F(1,381)=54.786)—a genuine balancing effect rather than an artifact of producing fewer arguments. Essay-revision analysis showed that 15 of 48 late-access participants reduced their bias during the access window; notably, 8 who initially presented only con arguments added an average of 2.22 pro arguments drawn from LLM responses. Under insufficient time, however, late access arrived too late to help (only 0.68 arguments added on average), and its lower Myside Bias simply mirrored lower argument quantity. Conversely, prolonged independent work entrenched one-sided reasoning: no-access participants' Myside Bias rose 1.92 points from insufficient to sufficient time. Late access thus functions as a post-hoc check that counteracts the natural myside tendency of solo deliberation.

Behavioral mechanisms

Interaction logs and self-reports illuminate why early access anchors deliberation. Direct copying was rare, but textual overlap with LLM responses was common (64–94% of participants with access), indicating subtle influence rather than offloading. Argument-overlap analysis showed that participants with early or continuous access were exposed to more valid arguments in LLM responses yet developed almost no additional unique arguments when given more time (+0.28 and +0.07 from insufficient to sufficient), whereas late- and no-access participants developed +1.80 and +2.10 unique arguments. Document-viewing behavior corroborated this: early/continuous participants iterated far less with sources during writing (e.g., 2.42 vs. 3.63 unique documents during writing for continuous vs. late access under sufficient time). Task-approach coding revealed that even seemingly benign uses—summarization and clarification, not content adoption—were sufficient to anchor subsequent deliberation about which documents, ideas, and stances to pursue. Notably, 38% of late-access participants under time pressure reported minimal or no AI use, explaining why late access provided little benefit there.

Self-reports diverge from performance

The Critical Thinking Self-Assessment Scale showed almost no variation across LLM access timings, despite substantial performance differences. Only time availability produced small self-report effects (interpretation, analysis, total score). The authors argue this dissociation demonstrates that self-report measures lack the resolution to capture how LLMs affect cognition under different temporal contexts, and that authentic performance assessment should be prioritized in human-AI collaboration evaluation.

Design implications

The findings motivate temporally-aware LLM design. When time is plentiful, tools could encourage independent work first—via nudges ("what are your initial thoughts before I help?"), friction-based delays, or automatic detection of critical-thinking tasks followed by suggestions of guided modes such as ChatGPT's study mode. Under time pressure, where early access demonstrably helps, designs should mitigate anchoring through metacognitive prompts targeting source diversity and reasoning balance, multiple candidate outputs grounded in different source subsets, or visualized diversity/credibility metrics. More broadly, applications could ask users upfront how much time they have and adapt assistance accordingly, and organizations should weigh immediate performance against longer-term maintenance of cognitive capacities such as memory internalization.

Limitations and open questions

The authors concede several constraints. The single iPAL civic decision-making task may not generalize to domains where prior expertise matters (e.g., debugging, medical diagnosis); the equal-footing design deliberately minimizes background knowledge, so extension to expertise-varying settings remains untested. Crowdworker recruitment and fixed laboratory-style timing raise ecological-validity concerns relative to naturalistic workplaces where time pressure emerges organically. The operationalization of timing (first/final third of task time) is one point in a larger space that includes graduated pressure, intermittent access, and user-controlled timing; single-session designs also cannot capture interleaved or longitudinal LLM use. Statistically, several key sufficient-time contrasts (e.g., late vs. early Essay score) are trends rather than individually significant pairwise differences, and the modest effect sizes for Evaluation warrant caution. Open questions include whether the reversal holds for domain-specific tasks, whether users can be reliably classified as needing scaffolding versus preservation of independent reasoning, and what thresholds define "sufficient" versus "insufficient" time for a given task.

Conclusion

This experiment establishes that the effect of LLM use on critical thinking performance is contingent on temporal context rather than uniformly positive or negative. Early LLM access boosts performance under time pressure but impairs essays, recall, and deliberation when time suffices; late access after independent work best preserves and augments reasoning when time allows, and counteracts myside bias. The results argue that evaluations of human-AI collaboration should systematically manipulate—and report—time constraints, and that LLM support for cognitively demanding tasks should be designed around when, not merely whether, assistance is available.

Paper to Video (Beta)

No one has generated a video about this paper yet.

Whiteboard

No one has generated a whiteboard explanation for this paper yet.

Open Problems

We haven't generated a list of open problems mentioned in this paper yet.

Tweets

Sign up for free to view the 11 tweets with 30 likes about this paper.