- The paper presents a 13-week, tool-agnostic course that develops AI-assisted literature review skills through progressive comprehension, organization, innovation, and synthesis modules.
- The preliminary study found significant self-reported gains across key competencies, including AI attribution (Cohen’s d = 2.40), hallucination detection (d = 1.45), and output verification (d = 1.22).
- The course emphasizes verified engagement, process documentation, and AI as a cognitive scaffold, while acknowledging that its single-cohort, self-reported design cannot yet establish causal effectiveness or durable skill transfer.
The curricular gap the course addresses
This paper reports the design, theoretical rationale, and preliminary outcome evidence for BSTA 495/395: Getting Started with AI-Assisted Research, a 13-week, prerequisite-free elective developed and delivered at Lehigh University in Spring 2026. The author positions the course against three existing forms of AI-related instruction—technical AI development courses, brief general-literacy interventions, and platform-specific tool tutorials—and argues that none develops the integrated competency set required for rigorous AI-assisted scholarship. Technical courses serve a narrow specialist audience; short workshops cannot build sustained practice-based skills; and tool tutorials produce platform-specific proficiency that degrades as tools evolve. The specific gap targeted is AI-assisted literature review, a competency expected of nearly all graduate students but rarely taught explicitly.
The paper is framed as curriculum design research in the design-based research tradition: it documents a designed artifact, articulates the principles motivating its design, and reports implementation observations plus preliminary outcome data. The author is explicit that formal inferential analysis and a dedicated outcomes study are ongoing.
Theoretical framework
Four theoretical commitments structure the design. First, Vygotsky's zone of proximal development grounds the principle of verified engagement: students must be able to independently explain, justify, and defend every claim they submit, regardless of how AI contributed to producing it. This operationalizes scaffolding while guarding against what Wood, Bruner, and Ross termed scaffold dependency—the risk that AI substitutes for skill development rather than accelerating it.
Second, the course extends responsible reliance theory from human factors research into academic practice. Verification effort should be calibrated to empirical failure profiles: citation generation by general-purpose LLMs carries hallucination rates reported in the range of 30–50% for bibliographic details, warranting full individual verification, whereas structured extraction by index-backed platforms (Elicit, Semantic Scholar) permits spot-checking of roughly 25% stratified samples. Uniform verification effort across tools and tasks is presented as both inefficient and pedagogically counterproductive.
Third, the paper argues that "AI as cognitive scaffold rather than cognitive replacement" functions as a threshold concept in Meyer and Land's sense, and demonstrates that it satisfies all five diagnostic criteria: transformative, irreversible, integrative, bounded, and troublesome. The course is organized around progression through this threshold, with each module's escalating AI demands forcing confrontation with the distinction between AI-assisted and AI-replaced cognition.
Fourth, ethics is treated as subject rather than policy. Rather than restricting AI use through rules, the course teaches students to analyze the ethical dimensions of their own practices, on the argument that principled understanding transfers across contexts where policy compliance does not.
Course architecture
The course meets weekly for a single 2-hour-40-minute session divided into didactic instruction (~45 minutes), guided hands-on laboratory work (~75 minutes), and structured debrief (~20 minutes). It enrolls upper-level undergraduates (395) and doctoral students (495) simultaneously with shared instruction and differentiated assessment—a deliberate choice justified by the comparative learning generated when graduate students' domain expertise surfaces how AI failure modes vary by field.
Instruction is organized into four sequential modules mapping onto the cognitive arc of literature review:
| Module |
Focus |
Unit of analysis |
Culminating artifact |
| A |
Comprehension |
Individual papers |
Structured reading log with verified citations |
| B |
Organization |
Paper collections |
Validated multi-dimensional taxonomy |
| C |
Innovation |
Field-level gaps |
Gap analysis report with frontier map |
| D |
Synthesis |
Complete reviews |
Original literature review |
Two architectural principles deserve emphasis. The cumulative input design makes each module's primary output the next module's primary input, reinforcing literature review as an integrated process. And the design is explicitly tool-agnostic: five platforms (ChatGPT, Claude, Elicit, Consensus, Semantic Scholar) are used as vehicles for transferable competencies—prompt construction, output verification, bias detection, attribution documentation—rather than as objects of instruction in themselves.
Module-level design decisions
Module A establishes four AI roles in research (cognitive scaffold, efficiency multiplier, conceptual bridge, pattern detector) alongside their limitations, introduces prompt engineering through five operational principles, and presents a six-step strategic reading procedure in which AI-augmented steps are bracketed within independent engagement at start and end. Its most technically demanding element is methodology evaluation via a four-stage framework in which discrepancies between AI characterization and independent researcher judgment are treated as investigative signals rather than resolved automatically.
Module B scales to collections, introducing typological, dimensional, and hierarchical taxonomies and a five-criterion validation procedure anchored by inter-rater reliability (Cohen's κ on a 25% stratified sample). Requiring quantitative validation makes AI-assisted organizational work falsifiable. A notable contribution is the distinction between genuine and spurious convergence: AI tools group papers by surface similarity, and students learn to check whether ostensibly convergent studies operationalize the same construct with comparable populations and independent samples.
Module C addresses gap identification via a six-type gap taxonomy paired with a five-dimension importance rubric, explicitly separating identification from importance assessment because AI tools tend to assert importance without substantiation. Frontier mapping uses temporal citation network analysis to distinguish genuine research frontiers from keyword-driven publication surges. The module closes with a systematic hallucination typology (citation, statistical, methodological, theoretical, synthesis) and six systematic output biases, deliberately sequenced after students have accumulated firsthand experience with these phenomena.
Module D integrates prior competencies into complete literature review production, distinguishing AI-assisted from AI-authored writing. Five bounded GenAI roles are identified (data extraction, pattern detection, draft generation for revision, structural suggestion, consistency checking), and two practical tests govern appropriate use: the revision depth test and the deletion test. Authorial voice is framed as an epistemological competency, not a stylistic preference. Notably, the methods section must be written without AI assistance.
Assessment architecture
Grading comprises pre/post surveys (5%), weekly assignments (60%), peer evaluation and participation (10%), and the final literature review (25%)—a distribution that operationalizes the commitment that process matters as much as product. Every assignment requires complete AI use documentation (exact prompts, outputs, verification steps, discrepancies found), maintained longitudinally in an AI use log. A four-dimension rubric applies to all major work: substantive quality, verification quality, attribution completeness, and authorial mastery—the last assessed primarily through whether students can defend their submissions under questioning. Undergraduates produce 3,000–4,000-word reviews; graduate students produce 4,000–6,000-word reviews with more extensive gap analysis. A standardized attribution template records tool and version, task purpose, prompts, outputs, verification method, error rate, and corrections, doubling as preparation for emerging journal disclosure requirements.
Implementation observations
Four findings from the inaugural offering bear on the design's validity. First, the experience-before-taxonomy sequencing effect: placing systematic hallucination instruction after weeks of hands-on tool use converted abstract warnings into retrospective explanations of phenomena students had already encountered, producing deeper engagement than early placement would have. Second, cross-disciplinary enrollment surfaced domain-dependent variation in hallucination rates and failure modes that was not planned content but was subsequently incorporated into instruction. Third, AI use logs showed consistent prompt sophistication gains across the semester, associated with fewer reported verification errors—supporting distributed rather than concentrated presentation of prompt engineering content. Fourth, requiring AI-free methods sections proved unexpectedly diagnostic: students who had managed their own research process produced accurate accounts, while those who had delegated process management produced generic or inaccurate descriptions.
Preliminary outcome evidence
Twenty-seven students completed the pre-course survey and 29 the post-survey; Wilcoxon signed-rank tests were computed on 26 matched pairs. The cohort comprised 33% traditional undergraduates, 48% accelerated 4+1 students registered at the undergraduate level, and 19% doctoral students. Notably, 89% had used generative AI at least three times in the preceding six months, yet the academic platforms central to the course were largely unfamiliar—confirming a genuine competency gap. Despite 41% having conducted more than three prior literature reviews, 82% rated themselves only slightly or moderately familiar with AI ethics guidelines.
Mean confidence across eight items rose from 3.41 to 4.13 on a five-point scale. The largest, statistically robust gains aligned precisely with the course's design emphases:
| Competency |
ΔM |
Cohen's d |
p |
| AI attribution practice |
+2.00 |
+2.40 |
< .001 |
| Detecting AI hallucinations |
+1.20 |
+1.45 |
< .001 |
| Using AI tools responsibly |
+1.03 |
+1.33 |
< .001 |
| Verifying AI outputs |
+1.00 |
+1.22 |
< .001 |
| Organizing papers into a taxonomy |
+0.91 |
+1.23 |
< .001 |
| Identifying research gaps |
+0.72 |
+0.77 |
.028 |
| Evaluating methodology / recognizing limitations |
+0.59 / +0.52 |
+0.67 each |
n.s. |
| Reading complex papers / formulating questions |
+0.40 / +0.41 |
+0.51 each |
n.s. |
The near-floor-to-ceiling shift in attribution practice (d=+2.40) directly validates the ethics-as-subject infrastructure and standardized template. Three general reading competencies showed moderate, non-significant gains; the author attributes this plausibly to a ceiling-approach effect, since these items had the highest pre-course means, while acknowledging that some downward recalibration among overconfident entrants would itself be pedagogically desirable. Post-course learning outcome ratings were uniformly strong, with four of six competency items achieving 100% agree-or-strongly-agree endorsement and "AI as scaffold" receiving the highest mean (M=4.58). Eighty-eight percent reported being likely or very likely to apply the skills to future research. Qualitative feedback identified workload as the dominant improvement concern.
The author states plainly that these are self-reported perceptions from a single-cohort pre–post design without a control condition; maturation, regression to the mean, and general AI exposure during the semester cannot be ruled out, and confidence is a proxy for—not a measure of—actual competency.
Limitations and open questions
Several limitations constrain the claims. The evidence base rests on one instructor, one institution, and one cohort, with no comparison condition. The 13-week format likely underdevelops Module D's synthesis competencies, particularly for students with limited prior research experience; a two-semester sequence is proposed as a possible remedy. The course omits quantitative synthesis methods entirely—meta-analysis, systematic review protocols, PRISMA procedures—which limits applicability in fields where systematic review is standard. Weekly assignment length was the most consistent student complaint, and future iterations must reduce it without sacrificing the documentation requirements underpinning process assessment. Finally, the five instructional platforms represent a snapshot of a rapidly changing landscape; the tool-agnostic principle mitigates but does not eliminate the need for ongoing curation. Whether the course's instructional inversion—centering tool use with methodology as the evaluative standard—produces superior methodological knowledge acquisition compared to traditional methods courses remains an open empirical question the ongoing outcome analysis will only partially address.
Conclusion
BSTA 495/395 demonstrates that AI research literacy can be taught as a coherent, discipline-agnostic competency within a single semester to mixed-level cohorts, organized around the authentic cognitive demands of literature review rather than AI technology itself. Five transferable design principles emerge: design for transfer rather than specific tools; organize instruction around the research workflow; sequence failure-mode instruction after accumulated tool experience; treat ethics as subject rather than policy; and assess process, not only product. The preliminary self-reported outcomes are consistent with the design's theoretical intentions, with the largest gains occurring exactly where the design concentrated its effort—but establishing causal effectiveness and durable transfer requires the controlled, multi-cohort study the author identifies as ongoing.