Academic Miscommunication: Causes & Remedies
- Academic Miscommunication is a breakdown in the clear transmission of scholarly meaning arising from ambiguous norms, compressed signals, and mismatches between intent and interpretation.
- It spans contexts from classroom misunderstandings and misaligned exam expectations to unclear AI usage policies, publication errors, and notation inconsistencies.
- Mitigation strategies focus on explicit communication, aligned assessment methods, and systematic verification to reduce ambiguity and enhance reproducibility.
Academic Miscommunication (AM) denotes breakdowns in how academic meaning is produced, transmitted, interpreted, or verified. In the literature, the term encompasses mismatches between instructor intent and student interpretation, loss of expert meaning in public science communication, ambiguity introduced by notation or verbal uncertainty expressions, submission conventions that render a simple methodological contribution unintelligible, and AI-mediated outputs whose fabricated citations or unsupported claims defeat routine checks (Scharlach, 29 Aug 2025, Cole, 2020, Galati, 2019, Oliver et al., 2022, Ansari, 5 Feb 2026). Rather than naming a single pathology, AM identifies a family of failure modes that emerge wherever scholarly communication depends on compressed signals, tacit norms, weakly specified rules, or superficial verification.
1. Conceptual range and domain-specific definitions
In classroom research, Academic Miscommunication is defined as the phenomenon in which “a student interprets the professor’s words, assignments, or other aspects of the curriculum in a manner that deviates from what the professor had intended.” In the physics study by Scharlach, the specific focus is mismatch between what students expected an exam to look like, in both style and content, and what the exam actually was (Scharlach, 29 Aug 2025). In software engineering education, de Souza Santos et al. characterize miscommunication less as misunderstanding of a single utterance than as ambiguity in the rule environment: course policies or rubrics may be vague, internally inconsistent, or silent about AI/LLM use, creating a “grey area” in which students negotiate what counts as allowed assistance (Santos et al., 17 Mar 2026).
In science communication, AM is framed as a translation failure between expert production and public uptake. The COVID-19 infodemic paper defines it as occurring when expert findings are lost in translation between researcher and public, leading to confusion, sensational headlines, or unintentional spread of false understandings. That scope explicitly includes jargon, over-rapid publication such as preprints, lack of plain-language summaries, and platform mismatches when academics address journalists or social media without adapting their communication (Cole, 2020). A basic distinction in this setting is between misinformation, defined as false information shared in the honest belief it is true, and disinformation, defined as deliberately false or misleading information spread maliciously; AM can participate in either process without being identical to either one (Cole, 2020).
Other literatures use the term or closely related formulations to describe failures internal to scholarship itself. In HCI, Oliver et al. argue that a “simple concept that has the potential for a methodological contribution to the field of HCI” can become “impossible to communicate…in a manner that is intelligible to the reader” because of submission conventions (Oliver et al., 2022). In statistical methodology for incomplete data, Galati identifies notational and framing choices that have “been present in the literature for a long time” and impede communication of core concepts such as observed versus missing components, Missing At Random (MAR), and ignorability (Galati, 2019). In uncertainty communication, Willems et al. show that verbal probability phrases have highly variable numerical interpretations, making them a direct risk for miscommunication even among statisticians (Willems et al., 2019).
This range of usage indicates that AM is not restricted to inaccurate content. It also includes accurate content that is underspecified, misframed, context-misaligned, linguistically unstable, or embedded in workflows that reward plausibility over verification.
2. Instructional alignment, assessment, and academic integrity
The most explicit quantitative operationalization of AM appears in the physics study. On a post-exam survey, students rated two statements indicating miscommunication and six statements indicating accurate communication, all on 1–5 Likert scales. The Academic Miscommunication score was defined as
A higher indicates a larger deviation between student expectations and the actual exam. The study compared two classes, Physics 42 for non-physics majors and Physics 71 for physics majors, and used linear regressions with Bonferroni correction for 20 tests, giving . In Physics 42, Exam Score vs. Academic Miscommunication yielded , , , and the interpretation reported in the paper is that students with greater miscommunication feelings performed worse. In Physics 71, none of the ten correlations met the Bonferroni threshold; for Exam vs. , , , (Scharlach, 29 Aug 2025).
The same broad phenomenon appears in software engineering education, but now around LLM use rather than exam format. In a cross-sectional survey of 116 undergraduate software engineering students, reported LLM cheating practices occurred primarily in programming assignments or labs, 82 (71%); regular classwork or weekly exercises, 53 (46%); software design or documentation tasks, 48 (41%); essays or written reflections, 40 (34%); group or capstone projects, 38 (33%); quizzes or short online tests, 34 (29%); and major exams or finals, 17 (15%). The top enabling factors were heavy workload or overlapping deadlines, 90 (78%); assignments completed remotely or online, 70 (60%); limited oversight or instructor checking, 43 (37%); unclear or missing rules about AI use, 39 (34%); belief that AI use would not be detected, 39 (34%); and peer norms encouraging similar behavior, 35 (30%). Only 43% of students had received any formal guidance or training on responsible AI use (Santos et al., 17 Mar 2026). The paper’s conclusion is not that cheating is reducible to student disposition, but that misuse is associated with assessment and instructional conditions and requires clearer alignment between assessment design, learning objectives, and expectations for LLM use (Santos et al., 17 Mar 2026).
Payne et al. extend this instructional picture by showing that even when “same rules” nominally govern a course, instructors, TAs, and students classify cheating scenarios differently. Across 13 scenarios, instructors averaged 37.2% “Not Cheating,” 18.0% “Trivial Cheating,” and 44.8% “Serious Cheating,” whereas TAs averaged 45.1%, 33.9%, and 21.0%, and students 37.6%, 37.2%, and 25.3%, respectively. Attribution of motives also diverged: students emphasized Prerequisite Knowledge (32.0%) and Time Management (19.5%); TAs emphasized Prerequisite Knowledge (38.1%) and Time Management (23.8%); instructors emphasized Grade Pressure (33.3%) and Laziness (33.3%) (Payne et al., 31 Mar 2026). This suggests that AM in integrity policy can be generated not only by unclear written rules but also by misaligned causal narratives about why students cross boundaries.
Across these education studies, AM is measurable, consequential, and heterogeneous. In one setting it is a mismatch between expected and actual exam form; in another it is policy ambiguity around LLMs; in a third it is divergence in perceived seriousness and motivation. The common structure is misalignment between institutional intent and participant interpretation.
3. Public science communication, infodemics, and uncertainty language
The COVID-19 infodemic literature treats AM as a structural risk of accelerated, high-volume science communication. The World Health Organization’s definition of an infodemic is “an overabundance of information—some accurate, some not—that occurs during an epidemic.” Within that ecology, academics function simultaneously as research producers, verifiers, and public communicators, so missteps can amplify both correct and incorrect messages (Cole, 2020). The same paper identifies advantages of academic engagement, including source credibility, rapid knowledge exchange through tools such as the Johns Hopkins CSSE dashboard and Worldometer, counter-misinformation “vaccines,” and the educator role in explaining transmission, mutation, and mitigation. It also identifies pitfalls: sensationalism, overconfidence in preliminary data, misinterpretation of projections, and communication ecology mismatch when academics speak on the wrong platform or in inaccessible language (Cole, 2020).
The failure pathways are concrete. One-sentence press releases can be distorted into alarming headlines; lack of plain-language summaries leads lay audiences to fill gaps with worst-case assumptions; expert threads on Twitter may never reach non-Twitter users; and reluctance to offer imperfect conclusions creates an information vacuum exploited by conspiracy theorists and anti-vaxxers (Cole, 2020). The paper situates mitigation within several frameworks: the WHO infodemic management framework, the Shannon–Weaver source–message–channel–receiver model extended by communicative ecologies, and epidemic-diffusion analogies in which information spread is modeled as
0
The point of the analogy is that correct information must propagate fast enough to outpace misinformation (Cole, 2020).
A complementary strand of AM research concerns uncertainty expressions themselves. Willems et al. studied 29 Dutch probability and frequency phrases using a survey of 881 Dutch/Flemish native speakers, including 226 statisticians and 655 non-statisticians. Participants interpreted each phrase as a point estimate in 1. The results showed large variability in interpretation, neutral contexts with no structural influence, and asymmetry in complementary phrases. For example, the mean interpretation of “likely” was approximately 75% and “unlikely” approximately 16%, summing to 91% rather than 100%; “very likely” and “very unlikely” summed to 96%; “almost always” and “almost never” to 98%. The paper reports no structural differences between statisticians and non-statisticians or between males and females, with maximum group mean differences of approximately 4 percentage points (Willems et al., 2019).
This line of work corrects a common assumption that domain expertise is sufficient to stabilize verbal uncertainty. It was not. The conclusion is therefore operational: avoid pure verbal expressions when precise risk communication is needed, prefer numerical formulations such as “There is at least 80 % chance of rain,” and, if verbal labels must be used, attach them to explicit numeric values or ranges (Willems et al., 2019).
4. Notation, methodology, and genre as sources of miscommunication
Some forms of AM arise not from disagreement over substance but from defects in representation. Galati’s analysis of incomplete-data methodology identifies three such defects. The first is the ambiguous use of 2 and 3 to denote both a formal partition of the random vector 4 induced by a missingness pattern 5 and a temporal device inside the marginal law of 6. The proposed remedy is to distinguish
7
for formal missingness from
8
for temporal overlaying. The second issue is the commonplace notation
9
used to communicate MAR. Galati argues that the right-hand side is undefined as written and encourages the stronger, and incorrect, conditional-independence reading 0. The third issue is the framing of ignorability by emulating complete-data methods rather than treating the question on its own merits. Galati’s reformulation is that, unless the mechanism is non-MAR or the investigator wants a specific joint restriction on 1, inference for 2 can proceed directly from the observed-data likelihood
3
without positing a missingness model that is then “ignored” (Galati, 2019).
In HCI, the obstacle is not notation but publication genre. Oliver et al. argue that page-count limits force authors to truncate background, context, and methodology; strict ACM/IEEE template rules discourage narrative or reflective writing styles; blind review rules obstruct the building of progressive context; and review panel turnover produces contradictory expectations across venues. The paper introduces two conceptual framings for this process. “Review Panel Ping-Pong” denotes repeated rewriting in response to successive, often contradictory reviews. “Matrioshka Unpacking” denotes the forced re-expansion of previously published ideas inside a new paper because anonymized self-citation and reviewer demands require context to be rebuilt from scratch (Oliver et al., 2022).
The significance of these two literatures is that AM can be generated by the scholarly infrastructure itself. In one case, ambiguous notation obscures central inferential claims; in the other, review conventions obstruct intelligibility of a conceptually simple contribution. A plausible implication is that clarity failures in academia are often systemic rather than merely stylistic.
5. AI-mediated fabrication, compound deception, and failure-mode taxonomies
The most direct recent analysis of AM in publication workflows concerns fabricated citations in accepted conference papers. The NeurIPS 2025 study examined 100 AI-generated hallucinated citations that appeared in 53 published papers, approximately 1% of all accepted papers, despite review by 3–5 expert researchers per paper. GPTZero Hallucination Check scanned 4,841 of 5,290 accepted papers, flagged 100 suspect citations across 53 papers, and human experts verified each flagged citation by searching Google Scholar, CrossRef, PubMed, arXiv, and DOI resolvers. The primary failure-mode distribution was Total Fabrication (66%), Partial Attribute Corruption (27%), Identifier Hijacking (4%), Placeholder Hallucination (2%), and Semantic Hallucination (1%). A critical finding was that every hallucination, 100%, exhibited compound failure modes. Secondary characteristics were dominated by Semantic Hallucination (63%) and Identifier Hijacking (29%). Among Total Fabrication primaries, 50 of 66, approximately 76%, also exhibited Semantic Hallucination as a secondary mode, and all 100 citations satisfied 4. The paper explains evasion of peer review through three heuristics: 5, “Title sounds domain-appropriate”; 6, “Link/ID resolves”; and 7, “Recognize author names/venues.” Compound layering lets a citation pass 8 unless every attribute is exhaustively checked (Ansari, 5 Feb 2026).
The distribution of contaminated papers was bimodal: 92% contained 1–2 hallucinations, interpreted in the paper as minimal AI use, while 8% contained 4–13 hallucinations, interpreted as heavy reliance. The risks identified include evidentiary collapse, reproducibility breakdown, citation-graph contamination, and training-data pollution through “Contamination Inheritance” (Ansari, 5 Feb 2026). The study therefore treats citation hallucination not as isolated sloppiness but as a multi-attribute deception structure that current peer review does not detect.
A related but distinct taxonomy appears in research-level mathematics. Auditing eight one-shot proofs generated by Gemini 2.5 Flash on Questions 1, 2, and 5 of the First Proof benchmark, the author identifies four failure modes: F1 citation fabrication, F2 premise smuggling, F3 silent problem reformulation, and F4 local-to-global compatibility gaps. The central empirical finding is that not one of the eight proofs contained a confirmed fabricated citation, yet every single one contained at least one load-bearing claim asserted as a “fundamental result” or “standard argument” with no justification attached. A two-stage premise-audit instrument, using a regex scan followed by a zero-temperature LLM judge, achieved 100% precision with 9, 0, and 50% proof-level recall with 1, yielding 2-score approximately 0.667. No proof was completely correct: final answer correct was 3 (Banerjee et al., 12 Jun 2026).
The mathematics study is especially important because it bounds the reach of citation verification. Retrieval-Augmented Generation can plausibly suppress F1 citation fabrication, but the author argues that it cannot address F2, F3, or F4 when the model asserts a false or load-bearing premise without citing anything. In that sense, citation verification catches only one class of AI-mediated AM. Unsupported premises, silent objective drift, and unverified global consistency remain invisible to purely bibliographic checks (Banerjee et al., 12 Jun 2026).
6. Mitigation strategies and institutional responses
The response to AM varies by domain, but a common theme is movement from tacit convention to explicit verification. For publication workflows, the NeurIPS citation study proposes mandatory automated citation verification at submission as an implementable solution. Its four-step pipeline consists of: Existence Check through exact string search in CrossRef API and Google Scholar; Metadata Consistency to verify that 4 form one unique record; Identifier Validation to fetch DOI or arXiv metadata and compare fields; and Semantic Plausibility Flags to detect out-of-domain or templated titles via a lightweight domain model. Implementation notes include batch queries to CrossRef or Unpaywall, fuzzy matching such as Levenshtein 5, threshold scoring with citations below 6 automatically flagged for author correction, and integration into CMT or OpenReview as a pre-review compliance check (Ansari, 5 Feb 2026).
For AI use in education, de Souza Santos et al. recommend explicit, activity-level AI policies in each syllabus or rubric, specifying how LLMs may or may not be used for code drafting versus final submission, documentation versus design artifact, and formative feedback versus summative output. They further recommend process-based and formative assessment, in-class or synchronous code walkthroughs, alignment of learning outcomes with AI use, formal guidance and training, and community norms beyond formal sanctions (Santos et al., 17 Mar 2026). In physics instruction, Scharlach recommends explicitly modeling multiple problem styles, providing a detailed exam blueprint, soliciting and correcting student expectations mid-semester, and fostering a classroom culture of belonging to reduce Impostor Syndrome (Scharlach, 29 Aug 2025). In computing education, Payne et al. recommend explicit integrity dialogues, transparent detection tools, automated variation, guidelines that frame generative AI as “akin to a study-buddy” with attribution and comprehension checks, scaffolded prerequisites, and flexible deadlines (Payne et al., 31 Mar 2026).
For public communication, the WHO infodemic management framework recommends scanning and verifying incoming evidence, explaining the science in accessible terms, amplifying accurate messages via trusted channels, measuring infodemic trends and response efficacy, and coordinating information technology interventions. The same literature recommends audience mapping, plain-language summaries, proactive relationships with reputable outlets such as Science Media Centre, participation in moderated forums such as r/coronavirus, cross-disciplinary collaboration, and WHO training programmes in infodemic management (Cole, 2020). For probabilistic language, Willems et al. recommend avoiding pure verbal expressions when precise risk communication is needed and supplementing any qualitative label with numeric anchors or ranges (Willems et al., 2019).
Finally, some proposed mitigations target the communication infrastructure itself. Oliver et al. recommend cross-venue reviewer continuity, template flexibility, reconsideration of anonymity rules to permit explicit self-citation or context linking, and greater openness to emerging paradigms in HCI methodology (Oliver et al., 2022). Galati’s remedies are even more basic: explicit dependence on missingness pattern 7, definition of previously undefined shorthand, and framing ignorability directly as a property of the true missingness mechanism rather than of a hypothetical model to be specified and ignored (Galati, 2019).
Taken together, these remedies suggest that AM is best addressed not by exhortation alone but by redesign of interfaces between producer and receiver: assessment interfaces, publication interfaces, notation systems, public-media interfaces, and inference-time pipelines. Where meaning is left to heuristic reconstruction, AM flourishes; where claims, rules, and references are made explicit, auditable, and context-matched, its scope narrows.