Papers
Topics
Authors
Recent
Search
2000 character limit reached

Impact of introducing "Informatics I" to the common university entrance examination in Japan: a longitudinal study on students' perceptions of their information-related knowledge and skills from 2006 to 2026

Published 13 Aug 2026 in cs.CY | (2608.12924v1)

Abstract: Despite the recent intensive development of secondary education curricula and assessments in informatics, the impact of assessments has not been well studied in this field. Since informatics education covers a diverse range of content, from computer science knowledge to ICT skills, careful consideration is needed to prevent assessments from distorting education. This study investigates the impact of introducing ``Informatics I'' into the Common Test for University Admissions in Japan, as an example of a large-scale, standardized, high-stakes assessment in 2025. As the data source for this analysis, this study uses a questionnaire that has been administered every year from 2006 to 2026 to all first-year students at the University of Tokyo. The questionnaire asks students for their self-perceptions of the information-related knowledge and skills they studied and acquired in high school. Using these data, we conduct a longitudinal study of the 2013 curriculum reform, the 2022 reform, and the introduction of the new entrance examination. We attempt to isolate the impact of the entrance examination through two comparisons: between the 2013 curriculum reform and the 2022 reform; and between direct-entry and gap-year students among those entering in 2025, who followed different curricula but took the new examination. We use the theoretical framework of the washback effect as a lens for interpreting these differences. We found that (1) the 2013 curriculum reform produced no discontinuity in students' perceptions, whereas (2) the introduction of the new entrance examination in 2025 produced a sharp change, particularly in the proportion of students reporting acquisition of computer-science topics; and (3) this change is too large to be interpreted as a gain in proficiency, and is better understood as a shift in students' criteria for judging acquisition.

Authors (1)

Summary

  • The paper uses a 2006–2026 longitudinal survey of University of Tokyo students and curriculum and cohort comparisons to identify the exam’s washback effects.
  • The 2025 exam produced an abrupt rise in reported computer science acquisition, including programming from below 20% to about 55%, but the pattern likely reflects changed self-assessment criteria rather than equivalent proficiency gains.
  • The findings show positive washback for tested CS content and negative washback for untested literacy skills, highlighting the need to validate self-reports with objective measures and monitor neglected competencies.

Context and motivation

Assessment has long been recognized as a driver of what is taught, yet studies of assessment impact in informatics education remain scarce. Most literature on informatics curricula—covering New Zealand, the UK, Poland, the USA, Wales, and Ireland—focuses on curriculum design and offers only limited discussion of assessment (2608.12924). The paper under review addresses this gap by examining Japan's 2025 introduction of "Informatics I" into the Common Test for University Admissions (CT), a large-scale, standardized, high-stakes examination taken by roughly half of all university applicants regardless of intended major. This makes the Japanese case distinctive: unlike the AP CSP test in the USA, NSI in France, or A-levels in the UK, where informatics testing is largely restricted to STEAM-bound students, the CT reaches a near-universal applicant population. In 2026, of approximately 620,000 students entering university, about 460,000 took the CT and more than 300,000 took CT "Informatics I" (2608.12924).

The study is grounded in the theoretical framework of the washback effect—the influence of tests on teaching and learning—as formulated by Hughes and refined by Alderson into fifteen hypotheses concerning participants, processes, and products. Prior washback research on learners' perceptions is limited, mostly confined to foreign-language testing, and the author identifies this work as the first substantial application of the framework to informatics education.

Data and identification strategy

The data source is the "Survey on the Study of High School Informatics," administered annually from 2006 to 2026 to all first-year students at the University of Tokyo (roughly 3,100 students per cohort). Students report whether they "studied" and whether they "acquired" each of thirteen information-related topics; the proportions reporting each are termed study rates and acquisition rates. The items are classified into eight "Literacy" items unlikely to be examined on the CT (e.g., word processing, spreadsheets, typing) and five "CS" items likely to be examined (programming, computer systems, simulation, databases, ethics), based on their coverage in the Courses of Study for Informatics I.

The central methodological challenge is disentangling the effect of the new examination from that of the 2022 curriculum reform, since students who took the CT had studied the new curriculum by construction. The study exploits two comparisons:

  • Curriculum reform without examination change: the 2013 reform (first to second period) affected cohorts entering from 2016 with no corresponding entrance examination change, providing an estimate of what a curriculum reform alone does to perceptions.
  • Same curriculum, different examination exposure: among students entering in 2025, gap-year (GY) students studied the pre-2022 curriculum but took the CT, while direct-entry (DE) students studied the post-2022 curriculum and took the CT. Comparing GY students with the 2024 cohort—who followed the same curriculum but took no informatics examination—isolates the examination's effect.

The author explicitly turns the perceptual nature of the data into an analytical asset: if students' reported perceptions shift while the underlying curriculum does not, that strongly indicates a washback effect rather than a change in actual instruction.

Findings across the first two periods

Over the first period (2003–2012) and second period (2013–2021), study rates followed continuous, gradual trends with no visible discontinuity at the 2016 curriculum transition. Literacy study rates rose in the first period and declined slowly thereafter—notably for word processing and presentation—with the decline beginning around 2019 rather than at the reform itself. CS study rates rose slowly until around 2020, then accelerated, which the author attributes to teachers anticipating the CT introduction. The conclusion for RQ1 is that the 2013 curriculum reform produced no observable discontinuity; changes appear driven by teachers' incremental instructional improvement and by topic familiarity in students' daily lives (e.g., messaging applications displacing email, Covid-era online classes raising typing familiarity).

A notable secondary finding is that acquisition rates are largely inconsistent with study rates: CS topics were increasingly studied over two decades while acquisition rates remained flat. This divergence foreshadows the paper's central interpretive claim about self-report criteria.

The discontinuity of 2025

In contrast to the continuity surrounding the 2013 reform, the 2025 introduction of the CT produced an abrupt change in CS-item rates. Study rates jumped for nearly all CS items, and acquisition rates rose sharply except for ethics. The magnitude is striking: the programming acquisition rate rose from below 20% in 2024 to about 55% in 2025—a single-year change exceeding the accumulated movement of nearly twenty years.

The author argues this change cannot plausibly represent a genuine proficiency gain. If taken literally, it would imply programming is easier to acquire than touch typing or word processing, which is counter-intuitive. Washback research consistently finds that tests have modest effects on learners' actual products but strong effects on learners' subjective perceptions—motivation, perceived difficulty, and self-efficacy. The preferred interpretation is therefore a shift in students' criteria for judging acquisition: before 2025, students may have required professional-level competence to claim they had "acquired" programming; after preparing for the CT, being able to solve examination questions sufficed. The ethics item supports this reading: its acquisition rate did not rise despite high study rates, plausibly because examination preparation exposed students to advanced legal and cryptographic content that made them judge the topic as harder than they had assumed.

The DE/GY comparison sharpens the causal attribution. GY students, who followed the same curriculum as the 2024 cohort but took the CT, showed significantly different study rates (chi-square tests with Bonferroni correction, significant at the 5% level for the five largest changes) moving in test-aligned directions—down for literacy, up for CS—and CS acquisition rates approaching DE levels. Because these students' actual curricula were identical to the 2024 cohort's, the differences must reflect perception change induced by the examination. Notably, some literacy rates for GY fell below those of the 2024 cohort even though GY students re-studied informatics for the examination, which the author interprets as negative washback on content unlikely to be examined—consistent with Scanlon et al.'s interview findings and prior washback studies. The answer to RQ3 is that study rates track the curriculum while acquisition rates track the examination, but the examination also distorts study-rate reports, indicating broad perception change; washback is positive for the CS domain and somewhat negative for literacy.

Implications

Three implications follow directly from the results. First, curriculum reform alone had no immediate measurable effect on student perceptions; its effects emerge slowly through teacher practice, and acquisition is dominated by everyday familiarity with technologies that shift rapidly. Second, high-stakes examinations produce drastic, rapid effects—but on perceptions more than proficiency—which complicates any claim that the examination improved informatics education. Third, and most consequentially for measurement methodology, self-reports of acquisition are unreliable indicators of proficiency across institutional changes, because respondents' internal criteria shift. The author recommends continuous measurement of unexamined areas whenever a high-stakes test is introduced, and calibration of self-reports against other instruments across institutional transitions. For large-scale paper-based tests like the CT, no obvious solution exists for assessing literacy-type competencies unsuited to paper formats; performance tasks as in AP CSP are feasible only at smaller scales.

Limitations

The paper is candid about several constraints. The data capture perceptions, not proficiency, so the criterion-shift interpretation, however well-motivated by washback theory, cannot be verified against objective measures within this study. Non-response bias is non-trivial precisely in the focal years: Cramér's V for the humanities/STEAM respondent composition is 0.16 in 2025 and 0.32 in 2026, and the STEAM skew may inflate the apparent improvements, though the author judges the 2025 change too large for bias to alter the overall picture. Item wording changed in 2025, adding supplementary keywords to the five CS items, which could itself have shifted impressions—for example, keywords like "conditional branching" might have lowered students' bar for claiming acquisition of programming. The single-institution sample (the most competitive university in Japan) likely amplifies the observed washback relative to other contexts, since high-performing schools' teachers and students engage more intensively with high-stakes tests. Finally, confounders such as the extra year since GY students last studied informatics, or societal changes including AI technology diffusion, cannot be fully excluded; the identification rests on the assumption that year-to-year variation in this survey is otherwise small. The author also acknowledges potential analyst bias, teaching the very population surveyed.

Conclusion

Drawing on a rare 21-year longitudinal dataset, this study demonstrates that Japan's 2013 curriculum reform left students' perceptions essentially unchanged, whereas the 2025 introduction of CT "Informatics I" produced a sharp, statistically significant discontinuity concentrated in computer-science topics—an effect better explained as a change in students' criteria for judging acquisition than as a proficiency gain, accompanied by negative washback on unexamined literacy content. As the first substantive application of washback theory to informatics education, the study establishes both the potency of high-stakes assessment in shaping the field's perceived content and the fragility of self-report measures across institutional change. The open question it leaves most pointedly is how to measure genuine proficiency gains from such reforms when the primary available instrument—student perception—is itself distorted by the intervention under study.

Paper to Video (Beta)

No one has generated a video about this paper yet.

Whiteboard

No one has generated a whiteboard explanation for this paper yet.

Tweets

Sign up for free to view the 1 tweet with 0 likes about this paper.