MDK12-Bench: Multimodal K-12 Reasoning Benchmark
- MDK12-Bench is a multi-discipline benchmark using real-world K-12 exam data to assess multimodal reasoning across six subjects from 2016 to 2025.
- It employs a dynamic evaluation framework with layered automated and manual curation, including GPT-based filtering and educator screening for robust quality assurance.
- Empirical results reveal notable weaknesses in math and physics, highlighting the challenges in achieving reliable multimodal reasoning in educational settings.
Searching arXiv for the latest MDK12-Bench papers and related benchmark descriptions. MDK12-Bench is a multi-discipline benchmark for evaluating reasoning in multimodal LLMs (MLLMs) using real-world K-12 examinations. It is presented as a large-scale resource for assessing multimodal reasoning through text-only and image-plus-text questions spanning multiple school subjects, difficulty levels, question formats, and exam years. Across the two associated papers, MDK12-Bench is characterized by six disciplines, full K-12 or near-full K-12 educational coverage, rich knowledge-point annotation, detailed answer explanations, and a dynamic evaluation framework intended to mitigate benchmark contamination by transforming question forms, textual surface realization, and image appearance during evaluation (Zhou et al., 8 Apr 2025, Zhou et al., 9 Aug 2025).
1. Definition and scope
MDK12-Bench is defined as a benchmark for evaluating multimodal reasoning in MLLMs through authentic K-12 exam questions collected from open-source exam repositories. The benchmark is explicitly motivated by four shortcomings in prior multimodal reasoning evaluation: limited data size, narrow domain coverage, weakly structured knowledge organization, and vulnerability to data contamination from static benchmark items appearing in model training corpora (Zhou et al., 8 Apr 2025).
The benchmark spans six disciplines: Mathematics, Physics, Chemistry, Biology, Geography, and Information Science. It supports both Text-only (T) and Image + text (I+T) settings, and includes several answer formats used in educational assessment. One paper reports four question formats—Multiple-choice, Fill-in-the-blank, True-or-false, and Open-ended—with multiple-choice further including single-answer and multi-answer variants (Zhou et al., 8 Apr 2025). A later paper states that MDK12-Bench covers five formats—SC, MC, Fill, T/F, and Open—which makes the single-choice and multiple-choice distinction explicit (Zhou et al., 9 Aug 2025). This suggests an evolution in the benchmark’s presentation rather than a change in its educational orientation.
The scale is large by benchmark standards. The benchmark is reported as containing 141,320 total instances, including 77,857 text-only instances, 63,463 multimodal instances, and 105,218 total images (Zhou et al., 8 Apr 2025). The later paper reports the same corpus at the rounded level as 141.3K instances and 105.2K images (Zhou et al., 9 Aug 2025). Both descriptions emphasize coverage from 2016 to 2025, enabling year-based analyses and contamination-oriented evaluation.
2. Data sources, organization, and annotation structure
MDK12-Bench is built from real-world K-12 examinations gathered from online open-source exam paper repositories. The source materials were originally in Chinese, and the benchmark was translated into English during processing, with manual verification by domain experts for technical fidelity (Zhou et al., 8 Apr 2025). The later paper describes an initial crawl of 5.8M multimodal exam instances, followed by several filtering stages that reduce the corpus first to 4.2M, then 0.6M, then 0.2M, and finally to the released 141.3K instances (Zhou et al., 9 Aug 2025).
Each question is stored with structured metadata. The earlier paper lists the core per-instance fields as Year, Question, Grade Level, Image, Difficulty Level, Answer, Question Type, Knowledge, Course, Analysis, and Modality (Zhou et al., 8 Apr 2025). The later paper describes similar metadata in expanded form, including difficulty level, exam year, question form, question, answer, text, image, grade level, curriculum, topic, knowledge points, and answer explanation (Zhou et al., 9 Aug 2025). These fields make the benchmark suitable not only for leaderboard-style evaluation but also for filtered analysis by topic, grade band, modality, or difficulty.
A central feature of MDK12-Bench is its hierarchical knowledge system. The benchmark is described as using a six-level or six-layer knowledge taxonomy. The operational hierarchy is given as discipline → grade → curriculum → topic → meta-knowledge → key knowledge point (Zhou et al., 8 Apr 2025). A later paper describes the six layers as Level 1 – Disciplines, Level 2 – Grade levels, Level 3 – Subfields, Level 4 – Curriculum, Level 5 – Topics, and Level 6 – Knowledge points (Zhou et al., 9 Aug 2025). The benchmark is thus not organized as a flat subject-labeled dataset; it is structured as an educational ontology.
The reported count of knowledge points differs across the two papers. One paper states 6,827 total knowledge points and attributes this to “instance-level knowledge point annotations based on a well-organized knowledge structure” (Zhou et al., 8 Apr 2025). The later paper reports 6,225 knowledge points in a six-layer taxonomy (Zhou et al., 9 Aug 2025). Because both values are explicitly reported in the source materials, the discrepancy is best understood as a version difference or taxonomy revision rather than something that can be resolved from the available text alone.
3. Construction pipeline and quality control
The benchmark is constructed through a staged curation process. The earlier paper describes a four-stage curation pipeline: Data Collection, Data Screening, Data Parsing, and Data Processing (Zhou et al., 8 Apr 2025). In that account, screening combines GPT-4o-based automated review, human inspection, and a predefined checklist, filtering out questions with low-quality images or without specific knowledge points. Parsing is rule-based, and the final processing step translates all Chinese text into English using the GPT-4o API, with domain experts manually reviewing translations and manually checking translated image text (Zhou et al., 8 Apr 2025).
The later paper provides a more granular five-stage account: large-scale collection, rule-based filtering, GPT-based filtering, educator filtering, and post-processing / final rule checks (Zhou et al., 9 Aug 2025). The rule-based stage evaluates criteria such as text-image correspondence, image resolution/clarity, content completeness, metadata accuracy, structural/format consistency, semantic coherence, duplication/redundancy, logical soundness, year coverage, non-educational content, encoding and unit consistency, and equation/symbol validity. GPT-4o is then used to assess semantic consistency, reasoning soundness, factual correctness, language clarity, completeness, grade/difficulty appropriateness, visual reference accuracy, and multi-step reasoning validity (Zhou et al., 9 Aug 2025).
These descriptions jointly indicate that MDK12-Bench is not a simple scrape of examination material. It is a processed benchmark with layered automated and human quality assurance. The paper from April 2025 further states that the dataset was built over about two months with 20+ researchers and several K-12 educators (Zhou et al., 8 Apr 2025). A plausible implication is that the benchmark’s emphasis on structured annotations and year metadata required substantial manual normalization beyond standard dataset cleaning.
4. Evaluation framework and benchmark dimensions
MDK12-Bench is designed to evaluate MLLMs along several axes beyond raw average accuracy. The earlier paper emphasizes coverage by discipline, difficulty, knowledge point, and cross-year partitions, as well as robustness under a dynamic evaluation framework (Zhou et al., 8 Apr 2025). The later paper formalizes four benchmark dimensions: difficulty levels, temporal (cross-year) shifts, contextual shifts, and knowledge-driven reasoning (Zhou et al., 9 Aug 2025).
Difficulty is explicitly annotated as Easy, Medium, and Hard (Zhou et al., 8 Apr 2025). Year metadata spans 2016–2025 and is retained to support breakdown analyses, cross-validations, and dynamic updates (Zhou et al., 8 Apr 2025). The benchmark therefore supports temporal diagnostics rather than only aggregate reporting.
The principal evaluation metric is accuracy, but the scoring procedure supports partial credit. The dynamic evaluation paper states that exact match gets full score 1.0, while otherwise GPT + predefined scoring rules are used for partial credit, with examples such as 0.5 when one of two blanks is correct, or for correct out of choices in a multi-select item (Zhou et al., 8 Apr 2025, Zhou et al., 9 Aug 2025). In practical terms, this means the reported “accuracy” is an averaged graded score rather than a strict all-or-nothing percentage for every item.
The benchmark also introduces MDK12-Mini for lighter-weight evaluation. One paper reports that MDK12-Mini contains 14,595 total instances, split into 4,951 easy, 4,692 medium, and 4,952 hard, with 10% of the data from each difficulty slice and uniform sampling over key knowledge points where possible (Zhou et al., 8 Apr 2025). A later paper reports 14,856 instances, balanced as 4,952 for each of easy, medium, and hard (Zhou et al., 9 Aug 2025). As with the knowledge-point counts, this numerical difference is explicit in the source record and is most conservatively read as reflecting benchmark revision.
5. Dynamic evaluation and contamination mitigation
A distinctive contribution of MDK12-Bench is its dynamic evaluation framework, introduced to mitigate the contamination risks associated with static benchmark questions (Zhou et al., 8 Apr 2025). The framework generates transformed versions of benchmark items that preserve the original answer while altering textual wording, response format, or image appearance. This is intended to reduce dependence on memorized surface forms and to test robustness under controlled distribution shifts (Zhou et al., 9 Aug 2025).
The framework consists of three modules: Image bootstrapping, Text bootstrapping, and Two-stage answer evaluation (Zhou et al., 8 Apr 2025). On the textual side, it applies Word Substitution, Sentence Paraphrasing, and Question Type Permutation, including examples such as multiple-choice → fill-in-the-blank (Zhou et al., 8 Apr 2025). On the visual side, it applies Spatial Transformation, Color Transformation, and Style Transformation. Spatial transformation pads the image with colors sampled from black, white, and grey, with padding widths sampled proportionally from 10% to 20% of side length (Zhou et al., 8 Apr 2025). Color transformation includes invert colors and salt-and-pepper noise with random density. Style transformation uses Flux-Dev while attempting to preserve semantic content (Zhou et al., 8 Apr 2025, Zhou et al., 9 Aug 2025).
Validity control is an explicit part of this framework. The paper states that “we apply a GPT-based judge to reject sampling wrong adapted instances” (Zhou et al., 8 Apr 2025). In the later paper’s notation, transformed samples are retained only if the semantic validity checker confirms that the transformed question or image remains consistent with the original answer (Zhou et al., 9 Aug 2025). This does not amount to a formal proof of semantic preservation, but it is a concrete safeguard against invalid perturbations.
Empirically, dynamic evaluation exposes substantial brittleness. On transformed items, Gemini2-thinking drops from 58.1 to 41.6, Gemini2-flash from 56.4 to 47.0, GPT-4o from 51.2 to 40.9, and Claude-3.7 from 46.7 to 31.4 (Zhou et al., 8 Apr 2025). The later paper summarizes the average reduction as 13.7% (Zhou et al., 9 Aug 2025). The text also states that text perturbations hurt more than image perturbations, and that compositional perturbations hurt most (Zhou et al., 8 Apr 2025). This suggests that, for many current MLLMs, academic multimodal reasoning remains tightly coupled to familiar linguistic presentation.
6. Empirical findings, diagnostic value, and significance
The benchmark’s experiments indicate that current MLLMs remain limited on multidisciplinary school-level reasoning even at large scale. On MDK12-Mini, one paper reports the best overall model as Gemini2-thinking: 59.4% overall, followed by Gemini2-flash: 57.2%, QVQ-72B: 53.2%, GPT-o1-mini: 53.1%, Qwen2.5-VL-72B: 51.9%, InternVL2.5-MPO: 51.7%, GPT-4o: 50.0%, and Claude-3.7: 49.8% (Zhou et al., 8 Apr 2025). A later paper reports stronger absolute numbers on a later evaluation table, with Gemini2-thinking: 67.8, Qwen2.5-VL-72B: 67.5, GPT-o1: 65.5, InternVL2.5-MPO: 65.2, InternVL2.5-78B: 64.6, and QVQ-72B: 64.4 (Zhou et al., 9 Aug 2025). The shared conclusion is that performance remains well below saturation.
Several consistent weaknesses are emphasized. Math and physics are persistently harder than the other disciplines: the later paper states that scores are 7.6% below the overall average in these areas, whereas chemistry, biology, geography, and information science are 3.7% above average (Zhou et al., 9 Aug 2025). Harder questions produce clear degradation; the later paper reports an average drop of 8.3% from easy to hard (Zhou et al., 9 Aug 2025). Cross-year evaluation also shows performance falling on newer exams, summarized as a 12.6% drop on newer material (Zhou et al., 9 Aug 2025). The earlier paper notes that higher performance on earlier exam data may indicate a contamination-related effect (Zhou et al., 8 Apr 2025).
The knowledge structure supports finer diagnosis than aggregate accuracy alone. Models perform better on frequently covered knowledge points and underperform on areas such as advanced geometry and biochemical processes (Zhou et al., 8 Apr 2025). A later paper additionally introduces KP-RAG, or knowledge-point reference-augmented generation, in which relevant knowledge-point references are appended to the prompt. Reported gains are +6.9% on easy, +6.0% on medium, and +2.1% on hard items (Zhou et al., 9 Aug 2025). This suggests that explicit knowledge support helps most when the task is primarily knowledge retrieval or concept recall, and less when the limiting factor is multi-step reasoning.
Relative to earlier multimodal benchmarks, MDK12-Bench’s distinctive contribution lies in the combination of scale, structured educational ontology, cross-year metadata, answer explanations, and dynamic robustness testing. One paper explicitly argues that future benchmarks should include structured knowledge systems, not just flat labels, and that evaluation should move beyond static test sets because contamination can distort conclusions (Zhou et al., 8 Apr 2025). The later paper extends this argument by positioning MDK12-Bench not merely as a static leaderboard dataset but as a diagnostic resource for robustness, generalization, knowledge use, and AI-assisted education (Zhou et al., 9 Aug 2025).