---
title: 'MDK12-Bench: Multimodal K-12 Reasoning Benchmark'
url: https://www.emergentmind.com/topics/mdk12-bench
type: topic
---

# MDK12-Bench: Multimodal K-12 Reasoning Benchmark

Searching arXiv for the latest MDK12-Bench papers and related benchmark descriptions.
MDK12-Bench is a multi-discipline benchmark for evaluating reasoning in multimodal large language models (MLLMs) using real-world K-12 examinations. It is presented as a large-scale resource for assessing multimodal reasoning through text-only and image-plus-text questions spanning multiple school subjects, difficulty levels, question formats, and exam years. Across the two associated papers, MDK12-Bench is characterized by six disciplines, full K-12 or near-full K-12 educational coverage, rich knowledge-point annotation, detailed answer explanations, and a dynamic evaluation framework intended to mitigate benchmark contamination by transforming question forms, textual surface realization, and image appearance during evaluation [2504.05782], [2508.06851].

## 1. Definition and scope

MDK12-Bench is defined as a benchmark for evaluating multimodal reasoning in MLLMs through authentic K-12 exam questions collected from open-source exam repositories. The benchmark is explicitly motivated by four shortcomings in prior multimodal reasoning evaluation: limited data size, narrow domain coverage, weakly structured knowledge organization, and vulnerability to data contamination from static benchmark items appearing in model training corpora [2504.05782].

The benchmark spans six disciplines: **Mathematics**, **Physics**, **Chemistry**, **Biology**, **Geography**, and **Information Science**. It supports both **Text-only (T)** and **Image + text (I+T)** settings, and includes several answer formats used in educational assessment. One paper reports four question formats—**Multiple-choice**, **Fill-in-the-blank**, **True-or-false**, and **Open-ended**—with multiple-choice further including **single-answer** and **multi-answer** variants [2504.05782]. A later paper states that MDK12-Bench covers five formats—**SC**, **MC**, **Fill**, **T/F**, and **Open**—which makes the single-choice and multiple-choice distinction explicit [2508.06851]. This suggests an evolution in the benchmark’s presentation rather than a change in its educational orientation.

The scale is large by benchmark standards. The benchmark is reported as containing **141,320 total instances**, including **77,857 text-only instances**, **63,463 multimodal instances**, and **105,218 total images** [2504.05782]. The later paper reports the same corpus at the rounded level as **141.3K instances** and **105.2K images** [2508.06851]. Both descriptions emphasize coverage from **2016 to 2025**, enabling year-based analyses and contamination-oriented evaluation.

## 2. Data sources, organization, and annotation structure

MDK12-Bench is built from **real-world K-12 examinations** gathered from **online open-source exam paper repositories**. The source materials were originally in **Chinese**, and the benchmark was translated into **English** during processing, with manual verification by domain experts for technical fidelity [2504.05782]. The later paper describes an initial crawl of **5.8M multimodal exam instances**, followed by several filtering stages that reduce the corpus first to **4.2M**, then **0.6M**, then **0.2M**, and finally to the released **141.3K instances** [2508.06851].

Each question is stored with structured metadata. The earlier paper lists the core per-instance fields as **Year**, **Question**, **Grade Level**, **Image**, **Difficulty Level**, **Answer**, **Question Type**, **Knowledge**, **Course**, **Analysis**, and **Modality** [2504.05782]. The later paper describes similar metadata in expanded form, including **difficulty level**, **exam year**, **question form**, **question**, **answer**, **text**, **image**, **grade level**, **curriculum**, **topic**, **knowledge points**, and **answer explanation** [2508.06851]. These fields make the benchmark suitable not only for leaderboard-style evaluation but also for filtered analysis by topic, grade band, modality, or difficulty.

A central feature of MDK12-Bench is its hierarchical knowledge system. The benchmark is described as using a **six-level** or **six-layer** knowledge taxonomy. The operational hierarchy is given as **discipline → grade → curriculum → topic → meta-knowledge → key knowledge point** [2504.05782]. A later paper describes the six layers as **Level 1 – Disciplines**, **Level 2 – Grade levels**, **Level 3 – Subfields**, **Level 4 – Curriculum**, **Level 5 – Topics**, and **Level 6 – Knowledge points** [2508.06851]. The benchmark is thus not organized as a flat subject-labeled dataset; it is structured as an educational ontology.

The reported count of knowledge points differs across the two papers. One paper states **6,827 total knowledge points** and attributes this to “instance-level knowledge point annotations based on a well-organized knowledge structure” [2504.05782]. The later paper reports **6,225 knowledge points** in a **six-layer taxonomy** [2508.06851]. Because both values are explicitly reported in the source materials, the discrepancy is best understood as a version difference or taxonomy revision rather than something that can be resolved from the available text alone.

## 3. Construction pipeline and quality control

The benchmark is constructed through a staged curation process. The earlier paper describes a **four-stage curation pipeline**: **Data Collection**, **Data Screening**, **Data Parsing**, and **Data Processing** [2504.05782]. In that account, screening combines **GPT-4o-based automated review**, **human inspection**, and a **predefined checklist**, filtering out questions with **low-quality images** or **without specific knowledge points**. Parsing is **rule-based**, and the final processing step translates all Chinese text into English using the **GPT-4o API**, with **domain experts manually reviewing** translations and manually checking translated image text [2504.05782].

The later paper provides a more granular five-stage account: **large-scale collection**, **rule-based filtering**, **GPT-based filtering**, **educator filtering**, and **post-processing / final rule checks** [2508.06851]. The rule-based stage evaluates criteria such as **text-image correspondence**, **image resolution/clarity**, **content completeness**, **metadata accuracy**, **structural/format consistency**, **semantic coherence**, **duplication/redundancy**, **logical soundness**, **year coverage**, **non-educational content**, **encoding and unit consistency**, and **equation/symbol validity**. GPT-4o is then used to assess **semantic consistency**, **reasoning soundness**, **factual correctness**, **language clarity**, **completeness**, **grade/difficulty appropriateness**, **visual reference accuracy**, and **multi-step reasoning validity** [2508.06851].

These descriptions jointly indicate that MDK12-Bench is not a simple scrape of examination material. It is a processed benchmark with layered automated and human quality assurance. The paper from April 2025 further states that the dataset was built over about **two months** with **20+ researchers** and several **K-12 educators** [2504.05782]. A plausible implication is that the benchmark’s emphasis on structured annotations and year metadata required substantial manual normalization beyond standard dataset cleaning.

## 4. Evaluation framework and benchmark dimensions

MDK12-Bench is designed to evaluate MLLMs along several axes beyond raw average accuracy. The earlier paper emphasizes coverage by **discipline**, **difficulty**, **knowledge point**, and **cross-year partitions**, as well as robustness under a **dynamic evaluation framework** [2504.05782]. The later paper formalizes four benchmark dimensions: **difficulty levels**, **temporal (cross-year) shifts**, **contextual shifts**, and **knowledge-driven reasoning** [2508.06851].

Difficulty is explicitly annotated as **Easy**, **Medium**, and **Hard** [2504.05782]. Year metadata spans **2016–2025** and is retained to support **breakdown analyses, cross-validations, and dynamic updates** [2504.05782]. The benchmark therefore supports temporal diagnostics rather than only aggregate reporting.

The principal evaluation metric is **accuracy**, but the scoring procedure supports partial credit. The dynamic evaluation paper states that **exact match gets full score 1.0**, while otherwise **GPT + predefined scoring rules** are used for partial credit, with examples such as **0.5** when one of two blanks is correct, or **\(m/n\)** for \(m\) correct out of \(n\) choices in a multi-select item [2504.05782], [2508.06851]. In practical terms, this means the reported “accuracy” is an averaged graded score rather than a strict all-or-nothing percentage for every item.

The benchmark also introduces **MDK12-Mini** for lighter-weight evaluation. One paper reports that MDK12-Mini contains **14,595 total instances**, split into **4,951 easy**, **4,692 medium**, and **4,952 hard**, with **10% of the data** from each difficulty slice and **uniform sampling over key knowledge points** where possible [2504.05782]. A later paper reports **14,856 instances**, balanced as **4,952** for each of easy, medium, and hard [2508.06851]. As with the knowledge-point counts, this numerical difference is explicit in the source record and is most conservatively read as reflecting benchmark revision.

## 5. Dynamic evaluation and contamination mitigation

A distinctive contribution of MDK12-Bench is its **dynamic evaluation framework**, introduced to mitigate the contamination risks associated with static benchmark questions [2504.05782]. The framework generates transformed versions of benchmark items that preserve the original answer while altering textual wording, response format, or image appearance. This is intended to reduce dependence on memorized surface forms and to test robustness under controlled distribution shifts [2508.06851].

The framework consists of three modules: **Image bootstrapping**, **Text bootstrapping**, and **Two-stage answer evaluation** [2504.05782]. On the textual side, it applies **Word Substitution**, **Sentence Paraphrasing**, and **Question Type Permutation**, including examples such as **multiple-choice → fill-in-the-blank** [2504.05782]. On the visual side, it applies **Spatial Transformation**, **Color Transformation**, and **Style Transformation**. Spatial transformation pads the image with colors sampled from **black, white, and grey**, with padding widths sampled proportionally from **10% to 20%** of side length [2504.05782]. Color transformation includes **invert colors** and **salt-and-pepper noise** with random density. Style transformation uses **Flux-Dev** while attempting to preserve semantic content [2504.05782], [2508.06851].

Validity control is an explicit part of this framework. The paper states that “**we apply a GPT-based judge to reject sampling wrong adapted instances**” [2504.05782]. In the later paper’s notation, transformed samples are retained only if the semantic validity checker confirms that the transformed question or image remains consistent with the original answer [2508.06851]. This does not amount to a formal proof of semantic preservation, but it is a concrete safeguard against invalid perturbations.

Empirically, dynamic evaluation exposes substantial brittleness. On transformed items, **Gemini2-thinking** drops from **58.1** to **41.6**, **Gemini2-flash** from **56.4** to **47.0**, **GPT-4o** from **51.2** to **40.9**, and **Claude-3.7** from **46.7** to **31.4** [2504.05782]. The later paper summarizes the average reduction as **13.7%** [2508.06851]. The text also states that **text perturbations hurt more than image perturbations**, and that **compositional perturbations hurt most** [2504.05782]. This suggests that, for many current MLLMs, academic multimodal reasoning remains tightly coupled to familiar linguistic presentation.

## 6. Empirical findings, diagnostic value, and significance

The benchmark’s experiments indicate that current MLLMs remain limited on multidisciplinary school-level reasoning even at large scale. On MDK12-Mini, one paper reports the best overall model as **Gemini2-thinking: 59.4% overall**, followed by **Gemini2-flash: 57.2%**, **QVQ-72B: 53.2%**, **GPT-o1-mini: 53.1%**, **Qwen2.5-VL-72B: 51.9%**, **InternVL2.5-MPO: 51.7%**, **GPT-4o: 50.0%**, and **Claude-3.7: 49.8%** [2504.05782]. A later paper reports stronger absolute numbers on a later evaluation table, with **Gemini2-thinking: 67.8**, **Qwen2.5-VL-72B: 67.5**, **GPT-o1: 65.5**, **InternVL2.5-MPO: 65.2**, **InternVL2.5-78B: 64.6**, and **QVQ-72B: 64.4** [2508.06851]. The shared conclusion is that performance remains well below saturation.

Several consistent weaknesses are emphasized. **Math and physics** are persistently harder than the other disciplines: the later paper states that scores are **7.6% below** the overall average in these areas, whereas **chemistry, biology, geography, and information science are 3.7% above average** [2508.06851]. Harder questions produce clear degradation; the later paper reports an average drop of **8.3%** from easy to hard [2508.06851]. Cross-year evaluation also shows performance falling on newer exams, summarized as a **12.6%** drop on newer material [2508.06851]. The earlier paper notes that higher performance on earlier exam data may indicate a contamination-related effect [2504.05782].

The knowledge structure supports finer diagnosis than aggregate accuracy alone. Models perform better on **frequently covered knowledge points** and underperform on areas such as **advanced geometry** and **biochemical processes** [2504.05782]. A later paper additionally introduces **KP-RAG**, or **knowledge-point reference-augmented generation**, in which relevant knowledge-point references are appended to the prompt. Reported gains are **+6.9%** on easy, **+6.0%** on medium, and **+2.1%** on hard items [2508.06851]. This suggests that explicit knowledge support helps most when the task is primarily knowledge retrieval or concept recall, and less when the limiting factor is multi-step reasoning.

Relative to earlier multimodal benchmarks, MDK12-Bench’s distinctive contribution lies in the combination of scale, structured educational ontology, cross-year metadata, answer explanations, and dynamic robustness testing. One paper explicitly argues that future benchmarks should include **structured knowledge systems**, not just flat labels, and that evaluation should move beyond **static test sets** because contamination can distort conclusions [2504.05782]. The later paper extends this argument by positioning MDK12-Bench not merely as a static leaderboard dataset but as a diagnostic resource for **robustness**, **generalization**, **knowledge use**, and **AI-assisted education** [2508.06851].

Source: https://www.emergentmind.com/topics/mdk12-bench