DCG-8K: Dynamic Chart Generation Dataset
- DCG-8K is a curated dataset of 8,000 dynamic chart examples that pairs executable HTML+JavaScript code with detailed textual instructions and dual QA annotations.
- It emphasizes chart fidelity and temporal accuracy by combining fine-grained animation details with both code-based and video-based quality evaluations.
- Designed as the core of DCG-Bench, the dataset supports three benchmark tasks—Detailed Text-to-Chart, Simple Text-to-Chart, and Video-to-Chart—for assessing MLLM performance.
DCG-8K is a curated corpus of 8,000 dynamic-chart examples introduced as the core dataset behind DCG-Bench, the Dynamic Chart Generation Benchmark, for evaluating multimodal LLMs on code-based dynamic chart generation (Li et al., 2 Oct 2025). In this setting, Dynamic Chart Generation (DCG) involves producing code-rendered animated visualizations as charts. Each DCG-8K sample pairs executable HTML+JavaScript chart code with its rendered animation, two levels of textual instruction, and question-answer annotations from both the code and video perspectives. The dataset is designed to support three benchmark tasks—Detailed Text-to-Chart, Simple Text-to-Chart, and Video-to-Chart—and to expose the gap between successful rendering and faithful recovery of fine-grained animation behavior (Li et al., 2 Oct 2025).
1. Definition and benchmark role
DCG-8K is defined as a high-quality dataset of 8,000 code-rendered animated chart examples designed to measure and drive progress in code-based dynamic chart generation (Li et al., 2 Oct 2025). Each sample includes an HTML+JavaScript snippet that, when executed, produces a smooth 2-second, 24 FPS animation of a chart; a detailed natural-language description ; a simple summary ; ten code-focused QA pairs ; and ten video-focused QA pairs .
The held-out test split of 700 samples is used as DCG-Bench to rigorously evaluate MLLMs on three complementary tasks: Detailed Text-to-Chart (D2C), Simple Text-to-Chart (S2C), and Video-to-Chart (V2C). Within the broader paper, DCG-8K is described as the heart of DCG-Bench, and the benchmark is positioned as the first benchmark evaluating MLLM capability on dynamic chart generation tasks from those three dimensions (Li et al., 2 Oct 2025).
The benchmark framing is important because recent advances in MLLMs are characterized as having significantly improved capability on static chart generation and comprehension, while dynamic chart generation and understanding remain underexplored. DCG-8K is therefore intended both as an evaluation substrate and as a training resource.
2. Dataset composition, splits, and sources
DCG-8K contains 8,000 dynamic chart instances. The splits are Train: 5,000, Validation: 2,300, and Test (DCG-Bench): 700. Within the train split, 4,000 samples are allocated for Supervised Fine-Tuning and 1,000 for GRPO.
All seed templates—175 of them across 18 different chart types from Apache ECharts—were crawled, manually cleaned, then randomly modified with 1–10 animation edits per template. The modifications were used to yield diversity in data elements, layout, colors, speeds, easing functions, and entry/exit effects. Non-renderable code was filtered out; the remaining code snippets average over 2,000 tokens in length and cover 20 chart categories, including bar, line, scatter, gauge, sankey, and map (Li et al., 2 Oct 2025).
| Statistic | Value |
|---|---|
| Total samples | 8,000 |
| Train / Val / Test splits | 5,000 / 2,300 / 700 |
| Avg. code tokens per sample | ≈2,100 tokens |
| Avg. detailed description length | ≈45 words |
| Avg. simple description length | ≈12 words |
| QA pairs per sample | 10 code + 10 video |
| Chart categories | 20 |
| Seed templates from ECharts | 175 (18 families) |
These statistics indicate that the dataset is not limited to a small number of canonical chart forms. A plausible implication is that the benchmark stresses both structural chart diversity and temporal animation diversity, since variation is introduced not only in visual appearance but also in motion parameters such as speeds, easing functions, and entry/exit effects.
3. Annotation schema and sample structure
Each data point consists of the 5-tuple . The detailed text specifies the animation at a fine level of granularity, while the simple text provides a shorter summary. The code is executable chart code, and the video 0 is the rendered animation derived from that code. The paired QA sets evaluate conformance from two distinct viewpoints: whether the code obeys the detailed specification, and whether the rendered animation exhibits the intended behavior.
The example in the source material illustrates the format. The detailed text is: “Create a bar chart with three bars whose heights rise sequentially: bar A from 0→50 in 0.8 s, bar B from 0→30 in 0.6 s after A finishes, bar C from 0→40 in 0.5 s; axes and grid lines appear instantly at start.” The corresponding simple text is: “Animate three bars rising one after another with axes visible immediately.” The abbreviated code snippet uses Apache ECharts with three bar-series entries configured by staggered animationDelay and animationDuration values, and the rendered video is described as a 2 s MP4 showing axes then bars A→B→C (Li et al., 2 Oct 2025).
The QA annotations are binary in form. An example 1 pair asks, “Does series 2 begin its animation exactly when series 1 ends?” with answer “Yes.” An example 2 pair asks, “Are gridlines visible before any bar movement?” with answer “Yes.” Because every sample includes ten code-focused and ten video-focused QA pairs, the dataset encodes supervision not only over final artifacts but over temporal and causal details of the animation sequence.
4. Task formulation in DCG-Bench
The benchmark defines three task categories. Given a sample 3, models must produce code 4 whose rendering 5 matches both the data sequence and the description.
Detailed Text-to-Chart (D2C) uses as input the data 6 plus the detailed text 7, which specifies chart type, object shapes, color transitions, easing, and precise timings. The requirement is to faithfully follow all low-level instructions.
Simple Text-to-Chart (S2C) uses as input the data 8 plus the simple summary 9, characterized as human-style shorthand. The requirement is to infer missing details, such as default durations, in order to produce a plausible, polished animation.
Video-to-Chart (V2C) uses as input the data 0 plus a reference animation video 1. The requirement is to visually parse motion—including order, speed, and easing—and reconstruct equivalent code (Li et al., 2 Oct 2025).
The distinction among these tasks is methodologically significant. D2C emphasizes literal adherence to explicit instructions; S2C evaluates underspecified generation; and V2C tests whether a model can infer executable animation logic from visual evidence. This suggests that DCG-8K is structured to probe complementary failure modes rather than a single notion of chart-generation quality.
5. Evaluation protocol and metrics
DCG-8K uses two complementary, QA-based metrics. The first is Execution Pass Rate, a Boolean indicator of whether candidate code 2 renders a nonblank animation. The second consists of QA-Based Scores over the code and the rendered video.
With 3 and 4, the benchmark defines
5
and
6
where 7 returns 1 if the generated artifact satisfies the question and 0 otherwise (Li et al., 2 Oct 2025).
This evaluation design separates renderability from fidelity. A model may generate code that executes successfully yet still fail to satisfy questions about temporal ordering, appearance timing, easing, or visibility conditions. The reported benchmark behavior underscores this distinction: execution success alone is not treated as sufficient evidence that a model has reproduced the intended dynamic chart.
6. Data curation pipeline and code transformation
The source material includes an illustrative code transformation from the data-curation pipeline, identified as Pipeline Stage 2 and Stage 3. In the example, Gemini-2.5-Pro is prompted with an ECharts bar-chart template and a set of animation edits such as “bars fade in over 1 s,” “x-axis drawn first,” and “y-axis dashed.” The model returns an HTML snippet 8 containing ECharts configuration fields such as animationDelay, animationDuration, animationEasing:'linear', and animationEasingUpdate:'cubicOut' (Li et al., 2 Oct 2025).
Stage 3 then headlessly renders this code into video 9 for subsequent QA extraction and human-readable description generation. In the dataset description, this pipeline is the mechanism by which instruction–code–video triplets are assembled: the template is edited into executable animated code, the animation is rendered into a video artifact, and annotations are attached for both code evaluation and video evaluation.
A plausible implication is that the dataset construction preserves a tight linkage between symbolic program structure and observable motion. That linkage is central to the benchmark’s dual evaluation scheme, because the same sample can be interrogated from the code side and from the rendered-output side.
7. Reported benchmark behavior and significance
The benchmark results reported for DCG-Bench show that strong execution performance does not necessarily imply strong fine-grained QA performance (Li et al., 2 Oct 2025). Among proprietary models, GPT-4.1 and Claude 3.7 achieve approximately 94–95% execution success in D2C and S2C and approximately 95–96% in V2C, but lag on fine-grained QA scores, with 0 in S2C and 1 in V2C.
For the best open-source MLLM prior to the reported method, Qwen2.5-VL-32B averages the following scores:
| Task | Model | Metrics |
|---|---|---|
| D2C | Qwen2.5-VL-32B | Exec 88.2%, 2, 3 |
| S2C | Qwen2.5-VL-32B | Exec ≈89.5%, 4, 5 |
| V2C | Qwen2.5-VL-32B | Exec 94.1%, 6, 7 |
By contrast, the authors’ lightweight 3 B model, described as Qwen2.5-DCG-3B after SFT plus joint-code-visual GRPO, achieves:
| Task | Model | Metrics |
|---|---|---|
| D2C | Qwen2.5-DCG-3B | Exec 91.95%, 8 7.45, 9 6.61 |
| S2C (zero-shot) | Qwen2.5-DCG-3B | Exec 89.37%, 0 5.77, 1 5.78 |
| V2C | Qwen2.5-DCG-3B | Exec 92.39%, 2 4.32, 3 5.66 |
The paper states that this amounts to an average 8.31% performance gain across the three tasks compared to the best prior open-source model, and that the 3 B model shows on-par performance against proprietary models. Within the narrative of the benchmark, this is presented as evidence that DCG-8K and its associated QA-based evaluation can both expose and help close the gap in dynamic chart generation performance (Li et al., 2 Oct 2025).
One common misconception in interpreting these results would be to treat execution pass rate as the primary indicator of competence. The reported numbers argue against that interpretation: even systems with execution success in the mid-90% range can remain weak on questions about precise visual timing or code-level compliance. DCG-8K is therefore notable less for renderability alone than for the way it operationalizes fidelity through paired code-focused and video-focused QA.