---
title: 'Qwen2.5-VL-DCG-3B: Dynamic Chart Generation Model'
url: https://www.emergentmind.com/topics/qwen2-5-vl-dcg-3b
type: topic
---

# Qwen2.5-VL-DCG-3B: Dynamic Chart Generation Model

Qwen2.5-VL-DCG-3B is a specialized 3B-parameter multimodal large language model for **code-based dynamic chart generation**, introduced in “OpusAnimation: Code-Based Dynamic Chart Generation” as a task-adapted derivative of **Qwen2.5-VL-3B**. It is trained to accept text prompts and videos and to generate long **HTML + JavaScript** programs that render animated charts. In the broader Qwen lineage, the generic **Qwen2.5-VL** technical report defines only the **3B, 7B, and 72B** base vision-language variants and does not define any “DCG” branch; accordingly, the “DCG” suffix denotes a downstream specialization rather than a separately specified base architecture [2510.03341] [2502.13923].

## 1. Provenance, naming, and task scope

Qwen2.5-VL-DCG-3B belongs to the Qwen2.5-VL family but is not part of the base model taxonomy established in the Qwen2.5-VL technical report. That report releases **Qwen2.5-VL-3B**, **Qwen2.5-VL-7B**, and **Qwen2.5-VL-72B**, and explicitly notes that the name **“Qwen2.5-VL-DCG-3B” does not appear anywhere** in the family definition. The DCG model therefore enters the literature as a **task-specific derivative** rather than as a separately documented architectural branch [2502.13923].

Within OpusAnimation, **dynamic chart generation (DCG)** is defined as producing **code-rendered animated visualizations as charts**. The model is intended to extend multimodal large language models beyond static chart generation and chart understanding toward **animation-rich visualizations** that require both temporal perception and long-form code synthesis. The paper frames DCG through three task dimensions: **Simple Text-to-Chart (S2C)**, **Detailed Text-to-Chart (D2C)**, and **Video-to-Chart (V2C)**, with Qwen2.5-VL-DCG-3B trained specifically to improve robustness on these settings [2510.03341].

A common misconception is to treat the model as if it were a generic small Qwen2.5-VL checkpoint. The OpusAnimation paper instead positions it as a **specialized MLLM** whose target output is not captions or QA responses but executable **dynamic chart code**, and whose core challenge is not only semantic correctness but also runtime validity and visual faithfulness under animation constraints [2510.03341].

## 2. Backbone and architectural basis

Qwen2.5-VL-DCG-3B is built on top of the **open-source Qwen2.5-VL-3B architecture**, and the DCG specialization keeps the **vision encoder frozen** while applying **LoRA** only to the **LLM part** during training. The underlying Qwen2.5-VL-3B family member is specified with a shared vision stack and a 3B-class language backbone: the common **Vision Transformer** uses hidden size **1280**, **32 layers**, **16 attention heads**, intermediate size **3456**, patch size **14**, window size **112**, and full-attention layers **{7, 15, 23, 31}**; the **Vision-Language Merger** outputs dimension **2048** for the 3B model; and the 3B **LLM** uses hidden size **2048**, **36 layers**, **2 KV heads**, head size **128**, intermediate size **4864**, shared embeddings, vocabulary size **151,646**, and **4.1T** trained tokens [2502.13923] [2510.03341].

Because the DCG model inherits the Qwen2.5-VL-3B base, it also inherits the family’s multimodal design assumptions: **dynamic resolution processing**, **absolute time encoding** for video, a unified visual-text token stream, and support for images and videos within one decoder-only framework. This matters for DCG because **V2C** requires not only reading chart appearance but also recovering temporal patterns such as staged object appearance, animation speed, and easing behavior [2502.13923].

OpusAnimation does not introduce extra detection heads, code-specific decoders, or separate video modules. The specialization is achieved entirely through post-training. During **SFT**, the LoRA configuration uses rank \(r=64\) and \(\alpha=128\) with learning rate \(1\times 10^{-4}\); during **RL**, LoRA uses \(r=32\), \(\alpha=64\), and learning rate \(1\times 10^{-5}\). The **vision encoder and adapter remain frozen** in both stages, and generated outputs are capped at **8192 tokens**, reflecting the length of HTML/JavaScript chart programs [2510.03341].

This design suggests a deliberate division of labor. The inherited Qwen2.5-VL backbone contributes general multimodal perception, while DCG post-training teaches the model to map those percepts into **fragile, long, execution-sensitive code artifacts**. A plausible implication is that the base model’s visual competence is treated as sufficient, and the principal bottleneck is the **decoder’s code-generation policy** rather than low-level vision.

## 3. Benchmark and dataset design

The DCG specialization is anchored in two resources: **DCG-8K**, the training corpus, and **DCG-Bench**, the benchmark derived from its test split. Each data point contains a chart program, its rendered video, textual descriptions at two granularities, extracted data sequences, and QA pairs for both code and rendered output [2510.03341].

| Task | Input | Target capability |
|---|---|---|
| D2C | Data sequence + detailed text | Fine-grained instruction following |
| S2C | Data sequence + short text | Zero-shot generalization from underspecified prompts |
| V2C | Data sequence + reference video | Video understanding and video-to-code translation |

DCG-8K contains **8,000 modified HTML code samples** derived from **175 seed templates** spanning **18 chart categories**, with **20 chart types overall** and reference code averaging **more than 2000 tokens**. The split is **5K training**—subdivided into **4K for SFT** and **1K for RL**—plus **2.3K validation** and **700 test** samples, the latter forming **DCG-Bench** [2510.03341].

The construction pipeline is deliberately programmatic. Seed JavaScript templates are crawled from **ECharts**, normalized into HTML snippets, then modified by **Gemini-2.5-Pro** through random changes in data, layout, style, speed, easing, and effects. Rendered videos are produced with **Timesnap** and **FFmpeg** as **2-second, 24 FPS** animations. The same Gemini model is then used to generate a **detailed description** \(t_d\), a **simple description** \(t_s\), a **data sequence** \(d\), and about **10 code QA pairs** plus **10 video QA pairs** per sample [2510.03341].

A crucial design choice is that **S2C is not used for training**. It functions as a **zero-shot generalization** test, probing whether a model trained on **D2C** and **V2C** can infer reasonable animation structure from short, realistic prompts. This setup makes S2C a sharper test of abstraction than of memorization [2510.03341].

## 4. Supervised and reinforcement training recipe

The training recipe is explicitly **two-stage**. First, the base Qwen2.5-VL-3B receives **multi-modal supervised fine-tuning** on **D2C** and **V2C**. Second, it undergoes **GRPO-based reinforcement learning** with a **Joint-Code-Visual Reward**. The intermediate checkpoint is denoted **Qwen2.5-DCG-3B\(_{s1}\)**, and the final checkpoint **Qwen2.5-DCG-3B\(_{s2}\)** is the model referred to as Qwen2.5-VL-DCG-3B [2510.03341].

The SFT stage is not merely a cold start for code generation; it is also a modality-balancing step. The paper reports that **D2C-only SFT** improves D2C most, **V2C-only SFT** improves V2C and even transfers somewhat to S2C, but the **mixed D2C+V2C setting** yields the **best overall performance** across **D2C, S2C, and V2C**. This indicates that multimodal supervision contributes directly to better chart-code abstraction, rather than merely improving video recognition [2510.03341].

The RL stage is centered on **Group Relative Policy Optimization (GRPO)** with a reward that combines code-level and video-level correctness. For each query, the paper defines QA-based scores
\[
S_{\text{code}}(q, c_g) = \frac{1}{N_c} \sum_{i=1}^{N_c} \mathrm{MLLM}_{\mathrm{eval}}(c_g, QA^{(i)}_{\text{code}})
\]
and
\[
S_{\text{video}}(q, v_g) = \frac{1}{N_v} \sum_{j=1}^{N_v} \mathrm{MLLM}_{\mathrm{eval}}(v_g, QA^{(j)}_{\text{video}}),
\]
then combines them as
\[
r_{jcv}(q, c, v) = w_{\text{code}} \cdot S_{\text{code}}(q,c) + w_{\text{video}} \cdot S_{\text{video}}(q,v).
\]
If rendering fails, both scores are set to **\(-0.25\)**, introducing an explicit penalty for non-executable code [2510.03341].

The GRPO objective is instantiated with group size **\(G=8\)**, clipping parameter **\(\epsilon=0.2\)**, and KL coefficient **\(\beta=0.04\)**. Advantages are normalized within the sample group:
\[
\hat{A}_{i,t} =
\frac{r_{jcv}^i - \mathrm{mean}(\mathbf{r}_{jcv})}
{\mathrm{std}(\mathbf{r}_{jcv})}.
\]
The important empirical finding is that reward decomposition matters. Pure code reward and pure video reward are both suboptimal; the best configuration is **\(8{:}2\)** in favor of **code reward**, which the paper interprets as evidence that **process-level correctness** dominates but **output-level visual fidelity** remains necessary [2510.03341].

The paper also argues that **RL generalizes where continued SFT overfits**. Continuing SFT on the same **1K D2C** subset for **8 epochs** leads to degradation on all three tasks, whereas **JCVR-GRPO trained only on D2C** improves D2C and also transfers to **S2C** and **V2C**. This is presented as evidence for the claim that, in this setting, **“SFT memorizes, RL generalizes.”** A plausible implication is that long chart code behaves like a brittle program synthesis problem, so supervision alone saturates quickly while reward-based optimization better captures executable structure and temporal coherence [2510.03341].

## 5. Empirical performance and observed behavior

The benchmark results show a large gap between the generic base model and the task-specialized derivative. On **D2C**, the base **Qwen2.5-VL-3B** achieves only **8.91%** execution rate, whereas the SFT model **Qwen2.5-DCG-3B\(_{s1}\)** reaches **87.36%** execution rate with \(S_{\text{code}}=6.73\) and \(S_{\text{video}}=5.78\). The final RL-enhanced model improves further [2510.03341].

| Task | Execution rate | \(S_{\text{code}}\) / \(S_{\text{video}}\) |
|---|---:|---:|
| D2C | **91.95%** | **7.45 / 6.61** |
| S2C | **89.37%** | **5.77 / 5.78** |
| V2C | **92.39%** | **4.32 / 5.66** |

These results have different interpretations across tasks. On **D2C**, the model benefits from explicit instruction detail and surpasses both open-source multimodal and code-oriented baselines. On **S2C**, it operates in zero-shot mode and still matches the **execution rate** of **Qwen2.5-VL-32B** almost exactly while slightly exceeding that larger model on **\(S_{\text{code}}\)**. On **V2C**, the model surpasses all open-source MLLMs and even exceeds **Gemini-2.5-Pro** on **code correctness**, though not on video score [2510.03341].

The paper states that the final model achieves an **average 8.31% performance gain** over the best open-source MLLM across the three task dimensions. In **D2C**, it beats **Qwen2.5-coder-32B** and **Qwen2.5-VL-32B** despite using only **3B parameters**. In **V2C**, it substantially outperforms **Qwen2.5-VL-32B** on \(S_{\text{code}}\), indicating that the specialization is not merely a small-model efficiency story but a task-aligned training story [2510.03341].

The qualitative analyses emphasize two consistent behaviors. First, the model tends to generate **well-structured, modular HTML/JS** with correct chart options and staged animations. Second, it captures **temporal ordering**—for example, axes before bars, or labels after growth completion—well enough that QA-based judges and human raters prefer its outputs to those of much larger baselines. In a **pairwise V2C user study**, it achieves a **42% win rate** against **Qwen2.5-coder-32B**. Error analysis also shows fewer **undefined property** and **syntax** errors than **Qwen2.5-VL-32B**, especially on D2C and S2C [2510.03341].

## 6. Limitations, ambiguities, and significance

The model’s limitations are stated directly. The current method still leaves **architecture and RL strategy underexplored**; the reward model depends on **Gemini-2.5-Pro**, raising reproducibility concerns; the dataset remains tied to **ECharts-derived templates**; and the authors explicitly note that a **CodeLLM backbone** could further reduce syntax errors. These points imply that Qwen2.5-VL-DCG-3B is best understood as a strong first specialization rather than a saturated endpoint for the task [2510.03341].

Another ambiguity concerns nomenclature. Since **Qwen2.5-VL Technical Report** does not define any **DCG** suffix, one should not infer a separate core architecture or deployment-grade compression scheme from the name alone. What is actually documented is a **post-trained, LoRA-adapted specialization of Qwen2.5-VL-3B** for dynamic chart generation. This distinction matters for model comparison: improvements stem from **task-specific data, multimodal SFT, and reward-guided optimization**, not from a newly described backbone family [2502.13923].

The broader significance of Qwen2.5-VL-DCG-3B lies in what it demonstrates about small multimodal models. It shows that a **3B** Qwen2.5-VL checkpoint, when trained on aligned **instruction-code-video** triplets and optimized with a **joint code/visual reward**, can become competitive with or better than much larger open-source and proprietary systems on a narrowly defined but technically demanding multimodal code-generation problem. This suggests that, for DCG-like workloads, **specialization and reward design** matter more than raw parameter count. A plausible implication is that future progress will come less from scaling alone and more from improving **video-to-code supervision**, **open reward models**, and **code-aware multimodal backbones** [2510.03341].

Source: https://www.emergentmind.com/topics/qwen2-5-vl-dcg-3b