Qwen2.5-VL-DCG-3B: Dynamic Chart Generation Model
- Qwen2.5-VL-DCG-3B is a specialized 3B-parameter multimodal model focused on dynamic chart generation through code-based training.
- Built on the Qwen2.5-VL-3B backbone, it employs LoRA adaptations and a two-stage (SFT and RL) training process to optimize execution-sensitive animated chart code.
- Empirical benchmarks indicate robust performance across D2C, S2C, and V2C tasks, outperforming larger models on code execution and visual fidelity.
Qwen2.5-VL-DCG-3B is a specialized 3B-parameter multimodal LLM for code-based dynamic chart generation, introduced in “OpusAnimation: Code-Based Dynamic Chart Generation” as a task-adapted derivative of Qwen2.5-VL-3B. It is trained to accept text prompts and videos and to generate long HTML + JavaScript programs that render animated charts. In the broader Qwen lineage, the generic Qwen2.5-VL technical report defines only the 3B, 7B, and 72B base vision-language variants and does not define any “DCG” branch; accordingly, the “DCG” suffix denotes a downstream specialization rather than a separately specified base architecture (Li et al., 2 Oct 2025, Bai et al., 19 Feb 2025).
1. Provenance, naming, and task scope
Qwen2.5-VL-DCG-3B belongs to the Qwen2.5-VL family but is not part of the base model taxonomy established in the Qwen2.5-VL technical report. That report releases Qwen2.5-VL-3B, Qwen2.5-VL-7B, and Qwen2.5-VL-72B, and explicitly notes that the name “Qwen2.5-VL-DCG-3B” does not appear anywhere in the family definition. The DCG model therefore enters the literature as a task-specific derivative rather than as a separately documented architectural branch (Bai et al., 19 Feb 2025).
Within OpusAnimation, dynamic chart generation (DCG) is defined as producing code-rendered animated visualizations as charts. The model is intended to extend multimodal LLMs beyond static chart generation and chart understanding toward animation-rich visualizations that require both temporal perception and long-form code synthesis. The paper frames DCG through three task dimensions: Simple Text-to-Chart (S2C), Detailed Text-to-Chart (D2C), and Video-to-Chart (V2C), with Qwen2.5-VL-DCG-3B trained specifically to improve robustness on these settings (Li et al., 2 Oct 2025).
A common misconception is to treat the model as if it were a generic small Qwen2.5-VL checkpoint. The OpusAnimation paper instead positions it as a specialized MLLM whose target output is not captions or QA responses but executable dynamic chart code, and whose core challenge is not only semantic correctness but also runtime validity and visual faithfulness under animation constraints (Li et al., 2 Oct 2025).
2. Backbone and architectural basis
Qwen2.5-VL-DCG-3B is built on top of the open-source Qwen2.5-VL-3B architecture, and the DCG specialization keeps the vision encoder frozen while applying LoRA only to the LLM part during training. The underlying Qwen2.5-VL-3B family member is specified with a shared vision stack and a 3B-class language backbone: the common Vision Transformer uses hidden size 1280, 32 layers, 16 attention heads, intermediate size 3456, patch size 14, window size 112, and full-attention layers {7, 15, 23, 31}; the Vision-Language Merger outputs dimension 2048 for the 3B model; and the 3B LLM uses hidden size 2048, 36 layers, 2 KV heads, head size 128, intermediate size 4864, shared embeddings, vocabulary size 151,646, and 4.1T trained tokens (Bai et al., 19 Feb 2025, Li et al., 2 Oct 2025).
Because the DCG model inherits the Qwen2.5-VL-3B base, it also inherits the family’s multimodal design assumptions: dynamic resolution processing, absolute time encoding for video, a unified visual-text token stream, and support for images and videos within one decoder-only framework. This matters for DCG because V2C requires not only reading chart appearance but also recovering temporal patterns such as staged object appearance, animation speed, and easing behavior (Bai et al., 19 Feb 2025).
OpusAnimation does not introduce extra detection heads, code-specific decoders, or separate video modules. The specialization is achieved entirely through post-training. During SFT, the LoRA configuration uses rank and with learning rate ; during RL, LoRA uses , , and learning rate . The vision encoder and adapter remain frozen in both stages, and generated outputs are capped at 8192 tokens, reflecting the length of HTML/JavaScript chart programs (Li et al., 2 Oct 2025).
This design suggests a deliberate division of labor. The inherited Qwen2.5-VL backbone contributes general multimodal perception, while DCG post-training teaches the model to map those percepts into fragile, long, execution-sensitive code artifacts. A plausible implication is that the base model’s visual competence is treated as sufficient, and the principal bottleneck is the decoder’s code-generation policy rather than low-level vision.
3. Benchmark and dataset design
The DCG specialization is anchored in two resources: DCG-8K, the training corpus, and DCG-Bench, the benchmark derived from its test split. Each data point contains a chart program, its rendered video, textual descriptions at two granularities, extracted data sequences, and QA pairs for both code and rendered output (Li et al., 2 Oct 2025).
| Task | Input | Target capability |
|---|---|---|
| D2C | Data sequence + detailed text | Fine-grained instruction following |
| S2C | Data sequence + short text | Zero-shot generalization from underspecified prompts |
| V2C | Data sequence + reference video | Video understanding and video-to-code translation |
DCG-8K contains 8,000 modified HTML code samples derived from 175 seed templates spanning 18 chart categories, with 20 chart types overall and reference code averaging more than 2000 tokens. The split is 5K training—subdivided into 4K for SFT and 1K for RL—plus 2.3K validation and 700 test samples, the latter forming DCG-Bench (Li et al., 2 Oct 2025).
The construction pipeline is deliberately programmatic. Seed JavaScript templates are crawled from ECharts, normalized into HTML snippets, then modified by Gemini-2.5-Pro through random changes in data, layout, style, speed, easing, and effects. Rendered videos are produced with Timesnap and FFmpeg as 2-second, 24 FPS animations. The same Gemini model is then used to generate a detailed description , a simple description , a data sequence , and about 10 code QA pairs plus 10 video QA pairs per sample (Li et al., 2 Oct 2025).
A crucial design choice is that S2C is not used for training. It functions as a zero-shot generalization test, probing whether a model trained on D2C and V2C can infer reasonable animation structure from short, realistic prompts. This setup makes S2C a sharper test of abstraction than of memorization (Li et al., 2 Oct 2025).
4. Supervised and reinforcement training recipe
The training recipe is explicitly two-stage. First, the base Qwen2.5-VL-3B receives multi-modal supervised fine-tuning on D2C and V2C. Second, it undergoes GRPO-based reinforcement learning with a Joint-Code-Visual Reward. The intermediate checkpoint is denoted Qwen2.5-DCG-3B, and the final checkpoint Qwen2.5-DCG-3B0 is the model referred to as Qwen2.5-VL-DCG-3B (Li et al., 2 Oct 2025).
The SFT stage is not merely a cold start for code generation; it is also a modality-balancing step. The paper reports that D2C-only SFT improves D2C most, V2C-only SFT improves V2C and even transfers somewhat to S2C, but the mixed D2C+V2C setting yields the best overall performance across D2C, S2C, and V2C. This indicates that multimodal supervision contributes directly to better chart-code abstraction, rather than merely improving video recognition (Li et al., 2 Oct 2025).
The RL stage is centered on Group Relative Policy Optimization (GRPO) with a reward that combines code-level and video-level correctness. For each query, the paper defines QA-based scores
1
and
2
then combines them as
3
If rendering fails, both scores are set to 4, introducing an explicit penalty for non-executable code (Li et al., 2 Oct 2025).
The GRPO objective is instantiated with group size 5, clipping parameter 6, and KL coefficient 7. Advantages are normalized within the sample group: 8 The important empirical finding is that reward decomposition matters. Pure code reward and pure video reward are both suboptimal; the best configuration is 9 in favor of code reward, which the paper interprets as evidence that process-level correctness dominates but output-level visual fidelity remains necessary (Li et al., 2 Oct 2025).
The paper also argues that RL generalizes where continued SFT overfits. Continuing SFT on the same 1K D2C subset for 8 epochs leads to degradation on all three tasks, whereas JCVR-GRPO trained only on D2C improves D2C and also transfers to S2C and V2C. This is presented as evidence for the claim that, in this setting, “SFT memorizes, RL generalizes.” A plausible implication is that long chart code behaves like a brittle program synthesis problem, so supervision alone saturates quickly while reward-based optimization better captures executable structure and temporal coherence (Li et al., 2 Oct 2025).
5. Empirical performance and observed behavior
The benchmark results show a large gap between the generic base model and the task-specialized derivative. On D2C, the base Qwen2.5-VL-3B achieves only 8.91% execution rate, whereas the SFT model Qwen2.5-DCG-3B0 reaches 87.36% execution rate with 1 and 2. The final RL-enhanced model improves further (Li et al., 2 Oct 2025).
| Task | Execution rate | 3 / 4 |
|---|---|---|
| D2C | 91.95% | 7.45 / 6.61 |
| S2C | 89.37% | 5.77 / 5.78 |
| V2C | 92.39% | 4.32 / 5.66 |
These results have different interpretations across tasks. On D2C, the model benefits from explicit instruction detail and surpasses both open-source multimodal and code-oriented baselines. On S2C, it operates in zero-shot mode and still matches the execution rate of Qwen2.5-VL-32B almost exactly while slightly exceeding that larger model on 5. On V2C, the model surpasses all open-source MLLMs and even exceeds Gemini-2.5-Pro on code correctness, though not on video score (Li et al., 2 Oct 2025).
The paper states that the final model achieves an average 8.31% performance gain over the best open-source MLLM across the three task dimensions. In D2C, it beats Qwen2.5-coder-32B and Qwen2.5-VL-32B despite using only 3B parameters. In V2C, it substantially outperforms Qwen2.5-VL-32B on 6, indicating that the specialization is not merely a small-model efficiency story but a task-aligned training story (Li et al., 2 Oct 2025).
The qualitative analyses emphasize two consistent behaviors. First, the model tends to generate well-structured, modular HTML/JS with correct chart options and staged animations. Second, it captures temporal ordering—for example, axes before bars, or labels after growth completion—well enough that QA-based judges and human raters prefer its outputs to those of much larger baselines. In a pairwise V2C user study, it achieves a 42% win rate against Qwen2.5-coder-32B. Error analysis also shows fewer undefined property and syntax errors than Qwen2.5-VL-32B, especially on D2C and S2C (Li et al., 2 Oct 2025).
6. Limitations, ambiguities, and significance
The model’s limitations are stated directly. The current method still leaves architecture and RL strategy underexplored; the reward model depends on Gemini-2.5-Pro, raising reproducibility concerns; the dataset remains tied to ECharts-derived templates; and the authors explicitly note that a CodeLLM backbone could further reduce syntax errors. These points imply that Qwen2.5-VL-DCG-3B is best understood as a strong first specialization rather than a saturated endpoint for the task (Li et al., 2 Oct 2025).
Another ambiguity concerns nomenclature. Since Qwen2.5-VL Technical Report does not define any DCG suffix, one should not infer a separate core architecture or deployment-grade compression scheme from the name alone. What is actually documented is a post-trained, LoRA-adapted specialization of Qwen2.5-VL-3B for dynamic chart generation. This distinction matters for model comparison: improvements stem from task-specific data, multimodal SFT, and reward-guided optimization, not from a newly described backbone family (Bai et al., 19 Feb 2025).
The broader significance of Qwen2.5-VL-DCG-3B lies in what it demonstrates about small multimodal models. It shows that a 3B Qwen2.5-VL checkpoint, when trained on aligned instruction-code-video triplets and optimized with a joint code/visual reward, can become competitive with or better than much larger open-source and proprietary systems on a narrowly defined but technically demanding multimodal code-generation problem. This suggests that, for DCG-like workloads, specialization and reward design matter more than raw parameter count. A plausible implication is that future progress will come less from scaling alone and more from improving video-to-code supervision, open reward models, and code-aware multimodal backbones (Li et al., 2 Oct 2025).