UI2V-Bench: Semantic I2V Evaluation
- UI2V-Bench is an evaluation benchmark for image-to-video generation that focuses on semantic understanding through defined dimensions such as spatial, attribute, category, and reasoning.
- It employs specialized MLLM-based evaluation pipelines alongside human studies to quantify instance-level subject tracking and causally coherent scene progression.
- Empirical findings reveal that while closed-source models score higher in semantic metrics, challenges remain in achieving robust attribute binding and commonsense consistency.
UI2V-Bench is an evaluation benchmark for image-to-video generation that targets semantic understanding of the input image and reasoning about world knowledge, rather than only low-level visual quality or temporal smoothness. It is designed for image-to-video diffusion models that take both a text prompt and a single input image, and it evaluates whether generated videos correctly identify specific subjects in the image, preserve their attributes and identities, and produce outcomes consistent with physical laws and human commonsense. The benchmark defines four primary dimensions—spatial understanding, attribute binding, category understanding, and reasoning—and uses MLLM-based automatic evaluation supplemented by human studies (Zhang et al., 29 Sep 2025).
1. Definition and motivation
UI2V-Bench, short for “Understanding-based I2V Bench,” is a benchmark suite comprising a dataset, metrics, and evaluation pipelines for image-to-video models. Its central premise is that conventional image-to-video evaluation is incomplete when it focuses mainly on image quality, aesthetic quality, motion smoothness, temporal consistency, or global alignment scores. In the image-to-video setting, the model is conditioned on an input image containing concrete subjects, so evaluation must also determine whether the model correctly understands which subject is referenced, whether attributes remain bound to the right instance, and whether the generated dynamics conform to commonsense and physical causality (Zhang et al., 29 Sep 2025).
This emphasis distinguishes UI2V-Bench from benchmarks that primarily assess coarse-grained perceptual quality. The benchmark treats semantic correctness as a first-class evaluation target. Two videos can therefore appear similarly plausible according to generic quality metrics while differing substantially in whether the correct object moves, whether the correct person is animated, or whether a cause–effect instruction is followed in a logically coherent way. A plausible implication is that UI2V-Bench operationalizes a shift from evaluating “how good a video looks” to evaluating “whether the generated video understands what should happen.”
2. Benchmark structure and evaluation dimensions
UI2V-Bench defines four primary evaluation dimensions: spatial understanding, attribute binding, category understanding, and reasoning (Zhang et al., 29 Sep 2025).
| Dimension | Target capability | Representative focus |
|---|---|---|
| Spatial understanding | Localize the correct subject from spatial descriptors | left/right, up/down |
| Attribute binding | Bind attributes to the correct instance | clothing, emotion, color, size, shape |
| Category understanding | Distinguish object categories in multi-object scenes | object classification |
| Reasoning | Infer plausible outcomes from causes or conditions | social, physical, temporal, natural |
Spatial understanding tests whether the model can parse spatial configurations of multiple same-category subjects in the input image and animate only the specified subject. The benchmark uses cases such as multiple dogs, cats, or balls arranged linearly, with prompts like “the dog on the left sits down” or “the second cat from the top jumps.” The number of subjects ranges from 2 to 4, and the arrangements in this version are linear, either horizontal or vertical. Failure modes include moving the wrong subject, moving multiple subjects, or failing to move any subject at all (Zhang et al., 29 Sep 2025).
Attribute binding tests whether the model can identify the target among multiple same-category subjects by using attributes rather than position. For persons, the benchmark includes age, height, build, dressing, makeup, emotion, pose, and object-holding. For objects, it includes color, size, shape, material, pattern or textures, and state or condition. The task is not merely to recognize attributes in isolation; it is to bind them to the correct instance and animate only that instance while leaving others unchanged. This design isolates a frequent failure mode in conditional generation: attribute confusion or partial matching across multiple candidates (Zhang et al., 29 Sep 2025).
Category understanding shifts from multiple subjects of the same category to scenes containing different categories. The model must recognize the specified category, localize it, and maintain its identity over time. Example scenarios include images containing vegetables and fruits, where only the carrot should disappear, or scenes with a man, a coffee cup, and a laptop, where the correct object must be manipulated. The benchmark reports this dimension as “Object Classification” accuracy (Zhang et al., 29 Sep 2025).
Reasoning is divided into four sub-dimensions: Human Society, Physical Interactions, Temporal Changes, and Natural Environment. Here the prompt typically specifies a cause or condition, and the generated video is expected to depict the implied effect. Examples include releasing a bowstring, letting go of a balloon, leaving bananas outside for a week, or an animal interacting with a flower. The benchmark therefore evaluates not only whether an action is animated, but whether the generated consequence is logically appropriate. This suggests that UI2V-Bench evaluates a form of causal and commonsense consistency that is not reducible to subject tracking alone (Zhang et al., 29 Sep 2025).
3. Evaluation pipelines and metrics
UI2V-Bench uses two MLLM-based evaluation pipelines: an instance-level pipeline for spatial understanding, attribute binding, and category understanding, and a feedback-based reasoning pipeline for the reasoning dimension (Zhang et al., 29 Sep 2025).
The instance-level pipeline has four components. A Prompt Analyzer extracts subject keywords and action keywords from the text prompt. A Subject Segmentor, implemented with GroundingDINO and SAM2, localizes the subject in the input image and produces masks. An Action Descriptor based on VideoRefer tracks the masked subject across frames and produces a description of its actions over time. A Result Judge then compares the original prompt with the subject-focused video description and assigns a score. The stated motivation for using VideoRefer is that generic MLLMs are weaker at precise subject tracking, so the benchmark introduces a more specialized component for spatial-temporal object understanding (Zhang et al., 29 Sep 2025).
The feedback-based reasoning pipeline is designed to mitigate prompt-induced hallucination during evaluation. First, a video description model such as Tarsier2 generates a detailed neutral description of what happens in the video. Then an LLM constructs a question chain: a sequence of observable yes/no questions ordered in a causal or logical progression. In a firearm example, the chain may progress from finger motion, to recoil, to muzzle flash, to the inference that a bullet was fired. The evaluator uses the video description and this question chain to determine whether the causal sequence is supported. This design is intended to reduce the risk that an evaluator simply assumes the correct answer from the prompt text alone (Zhang et al., 29 Sep 2025).
The benchmark defines the overall “Image Understanding (Ours)” score as the average of the four major dimensions:
For human alignment analysis, the benchmark reports Kendall’s tau and Spearman’s rho between automatic rankings and human rankings. The reported average correlation over all dimensions is and . Dimension-specific correlations are higher for Spatial Understanding and Reasoning than for Attribute Binding, which provides an explicit indication that attribute binding is the hardest of the automatic metrics to align with human judgment (Zhang et al., 29 Sep 2025).
4. Dataset construction and experimental setup
UI2V-Bench contains approximately 500 carefully constructed text-image pairs. The majority of the images come from open-access stock sources including Unsplash and Pexels. A minority are synthetic images generated by GPT-4 for configurations that are difficult to obtain in real photographs, such as specific arrangements of objects. The benchmark does not rely on manual per-frame annotations; instead, it depends on careful case design and later validation through correlation with human evaluation (Zhang et al., 29 Sep 2025).
The data construction varies by dimension. Spatial understanding and attribute binding cases use multiple same-category subjects so that spatial or attribute cues are necessary for correct disambiguation. Category understanding uses images with multiple common objects from different categories. Reasoning cases are formulated as cause–effect scenarios, where the image provides an initial state and the prompt specifies an action or condition whose consequence should unfold in the generated video (Zhang et al., 29 Sep 2025).
The benchmark evaluates both open-source and closed-source image-to-video models. The open-source models are Wan2.1, HunyuanVideo, and CogVideoX 1.5-5B. The closed-source models are SeedDance 1.0 Pro and MinMax-Hailuo-02. One video is generated per prompt for each model using fixed settings chosen to reflect each model’s native configuration. The reported settings are 5-second 480p 24 fps for Wan2.1, 3-second 720p 24 fps for HunyuanVideo, 6-second 768×1360 8 fps for CogVideoX 1.5-5B, 5-second 580p 25 fps for SeedDance-1.0-pro, and 6-second 768p 24 fps for MinMax-Hailuo-02 (Zhang et al., 29 Sep 2025).
This fixed-setting protocol means that the benchmark compares models under their intended operating configurations rather than forcing a single standardized output format. A plausible implication is that the comparison prioritizes ecological validity over strict resolution normalization.
5. Empirical findings and human evaluation
The reported results show that semantic understanding and reasoning remain substantially weaker than generic quality metrics would suggest. On the overall “Image Understanding (Ours)” metric, HunyuanVideo scores 28.49, CogVideoX 32.93, Wan2.1 38.68, SeedDance 41.67, and Hailuo 46.30. On the Reasoning average, the corresponding scores are 34.81, 37.80, 43.18, 52.77, and 67.04. On Spatial understanding average, they are 24.24, 33.59, 37.76, 44.41, and 52.57. On Attribute Binding average, they are 34.91, 40.05, 43.74, 39.33, and 44.69. On Category understanding, reported as Object Classification accuracy, they are 20.00, 20.26, 30.04, 30.15, and 20.89 (Zhang et al., 29 Sep 2025).
These numbers establish several patterns. Closed-source models outperform open-source models overall, with Hailuo as the strongest reported model on Reasoning, Spatial understanding, and the overall Image Understanding score. Wan2.1 is the strongest among the open-source models across several dimensions. At the same time, absolute scores remain moderate: even the best reported Reasoning score is 67.04, and Category understanding remains around 20–30% across models. This indicates that robust instance-level semantic control and causal reasoning are not solved even when overall video quality is comparatively strong (Zhang et al., 29 Sep 2025).
The benchmark explicitly contrasts these results with conventional quality metrics. Image Quality scores range from 0.7066 to 0.7321, Aesthetic quality from 0.5642 to 0.6080, and Motion Smoothness from 0.9865 to 0.9943 across the evaluated models. These values are relatively similar across models and do not track the larger spread in Image Understanding scores. The benchmark therefore argues that visual quality and smoothness do not correlate with semantic correctness or reasoning alignment in image-to-video generation (Zhang et al., 29 Sep 2025).
Human evaluation uses a random subset of benchmark samples, with at least 8 human annotators per sample. For semantic understanding dimensions, annotators use a 1–5 scale ranging from “Very Poor” to “Excellent.” For reasoning, evaluators are shown the image, prompt, and expected target outcome so that they can assess both explicit prompt depiction and causal correctness. The automatic metrics show strong positive correlations with human rankings, particularly for Spatial Understanding, Category Understanding, and Reasoning. The reported values are , for Spatial Understanding; , for Attribute Binding; , for Category Understanding; and , 0 for Reasoning (Zhang et al., 29 Sep 2025).
6. Position in the benchmark landscape, limitations, and implications
UI2V-Bench is situated against benchmarks that emphasize video quality, temporal smoothness, or global alignment. In the broader video-generation literature, VBench++ provides a comprehensive and versatile benchmark suite for text-to-video and image-to-video, with dimensions such as subject consistency, motion smoothness, spatial relationship, and I2V-specific consistency measures (Huang et al., 2024). UI2V-Bench differs by centering its evaluation on fine-grained subject-specific semantic understanding and explicit reasoning. Where broader video benchmarks decompose video quality and alignment, UI2V-Bench targets whether the correct subject in the conditioning image is manipulated and whether the generated consequence follows commonsense or physical law (Zhang et al., 29 Sep 2025).
The benchmark also states several limitations. The diversity of evaluated image-to-video models is still limited, especially on the open-source side. The current spatial understanding tasks cover only linear arrangements such as left–right and up–down, leaving more complex layouts such as circular or grid configurations for future work. The semantic evaluation pipeline depends on open-vocabulary detection and segmentation through GroundingDINO and SAM2, so difficult references such as “the second knife from the top” or “the sitting man” may affect reliability. The MLLM-based evaluators can still exhibit bias or hallucination, although the benchmark attempts to mitigate this with question chains and specialized models, and the lower human correlation for Attribute Binding makes that limitation visible rather than implicit (Zhang et al., 29 Sep 2025).
The benchmark’s main implication is diagnostic rather than merely comparative. It provides a way to determine whether an image-to-video model is truly grounding prompts in the conditioning image and whether it can maintain correct instance-level semantics over time. It also exposes a gap between general perceptual realism and understanding-based generation. This suggests that future progress in image-to-video synthesis will require more than improvements in fidelity and temporal coherence; it will also require stronger instance-centric conditioning, better identity and attribute preservation, and mechanisms that support causal and commonsense reasoning during generation (Zhang et al., 29 Sep 2025).