Papers
Topics
Authors
Recent
Search
2000 character limit reached

Aura: Consistent Multi-Subject Video Generation via VLM-Grounded Semantic Alignment

Published 5 Jul 2026 in cs.CV | (2607.04311v2)

Abstract: Subject-driven and multi-element video generation are central to controllable video synthesis, but existing methods still struggle to preserve identity consistency and model complex relationships among multiple subjects. In this paper, we propose Aura, a unified framework for high-fidelity and identity-consistent video generation. To better capture scene dynamics and subject interactions, we introduce AI director-level captions that provide dense and structured descriptions of video content. We further leverage a vision-LLM (VLM) with learnable queries to extract multimodal semantic features from textual and visual references, covering both global semantics and fine-grained visual cues. To bridge the representational gap between the VLM and the Diffusion Transformer (DiT), we design a two-stage alignment strategy that progressively maps VLM features into the DiT feature space. For visual conditioning, we adopt token concatenation to inject reference information directly into the generation process. To distinguish heterogeneous subject types and reduce common copy-paste artifacts, we develop a subject-aware RoPE-Shift mechanism. To further differentiate reference images of different categories, we introduce subject-aware learnable tokens. In addition, we introduce Memory Tokens to balance the training signal across examples with different numbers of reference subjects. During inference, Progressive-APG (Adaptive Prompt Guidance) further alleviates oversaturation and improves semantic alignment with user prompts. Finally, we build a high-quality video-subject image dataset through a dedicated data construction pipeline. Extensive experiments show that our method achieves state-of-the-art performance on both single-subject generation and more challenging multi-element scenarios.

Summary

  • The paper introduces a novel diffusion transformer architecture that decouples identity, scene, and motion for multi-subject video generation.
  • The framework incorporates token-concatenated reference injection and subject-aware RoPE shift to ensure semantic alignment and robust identity preservation.
  • Empirical results demonstrate state-of-the-art performance with superior naturalness, motion amplitude, and user preference over existing T2V methods.

Aura: Consistent Multi-Subject Video Generation via VLM-Grounded Semantic Alignment

Motivation and Problem Statement

Controllable video generation demands precise appearance preservation for specified subjects, compositional scene generation, and consistent motion, all guided by structured prompts. Conventional text-to-video (T2V) diffusion models have made substantial progress, yet subject-driven and multi-element generation remains challenging due to weak identity consistency, copy-paste artifacts, and ambiguous attribute binding in multi-subject scenarios. Current approaches fail to robustly maintain subject identity, especially with in-the-wild references, and struggle with semantic binding and cross-modal alignment. The absence of explicit reference-aware supervision and modality alignment inevitably results in attribute leakage and degraded generative controllability.

Aura Architecture and Core Innovations

Aura introduces a unified framework based on a diffusion transformer (DiT) backbone, integrating a dual-stream semantic conditioning architecture. The system is engineered for arbitrary multi-reference controllability, explicitly decoupling identity, scene, and motion.

Aura’s architecture comprises several core modules:

  • Token-Concatenated Reference Injection: All references (human, object, scene) are encoded, mapped to DiT’s hidden dimensions, and concatenated along the sequence axis, enabling joint full-attention across video and reference tokens. Fixed slot counts and per-category learnable tokens (feature level) with asymmetric clean-timestep embeddings disambiguate reference roles and categories.
  • Subject-Aware RoPE Shift: Per-category spatial offsets on a 3D rotary position grid (RoPE) allocate disjoint “quadrants” for each reference category, resolving coordinate and collision ambiguity across heterogeneous references.

Figure 1

Figure 1: Aura architecture with token-concatenated reference injection, learnable meta-queries, and disjoint RoPE shifts for robust category separation.

  • Dual-Stream Semantic Conditioning: References and prompts are jointly processed via a frozen Qwen2.5-VL VLM and structured T5 embeddings. This multimodal extraction pathway uses learnable meta-queries and a two-stage alignment (sentence-level InfoNCE plus token-level Hungarian matching), allowing parameter-free shared-KV cross-attention in DiT.
  • Training Curriculum: A four-phase schedule—Coarse-Align, Fine-Align, Ref-Only, Joint-Mix—progressively adapts the DiT to reference-grounded conditioning and aligns VLM features with T5 for coherent multi-modality.
  • Inference with Norm-Only Progressive Adaptive Prompt Guidance (APG): Dual CFG axes (text, reference) are independently rescaled with schedule-dependent norm clipping, ensuring stable semantic and identity injection without guidance-saturation artifacts.

Data Pipeline and Curation

Aura trains on a high-quality, ∼\sim15M-scale dataset constructed through a dedicated pipeline operating on raw video sources. This pipeline includes:

  • Shot segmentation, quality filtering, and director-style re-captioning.
  • Three reference streams for each sample: human (identity-preserving, edited via I2I, filtered by ArcFace), object (pose-completed, inpainted, BLIP-2 filtered), and scene (foreground-erased, multi-view reconstructed).
  • Semantic editing and identity similarity filtering, breaking shortcut correlations and enforcing compositional generalization.

Figure 2

Figure 2: Data pipeline with shot segmentation, structured captions, and curated (clip, caption, reference set) tuples.

Curated reference samples demonstrate controlled variance in background, illumination, pose, viewpoint, and occlusion without sacrificing semantic fidelity.

Figure 3

Figure 3: Curated examples illustrating robust identity/scene preservation under significant low-level edits.

Figure 4

Figure 4: Diverse reference samples showing effective occlusion completion, appearance edits, and scene perturbations.

Empirical Evaluation

Quantitative experiments utilize OpenS2V-Eval and VLM-based metrics. Aura attains state-of-the-art Total scores as well as top marks in NaturalScore and AES, reflecting strong physical plausibility, subject fidelity, and aesthetic quality. Notably, Aura achieves maximal motion amplitude, outperforming near-static S2V baselines and demonstrating robust compositional scene editing. The dual-stream conditioning is critical: removing the VLM pathway degrades motion and naturalness, validating the necessity of multimodal grounding.

Qualitative results showcase Aura’s capacity for multi-element, multi-subject synthesis with consistent identity and prompt adherence.

Figure 5

Figure 5: Qualitative comparison showing Aura’s superior compositional and identity-consistent video generation.

User studies using the GSB protocol corroborate the automatic metrics, with Aura consistently preferred over all baselines, both T2V and S2V.

Figure 6

Figure 6: User preference statistics under GSB, with Aura achieving dominant Good rates over every competitor.

Ablation Study and Alignment Objective

Ablation experiments reveal that Hungarian matching in the alignment loss is pivotal for token-level discriminative binding between Qwen2.5-VL and T5, preventing collapse to uniform or modality-agnostic representations.

Figure 7

Figure 7: Hungarian matching dramatically improves reference-specific facial fidelity in plug-and-play alignment.

Figure 8

Figure 8: Pair-wise token similarity matrices showing grid structure only with Hungarian matching.

Assessment of progressive APG confirms that schedule-dependent norm-only clipping is superior to conventional projection-based guidance, directly mitigating late-stage artifacts and over-saturation.

Limitations and Future Directions

Aura relies on post-hoc VLM-T5 conditioning, which only approximates true joint distributions. Hyper-parameters for identity-hard-copy balance and curation pipeline throughput remain hand-tuned; adaptive controllers and faster I2I editors are promising lines of work. Risks regarding deepfakes, demographic bias, and IP imitation persist from the T2V backbone and require deployment safeguards.

Conclusion

Aura establishes a robust paradigm for compositional controllable video generation, combining token-concatenated references, subject-aware spatial separation, aligned multimodal conditioning, and norm-controlled APG. Large-scale data curation and explicit semantic alignment yield superior results across fidelity, naturalness, and identity metrics, with demonstrated generalization to complex, multi-subject scenarios. The architectural and training strategies in Aura constitute a technical benchmark for future video generation systems combining structured multimodal prompts, reference injection, and compositional reasoning.

Paper to Video (Beta)

No one has generated a video about this paper yet.

Whiteboard

No one has generated a whiteboard explanation for this paper yet.

Open Problems

We haven't generated a list of open problems mentioned in this paper yet.