Papers
Topics
Authors
Recent
Search
2000 character limit reached

4Director: 3D Video Generation with Camera and Object Control

Updated 5 October 2026
  • 4Director is a video-generation system that combines explicit 4D scene representation with a motion adapter to enable precise control of camera and object motion from a single static image, ensuring consistent and clear geometric information across frames.
  • The system reconstructs a background point cloud and complete canonical object meshes, then renders the resulting scene as a depth video to accurately capture rigid 3D geometry, orientation, and occlusion, streamlining its Motion Adapter to add RGB coloration, texture, and non-rigid motion.
  • Evaluation shows 4Director achieves the lowest FID, FVD, and translation errors, as well as a highly accurate identity preservation through Identity-Gated IoU modelling

4Director is a video-generation system for precise camera and object-motion control using an explicit rigid 4D scene representation. Given a single image, the system reconstructs a static background point cloud and complete canonical meshes for marked objects, assigns one rigid SE(3)\mathrm{SE}(3) transformation to each object at every frame, renders the resulting scene as a depth video, and uses a Motion Adapter to generate RGB video with appearance, illumination, background completion, and non-rigid dynamics. Its central contribution is the use of complete 3D geometry rather than image-plane cues, sparse 3D points, bounding boxes, or incomplete object proxies (Cao et al., 1 Oct 2026).

1. Motivation and control paradigm

Video-generation systems commonly control motion through 2D points, masks, bounding boxes, or image-plane paths. These controls are ambiguous in depth and orientation. A shrinking bounding box, for example, may represent an object moving away from the camera or an object becoming smaller. A 2D trajectory does not distinguish object motion from camera-induced parallax, determine object orientation, or specify front–back occlusion.

Some systems introduce 3D trajectories, 3D boxes, tracked points, spheres, or one Gaussian per object. These representations reduce depth ambiguity but generally lack complete surface geometry. When a camera moves around an object or the object rotates, unseen surfaces must be regenerated, potentially causing view-dependent identity drift and inconsistent geometry.

4Director separates directorial control from generative completion:

  • Explicitly controlled variables: camera motion and rigid 3D object motion.
  • Synthesized variables: RGB appearance, illumination, background completion, view-consistent texture, non-rigid dynamics, and secondary motion.

The system treats a scene as a shared 4D representation rather than as independently generated frames. Complete object meshes are reconstructed once from the input image, and their per-frame transformations determine their spatial placement and visibility. This provides a single object-level control interface: one trajectory controls the complete object instead of requiring independent trajectories for many points.

The reference coordinate system is the camera coordinate frame of the input image. The first-frame camera transformation is set to the identity, E1=I\mathbf E^1=\mathbf I. The scene contains a background point cloud P\mathcal P, a complete canonical mesh Mo\mathcal M_o for each marked object oo, a camera intrinsic matrix K\mathbf K, a camera trajectory {Et}t=1F\{\mathbf E^t\}_{t=1}^{F}, and an object trajectory {Tot}t=1F\{\mathbf T_o^t\}_{t=1}^{F} for every object.

2. Explicit rigid 4D scene representation

The complete scene state is

S=(P,{Mo},{Tot},K,{Et}).\mathcal S= \left( \mathcal P, \{\mathcal M_o\}, \{\mathbf T_o^t\}, \mathbf K, \{\mathbf E^t\} \right).

The background point cloud and object meshes are canonical and static. Camera and object transformations vary over time.

For object oo, the rigid transformation is

E1=I\mathbf E^1=\mathbf I0

where E1=I\mathbf E^1=\mathbf I1 and E1=I\mathbf E^1=\mathbf I2. A canonical mesh vertex E1=I\mathbf E^1=\mathbf I3 is transformed at frame E1=I\mathbf E^1=\mathbf I4 according to

E1=I\mathbf E^1=\mathbf I5

The initial object transformation is

E1=I\mathbf E^1=\mathbf I6

For training clips, the observed position of a surface point is modeled as

E1=I\mathbf E^1=\mathbf I7

where E1=I\mathbf E^1=\mathbf I8 is the canonical position, E1=I\mathbf E^1=\mathbf I9 is the rigid component, and P\mathcal P0 is a non-rigid residual such as a swinging limb. Only the rigid component is rendered into the depth control. The Motion Adapter learns to synthesize the residual dynamics.

This decomposition makes the geometric control explicit while retaining the expressive capacity of a pretrained video model. A rigidly controlled person can still exhibit synthesized cloth motion, hair movement, or other non-rigid behavior, but those effects are not independently prescribed by the user.

Camera projection and occlusion

The camera transformation P\mathcal P1 is a world-to-camera matrix. A world-space point P\mathcal P2 is mapped to camera coordinates by

P\mathcal P3

For camera coordinates P\mathcal P4, perspective projection is

P\mathcal P5

At each frame, the background point cloud and transformed object meshes are projected using P\mathcal P6. The renderer selects the nearest surface depth for every pixel. Pixels without geometric support are marked invalid. The resulting sequence is a depth video

P\mathcal P7

The depth representation encodes camera viewpoint, object translation and rotation, occlusion ordering, object disappearance and reappearance, and the projected geometry of complete objects. It does not directly encode RGB appearance, illumination, non-rigid motion, or unseen background regions.

In the implementation, depth is converted to an 8-bit grayscale control image. If P\mathcal P8 is the nearest camera-axis depth, then

P\mathcal P9

where Mo\mathcal M_o0 is computed from the inverse-depth range across scene entities. Nearer surfaces are brighter, and invalid pixels receive the constant value 128.

3. Reconstruction and trajectory authoring

4Director reconstructs a scene from a single input image through a sequence of geometric and segmentation operations.

First, MoGe-2 estimates monocular depth and camera intrinsics Mo\mathcal M_o1. The user marks objects with clicks, and SAM 2 produces object masks. Pixels outside the masks are back-projected into a static background point cloud Mo\mathcal M_o2.

For each marked object, Pixal3D reconstructs a complete canonical textured mesh Mo\mathcal M_o3, including surfaces not visible in the input image. The reconstructed mesh is aligned to the observed object mask and depth using a similarity transform. Complete geometry is essential when an object rotates or the camera moves to reveal an unseen side: the newly visible surface originates from the same reconstructed mesh rather than being independently regenerated in each frame.

The user edits the scene in a 3D viewer and specifies camera and object trajectories by keyframes:

Mo\mathcal M_o4

Because the camera and objects share a coordinate system, translations and rotations are expressed in 3D rather than inferred from projected silhouettes. A new object may be reconstructed from a separate reference image, inserted into the common scene, and assigned its own trajectory. Its textured mesh is rendered into the input image so that the composite becomes the first frame.

The authoring representation also specifies visibility. An object can move behind another object, leave the camera frustum, or later re-enter the frame while retaining the same canonical geometry and identity. These cases are difficult for controls based solely on points, boxes, or masks because those representations do not explicitly encode the complete surface or depth ordering.

The recovered scene uses an arbitrary-scale shared coordinate system rather than an absolutely metric reconstruction. MegaSaM estimates per-frame depth, a clip-level intrinsic matrix, and camera poses for training data, with MoGe-2 used as a depth prior and UniDepthV2 providing metric-depth cues.

4. Motion Adapter and video synthesis

The geometric scene is a control scaffold rather than the final RGB video. 4Director uses a Motion Adapter built on Wan2.1-VACE-14B. The VAE, text encoder, and main diffusion transformer remain unchanged. The Motion Adapter is a DiT-style conditioning branch initialized from the released VACE branch.

The VAE encodes the rendered depth-control video, the input image, and the text prompt. Eight context blocks, inserted at every fifth DiT block, process the control stream. Their outputs are projected and injected into the main video transformer through cross-attention and residual hints.

The conditional generation process can be summarized as

Mo\mathcal M_o5

where Mo\mathcal M_o6 constrains the spatial-temporal geometry while the pretrained video prior supplies visual content not specified by the depth representation.

The generated video is expected to preserve the geometry-induced motion while synthesizing:

  • realistic RGB appearance;
  • view-consistent texture and shape;
  • illumination changes;
  • realistic background completion;
  • non-rigid dynamics;
  • temporal continuity.

The training target contains rigid and non-rigid motion, whereas the depth control contains only the rigid component. This enables the adapter to learn residual dynamics without replacing the user-specified rigid transformation.

Flow-matching objective

Let Mo\mathcal M_o7 be the latent representation of the target video and let

Mo\mathcal M_o8

be its flow-matching noisy state, where Mo\mathcal M_o9 is Gaussian noise. The Motion Adapter predicts a velocity field

oo0

conditioned additionally on the input image and text. The training objective is

oo1

At inference, the reconstructed scene is rendered into a depth video, and the adapter generates the final RGB sequence from the depth control, input image, and text prompt. New trajectories do not require retraining because authoring and training share the same depth-video interface.

5. RealCOD-Rigid dataset and training data

Training requires video clips paired with rigid 3D scene annotations. 4Director introduces RealCOD-Rigid, constructed from RealCOD-25K.

The source collection contains 25,318 monocular clips, each with 81 frames at oo2 resolution and 16 frames per second, together with a text prompt and SAM3 masks for one or two objects. The automatic annotation pipeline produces 20,774 usable training clips, covering 29,338 objects in the source collection before filtering.

Each retained clip contains:

  • a shared scene coordinate frame;
  • camera intrinsics;
  • per-frame camera poses;
  • a background point cloud;
  • complete canonical object meshes;
  • per-frame rigid transformations;
  • a rendered rigid depth-control video.

Automatic annotation pipeline

Camera and depth estimation: MegaSaM estimates per-frame depth, a clip-level intrinsic matrix, and camera poses. MoGe-2 supplies a depth prior, and UniDepthV2 provides metric-depth cues. The first frame, with object masks removed, is back-projected into the background point cloud.

Object reconstruction: Pixal3D reconstructs each object from its first-frame crop. A robust similarity alignment places the canonical mesh into the scene coordinate system using pixel-to-mesh correspondences and depth. The implementation uses a 100k-face textured mesh, an Umeyama similarity fit, robust outlier rejection, and geometry and silhouette quality gates.

Rigid-body tracking: TAPIP3D tracks 3D points on each object. Because articulated regions may not move rigidly, the method estimates a stable core oo3 and fits a robust rigid transformation per frame:

oo4

The implementation uses RANSAC, Tukey robust weighting, weighted Kabsch rotation updates, a robust geometric median for translation, and spatial balancing. The estimated transformation follows the stable object body rather than a moving limb.

Rendering: the aligned meshes are transformed according to the estimated rigid motions and rendered with the background point cloud. The original video is the target, while the rigid depth rendering is the training condition.

This construction formalizes the central learning problem: generate realistic video from complete geometry and rigid motion while supplying appearance, illumination, and non-rigid motion through the video prior.

6. Evaluation, results, and limitations

4Director is evaluated on 100 held-out RealCOD-25K clips. The compared systems are MotionCtrl, Perception-as-Control, SymphoMotion, and VerseCrafter. All methods receive the same underlying recovered camera and object motion, converted into the control format required by each baseline.

The reported metrics include FID, FVD, CLIP similarity, VBench-I2V dimensions, camera rotation and translation errors, recognition rate, and Identity-Gated IoU.

Method FID ↓ FVD ↓ CLIP ↑ RotErr ↓ TransErr ↓ Recognition ↑ IG-IoU ↑
MotionCtrl 115.6 1653.3 27.5 4.74° 0.231 32.6 18.8
Perception-as-Control 55.7 1454.6 30.2 7.00° 0.228 80.5 40.8
VerseCrafter 55.0 484.7 31.7 6.16° 0.174 94.8 54.8
SymphoMotion 45.8 405.4 30.9 3.89° 0.132 90.2 52.9
4Director 44.1 370.4 31.8 3.65° 0.122 94.9 60.4

The paper reports that 4Director reduces FID by 3.7% relative to SymphoMotion and FVD by 8.6%. It obtains the lowest camera rotation and translation errors, improves IG-IoU by 10.2% relative to VerseCrafter, and achieves a recognition rate comparable to VerseCrafter while providing stronger spatial and pose adherence.

The distinction between recognition rate and IG-IoU is significant. Recognition rate measures whether the generated object remains identity-consistent. IG-IoU additionally measures whether that identity-consistent object occupies the prescribed location and pose.

Identity-Gated IoU

Standard mask IoU can reward spatial overlap even when the generated object is the wrong object, structurally deformed, or identity-inconsistent. 4Director introduces Identity-Gated IoU (IG-IoU) to jointly evaluate prescribed motion and identity preservation.

For oo5 sampled visible frames, let oo6 be the generated-to-reference object-mask IoU and let oo7 indicate whether the object remains identity-consistent and free of major structural failure. The recognition rate is

oo8

IG-IoU is

oo9

An invalid frame contributes zero even if its segmentation overlaps the prescribed position. SAM3 supplies masks, and Qwen3-VL judges identity consistency, structural failure, both, or uncertainty. The evaluation samples 16 uniformly spaced visible frames when possible.

Ablation of control representations

Control FID ↓ FVD ↓ RotErr ↓ TransErr ↓ Recognition ↑ IG-IoU ↑
2D box 58.3 608.8 9.61° 0.233 95.0 48.5
3D box 54.7 556.0 5.00° 0.142 93.3 50.3
3D mesh without rotation 53.0 485.6 4.84° 0.147 91.9 53.0
Rigid 3D geometry 44.1 370.4 3.65° 0.122 94.9 60.4

The ablation attributes the performance difference to the successive addition of depth, orientation, surface geometry, and per-frame occlusion. A 2D box lacks depth and orientation; a 3D box encodes pose but not surface geometry; and a non-rotating mesh cannot expose the correct surface under orientation changes.

Computational cost and limitations

The full Motion Adapter contains approximately 3.049 billion trainable parameters. Training uses three epochs on 24 GPUs with global batch size 24, K\mathbf K0 resolution, 81 frames, AdamW, peak learning rate K\mathbf K1, and 25 warmup steps. Inference uses 20 flow-matching sampling steps, classifier-free guidance scale 5.0, 81 frames, 16 frames per second, K\mathbf K2 resolution, fixed seed 42, and hint scale 1.0. A rank-16 LoRA variant trains approximately 15.3 million parameters and remains competitive, although the full adapter performs better.

The principal limitation is rigid per-object control. A single transformation K\mathbf K3 cannot prescribe articulated motion. A breakdancer represented as one rigid object may retain the first-frame handstand pose rather than performing independently controlled limb movements. The generator can synthesize non-rigid motion, but the user cannot directly control joints, limbs, or articulated parts.

Additional limitations include dependence on single-image reconstruction quality, arbitrary rather than metric scene scale, learned background completion when motion reveals unseen regions, evaluator dependence in IG-IoU, and the inability of silhouette overlap to detect every orientation error for approximately front–back-symmetric objects. Automatic annotation can also fail because of poor depth, unstable tracking, incomplete reconstruction, or an unreliable rigid core.

4Director therefore occupies a specific position in controllable video generation: it provides explicit, object-level, rigid 3D direction while delegating visual completion and residual dynamics to a pretrained video world model. Its principal future direction is the extension of the scene representation with articulated parts, enabling finer-grained motion control beyond whole-object rigid trajectories.

Definition Search Book Streamline Icon: https://streamlinehq.com
References (1)

Topic to Video (Beta)

No one has generated a video about this topic yet.

Whiteboard

No one has generated a whiteboard explanation for this topic yet.

Follow Topic

Get notified by email when new papers are published related to 4Director.