---
title: '4Director: 3D Video Generation with Camera and Object Control'
url: https://www.emergentmind.com/topics/4director
type: topic
---

# 4Director: 3D Video Generation with Camera and Object Control

4Director is a video-generation system for precise camera and object-motion control using an explicit rigid 4D scene representation. Given a single image, the system reconstructs a static background point cloud and complete canonical meshes for marked objects, assigns one rigid $\mathrm{SE}(3)$ transformation to each object at every frame, renders the resulting scene as a depth video, and uses a Motion Adapter to generate RGB video with appearance, illumination, background completion, and non-rigid dynamics. Its central contribution is the use of complete 3D geometry rather than image-plane cues, sparse 3D points, bounding boxes, or incomplete object proxies [2610.02160].

## 1. Motivation and control paradigm

Video-generation systems commonly control motion through 2D points, masks, bounding boxes, or image-plane paths. These controls are ambiguous in depth and orientation. A shrinking bounding box, for example, may represent an object moving away from the camera or an object becoming smaller. A 2D trajectory does not distinguish object motion from camera-induced parallax, determine object orientation, or specify front–back occlusion.

Some systems introduce 3D trajectories, 3D boxes, tracked points, spheres, or one Gaussian per object. These representations reduce depth ambiguity but generally lack complete surface geometry. When a camera moves around an object or the object rotates, unseen surfaces must be regenerated, potentially causing view-dependent identity drift and inconsistent geometry.

4Director separates directorial control from generative completion:

- **Explicitly controlled variables**: camera motion and rigid 3D object motion.
- **Synthesized variables**: RGB appearance, illumination, background completion, view-consistent texture, non-rigid dynamics, and secondary motion.

The system treats a scene as a shared 4D representation rather than as independently generated frames. Complete object meshes are reconstructed once from the input image, and their per-frame transformations determine their spatial placement and visibility. This provides a single object-level control interface: one trajectory controls the complete object instead of requiring independent trajectories for many points.

The reference coordinate system is the camera coordinate frame of the input image. The first-frame camera transformation is set to the identity, $\mathbf E^1=\mathbf I$. The scene contains a background point cloud $\mathcal P$, a complete canonical mesh $\mathcal M_o$ for each marked object $o$, a camera intrinsic matrix $\mathbf K$, a camera trajectory $\{\mathbf E^t\}_{t=1}^{F}$, and an object trajectory $\{\mathbf T_o^t\}_{t=1}^{F}$ for every object.

## 2. Explicit rigid 4D scene representation

The complete scene state is

$$
\mathcal S=
\left(
\mathcal P,
\{\mathcal M_o\},
\{\mathbf T_o^t\},
\mathbf K,
\{\mathbf E^t\}
\right).
$$

The background point cloud and object meshes are canonical and static. Camera and object transformations vary over time.

For object $o$, the rigid transformation is

$$
\mathbf T_o^t=(\mathbf R_o^t,\mathbf t_o^t)\in\mathrm{SE}(3),
$$

where $\mathbf R_o^t\in\mathrm{SO}(3)$ and $\mathbf t_o^t\in\mathbb R^3$. A canonical mesh vertex $\mathbf v\in\mathcal M_o$ is transformed at frame $t$ according to

$$
\mathcal M_o^t=
\left\{
\mathbf R_o^t\mathbf v+\mathbf t_o^t
\mid
\mathbf v\in\mathcal M_o
\right\}.
$$

The initial object transformation is

$$
\mathbf T_o^1=(\mathbf I,\mathbf 0).
$$

For training clips, the observed position of a surface point is modeled as

$$
\mathbf X_i^t
=
\mathbf R_o^t\mathbf Y_i
+
\mathbf t_o^t
+
\mathbf r_i^t,
$$

where $\mathbf Y_i$ is the canonical position, $\mathbf R_o^t\mathbf Y_i+\mathbf t_o^t$ is the rigid component, and $\mathbf r_i^t$ is a non-rigid residual such as a swinging limb. Only the rigid component is rendered into the depth control. The Motion Adapter learns to synthesize the residual dynamics.

This decomposition makes the geometric control explicit while retaining the expressive capacity of a pretrained video model. A rigidly controlled person can still exhibit synthesized cloth motion, hair movement, or other non-rigid behavior, but those effects are not independently prescribed by the user.

### Camera projection and occlusion

The camera transformation $\mathbf E^t$ is a world-to-camera matrix. A world-space point $\mathbf X$ is mapped to camera coordinates by

$$
\tilde{\mathbf x}^{\,t}
=
\mathbf E^t
\begin{bmatrix}
\mathbf X\\
1
\end{bmatrix}.
$$

For camera coordinates $(X_c,Y_c,Z_c)$, perspective projection is

$$
\begin{bmatrix}
u\\
v\\
1
\end{bmatrix}
\sim
\mathbf K
\begin{bmatrix}
X_c/Z_c\\
Y_c/Z_c\\
1
\end{bmatrix}.
$$

At each frame, the background point cloud and transformed object meshes are projected using $(\mathbf K,\mathbf E^t)$. The renderer selects the nearest surface depth for every pixel. Pixels without geometric support are marked invalid. The resulting sequence is a depth video

$$
\mathbf C=\{\mathbf C_t\}_{t=1}^{F}.
$$

The depth representation encodes camera viewpoint, object translation and rotation, occlusion ordering, object disappearance and reappearance, and the projected geometry of complete objects. It does not directly encode RGB appearance, illumination, non-rigid motion, or unseen background regions.

In the implementation, depth is converted to an 8-bit grayscale control image. If $Z_t$ is the nearest camera-axis depth, then

$$
\mathbf C_t
=
\left\lfloor
255\,
\operatorname{clip}
\left(
\frac{1/Z_t-\alpha}{\beta-\alpha},
0,1
\right)
\right\rceil,
$$

where $[\alpha,\beta]$ is computed from the inverse-depth range across scene entities. Nearer surfaces are brighter, and invalid pixels receive the constant value 128.

## 3. Reconstruction and trajectory authoring

4Director reconstructs a scene from a single input image through a sequence of geometric and segmentation operations.

First, **MoGe-2** estimates monocular depth and camera intrinsics $\mathbf K$. The user marks objects with clicks, and **SAM 2** produces object masks. Pixels outside the masks are back-projected into a static background point cloud $\mathcal P$.

For each marked object, **Pixal3D** reconstructs a complete canonical textured mesh $\mathcal M_o$, including surfaces not visible in the input image. The reconstructed mesh is aligned to the observed object mask and depth using a similarity transform. Complete geometry is essential when an object rotates or the camera moves to reveal an unseen side: the newly visible surface originates from the same reconstructed mesh rather than being independently regenerated in each frame.

The user edits the scene in a 3D viewer and specifies camera and object trajectories by keyframes:

$$
\{\mathbf E^t\}_{t=1}^{F},
\qquad
\{\mathbf T_o^t\}_{t=1}^{F}.
$$

Because the camera and objects share a coordinate system, translations and rotations are expressed in 3D rather than inferred from projected silhouettes. A new object may be reconstructed from a separate reference image, inserted into the common scene, and assigned its own trajectory. Its textured mesh is rendered into the input image so that the composite becomes the first frame.

The authoring representation also specifies visibility. An object can move behind another object, leave the camera frustum, or later re-enter the frame while retaining the same canonical geometry and identity. These cases are difficult for controls based solely on points, boxes, or masks because those representations do not explicitly encode the complete surface or depth ordering.

The recovered scene uses an arbitrary-scale shared coordinate system rather than an absolutely metric reconstruction. MegaSaM estimates per-frame depth, a clip-level intrinsic matrix, and camera poses for training data, with MoGe-2 used as a depth prior and UniDepthV2 providing metric-depth cues.

## 4. Motion Adapter and video synthesis

The geometric scene is a control scaffold rather than the final RGB video. 4Director uses a **Motion Adapter** built on **Wan2.1-VACE-14B**. The VAE, text encoder, and main diffusion transformer remain unchanged. The Motion Adapter is a DiT-style conditioning branch initialized from the released VACE branch.

The VAE encodes the rendered depth-control video, the input image, and the text prompt. Eight context blocks, inserted at every fifth DiT block, process the control stream. Their outputs are projected and injected into the main video transformer through cross-attention and residual hints.

The conditional generation process can be summarized as

$$
p(\text{video}\mid
\mathbf C,
\text{input image},
\text{text}),
$$

where $\mathbf C$ constrains the spatial-temporal geometry while the pretrained video prior supplies visual content not specified by the depth representation.

The generated video is expected to preserve the geometry-induced motion while synthesizing:

- realistic RGB appearance;
- view-consistent texture and shape;
- illumination changes;
- realistic background completion;
- non-rigid dynamics;
- temporal continuity.

The training target contains rigid and non-rigid motion, whereas the depth control contains only the rigid component. This enables the adapter to learn residual dynamics without replacing the user-specified rigid transformation.

### Flow-matching objective

Let $\mathbf z$ be the latent representation of the target video and let

$$
\mathbf z_k=(1-\sigma_k)\mathbf z+\sigma_k\boldsymbol\epsilon
$$

be its flow-matching noisy state, where $\boldsymbol\epsilon$ is Gaussian noise. The Motion Adapter predicts a velocity field

$$
\boldsymbol\nu
\left(
\mathbf z_k,k;\mathbf C
\right),
$$

conditioned additionally on the input image and text. The training objective is

$$
\mathbb E_{\mathbf z,k,\boldsymbol\epsilon}
\left[
\lambda_k
\left\|
\boldsymbol\nu(\mathbf z_k,k;\mathbf C)
-
(\boldsymbol\epsilon-\mathbf z)
\right\|_2^2
\right].
$$

At inference, the reconstructed scene is rendered into a depth video, and the adapter generates the final RGB sequence from the depth control, input image, and text prompt. New trajectories do not require retraining because authoring and training share the same depth-video interface.

## 5. RealCOD-Rigid dataset and training data

Training requires video clips paired with rigid 3D scene annotations. 4Director introduces **RealCOD-Rigid**, constructed from **RealCOD-25K**.

The source collection contains 25,318 monocular clips, each with 81 frames at $832\times480$ resolution and 16 frames per second, together with a text prompt and SAM3 masks for one or two objects. The automatic annotation pipeline produces 20,774 usable training clips, covering 29,338 objects in the source collection before filtering.

Each retained clip contains:

- a shared scene coordinate frame;
- camera intrinsics;
- per-frame camera poses;
- a background point cloud;
- complete canonical object meshes;
- per-frame rigid transformations;
- a rendered rigid depth-control video.

### Automatic annotation pipeline

**Camera and depth estimation**: MegaSaM estimates per-frame depth, a clip-level intrinsic matrix, and camera poses. MoGe-2 supplies a depth prior, and UniDepthV2 provides metric-depth cues. The first frame, with object masks removed, is back-projected into the background point cloud.

**Object reconstruction**: Pixal3D reconstructs each object from its first-frame crop. A robust similarity alignment places the canonical mesh into the scene coordinate system using pixel-to-mesh correspondences and depth. The implementation uses a 100k-face textured mesh, an Umeyama similarity fit, robust outlier rejection, and geometry and silhouette quality gates.

**Rigid-body tracking**: TAPIP3D tracks 3D points on each object. Because articulated regions may not move rigidly, the method estimates a stable core $\Omega_o$ and fits a robust rigid transformation per frame:

$$
(\mathbf R_o^t,\mathbf t_o^t)
=
\arg\min_{\mathbf R,\mathbf t}
\sum_{i\in\Omega_o}
\rho
\left(
\left\|
\mathbf R\mathbf Y_i+\mathbf t-\mathbf X_i^t
\right\|
\right).
$$

The implementation uses RANSAC, Tukey robust weighting, weighted Kabsch rotation updates, a robust geometric median for translation, and spatial balancing. The estimated transformation follows the stable object body rather than a moving limb.

**Rendering**: the aligned meshes are transformed according to the estimated rigid motions and rendered with the background point cloud. The original video is the target, while the rigid depth rendering is the training condition.

This construction formalizes the central learning problem: generate realistic video from complete geometry and rigid motion while supplying appearance, illumination, and non-rigid motion through the video prior.

## 6. Evaluation, results, and limitations

4Director is evaluated on 100 held-out RealCOD-25K clips. The compared systems are **MotionCtrl**, **Perception-as-Control**, **SymphoMotion**, and **VerseCrafter**. All methods receive the same underlying recovered camera and object motion, converted into the control format required by each baseline.

The reported metrics include FID, FVD, CLIP similarity, VBench-I2V dimensions, camera rotation and translation errors, recognition rate, and Identity-Gated IoU.

| Method | FID ↓ | FVD ↓ | CLIP ↑ | RotErr ↓ | TransErr ↓ | Recognition ↑ | IG-IoU ↑ |
|---|---:|---:|---:|---:|---:|---:|---:|
| MotionCtrl | 115.6 | 1653.3 | 27.5 | 4.74° | 0.231 | 32.6 | 18.8 |
| Perception-as-Control | 55.7 | 1454.6 | 30.2 | 7.00° | 0.228 | 80.5 | 40.8 |
| VerseCrafter | 55.0 | 484.7 | 31.7 | 6.16° | 0.174 | 94.8 | 54.8 |
| SymphoMotion | 45.8 | 405.4 | 30.9 | 3.89° | 0.132 | 90.2 | 52.9 |
| **4Director** | **44.1** | **370.4** | **31.8** | **3.65°** | **0.122** | **94.9** | **60.4** |

The paper reports that 4Director reduces FID by 3.7% relative to SymphoMotion and FVD by 8.6%. It obtains the lowest camera rotation and translation errors, improves IG-IoU by 10.2% relative to VerseCrafter, and achieves a recognition rate comparable to VerseCrafter while providing stronger spatial and pose adherence.

The distinction between recognition rate and IG-IoU is significant. Recognition rate measures whether the generated object remains identity-consistent. IG-IoU additionally measures whether that identity-consistent object occupies the prescribed location and pose.

### Identity-Gated IoU

Standard mask IoU can reward spatial overlap even when the generated object is the wrong object, structurally deformed, or identity-inconsistent. 4Director introduces **Identity-Gated IoU (IG-IoU)** to jointly evaluate prescribed motion and identity preservation.

For $N$ sampled visible frames, let $\mathrm{MaskIoU}_f$ be the generated-to-reference object-mask IoU and let $v_f\in\{0,1\}$ indicate whether the object remains identity-consistent and free of major structural failure. The recognition rate is

$$
\mathrm{RecognitionRate}
=
\frac{1}{N}\sum_{f=1}^{N}v_f.
$$

IG-IoU is

$$
\mathrm{IG\text{-}IoU}
=
\frac{1}{N}
\sum_{f=1}^{N}
v_f\,\mathrm{MaskIoU}_f.
$$

An invalid frame contributes zero even if its segmentation overlaps the prescribed position. SAM3 supplies masks, and Qwen3-VL judges identity consistency, structural failure, both, or uncertainty. The evaluation samples 16 uniformly spaced visible frames when possible.

### Ablation of control representations

| Control | FID ↓ | FVD ↓ | RotErr ↓ | TransErr ↓ | Recognition ↑ | IG-IoU ↑ |
|---|---:|---:|---:|---:|---:|---:|
| 2D box | 58.3 | 608.8 | 9.61° | 0.233 | 95.0 | 48.5 |
| 3D box | 54.7 | 556.0 | 5.00° | 0.142 | 93.3 | 50.3 |
| 3D mesh without rotation | 53.0 | 485.6 | 4.84° | 0.147 | 91.9 | 53.0 |
| **Rigid 3D geometry** | **44.1** | **370.4** | **3.65°** | **0.122** | **94.9** | **60.4** |

The ablation attributes the performance difference to the successive addition of depth, orientation, surface geometry, and per-frame occlusion. A 2D box lacks depth and orientation; a 3D box encodes pose but not surface geometry; and a non-rotating mesh cannot expose the correct surface under orientation changes.

### Computational cost and limitations

The full Motion Adapter contains approximately 3.049 billion trainable parameters. Training uses three epochs on 24 GPUs with global batch size 24, $832\times480$ resolution, 81 frames, AdamW, peak learning rate $5\times10^{-5}$, and 25 warmup steps. Inference uses 20 flow-matching sampling steps, classifier-free guidance scale 5.0, 81 frames, 16 frames per second, $832\times480$ resolution, fixed seed 42, and hint scale 1.0. A rank-16 LoRA variant trains approximately 15.3 million parameters and remains competitive, although the full adapter performs better.

The principal limitation is rigid per-object control. A single transformation $\mathbf T_o^t$ cannot prescribe articulated motion. A breakdancer represented as one rigid object may retain the first-frame handstand pose rather than performing independently controlled limb movements. The generator can synthesize non-rigid motion, but the user cannot directly control joints, limbs, or articulated parts.

Additional limitations include dependence on single-image reconstruction quality, arbitrary rather than metric scene scale, learned background completion when motion reveals unseen regions, evaluator dependence in IG-IoU, and the inability of silhouette overlap to detect every orientation error for approximately front–back-symmetric objects. Automatic annotation can also fail because of poor depth, unstable tracking, incomplete reconstruction, or an unreliable rigid core.

4Director therefore occupies a specific position in controllable video generation: it provides explicit, object-level, rigid 3D direction while delegating visual completion and residual dynamics to a pretrained video world model. Its principal future direction is the extension of the scene representation with articulated parts, enabling finer-grained motion control beyond whole-object rigid trajectories.

Source: https://www.emergentmind.com/topics/4director