---
title: 'DyCrowd: 3D Crowd Reconstruction'
url: https://www.emergentmind.com/topics/dycrowd
type: topic
---

# DyCrowd: 3D Crowd Reconstruction

DyCrowd is a multi-stage optimization framework for **spatio-temporally consistent 3D crowd reconstruction** from **large-scene video**, designed to recover the **global positions, body poses, and shapes of hundreds of people** in a shared global coordinate system [2508.12644]. It was introduced to address two limitations attributed to prior large-scene crowd reconstruction settings: reconstruction from a **static image**, which lacks temporal consistency, and insufficient robustness to **severe and repeated occlusions**. Its defining mechanism is a **coarse-to-fine group-guided motion optimization** strategy that leverages **collective crowd behavior**, a **VAE-based human motion prior**, and the **Asynchronous Motion Consistency (AMC)** loss so that high-quality unoccluded motion segments can guide the recovery of occluded ones, even under **temporal desynchronization** and **rhythmic inconsistencies**. The framework is accompanied by **VirtualCrowd**, a virtual benchmark dataset for evaluating dynamic crowd reconstruction from large-scene videos [2508.12644].

## 1. Problem formulation and representation

DyCrowd takes as input a large-scene video with \(T\) frames and \(N\) people, and estimates for each person \(n\) a temporal sequence of SMPL parameters in a **global scene coordinate system**:
\[
\mathbf{Q}_n = \{Q_{n,t}\}_{t=1}^{T} = \{\boldsymbol{\theta}_{n,t}, \boldsymbol{\beta}_{n,t}, \boldsymbol{\gamma}_{n,t}, \boldsymbol{\tau}_{n,t}\}_{t=1}^{T}.
\]
Here, \(\boldsymbol{\theta}_{n,t}\) denotes **body pose parameters**, \(\boldsymbol{\beta}_{n,t}\) the **body shape parameters**, \(\boldsymbol{\gamma}_{n,t}\) the **global/root rotation**, and \(\boldsymbol{\tau}_{n,t}\) the **global/root translation**. These parameters are mapped by SMPL to mesh vertices and joints:
\[
[\mathbf{V}_{n,t}, \mathbf{J}^{\text{SMPL}}_{n,t}] = \mathcal{M}(\boldsymbol{\gamma}_{n,t}, \boldsymbol{\theta}_{n,t}, \boldsymbol{\beta}_{n,t}) + \boldsymbol{\tau}_{n,t}.
\]

The technical difficulty of this formulation lies in the simultaneous presence of **extreme scale variation**, **hundreds of people**, **frequent and long-duration occlusions**, and the need to preserve both **scene-scale geometric consistency** and **temporal continuity** across the full video. DyCrowd therefore does not treat frames independently. Instead, it performs optimization at both **frame level** and **segment level**, and it incorporates crowd-level regularities rather than relying solely on person-wise image evidence.

A central premise of the framework is that large scenes must be reconstructed in a **shared global coordinate system** with plausible inter-person spacing and stable long-range motion. This distinguishes the task from local person reconstruction or frame-wise 3D pose lifting. A plausible implication is that the method is as much a **scene-consistent motion recovery system** as a person-level mesh recovery pipeline.

## 2. Global crowd motion initialization

The first stage, **global crowd motion initialization**, produces an initial global motion estimate for each person. Because scale variation is substantial, each frame is split into multiple regions, and a top-down detector, **VitDet**, is run region-wise. From detections, 2D keypoints are extracted using **DWPose**, while initial local SMPL pose and shape are estimated using **HMR2.0** [2508.12644].

Scene calibration is then performed by estimating the **ground plane** and **camera intrinsics** using walking and standing priors, following a Crowd3D-style calibration procedure. DyCrowd also estimates **Human-scene Virtual Interaction Points (HVIPs)** in 2D and 3D. HVIP is defined as the projection of the body’s torso center onto the ground plane, and it serves as a global localization cue.

For temporal association, DyCrowd uses **PHALP** but modifies it in two ways: it replaces PHALP’s default position representation with **3D HVIP**, and it restricts matching to **spatially adjacent people** for efficiency. This produces a tracked sequence \(\bar{\mathbf{Q}}_n\) for each person, which is used as the starting point for later optimization.

This initialization stage is not yet the final reconstruction. Its role is to make subsequent optimization feasible in large scenes where ordinary frame-by-frame association and local-body fitting would be unstable. A common misconception is that DyCrowd directly regresses the final 3D crowd state from video. In fact, its design is explicitly **optimization-based** and depends on progressively refined intermediate estimates.

## 3. Fundamental individual optimization

After initialization, DyCrowd performs **fundamental individual optimization** to stabilize each person’s global position, root orientation, and body configuration. This stage has two subcomponents: **root optimization** and **SMPL optimization**.

Root optimization adjusts \(\boldsymbol{\tau}_{n,t}\) and \(\boldsymbol{\gamma}_{n,t}\) using the objective
\[
E_{\text{root}} = E_{\text{2D}} + E_{\text{hvip2D}} + E_{\text{contact}} + E_{\text{t-trans}}.
\]
The **2D reprojection loss** \(E_{\text{2D}}\) constrains projected 3D joints to agree with detected 2D keypoints. The **HVIP 2D loss** \(E_{\text{hvip2D}}\) enforces plausible human-ground interaction through the projected virtual interaction point. The **ground contact loss** \(E_{\text{contact}}\) penalizes unrealistic mesh-ground offsets, and the **temporal translation smoothness** term \(E_{\text{t-trans}}\) reduces jitter in root motion [2508.12644].

SMPL optimization then refines pose and shape using
\[
E_{\text{smpl}} = E_{\text{root}} + E_{\text{pose}} + E_{\text{shape}} + E_{\text{t-pose}}.
\]
Here, \(E_{\text{pose}}\) uses a **VPoser latent prior**, \(E_{\text{shape}}\) regularizes body shape, and \(E_{\text{t-pose}}\) imposes temporal smoothness on 3D joint trajectories.

The function of this stage is to anchor each person to the calibrated scene before segment-level motion reasoning begins. It addresses the ambiguity inherent in local per-frame estimates, especially for **global location** and **root orientation**. This suggests that DyCrowd treats scene grounding as a prerequisite for motion completion rather than as a byproduct of pose fitting.

## 4. Coarse-to-fine group-guided motion optimization

The core contribution of DyCrowd is its **coarse-to-fine group-guided motion optimization**, introduced specifically to handle **temporal instability** and **long-term occlusion** [2508.12644]. The framework exploits two empirical regularities: people with similar trajectories often exhibit similar motion patterns, and within a local crowd group there may be both **occluded** and **unoccluded** subjects.

### VAE-based Human Motion Prior Optimization

The first part of this stage is **VAE-based Human Motion Prior Optimization (VHMP-Optim)**. DyCrowd trains a **transformer-based VAE** as a motion prior on segments of length **64 frames**. Each segment is encoded into a latent variable
\[
\boldsymbol{z}_{n,s}^{\vartheta} = (\boldsymbol{z}_{n,s}^{p}, \boldsymbol{z}_{n,s}^{r}) \in \mathbb{R}^{256+128}.
\]
The local state includes joints, velocities, and pose parameters, while the global state includes root position, translation, and rotation, together with contact probabilities. The VAE is trained on **AMASS** using **KL regularization**, **\(L_2\) reconstruction losses**, and **occlusion augmentation**, including random joint occlusion, frame occlusion, partial body occlusion, and consecutive occlusion.

The optimization objective is
\[
E_{\text{motion}} = E_{\text{smpl}} + E_{\text{VAE}} + E_{\text{env}} + E_{\text{connect}}.
\]
The **VAE latent prior** regularizes motion segments, the **environment/contact term** encourages plausible contact heights and zero velocity at contact, and the **connection term** enforces continuity between adjacent segments.

This stage improves short-term plausibility and continuity, but the paper states that it is still insufficient for **long-term occlusion**. That limitation motivates the next component.

### Segment-level Group-guided Optimization

The second part is **Segment-level Group-guided Optimization (SG-Optim)**. Each person’s motion is partitioned into segments, and segments are clustered by **relative trajectory similarity** using **symmetric segment path distance** and **affinity propagation**. Because affinity propagation is used, the number of groups need not be fixed in advance.

DyCrowd computes a **segment confidence** to determine which segments are reliable and which require repair. Segments are labeled as **optimal** (\(\xi=1\)), **poor** (\(\xi=0\)), or not requiring optimization (\(\xi=-1\)). High-quality segments are then used to guide poor segments in two settings: **in-sequence**, where a person’s own reliable motion guides its occluded segment, and **cross-sequence**, where guidance comes from another person with similar motion.

The key loss in this stage is the **Asynchronous Motion Consistency (AMC)** loss:
\[
E_{\mathrm{AMC}} = \lambda_{\text{AMC}} \sum_{n,s} w\,\mathcal{L}_s(\tilde{\boldsymbol{\theta}}_{n,s}, \tilde{\boldsymbol{\theta}}_{n',s'}) \cdot \mathds{1}(\xi_{n,s}=0, \xi_{n',s'}=1).
\]
AMC uses **soft dynamic time warping (soft-DTW)** to compare motion segments at the sequence level rather than frame-by-frame. This is critical because similar crowd members may not be synchronized: one may step earlier, walk with a different rhythm, or exhibit a phase shift relative to another.

The final group-level objective is
\[
E_{\text{group}} = E_{\text{motion}} + E_{\text{AMC}}.
\]

AMC corrects another possible misunderstanding: DyCrowd’s group guidance does **not** assume that similar motions are temporally aligned. The method is expressly designed for **temporal desynchronization and rhythmic inconsistencies**, and soft-DTW is the mechanism that makes such guidance differentiable and optimization-compatible [2508.12644].

## 5. VirtualCrowd benchmark and evaluation protocol

DyCrowd is accompanied by **VirtualCrowd**, a synthetic benchmark for large-scene dynamic crowd reconstruction [2508.12644]. It was built using **Blender** with the **iCity3D** plugin for scene generation, **SynBody** human models, **DIMOS** motion generation, and Blender rendering. The main dataset contains **4 scenes**, each covering **more than 2500 square meters**, with **two configurations** per scene—**high-angle view** and **low-angle view**—for a total of **8 validation videos**. The videos are rendered at **7680 × 4320 (8K)** and **30 fps**, with a **total duration of 1600 frames**, **crowd size from 60 to 200 people**, **931 motion sequences**, and **186,200 poses**.

VirtualCrowd provides **2D joints**, **MOT tracking annotations**, **3D joints**, **3D positions**, and **SMPL-X parameters**. The paper also mentions **two additional sloped-scene videos** as an extension.

Evaluation uses several metrics with distinct roles. **PA-PPDS** measures spatial crowd distribution consistency in global space, **PCOD** evaluates ordinal depth correctness, **MPJPE** and **PA-MPJPE** measure 3D pose accuracy, **WA-MPJPE** and **W-MPJPE** evaluate sequence-level global motion accuracy under different alignment protocols, and **ACCEL** measures motion smoothness. Occlusion-specific evaluation additionally reports **MPJPE** and **PA-MPJPE** on **severely occluded instances**.

The reported implementation uses **PyTorch**, **RMSprop**, a learning rate of **0.01**, and stage-specific optimization iterations of **100** for root optimization, **150** for SMPL optimization, **200** for VAE motion prior optimization, and **200** for group-guided optimization. Runtime is reported as roughly **4 hours** for a scene with **100 people and 200 frames** on an **NVIDIA RTX 3090 with 128 GB memory**. This establishes that DyCrowd is a high-cost offline optimization pipeline rather than a real-time system.

## 6. Empirical performance, related distinctions, and limitations

On VirtualCrowd, DyCrowd is compared against **Crowd3D**, **GroupRec**, and **SLAHMR-Large**. In the main quantitative comparison, DyCrowd reports **PA-PPDS 89.10**, **PCOD 92.20**, **MPJPE 69.74**, **PA-MPJPE 48.57**, **WA-MPJPE 68.99**, **W-MPJPE 83.39**, and **ACCEL 15.72**; with ground-truth tracking, **DyCrowd\(^\star\)** reports **PA-PPDS 91.23**, **PCOD 95.38**, **MPJPE 68.81**, **PA-MPJPE 45.34**, **WA-MPJPE 65.91**, **W-MPJPE 80.34**, and **ACCEL 15.53** [2508.12644]. The paper states that DyCrowd significantly improves **pose accuracy** over the baselines and achieves the best or near-best **global arrangement metrics**. It also notes that **SLAHMR-Large** is smoother in terms of **ACCEL**, but may produce implausible sliding under occlusion.

Ablation results indicate that removing the **coarse-to-fine group-guided motion optimization** substantially worsens pose metrics and occlusion recovery. Removing **AMC** causes only a slight drop in overall global metrics but weakens recovery on occluded instances. The paper also compares its motion prior to **NeMF** and **DMMR-VAE**, reporting that its own prior gives the best reconstruction quality overall.

The framework has several stated limitations. It depends on **2D detection and tracking**; catastrophic failures at that stage can propagate into motion recovery. If a person’s **tracklet is interrupted by occlusion**, the method cannot recover that person’s motion. It can handle **multi-ground scenes** using region-level processing but struggles with **complex terrain interactions like stairs**. It is **not real-time**, and its group-guided recovery assumes that people in the same group have sufficiently similar motion to guide one another. The benchmark, **VirtualCrowd**, is **synthetic**, so real-world generalization is evaluated mainly **qualitatively on PANDA** rather than through equivalent full 3D ground truth.

Within the broader crowd-analysis literature, DyCrowd occupies a distinct position. Earlier dense-crowd work such as **flow segmentation** based on **FTLE**, **LCS**, and watershed transforms partitions scenes into coherent motion regions rather than reconstructing individual 3D humans [1506.04608]. Drone crowd-flow methods based on centroid density maps and inter-frame centroid matching focus on dense-group motion rather than body pose and shape [2301.04937]. Depth-guided counting systems such as **DigCrowd** divide scenes into **far-view** and **near-view** regions for counting in **EDOF** imagery rather than recovering dynamic 3D crowds [1803.02256]. This suggests that DyCrowd extends crowd analysis from **counting**, **segmentation**, and **flow estimation** to **scene-scale temporally consistent 3D reconstruction** of many individuals, while remaining constrained by the reliability of upstream detection, tracking, and scene calibration.

Source: https://www.emergentmind.com/topics/dycrowd