- The paper introduces MMFA, a one-shot face-animation framework that uses unsupervised 3D keypoints, explicit scale modeling, and self-supervised expression features to independently control pose, translation, scale, and expression.
- MMFA achieves the best reported FID in same-identity and cross-identity tests—13.265 and 77.445, respectively—while matching the strongest same-identity CSIM score of 0.974, though it does not lead every pixel-level metric.
- The paper adds a VAE-based expression space for smooth interpolation without a driving video, while acknowledging higher computational cost and weaker eye-gaze tracking than some 2D keypoint methods.
Overview and motivation
MMFA (Motion Manipulation via unsupervised keypoint positioning in Face Animation) is a one-shot face animation framework that combines 3D unsupervised keypoint estimation with self-supervised representation learning to achieve explicit, independent control over facial motion attributes (2603.04302). The authors identify two shortcomings of prior keypoint-based pipelines. First, methods such as FOMM and MRAA estimate keypoints without semantic structure, so identity semantics remain entangled with motion information, preventing separate manipulation of facial movement. Second, Face-vid2vid's canonical-keypoint decomposition ignores camera perspective projection; because faces at different depths in VoxCeleb videos appear at different scales, the expression deformation δ is forced to absorb scale changes, coupling expression with facial scaling and rotation. The paper's central claim is that pose, translation, scaling, and expression can be decoupled under reasonable geometric assumptions while retaining the detail-transfer capability of unsupervised keypoints.
Keypoint decomposition pipeline
MMFA builds on three assumptions: the object centroid lies at the world/camera origin, the camera-to-image mapping follows orthographic projection, and the face is rigid under scale and rotation transforms. Under these assumptions, each face is modeled by K canonical keypoints pCk∈R3 (identity anchors), a rotation matrix R, a 2D translation t, a scalar scale factor f, and per-keypoint expression deformations δk. Target keypoints are computed as:
pSk=RSfS(pCk+δSk)+tS
The explicit scale factor f is the key departure from Face-vid2vid: it absorbs depth-induced size variation so that δ encodes only genuine expression deformation. Rotation is obtained from a pre-trained pose estimator rather than learned jointly, avoiding contamination of K0 by K1. Expression deformations are predicted by an encoder-decoder: a Bottleneck-based encoder compresses the input image into a 256-dimensional feature vector K2, which is concatenated with the canonical keypoints and decoded into K3. Because K4 depends on both the expression feature and the target identity's canonical keypoints, cross-identity reenactment transfers expressions adapted to the source face's geometry.
Three auxiliary losses support this decomposition. A self-supervised representation loss maximizes cosine similarity between expression features of an image and its augmented version (rotation, scaling, translation), enforcing invariance of K5 to non-expression factors. An identity latent consistency loss aligns canonical keypoints between source and driving frames in canonical space, stabilizing identity representation across poses. A 2D landmark loss matches 145 detected landmarks (face, mouth, pupils) between generated and driving images. Generation uses a multi-scale generator producing outputs at 64, 128, and 256 resolutions with corresponding multi-scale perceptual losses.
VAE latent space for expression interpolation
To enable continuous expression control without a driving source, MMFA trains a variational autoencoder that maps K6 to a Gaussian latent variable K7. The authors report a practical training failure mode: the KL term converges faster than reconstruction, causing posterior collapse to a constant average expression before distinguishable representations form. They mitigate this with an adversarial loss on reconstructed features, which they state preserves diversity in the expression feature distribution. This VAE is trained independently after MMFA and enables linear interpolation of expressions (K8), which the authors claim is the first expression interpolation capability within an unsupervised face animation framework. Qualitative results show smooth, coherent expression transitions in both same-identity and cross-identity settings.
Experimental results
Training uses VoxCeleb with same-video frame pairs for self-supervision; evaluation covers 80 image–video pairs built from CelebA, FFHQ, and the VoxCeleb test set, yielding roughly 25K synthesized images per method. Baselines include FOMM, MRAA, Face-vid2vid, DaGAN, LIA, and DPE.
| Method |
Same-ID FID ↓ |
Same-ID CSIM ↑ |
Cross-ID CSIM ↑ |
Cross-ID APD ↓ |
Cross-ID FID ↓ |
| FOMM |
22.755 |
0.965 |
0.903 |
0.027 |
105.216 |
| MRAA |
21.397 |
0.970 |
0.901 |
0.026 |
87.276 |
| Face-vid2vid |
20.092 |
0.972 |
0.927 |
0.033 |
86.054 |
| DaGAN |
14.463 |
0.974 |
0.907 |
0.023 |
86.012 |
| LIA |
23.992 |
0.967 |
0.936 |
0.141 |
83.065 |
| DPE |
29.620 |
0.967 |
0.887 |
0.043 |
128.904 |
| MMFA |
13.265 |
0.974 |
0.925 |
0.042 |
77.445 |
MMFA achieves the lowest FID in both same-identity (13.265 vs. DaGAN's 14.463) and cross-identity (77.445 vs. LIA's 83.065) settings, indicating superior generation authenticity. It ties DaGAN for best same-identity CSIM (0.974). In cross-identity reenactment, LIA attains the highest CSIM but degrades sharply on APD (0.141), indicating poor pose transfer; MMFA achieves the best AED and second-best CSIM, which the authors interpret as stable operation in unconstrained conditions. Notably, MMFA does not dominate all metrics — its PSNR (22.898) trails MRAA and DaGAN, and its AKD (1.476) trails DaGAN — so the quantitative advantage is concentrated in perceptual quality (FID) rather than pixel-level fidelity.
For motion attribute editing, comparison against DPE shows that DPE suffers identity loss and background distortion (e.g., clothing artifacts) under large pose or expression changes, whereas MMFA's explicit keypoints confine edits to the face region and permit direct specification of pose, scale, and position. Visualization of canonical faces confirms that identities are nearly identical in neutral space regardless of input pose and expression, supporting the feasibility of cross-identity editing. The ablation study shows each proposed loss contributes: K9 improves cross-identity identity preservation, pCk∈R30 further improves it and reduces canonical-face distortion under large pose variation, and pCk∈R31 yields the largest gains, bringing FID from 16.781 (base) to 13.265 and improving LPIPS, PSNR, AKD, and AED simultaneously.
Limitations and open questions
The authors acknowledge several constraints. The 3D keypoint estimator and 3D convolutions in dense motion estimation make MMFA more resource-intensive to train than 2D keypoint baselines such as FOMM. In same-identity expression driving, 3D keypoint methods do not outperform 2D counterparts despite better cross-identity preservation and image quality; supplementary results further note that MMFA tracks eyeball position less accurately than 2D-keypoint methods like FOMM and MRAA. The framework also rests on the orthographic projection and rigidity assumptions, which approximate real camera behavior and non-rigid facial deformation; how far these assumptions hold under extreme perspective or wide-baseline capture is not quantified. Open questions include whether 2D and 3D keypoints can be combined for more accurate fine-grained driving (e.g., eye gaze), and whether the network can be simplified for faster training and inference.
Conclusion
MMFA extends unsupervised 3D keypoint-based face animation with an explicit scale-aware decomposition and self-supervised expression feature learning, achieving independent control of rotation, translation, scaling, and expression. Its VAE-based expression latent space adds continuous interpolation without driving sources. Empirically, the method delivers the best FID among compared baselines in both reenactment regimes and competitive identity preservation, though it does not lead on pixel-level reconstruction metrics and incurs higher training cost than 2D alternatives.