Papers
Topics
Authors
Recent
Search
2000 character limit reached

Motion Manipulation via Unsupervised Keypoint Positioning in Face Animation

Published 4 Mar 2026 in cs.CV | (2603.04302v1)

Abstract: Face animation deals with controlling and generating facial features with a wide range of applications. The methods based on unsupervised keypoint positioning can produce realistic and detailed virtual portraits. However, they cannot achieve controllable face generation since the existing keypoint decomposition pipelines fail to fully decouple identity semantics and intertwined motion information (e.g., rotation, translation, and expression). To address these issues, we present a new method, Motion Manipulation via unsupervised keypoint positioning in Face Animation (MMFA). We first introduce self-supervised representation learning to encode and decode expressions in the latent feature space and decouple them from other motion information. Secondly, we propose a new way to compute keypoints aiming to achieve arbitrary motion control. Moreover, we design a variational autoencoder to map expression features to a continuous Gaussian distribution, allowing us for the first time to interpolate facial expressions in an unsupervised framework. We have conducted extensive experiments on publicly available datasets to validate the effectiveness of MMFA, which show that MMFA offers pronounced advantages over prior arts in creating realistic animation and manipulating face motion.

Summary

  • The paper introduces MMFA, a one-shot face-animation framework that uses unsupervised 3D keypoints, explicit scale modeling, and self-supervised expression features to independently control pose, translation, scale, and expression.
  • MMFA achieves the best reported FID in same-identity and cross-identity tests—13.265 and 77.445, respectively—while matching the strongest same-identity CSIM score of 0.974, though it does not lead every pixel-level metric.
  • The paper adds a VAE-based expression space for smooth interpolation without a driving video, while acknowledging higher computational cost and weaker eye-gaze tracking than some 2D keypoint methods.

Overview and motivation

MMFA (Motion Manipulation via unsupervised keypoint positioning in Face Animation) is a one-shot face animation framework that combines 3D unsupervised keypoint estimation with self-supervised representation learning to achieve explicit, independent control over facial motion attributes (2603.04302). The authors identify two shortcomings of prior keypoint-based pipelines. First, methods such as FOMM and MRAA estimate keypoints without semantic structure, so identity semantics remain entangled with motion information, preventing separate manipulation of facial movement. Second, Face-vid2vid's canonical-keypoint decomposition ignores camera perspective projection; because faces at different depths in VoxCeleb videos appear at different scales, the expression deformation δ\delta is forced to absorb scale changes, coupling expression with facial scaling and rotation. The paper's central claim is that pose, translation, scaling, and expression can be decoupled under reasonable geometric assumptions while retaining the detail-transfer capability of unsupervised keypoints.

Keypoint decomposition pipeline

MMFA builds on three assumptions: the object centroid lies at the world/camera origin, the camera-to-image mapping follows orthographic projection, and the face is rigid under scale and rotation transforms. Under these assumptions, each face is modeled by KK canonical keypoints pCk∈R3p_C^k \in \mathbb{R}^3 (identity anchors), a rotation matrix RR, a 2D translation tt, a scalar scale factor ff, and per-keypoint expression deformations δk\delta^k. Target keypoints are computed as:

pSk=RSfS(pCk+δSk)+tSp_S^k = R_S f_S (p_C^k + \delta_S^k) + t_S

The explicit scale factor ff is the key departure from Face-vid2vid: it absorbs depth-induced size variation so that δ\delta encodes only genuine expression deformation. Rotation is obtained from a pre-trained pose estimator rather than learned jointly, avoiding contamination of KK0 by KK1. Expression deformations are predicted by an encoder-decoder: a Bottleneck-based encoder compresses the input image into a 256-dimensional feature vector KK2, which is concatenated with the canonical keypoints and decoded into KK3. Because KK4 depends on both the expression feature and the target identity's canonical keypoints, cross-identity reenactment transfers expressions adapted to the source face's geometry.

Three auxiliary losses support this decomposition. A self-supervised representation loss maximizes cosine similarity between expression features of an image and its augmented version (rotation, scaling, translation), enforcing invariance of KK5 to non-expression factors. An identity latent consistency loss aligns canonical keypoints between source and driving frames in canonical space, stabilizing identity representation across poses. A 2D landmark loss matches 145 detected landmarks (face, mouth, pupils) between generated and driving images. Generation uses a multi-scale generator producing outputs at 64, 128, and 256 resolutions with corresponding multi-scale perceptual losses.

VAE latent space for expression interpolation

To enable continuous expression control without a driving source, MMFA trains a variational autoencoder that maps KK6 to a Gaussian latent variable KK7. The authors report a practical training failure mode: the KL term converges faster than reconstruction, causing posterior collapse to a constant average expression before distinguishable representations form. They mitigate this with an adversarial loss on reconstructed features, which they state preserves diversity in the expression feature distribution. This VAE is trained independently after MMFA and enables linear interpolation of expressions (KK8), which the authors claim is the first expression interpolation capability within an unsupervised face animation framework. Qualitative results show smooth, coherent expression transitions in both same-identity and cross-identity settings.

Experimental results

Training uses VoxCeleb with same-video frame pairs for self-supervision; evaluation covers 80 image–video pairs built from CelebA, FFHQ, and the VoxCeleb test set, yielding roughly 25K synthesized images per method. Baselines include FOMM, MRAA, Face-vid2vid, DaGAN, LIA, and DPE.

Method Same-ID FID ↓ Same-ID CSIM ↑ Cross-ID CSIM ↑ Cross-ID APD ↓ Cross-ID FID ↓
FOMM 22.755 0.965 0.903 0.027 105.216
MRAA 21.397 0.970 0.901 0.026 87.276
Face-vid2vid 20.092 0.972 0.927 0.033 86.054
DaGAN 14.463 0.974 0.907 0.023 86.012
LIA 23.992 0.967 0.936 0.141 83.065
DPE 29.620 0.967 0.887 0.043 128.904
MMFA 13.265 0.974 0.925 0.042 77.445

MMFA achieves the lowest FID in both same-identity (13.265 vs. DaGAN's 14.463) and cross-identity (77.445 vs. LIA's 83.065) settings, indicating superior generation authenticity. It ties DaGAN for best same-identity CSIM (0.974). In cross-identity reenactment, LIA attains the highest CSIM but degrades sharply on APD (0.141), indicating poor pose transfer; MMFA achieves the best AED and second-best CSIM, which the authors interpret as stable operation in unconstrained conditions. Notably, MMFA does not dominate all metrics — its PSNR (22.898) trails MRAA and DaGAN, and its AKD (1.476) trails DaGAN — so the quantitative advantage is concentrated in perceptual quality (FID) rather than pixel-level fidelity.

For motion attribute editing, comparison against DPE shows that DPE suffers identity loss and background distortion (e.g., clothing artifacts) under large pose or expression changes, whereas MMFA's explicit keypoints confine edits to the face region and permit direct specification of pose, scale, and position. Visualization of canonical faces confirms that identities are nearly identical in neutral space regardless of input pose and expression, supporting the feasibility of cross-identity editing. The ablation study shows each proposed loss contributes: KK9 improves cross-identity identity preservation, pCk∈R3p_C^k \in \mathbb{R}^30 further improves it and reduces canonical-face distortion under large pose variation, and pCk∈R3p_C^k \in \mathbb{R}^31 yields the largest gains, bringing FID from 16.781 (base) to 13.265 and improving LPIPS, PSNR, AKD, and AED simultaneously.

Limitations and open questions

The authors acknowledge several constraints. The 3D keypoint estimator and 3D convolutions in dense motion estimation make MMFA more resource-intensive to train than 2D keypoint baselines such as FOMM. In same-identity expression driving, 3D keypoint methods do not outperform 2D counterparts despite better cross-identity preservation and image quality; supplementary results further note that MMFA tracks eyeball position less accurately than 2D-keypoint methods like FOMM and MRAA. The framework also rests on the orthographic projection and rigidity assumptions, which approximate real camera behavior and non-rigid facial deformation; how far these assumptions hold under extreme perspective or wide-baseline capture is not quantified. Open questions include whether 2D and 3D keypoints can be combined for more accurate fine-grained driving (e.g., eye gaze), and whether the network can be simplified for faster training and inference.

Conclusion

MMFA extends unsupervised 3D keypoint-based face animation with an explicit scale-aware decomposition and self-supervised expression feature learning, achieving independent control of rotation, translation, scaling, and expression. Its VAE-based expression latent space adds continuous interpolation without driving sources. Empirically, the method delivers the best FID among compared baselines in both reenactment regimes and competitive identity preservation, though it does not lead on pixel-level reconstruction metrics and incurs higher training cost than 2D alternatives.

Paper to Video (Beta)

No one has generated a video about this paper yet.

Whiteboard

No one has generated a whiteboard explanation for this paper yet.

Open Problems

We haven't generated a list of open problems mentioned in this paper yet.