Papers
Topics
Authors
Recent
Search
2000 character limit reached

Follow-Your-Emoji-Faster: Efficient Portrait Animation

Updated 12 July 2026
  • Follow-Your-Emoji-Faster is a diffusion-based framework that uses expression-aware landmarks to drive realistic and controlled portrait animations.
  • It enhances generation efficiency with a Taylor-interpolated cache, achieving 2.6X acceleration while maintaining identity and expression fidelity.
  • The method integrates a Stable Diffusion pipeline with temporal attention and fine-grained facial supervision to ensure long-term temporal consistency.

Follow-Your-Emoji-Faster is a diffusion-based framework for freestyle portrait animation in which a static reference portrait is animated by a target sequence of facial landmarks while preserving identity, transferring subtle and exaggerated expressions, maintaining long-term temporal consistency, and improving generation efficiency. In the reported formulation, the method is built on a Stable Diffusion-style latent diffusion pipeline, retains the expression-aware landmark conditioning and facial fine-grained supervision introduced by the earlier Follow-Your-Emoji system, and adds a Taylor-interpolated cache together with the EmojiBench++ benchmark to address the efficiency bottleneck of long-form portrait animation (Ma et al., 20 Sep 2025, Ma et al., 2024).

1. Task formulation and research lineage

The task addressed by Follow-Your-Emoji-Faster is freestyle portrait animation driven by facial landmarks. The input is a reference portrait image I0I_0 and a target motion sequence represented as landmarks {Lt}t=1..N\{L_t\}_{t=1..N}, and the output is a synthesized video that preserves the appearance of I0I_0 while following the target expressions and head dynamics. The paper identifies four central difficulties: identity preservation, accurate expression retargeting, long-term temporal consistency, and efficiency (Ma et al., 20 Sep 2025).

The framework is positioned as a direct continuation of the Follow-Your-Emoji line. The earlier system already combined a Stable Diffusion or Latent Diffusion backbone with expression-aware landmarks, a facial fine-grained loss, and a progressive generation strategy for long-term animation (Ma et al., 2024). Follow-Your-Emoji-Faster retains those core ideas but explicitly targets efficient sampling, reporting a Taylor-interpolated cache and a “2.6X lossless acceleration,” while also expanding evaluation from EmojiBench to EmojiBench++ (Ma et al., 20 Sep 2025).

A plausible implication is that the contribution is not merely incremental acceleration. The faster variant reframes efficiency as part of controllability: the same conditioning stack that preserves identity and expression is also used to decide where feature reuse is safe and where interpolation must remain facially precise.

2. Backbone architecture and conditioning pathways

The base model is a Stable Diffusion-style latent diffusion model with a frozen VAE encoder-decoder and a UNet denoiser operating in latent space. The architecture includes four plug-ins: image prompt injection, an appearance network, temporal attention, and a landmark-driven motion-control pathway (Ma et al., 20 Sep 2025).

Image prompt injection is implemented with a CLIP image encoder whose identity tokens are refined through a 4-layer Q-Former and injected through cross-attention. This pathway provides identity semantics from the reference portrait. In parallel, an appearance network with an SD-UNet-like structure extracts appearance features from I0I_0 and injects them into self-attention blocks of the main UNet, explicitly supporting identity and background retention. Temporal transformer layers, initialized from AnimateDiff, are inserted into the UNet to model cross-frame coherence. Motion control is handled by a landmark encoder that maps the expression-aware landmark sequence into spatial control features and fuses them directly into UNet inputs; the reported implementation does not require ControlNet or LoRA (Ma et al., 20 Sep 2025).

This conditioning design is technically notable because it separates three forms of guidance that are often entangled in portrait animation: identity tokens for cross-attention, appearance features for self-attention, and motion features for direct latent-space conditioning. The earlier Follow-Your-Emoji system described the same overall logic in terms of a landmark encoder, CLIP image conditioning, appearance-feature injection, and AnimateDiff-based temporal attention (Ma et al., 2024), and the faster variant preserves that decomposition.

3. Expression-aware landmarks and facially localized supervision

A defining element of the method is the use of expression-aware landmarks as explicit motion signals. The pipeline extracts MediaPipe 3D facial keypoints from the driving sequence, projects them to 2D, removes facial contour points, and includes pupil points with relative iris positions. Excluding contour points is intended to avoid contour instability and identity distortion, while including pupils preserves fine eye motion and exaggerated gaze behavior (Ma et al., 20 Sep 2025).

The target landmarks are aligned to the canonical or reference coordinate frame by a similarity transform. For 2D driver and reference landmarks D,S∈R2×KD,S \in \mathbb{R}^{2 \times K}, the paper gives a Procrustes-like alignment: center both sets, compute SVD(D~⊤S~)=UΣV⊤\mathrm{SVD}(\tilde D^\top \tilde S)=U\Sigma V^\top, set R=VU⊤R=VU^\top, s=tr(Σ)/∥D~∥2s=\mathrm{tr}(\Sigma)/\|\tilde D\|^2, τ=μS−sRμD\tau=\mu_S-sR\mu_D, and obtain warped landmarks L^=sRD+τ\hat L=sRD+\tau (Ma et al., 20 Sep 2025). This construction is used to reduce identity leakage during motion transfer.

The diffusion objective remains the standard latent denoising loss,

{Lt}t=1..N\{L_t\}_{t=1..N}0

but the model adds a facial fine-grained loss

{Lt}t=1..N\{L_t\}_{t=1..N}1

with total loss

{Lt}t=1..N\{L_t\}_{t=1..N}2

Here {Lt}t=1..N\{L_t\}_{t=1..N}3 is an expression mask formed by dilated landmark disks, and {Lt}t=1..N\{L_t\}_{t=1..N}4 is a facial mask derived from contour projection. The reported role of this masked residual term is to emphasize subtle expression regions such as eyes and lips while also reinforcing facial identity reconstruction (Ma et al., 20 Sep 2025).

A common misconception is that identity preservation in such systems must rely on an explicit recognition loss. In this framework, ArcFace cosine similarity is used for evaluation rather than as a training objective; the paper explicitly states that no ArcFace identity loss is included in training (Ma et al., 20 Sep 2025).

4. Training data and optimization regime

The reported training corpus combines HDTF and VFHQ with a newly collected expression dataset. That expression set contains 115 subjects, each with a 20-minute indoor video, together with 18 posed or exaggerated expressions. Outdoor scenes are covered by VFHQ, and stylized training targets include cartoons, animals, and sculptures from public sources. EmojiBench++ is reserved for evaluation and does not overlap with training (Ma et al., 20 Sep 2025).

Training proceeds in two phases. Phase I emphasizes appearance and identity with single-frame supervision for 30,000 iterations at batch size 32 and resolution {Lt}t=1..N\{L_t\}_{t=1..N}5. Phase II introduces temporal attention with 16-frame clips for 10,000 steps, again at batch size 32. Optimization uses Adam with a constant learning rate of {Lt}t=1..N\{L_t\}_{t=1..N}6, a frozen VAE, and temporal attention initialized from AnimateDiff. The training setup uses 32 NVIDIA A800 GPUs and is reported to require approximately 68 hours (Ma et al., 20 Sep 2025).

The earlier Follow-Your-Emoji system followed a similar two-stage regimen—30,000 image-level steps followed by 10,000 video-level steps at batch size 32 and learning rate {Lt}t=1..N\{L_t\}_{t=1..N}7—which suggests continuity in the optimization design even as the faster model changes the inference-time efficiency profile (Ma et al., 2024).

5. Progressive long-term generation and Taylor-interpolated cache

Long-term temporal stability is handled by a progressive generation strategy. During training, the model alternates between two masking schemes with equal probability: keyframe masking, which keeps only the first and last latents and masks all middle frames, and random masking, in which each frame latent is masked independently with probability {Lt}t=1..N\{L_t\}_{t=1..N}8. At inference time, the model first generates the endpoints and then recursively fills intermediate frames by bisection with anchor-aware sampling (Ma et al., 20 Sep 2025).

The inference logic is summarized by the reported pseudo-code: I0I_07 This anchor-conditioned interpolation is meant to constrain mid-sequence content to the identity and appearance established by the keyframes (Ma et al., 20 Sep 2025).

The principal efficiency contribution is the Taylor-interpolated cache, a training-free mechanism that caches feature maps in spatial attention, cross-attention, and temporal attention blocks and predicts later-step features by finite-difference Taylor interpolation. For a module {Lt}t=1..N\{L_t\}_{t=1..N}9 and stride I0I_00, the interpolation rule is

I0I_01

with

I0I_02

and higher-order differences defined recursively. The paper applies interpolation after a timestep boundary I0I_03 of total denoising steps, or inside landmark-mask regions, while earlier steps and background tokens rely more heavily on reuse (Ma et al., 20 Sep 2025).

The reported effect is a “2.6X lossless acceleration,” and the detailed evaluation gives latency I0I_04 versus I0I_05, corresponding to I0I_06 speedup, for the TIC configuration with I0I_07 and Taylor order up to I0I_08 (Ma et al., 20 Sep 2025). This suggests that the acceleration is achieved not by uniformly reducing computation, but by reallocating computation toward landmark-relevant regions and later denoising stages.

6. Benchmarking and empirical performance

Follow-Your-Emoji-Faster introduces EmojiBench++, described as an extension of EmojiBench to 500 diverse portrait animation sequences spanning real faces, cartoons, sculptures, and animals. Evaluation is divided into self reenactment and cross reenactment. Self reenactment uses L1, SSIM, LPIPS, and FVD. Cross reenactment uses ArcFace cosine identity similarity, HyperIQA image quality, and landmark accuracy. A user study with 30 participants over 45 cases ranks methods on expression fidelity, identity preservation, and overall quality (Ma et al., 20 Sep 2025).

The paper reports best-in-table performance across all headline metrics. The summary values are as follows.

Setting Metric Follow-Your-Emoji-Faster
Self reenactment L1 / SSIM / LPIPS / FVD 0.027 / 0.829 / 0.137 / 115.4
Cross reenactment ID / Image Quality / Landmark Accuracy 0.761 / 60.147 / 38.26
User study Expression / Identity / Overall 1.17 / 1.42 / 1.67

The acceleration study also compares the proposed cache with Half Steps, TokenPrune, TokenMerge, and DeepCache. The TIC configuration is reported to preserve the strongest quality metrics among accelerated variants, while the “Half Steps” setting increases speed but degrades quality more noticeably. In the same table, TIC is associated with SSIM I0I_09, LPIPS I0I_00, FVD I0I_01, identity I0I_02, image quality I0I_03, landmark accuracy I0I_04, latency I0I_05, and speed I0I_06 (Ma et al., 20 Sep 2025).

Ablation results attribute performance gains to several design choices. Replacing expression-aware landmarks with 2D landmarks degrades identity and accuracy; including contour points in the motion signal harms identity on non-human styles; removing pupil points weakens gaze fidelity; removing either facial or expression masks from the fine-grained loss degrades identity preservation and subtle expression quality; and omitting the progressive strategy worsens long-term coherence (Ma et al., 20 Sep 2025).

7. Limitations, ethical considerations, and broader significance

The method is explicitly intended to work across real faces, cartoons, sculptures, and animals, but the paper also states several failure modes. Animal portraits with atypical mouth, tongue, or teeth details may be underrepresented in training. Landmark detection failures, extreme occlusions, extreme head poses, or heavy stylization can degrade retargeting and identity stability. Although the framework can be used at higher resolutions, robustness there remains data-dependent (Ma et al., 20 Sep 2025).

The paper also foregrounds misuse risk. Because the system performs high-fidelity portrait animation, it raises standard deepfake concerns; the stated recommendations are visible watermarking, usage agreements in the repository, respect for portrait rights, and restriction to legitimate use cases (Ma et al., 20 Sep 2025).

In the broader research trajectory, Follow-Your-Emoji-Faster can be understood as the efficiency-oriented consolidation of the earlier Follow-Your-Emoji design. The 2024 system established that expression-aware landmarks, facial masks, and progressive generation improved controllability, identity preservation, and long-term stability on EmojiBench (Ma et al., 2024). The faster variant preserves that formulation, adds the Taylor-interpolated cache and EmojiBench++, and is reported to improve all headline quality metrics while substantially reducing denoising latency (Ma et al., 20 Sep 2025). This suggests that, in this line of work, efficiency is not treated as a separate deployment concern but as a model-design problem coupled to facial control, temporal consistency, and spatially selective reuse.

Topic to Video (Beta)

No one has generated a video about this topic yet.

Whiteboard

No one has generated a whiteboard explanation for this topic yet.

Follow Topic

Get notified by email when new papers are published related to Follow-Your-Emoji-Faster.