- The paper introduces InterTalk, a motion-driven framework that unifies context encoding, interactive motion generation, and efficient rendering for realistic conversational avatar synthesis.
- The paper demonstrates superior performance with state-of-the-art results in lip synchronization, visual quality, and temporal consistency across single and multi-person scenarios.
- The paper validates its iterative refinement and 3D data augmentation through ablation studies, underlining their critical roles in enhancing feedback-driven natural interactions.
InterTalk: Flexible, Natural, and Efficient Conversational Talking Face Generation
Conversational talking face generation demands solutions for multi-round dialogues, arbitrary participant numbers, natural interaction including non-verbal feedback, and real-time performance. Prior art either restricts interactivity to isolated roles or scales computationally via large video diffusion models, which underscores a trade-off between flexibility, naturalness, and efficiency. The proposed InterTalk framework directly targets these requirements, establishing a motion-driven architecture that unifies participant dynamics, enables role switching, and delivers efficient multi-person interactive video synthesis.

Figure 1: InterTalk generates conversational talking face videos with flexible, natural, and real-time group interaction.
Architecture: Responsive Context Encoding, Motion Generation, and Rendering
InterTalk decomposes the task into three modules:
- Responsive Context Encoder (RCE): Encodes environmental context by fusing each participant's audio (via Wav2Vec 2.0) and disentangled facial motion (lip, eye, and head pose) into a latent interactive representation. A cross-attention block aligns modalities, followed by a bi-directional LSTM for temporal coherence and adaptive attention pooling across participants, yielding a unified context vector.
- Interactive Motion Generator (IMG): Generates refined facial motion conditioned on RCE outputs and self audio. A multi-stage pipeline employs a fused Transformer Encoder, temporal alignment attention, Transformer Decoder, and feedforward refinement for facial components. Further enhancement leverages dedicated modules for lip-sync (trained on single-speaker datasets), eye-blinking, and head pose initialization for gaze alignment. The iterative generation strategy enables feedback-driven refinement of interactive motions.
- Rendering Pipeline: Uses implicit 3D keypoints derived from pose and expression estimators, with a warping decoder and spatial blending mask to animate each participant and composite them into a cohesive scene.

Figure 2: InterTalk consists of Responsive Context Encoder, Interactive Motion Generator, and a Rendering Pipeline.
Iterative Generation and Motion Feedback
The iterative generation strategy sequentially updates each participant’s motion, integrating feedback across cycles. The approach progressively incorporates environmental signals and responses, significantly enhancing realism and interaction diversity.

Figure 3: Iterative generation strategy progressively updates participant motions for mutual feedback and natural interactivity.
Dataset Construction and 3D Data Augmentation
Due to scarcity of suitable public datasets, InterTalk introduces a new multi-person conversational dataset. Multi-track audio is separated via visual-audio networks, ensuring clean individual streams. To augment training diversity and realism, 3D FLAME coefficients are converted to motion signals through specialized converters with region masks for lip and eye expressions. The integration of 3D data into InterTalk’s pipeline delivers robust supervision and expanded conversational coverage.

Figure 4: Dataset samples, multi-track audio separation, and 3D augmentation pipeline utilizing FLAME coefficients.
Empirical Evaluation and Numerical Results
InterTalk demonstrates superior performance across single-person and multi-participant conversational scenarios:
- Talking Face (Single-Speaker): Outperforms state-of-the-art models—including SadTalker, EchoMimic, Hallo2, and Sonic—on HDTF with Sync-C (8.30), FID (23.07), and FVD (178.73), indicating excellent lip synchronization, visual quality, and temporal consistency.
- Interactive Head Generation: Achieves minimum Fréchet Distance (FD = 18.33), maximum diversity (SID = 4.85), and lowest MSE across ViCo, outperforming DIM, INFP, ARIG, and DualTalk.
- Multi-Participant Conversations: On the custom dataset, InterTalk produces substantially lower FD and higher FPS compared to MultiTalk, supporting arbitrary participant counts and real-time throughput.
- Ablation Study: Disabling motion feedback, iterative generation, motion disentanglement, and 3D augmentation each notably degrades interaction quality, highlighting necessity of these components.

Figure 5: Comparative results on interactive head generation highlighting responsive agent behaviors.

Figure 6: InterTalk generates avatars supporting arbitrary users and agents with multi-track audio inputs.

Figure 7: InterTalk achieves unconstrained participant numbers and highly interactive behaviors in group conversations.
Analysis of Iterative Refinement
FD and SID trends across generation iterations confirm that iterative feedback elevates realism and interaction diversity, as seen in decreasing FD and rising SID. This mechanism is pivotal for capturing lifelike mutual participant responses.

Figure 8: Iterative refinement reduces FD and increases SID, evidencing improvement in motion realism and diversity.
Implications and Future Directions
Practically, InterTalk enables real-time, flexible, and natural conversational talking face generation for applications in remote communication, digital companionship, and online education. Theoretically, explicit modeling of interaction context, facial motion disentanglement, and iterative feedback set a new benchmark for controllable multi-modal generative systems. Ongoing directions include expanding the scale and diversity of conversational datasets, integrating emotional cues, hierarchical modeling of group dynamics, and adaptation to open-domain dialog systems.
Conclusion
InterTalk delivers a unified, efficient, and interactive framework for conversational talking face generation, satisfying requirements for flexibility, naturalness, and scalability. Strong empirical results validate each architectural component and the overall motion-driven paradigm. The introduction of a new dataset and a principled 3D augmentation approach propels supervised conversational modeling. InterTalk’s modular design and real-time performance establish it as a robust backbone for future research in conversational agents and interactive AI-driven facial animation (2606.31088).