- The paper introduces an automated pipeline that synthesizes facial mechanisms using parametric templates and modular linkages for scalable animatronic designs.
- It employs a hierarchical design process with kinematic optimization and collision-free assembly refinement to ensure anatomical feasibility and efficient manufacturing.
- Quantitative results show a success rate of 66.7% with reduced convergence times and enhanced interaction performance via dual-speaker audio-to-actuator mapping.
Automated Synthesis of Facial Mechanisms for Conversational Animatronic Robots
Parametric Face Template and Modular Actuation
The paper presents an automated pipeline for scalable hardware generation of animatronic robotic faces through parametric templates and modular linkage-driven actuation. The mechanical face template consists of four principal modules—eyebrow, eyes, mouth, and jaw—each modeled as spatial linkages with explicit DoF allocation. The eyebrow mechanism employs a symmetric six-bar linkage for vertical and brow-center motions, the eye integrates four-bar linkages for eyelid and eyeball actuation, the mouth module realizes lip and mouth-corner articulation via four- and five-bar mechanisms, and the jaw uses a four-bar linkage for opening/closing. This modularity supports systematic scaling and retargeting across diverse geometries, enabling efficient adaptation for non-human or stylized morphologies Figure 1.

Figure 1: Mechanical face template with modular linkage-driven actuation and a 3-DoF neck, supporting scalable synthesis.
Hierarchical Automated Design Pipeline
The proposed pipeline operates hierarchically: given a 2D portrait, it reconstructs a metric-accurate 3D mesh with semantic facial landmarks. Coarse initialization assigns module base poses, while inner-loop kinematic synthesis optimizes linkage parameters under anatomy-guided feasible motion volumes and AU-derived trajectory primitives, maximizing expressiveness while enforcing manufacturability. Expressiveness is formalized through trajectory amplitude scaling, constrained by anatomical bounds and collision detection. Outer-loop QP-based assembly refinement resolves spatial interferences via minimum translation vectors, iteratively adjusting base poses or amplitude limits until a collision-free assembly is achieved Figure 2.

Figure 2: Hierarchical design pipeline: initialization, kinematic synthesis under anatomical and AU trajectory constraints, and outer-loop collision-driven refinement.
Extensive kinematic modeling for each module ensures precise spatial mapping and manipulability analysis, with detailed nonlinear optimization systems solved for robust forward kinematics and workspace boundary estimation (see supplement: Figure 3).
Interaction Synthesis and Semantic Mapping
Beyond hardware synthesis, conversational interaction is achieved through a dual-speaker talking head model. The framework joins audio from both interlocutors using Wav2Vec feature extraction, temporal encoding, and turn-aware gating. A Transformer backbone with turn-conditioned attention models speaker-listener dynamics, outputting temporally coherent blendshape parameters and head pose for both roles. The mapping from synthesized facial motion to actuator commands exploits explicit semantic region decomposition: lightweight MLPs regress commands for each region (eyebrow, eyes, mouth, jaw), trained on calibration datasets generated via Latin hypercube sampling and MediaPipe-based facial parameter extraction. This region-wise mapping ensures efficient, high-fidelity translation of digital animation to physical actuation Figure 4.

Figure 4: Interaction synthesis and region-wise mapping from dual-speaker audio to robot motor commands for multi-round conversational expressions.
Quantitative Evaluation and Algorithm Efficacy
The pipeline achieves a success rate of 66.7% in synthesizing collision-free mechanisms across 15 diverse facial geometries, outperforming baselines including local-only optimization (20%), global joint optimization (0%), and heuristic repulsion (33.3%). Convergence time is notably reduced (591.3 s vs. 1098.2 s heuristic), with expressiveness scores (∑k​αk​) comparable to manual design but with substantially accelerated runtime (11.7 min vs. 22.8 h manual). The methodology generalizes to non-human morphologies—illustrated by feasible assemblies for "Yoda" and "Jack"—and delivers volume-efficient layouts where template deviation is substantial Figure 5.

Figure 5: (a) Algorithmic assembly for morphologically divergent characters ("Jack" and "Yoda"). (b) Comparison of mouth-corner mechanism, algorithmic versus manual on compact geometry.
In conversational benchmarking, the system achieves the lowest MSE (0.38), lowest pose dynamic deviation (5.63), and robust speaker-listener performance across metrics (LVE, SID, FDD). Unlike prior dyadic systems requiring partner's future motion and restricting real-time deployment, the introduced model infers behaviors exclusively from audio, supporting online interaction. Region-wise semantic mapping further improves audio-lip synchronization (LSE-D 10.61, LSE-C 3.19), outpacing landmark- and FLAME-based baselines and sustaining high frame rates on embedded platforms, which is critical for real-time robotics.
End-to-End Demonstration and Perceptual Validation
The pipeline enables rapid physical realization, compressing mechanism synthesis and mapping stages to minutes, with the total build time for a new personalized head under 26 h (main bottleneck: physical manufacturing, inherently parallelizable). End-to-end deployment demonstrates bidirectional interaction in both human-robot and robot-robot dialogue re-enactments, achieving contextually congruent reactive expressions and fluid multi-round gestural engagement Figure 6.

Figure 6: End-to-end system visualization; (a) real-time human-robot conversation, (b) dyadic role-play with context-aware expressions.
User studies with 100 participants confirm highest perceived performance for the proposed pipeline, with explicit listener behaviors and coordinated neck motion producing appreciably more natural interaction than speaker-only, random neck, or mouth-only mapping variants Figure 7.

Figure 7: User study results: (a) Overall interaction performance, (b) Mapping quality for short sentences.
Morphological Diversity and Extended Scenarios
Automated mechanism synthesis is validated across 8 distinct identities spanning human, stylized, and folk characters—demonstrating robust volumetric adaptability and kinematic integrity Figure 8. Additional visuals illustrate successful cross-modal interaction scenarios including physical-virtual dyads and Mandarin mythological dialogues, evidencing the framework's extensibility Figure 9.

Figure 8: Synthesis of optimized internal CAD assemblies for eight distinct identities, emphasizing volumetric diversity.

Figure 9: Diverse interaction scenarios: Yoda-Luke, elf-virtual avatar, and Mandarin myth.
Practical and Theoretical Implications
Practically, the framework offers scalable hardware generation for animatronic faces, compressing customization timelines and reducing expert dependency in prototyping. Theoretically, it raises the prospect of treating embodied social robotics as an automatic design and computational optimization challenge, bridging mechanistic expressiveness, anatomical validity, and interactive coherence. The integration of anatomy- and AU-guided constraints, semantic mapping, and dual-speaker temporal modeling represents a step toward unified frameworks for personalized, conversational agents.
Future directions include enhancing mechanical reconfigurability, fully end-to-end differentiable learning for direct audio-to-actuator mapping, high-fidelity differentiable simulation for sim-to-real transfer, and synthesis of full-body embodiment for richer multimodal interaction.
Conclusion
The presented pipeline demonstrates automated scalable synthesis and real-time deployment of high-fidelity animatronic faces for conversational robots, combining hierarchical mechanical design with neural dual-speaker expression modeling and semantic motion mapping. This architecture supports large-scale personalization, robust interaction, and efficient manufacturing, laying foundational work for embodied AI agents with distinct identities and expressive conversational capabilities (2607.11688).