Paula Signing Avatar: Linguistic & Emotional Control
- Paula Signing Avatar is a digital signing system that integrates coordinated manual and non-manual signals through layered facial animation for naturalistic communication.
- It employs a two-parameter Pleasure–Arousal model with EASIER notation to achieve precise, culturally adaptable emotional expression.
- The architecture blends sign language grammar with fine-grained facial controls, enabling real-time, multilingual applications and mixed-reality interactions.
Paula Signing Avatar is a signing avatar designed through collaboration with Deaf communities for natural, legible signing. In the research literature, Paula is presented as an avatar whose manual and non-manual channels must be coordinated so that facial expressions, head and torso posture, eye gaze, and mouth actions can convey grammatical, pragmatic, and affective information. Recent work on Paula focuses especially on emotional non-manual signals, using a two-parameter representation integrated with the EASIER notation to make linguistic specification of emotional facial expression more coherent than earlier label-based or low-level control schemes (McDonald et al., 11 Aug 2025).
1. Linguistic function and the problem of non-manual signals
In signed languages, non-manual signals carry grammatical, pragmatic, and affective information. Reported examples include eyebrow raise or lower for question types, lexical mouth actions that accompany particular signs, and subtle cheek and eye changes that occur even during obligatory mouth gestures such as PAH in ASL. For generated sign language and avatars, these channels are difficult because multiple processes can affect the same facial regions asynchronously. Raised brows can simultaneously encode a polar question or surprise depending on context; mouthing and emotional expressions share the lower face; and grammatical non-manuals and emotive signals may start and end at different times (McDonald et al., 11 Aug 2025).
This problem is not merely one of adding facial animation. Emotional content is described as difficult to specify consistently because discrete emotion labels such as Ekman’s seven are too coarse for the richness of sign language discourse and do not blend well with other facial processes. Low-level FACS or AU rigs, and very high-control rigs such as Metahuman, are reported as non-intuitive for linguistic annotators and difficult to combine without conflicts. Cultural and language variation further modifies the facial realization of similar communicative intents, making label-based approaches brittle across communities. A recurring implication is that a signing avatar cannot treat affect as an isolated overlay; it must reconcile affect with linguistic facial structure at the level of concurrent control over facial subregions (McDonald et al., 11 Aug 2025).
2. Avatar architecture and layered facial animation
Paula’s facial system is described as bone- and skin-based facial animation originally inspired by FACS, but expanded beyond AUs to enable fine-grained motions, including eyelid shape control not covered by the original FACS. The artist-facing interface provides over 60 controls for the face, including eyebrow movement, eyelid openness and shape, lip corner up or down, lip pucker or spread, and jaw opening. This high-control rig is intended for precision, but the animation pipeline places those controls within a layered approach that organizes concurrent linguistic and emotive processes and allows controlled blending on shared facial regions (McDonald et al., 11 Aug 2025).
The layered design is particularly important for the lower face. Prior work on Paula is described as outlining methods to reconcile conflicts between mouth actions, including mouthing and lexical mouth gestures, and other facial expressions. Manual and non-manual channels are coordinated via layering, and onset and duration can differ across processes. Facial emotion must blend with grammatical non-manuals and lexical mouth actions; accordingly, the system avoids a single “brows tier” and instead manages multiple influences on each facial subregion. This architecture positions Paula as a linguistically driven avatar rather than a facial rig driven by one monolithic control stream (McDonald et al., 11 Aug 2025).
3. Pleasure–Arousal modelling and the EASIER notation
The emotional model applied to Paula uses the core affect dimensions of PAD, specifically Pleasure and Arousal , while noting Dominance for future work. These parameters represent affective valence and intensity or activation. For ease of linguistic annotation, values of each of and are restricted to the set , yielding nine “corner cases” in a grid. An animator trained in sign language facial expression created nine prototype facial poses for Paula, one for each . AffectNet images were used as visual references: for each corner case, the 10 closest images, measured by Euclidean distance in valence-arousal space, were selected to guide the artist (McDonald et al., 11 Aug 2025).
At runtime, Paula uses a discrete selection scheme rather than an explicit parametric formula. The system chooses the prototype pose indexed by and blends it with other facial layers, including lexical mouth gestures and grammatical non-manuals, using Paula’s layered blending system. The paper does not provide explicit formulas for mapping and 0 to individual facial controls such as eyebrow angle or mouth-corner weights, and it does not specify exact interpolation formulas or ramp times. Intermediate expressions are identified as future work through interpolation of bone rotations; the current implementation uses discrete extremes with blending to combine layers (McDonald et al., 11 Aug 2025).
EASIER is the textual representation that enables annotators to specify Paula’s emotive non-manuals via 1 and 2. The paper states that EASIER lets users control Paula’s emotional expressions with two numerical parameters, but it does not provide a formal grammar or example lines. The operational model is per-segment or per-clause assignment of 3 and 4, with linguistic annotations providing segment timing and EASIER adding emotional settings. Neutral is treated as the baseline at 5, and current recommendations are to use extreme values for clarity, default to 6 when emotion is not relevant, and rely on the blend system for transitions (McDonald et al., 11 Aug 2025).
4. Multilingual evaluation and annotation practice
The reported evaluation animated five common service sentences in English translated into Greek, German, and French sign languages on Paula. Cross-language differences in emotive non-manuals were evident. For “Could you repeat that?” DGS and LSF used furrowed brows, whereas GSL used raised brows. Apology expressions differed across the three languages, and smile intensity varied in greetings. These observations are central because they show that a single named emotion category is not sufficient to capture the realization of affective and grammatical facial behavior across signing communities (McDonald et al., 11 Aug 2025).
The outcome is described qualitatively rather than quantitatively. Using the 7 and 8 dimensions, described as descriptive rather than label-based, helped annotators specify appropriate facial animation across four language communities, and the approach was judged effective for producing facial expressions deemed appropriate by the groups. The paper does not report quantitative metrics. It does, however, characterize Paula’s Pleasure–Arousal approach as parsimonious and intuitive for annotators, more flexible than discrete labels, and culturally adaptable because it describes form via valence and arousal rather than naming “anger” or “sadness.” The corresponding trade-offs are also explicit: the current system is limited to discrete extremes, Dominance is absent, the mapping is face-centric, explicit head and torso cues for dominance are not yet integrated, and interaction with lexical mouthing remains an active area (McDonald et al., 11 Aug 2025).
5. Scripting, notation, and prompt-driven production
Paula belongs to a longer tradition of avatar-independent sign animation in which movement is specified abstractly rather than as raw motion clips. SiGML, based on HamNoSys, provides a compact, human-readable, avatar-independent description of posture and movement from which concrete motion data are synthesized in real time for any target avatar. HamNoSys records handshape, hand location, hand orientation, movement, and non-manual signals in tiers for shoulders, torso, head, eyes, facial expression, and mouth. The SiGML runtime resolves named locations to avatar-specific coordinates, synthesizes wrist orientation and path, applies inverse kinematics, maps handshape descriptions to joint parameters, and schedules overlapping non-manual tiers with attack–sustain–release envelopes. The paper does not name Paula specifically, but it states that SiGML is avatar-independent: as long as Paula’s rig exposes the required skeleton data, morphs, and named locations, the same SiGML drives her in real time (Kennaway, 2015).
This notation-centered view now coexists with prompt-driven production systems. SignAvatars introduces a large-scale, multi-prompt 3D sign language motion dataset comprising 70,000 videos from 153 signers, totaling 8.34 million frames, covering both isolated signs and continuous, co-articulated signs, with prompts including HamNoSys, spoken language, words, and glosses. Its benchmark supports 3D sign language production from text scripts, individual words, and HamNoSys notation, using SMPL-X and MANO annotations and a SignVAE baseline with semantic and motion codebooks. The paper further states that the resulting SMPL-X motion can be retargeted to a custom avatar such as Paula by mapping joints, converting 6D rotations, and driving hand and facial channels, although exact rig-mapping procedures are not provided (Yu et al., 2023).
6. Reconstruction and generative modelling around Paula
Data-driven methods extend Paula beyond rule-based scripting. SGNify addresses sign language capture from in-the-wild monocular videos by fitting SMPL-X with linguistic priors grounded in sign phonology and Battison’s symmetry and dominance conditions. It reconstructs hand pose, facial expression, and body movement from isolated signs and reports quantitative gains over FrankMocap, PIXIE, PyMAF-X, and an adapted SMPLify-X baseline on sign-language videos. On the reported dataset, SGNify achieves mean translationally-aligned vertex-to-vertex error of 55.63 mm for the upper body, 19.22 mm for the left hand, and 17.50 mm for the right hand, with perceptual recognition rates that are significantly more comprehensible and natural than those of previous methods and not significantly different from real video. The practical pipeline explicitly describes retargeting reconstructed SMPL-X motion to a stylized avatar such as Paula (Forte et al., 2023).
Generative motion models address the converse problem: producing sign motion from semantic inputs. SignAvatar uses a transformer-based conditional variational autoencoder conditioned by CLIP embeddings of gloss words or images and targets isolated word-level signs with SMPL-X body, hand, and face parameters. Reported results include reconstruction accuracy improvements from 0.818 to 0.952 on the full 103-word setting with curriculum learning, and generation accuracy improvements from 0.511 to 0.733 on the same setting. The framework is presented as suitable for retargeting to a Paula Signing Avatar in real time, particularly for dictionary-like isolated-word generation (Dong et al., 2024).
A later strand of work shifts from rig-level motion generation to photorealistic signer synthesis. “Diverse Signer Avatars with Manual and Non-Manual Feature Modelling for Sign Language Production” proposes a Stable Diffusion 2.1 backbone with ControlNet-like conditioning, a sign feature aggregation module for manual and non-manual cues, Sapiens foundation features, and a reference-image branch for signer identity. The paper does not reference Paula by name, but it states that the identity-agnostic pipeline readily generalizes to any named avatar, including Paula. On YouTube-SL-25, it reports PSNR 22.71, SSIM 0.8635, and LPIPS 0.1139, along with a user study in which native BSL interpreters rated the method higher than AnimateAnyone and PENet for both accuracy and realism. A central observation is that perceived aesthetics by non-signers do not correlate with sign correctness, so native signer feedback remains essential (Lakhal et al., 21 Aug 2025).
7. Mixed-reality customization, controversy, and open directions
Paula has also been framed as a mixed-reality interpreter for deaf–hearing communication. In a participatory design study with 15 DHH college students, three spatial blend patterns were identified: no overlay, partial overlay, and complete overlay. No overlay was valued for preserving clear human-human eye contact and picture-in-picture control. Partial overlay was preferred by most because it balanced intelligibility of generated content with “human touch” and authenticity. Complete overlay was controversial; several participants described it as “dishonest” because it can falsely imply that the hearing person is signing or obscure the real person’s identity. Across conditions, participants emphasized keeping the face, especially the eyes, as the locus of real-person contact and selectively overlaying facial movements and hands for language rather than replacing the torso (Chen et al., 2024).
The same study specifies the customization controls expected of a system such as Paula. These include adjustable signing speed, sign duration, transition timing, and signing box size; region-specific and community-accepted variants; correct handshapes and smooth transitions; accurate modelling of eyebrows, head tilts, eye gaze, and facial grammar; optional mouthings; controllable voice timbre, pitch, and rate; and tight synchronization of hands, face, and voice. The proposed interface includes per-region overlay sliders, style presets such as “Neutral,” “Professional,” and “Expressive,” caption-only mode, and explicit consent gates for complete overlay. The design recommendations also insist that eye contact remain primary, that complete overlay require mutual consent, and that generated voices avoid mimicking “deaf accent” (Chen et al., 2024).
Open research directions for Paula follow directly from the current limitations of the emotional model. The reported agenda is to support intermediate, continuous values of 9 and 0 via interpolation of bone rotations; add Dominance and expand the prototypes from 9 to 27; integrate torso, shoulder, and potentially head and neck postures tied to dominance; refine lower-face blending with mouthing and lexical mouth gestures; model eye gaze and head movements in tandem with emotion; couple automatic emotion recognition from spoken or text input to EASIER parameters; and continue iterative testing with Deaf researchers and communities to judge legibility, naturalness, and ease of specification. This suggests that Paula’s future development lies in combining annotation-friendly control, layered linguistic blending, community-led validation, and stronger coupling between manual and non-manual channels rather than in replacing one control vocabulary with another (McDonald et al., 11 Aug 2025).