- The paper introduces PuppetAI, a platform for expressive human-robot interaction using a customizable puppet-inspired robot and a four-layer control system.
- The robot's modular design allows for scalable bending angles and economic actuation, while the software stack decouples speech perception, emotion synthesis, gesture scheduling, and motor control.
- The platform demonstrates a loop where user speech triggers affective gestures via an LLM, but lacks empirical validation and user studies.
PuppetAI is a modular soft robot interaction platform that couples a scalable cable-driven actuation system with a puppet-inspired continuum-robot gesture framework, together with a four-layer decoupled software stack. The authors position it as an accessible, customizable foundation for tactile-based expressive HRI research, and validate it through a demonstration loop in which an LLM translates human vocal input into parameterized affective gestures (2602.04787).
Motivation and positioning
The paper identifies two gaps. First, prior robotic puppetry systems—such as HinHRob for glove puppetry, PINOKY's ring-based plush-toy animation, and telepresence puppetry interfaces—are typically built on rigid internal structures with specialized manipulation mechanisms. They accommodate only limited morphological variation across puppet sizes and structures, and some employ complex or closed-source designs that hinder reproducibility. Second, cable-driven continuum robots offer compliance, remote actuator placement (reducing moving mass), and modular segment-based configurability, but the literature has focused on mechanical modeling and manipulation rather than emotion-oriented HRI; such systems have rarely been deployed as functional platforms for interaction studies.
Mechanical design
The platform consists of a modular actuation base driving cables routed through a TPU-95 soft framework housed in a plush puppet body. The current embodiment has a torso and two arms, each supporting spatial motion via two orthogonal bending planes separated by 90°. Non-deformable sections connect to deformable ones via mortise-and-tenon joints.
Customization follows a simple geometric scheme: a deformable section of length Lflex is divided into n cutting units, each contributing a bending angle θ=θmax/n. In the implemented design, each arm uses a 122 mm deformable section with five cutting units at 30° per unit, achieving up to 150° of vertical or forward swing; the body uses a 100 mm deformable section with three units at 15° per unit, supporting bending up to 45°. Because Lflex and n are free parameters, the same fabrication approach scales to other puppet sizes and gesture ranges without added structural complexity.
Actuation is deliberately economical: within each plane, one motor-driven cable produces bending while an elastic rope on the opposite side supplies passive restoring force, halving the number of actively controlled inputs per DoF. Motors are mounted at the base, keeping moving mass low—an arrangement the authors argue improves safety in human-centered settings. The number of motors is scalable to match research requirements.
Software architecture
Control software runs on a host computer connected over USB-B and is organized into four decoupled layers:
- Perceptual processing: SenseVoiceSmall performs speech-to-text and vocal-emotion estimation.
- Affective modeling: a prompted ChatGPT-based LLM synthesizes action sequences by retrieving motions from a predefined gesture library, emitting a structured format such as
[Waving] [2] [Happy] [2] where numeric fields encode inter-gesture intervals or continuous-motion durations.
- Action sequence scheduler: manages execution ordering of motion sequences.
- Low-level actuation control: handles phase-shifted multi-motor channel control, sinusoidal displacement profiles, and torque limit detection.
The decoupling allows any layer to be refined or replaced independently—for example, swapping the LLM backend or the perceptual model without touching actuation code—which the authors present as the platform's principal value as a research testbed.
Expressive gestures and the affective expression loop
Gesture design is grounded in professional puppetry practice: two researchers performed open coding on video and text archives of puppet performances, then compiled a motion library split into discrete gestures (waving, joy, sadness, hug, confusion) and continuous states (dancing). The neutral idle pose is bilateral arm extension slightly below shoulder height; when unengaged, the robot performs subtle randomized lateral swaying to preserve liveliness.
The demonstrated interaction loop proceeds from user speech through transcription and emotion analysis to LLM-generated action sequences executed as gestural feedback, which in turn shapes subsequent user responses. The behavioral model implements empathetic resonance: positive user affect elicits celebratory sequences ([Joy] [1] [Dancing] [3]), negative affect sympathetic ones ([Sadness] [1] [Hug] [3]). Interaction is currently triggered by push-to-talk button presses rather than hands-free turn-taking.
Assessment
The paper's contribution is architectural and infrastructural rather than empirical. Its strengths lie in the parametric bend-segment customization (θmax/n per unit), single-motor-per-plane actuation with elastic restoring force, interchangeable plush exteriors intended for tactile studies, and open modularity aimed at lowering cost barriers for smaller laboratories.
Several limitations should be noted plainly. No user study or quantitative evaluation of perceived emotional fidelity, gesture legibility, or interaction quality is reported; validation rests entirely on a small set of illustrative dialogue scenarios. The affective behavior is reactive resonance only—the system does not attempt active emotion regulation—and the mapping from LLM output to gesture sequences depends on a proprietary ChatGPT backend, which partially conflicts with the reproducibility concerns the authors raise about closed-source platforms. Voice input requires manual push-to-talk, and the tactile dimension central to the platform's motivation ("pleasant-to-touch" exteriors) remains unexploited in the demonstrated loop. Open questions left by the paper include whether puppetry-derived discrete/continuous gesture taxonomies generalize across embodiments, how LLM-selected sequences perform against hand-authored baselines in perception studies, and how latency and turn-taking behave once voice activity detection replaces push-to-talk. Planned extensions include subdermal pneumatic tactile systems and spatialized synthetic speech.
Conclusion
PuppetAI contributes a configurable hardware–software substrate for expressive social robotics: a scalable cable-driven continuum framework with parametrically adjustable bending segments, paired with a four-layer decoupled control pipeline integrating speech perception, LLM-based emotion synthesis, gesture scheduling, and motor execution. Its demonstrated affective expression loop shows end-to-end functionality but not yet validated effectiveness. The platform's main claim—that reduced complexity and cost can broaden access to tactile-rich affective HRI research—is plausible and well-motivated, though its evidentiary support is currently limited to system description and anecdotal scenarios.