- The paper demonstrates a low-cost, autonomous storytelling system that synchronizes audio-driven puppet mouth motion, neck articulation, narration, and ProMP-based arm gestures on a Baxter robot.
- The paper reports more favorable social-affective ratings and higher, less variable story recall for puppet-based delivery than a gesture-only baseline in a preliminary pilot of 10 adults, without inferential statistical testing.
- The paper shows practical portability by transferring the system to an ABB IRB 120 with a minor adapter, approximately 45 minutes of hardware work, reconfigured motion parameters, and a CSV-to-RAPID software bridge.
RoboTales is a low-cost robotic storytelling system that uses expressive sock puppetry to deliver character-driven narratives autonomously. Implemented on a Baxter robot as a demonstration platform, the system synchronizes narration, arm gestures, and puppet mouth articulation, and was evaluated in a small pilot study against a gesture-only baseline. The work contributes a reusable design pattern for embodied robotic storytelling rather than a novel learning-theoretic claim, and its empirical evidence is correspondingly preliminary.
Motivation and positioning
The authors situate RoboTales at the intersection of two largely separate traditions: social assistive robotics (SAR) and puppetry. Prior robotic storytellers typically rely on screen-based avatars or rigid gesture libraries, and often require human supervision; such systems yield engagement but inconsistent gains in comprehension and recall. Puppetry research, by contrast, shows benefits for attention, emotional engagement, and memory in children. The paper's central design hypothesis is that physically embodied puppets mounted on general-purpose manipulators can bridge this gap: sock puppets require only minimal actuation (mouth and neck), are mechanically robust, and provide intuitive anthropomorphic cues, unlike the complex marionettes that dominate prior robotic puppetry work.
Design goals and system architecture
Three design goals structure the work: G1, synchronization of narration with physical embodiment (jaw, neck, gestures) to convey character identity and emotion; G2, use of affordable, maintainable components fabricable by non-experts; and G3, modularity enabling multiple stories/characters and transfer across robot platforms. Each goal is paired with a success criterion (S1–S3) covering engagement/comprehension, autonomous operation, and deployability.
Hardware. Each puppet end-effector has two degrees of freedom: an MX-64 Dynamixel servo drives a nodding neck and an MX-28 drives jaw articulation. Parts are CAD-modeled in SolidWorks and 3D-printed in PLA. Two distinct characters—Barneby (energetic, scatterbrained) and Fitzwilliam (cautious, organized)—provide contrasting visual and behavioral cues intended to aid character recognition.
Software. Mouth motion is audio-driven without speech recognition: narration is segmented into RMS windows, filtered with a deadband and exponential moving average to suppress noise, then mapped to pre-calibrated jaw angle limits. Arm gestures are generated with probabilistic movement primitives (ProMPs), which introduce controlled variation and smoothness so that each character's personality is expressed through distinct motion styles; because ProMP trajectories are time-agnostic, they blend naturally with narration timing. In the non-puppet baseline mode, Baxter instead uses viseme-based lip sync on its screen face, with ambient conversational gestures triggered at 60% probability per opportunity to avoid over-animation.
Pilot study
A between-subjects pilot with ten adult participants (Mage​=23.20, SD=2.90) compared the puppet condition against a gesture-only condition on the same five-minute story. Measures were the Human–Robot Interaction Evaluation Scale (HRIES, 7-point Likert across 16 social/affective items) and a weighted story-recall rubric over six questions of graded difficulty.
Participants in the puppet condition gave more favorable HRIES affective and social perception ratings and achieved higher story recall with lower variance than those in the gesture-only condition. The authors explicitly state the pilot was not designed for statistical inference, so these results should be read as consistent directional trends rather than confirmed effects—a limitation compounded by the small sample and the fact that participants were adults, whereas the intended users are children. On the systems side, all sessions ran fully autonomously without operator intervention (S2). Portability was demonstrated by mounting Barneby on an ABB IRB 120: the transition required only a minor adapter modification, roughly 45 minutes of hardware workflow, reconfigured six-DoF ProMP parameters, and a CSV parser bridging to ABB's RAPID environment (S3).
Limitations and open questions
The paper concedes several constraints plainly. The pilot sample of ten adults cannot support statistical claims about engagement or recall, and no inferential tests are reported; validation with children in actual deployment contexts (classrooms, libraries, community settings) awaits IRB approval and age-appropriate protocols. Mouth-timing alignment remains coarse, since the amplitude-driven jaw mapping does not perform true phoneme-level synchronization. The comparison space is narrow—a single puppet-versus-gesture contrast—leaving open how specific movement qualities (timing, amplitude, motion style drawn from traditional puppetry) causally contribute to expressiveness. Finally, the HRIES differences favoring the puppet condition are reported as trends; which of the sixteen perception dimensions drive the effect is not established.
Conclusion
RoboTales demonstrates that simple two-DoF sock-puppet end-effectors, combined with audio-driven jaw actuation and ProMP-based gesture generation, suffice for autonomous, expressive robotic storytelling at very low cost and with demonstrated cross-platform portability. Its preliminary findings—that puppet embodiment improves perceived social-affective qualities and story recall relative to gesture-only delivery—are promising but underpowered, and the system's value for its target population of children remains to be validated in situ.