- The paper shows that personalised prompts increased response duration by 70%, self-referential language from 29.2% to 67.8%, and positive sentiment from 13.5% to 29.3% compared with generic prompts.
- Conversational output declined during the second half of sessions, while higher first-session latency and hesitation were associated with participation in three or fewer sessions, supporting adaptive pacing and shorter activities.
- People living alone gave longer, faster, and more self-referential responses but also showed more self-corrections and cognitive breakdowns, highlighting the need to tailor robot interactions to social context.
Study context and motivation
Cognitive Stimulation Therapy (CST) is among the best-evidenced non-pharmacological interventions for mild-to-moderate dementia, but its delivery at scale is constrained by the need for trained facilitators and structured group settings. Individual CST (iCST) was developed to address this, yet adherence remains a persistent problem in home-based delivery. This paper by Akinrintoyo and Salomons examines the fine-grained conversational dynamics that arise when a socially assistive robot (SAR) delivers iCST in participants' own homes, moving beyond the acceptance and usability outcomes that have dominated prior SAR-for-dementia work (2607.07998).
The study addresses one hypothesis — that personalised prompts elicit higher conversational engagement than generic prompts (H1) — and three research questions: whether conversational behaviour changes within a session (RQ1), whether first-session conversational metrics predict long-term participation (RQ2), and whether living situation shapes engagement (RQ3).
System, deployment, and analysis pipeline
The Co-STAR system (Cognitive Stimulation Therapy by an Autonomous Robot) is built on a Misty platform with a tablet for visual prompts, a microphone, and a local mini-computer running a private cloud architecture; all speech processing and storage occurred on-device with no third-party cloud involvement, and no video data were collected. All robot utterances were pre-generated before deployment, and an adaptive response-latency mechanism adjusted turn-taking timing to each user's speaking pace.
Eight people with clinical dementia diagnoses (mean age 79.67 years) received 30-minute sessions over one week of in-home deployment, completing between one and eight sessions each (31 sessions total; four recordings lost to technical issues, leaving 27 analysed). Five turn-based activities were used: Word Association, Popular Places, Object Categorisation, Famous Faces, and Common Sayings, interleaved with open-ended, reminiscence-based follow-up prompts.
Transcription combined WhisperX for time-aligned ASR with WhisperD, a dementia-specific recogniser that preserves filler words ("uh", "um") with diagnostic relevance (2607.07998). Speaker diarisation used a pyannote-based pipeline, with Co-STAR's prompts as ground truth. A locally deployed Llama-3 model classified conversational markers (hesitation, self-correction, cognitive breakdown, self-reference) and categorised prompts as personalised or generic; all LLM outputs were manually verified. Metrics spanned speech production (words per turn, response duration, speech production rate, response latency), fluency/repair markers, and engagement markers (self-reference rate, positive sentiment). Non-parametric tests (Mann–Whitney U, chi-squared, Kruskal–Wallis) were used throughout, appropriate given the cohort size.
Personalised versus generic prompts
Personalisation produced the strongest effects in the study. With roughly balanced turn counts across conditions (174 vs 185), responses to personalised prompts were 70% longer than responses to generic prompts (17.3 vs 10.07 seconds, p<0.01). Self-referential language more than doubled (67.8% vs 29.2%, p<0.001), and positive sentiment also rose sharply (29.3% vs 13.5%, p<0.001).
Two nuances qualify this picture. First, words per turn showed only a non-significant trend (24.29 vs 17.32, p=0.113), so the duration effect is not fully mirrored in word count. Second, speech production rate was lower for personalised prompts (1.92 vs 2.08 words/s, p<0.05) while response latency did not differ (p=0.167). The authors interpret the slower pace as reflective autobiographical recall rather than retrieval difficulty, supported by the absence of significant differences in hesitation, self-correction, or cognitive breakdown rates. This interpretation is plausible but rests on null results from a small sample; it cannot be ruled out that personalisation does impose some processing cost. The design implication is direct: robots should prioritise prompts anchored in a PwD's autobiographical experience, since engagement gains occur without increased initiation delay.
Within-session dynamics
Comparing the first and second 15-minute halves of sessions revealed consistent decline: words per turn fell 21% (18.83 vs 14.74, p<0.001), speech production rate dropped (2.36 vs 2.12 words/s, p<0.001), and self-reference rate declined (42.8% vs 32.5%, p<0.001). Response latency and all fluency/repair markers remained stable. Qualitatively, early-session narratives were elaborated and temporally structured, whereas late-session responses became sparse and procedural ("Yes, it was good").
The authors attribute this pattern to cognitive fatigue, though they acknowledge it may equally reflect reduced engagement; the stable latency and repair rates argue against increased retrieval difficulty specifically. Either way, the result supports phase-adaptive session design: high-cognitive-load activities such as autobiographical recall should be front-loaded, with lower-load content later, and shorter, more frequent sessions may outperform single 30-minute blocks.
First-session metrics as predictors of sustained participation
Splitting participants into low-engagement (≤3 sessions, N=4) and high-engagement (>3 sessions, N=4) groups yielded a distinct first-session profile for those who disengaged: higher response latency (2.70 vs 1.63 s, p<0.001), slower speech production rate (2.066 vs 2.270 words/s, p<0.0010), and higher hesitation rates (8.2% vs 3.7%, p<0.0011).
A notable and somewhat counterintuitive finding is that low-engagement participants expressed more positive sentiment (21.5% vs 14.7%, p<0.0012), indicating that affective expression does not align directly with sustained engagement — a caution against using sentiment alone as an engagement proxy. Self-reference rates did not differentiate groups, unlike in the personalisation and session-phase analyses, suggesting autobiographical framing reflects the interaction itself rather than disposition toward continued use. With four participants per group, these associations are preliminary; the paper frames them as evidence that real-time detection of elevated latency could trigger adaptive simplification of prompts, but predictive validity has not been established.
Influence of living situation
Participants living alone produced significantly longer responses (21.9 vs 15.5 WPT, p<0.0013; 16.8 vs 9.1 s per turn, p<0.0014), responded faster (latency 1.55 vs 2.24 s, p<0.0015), and used more self-referential language (47.9% vs 32.2%, p<0.0016) than those living with a family caregiver. However, this greater narrative engagement came with significantly more self-corrections (7.6% vs 2.3%, p<0.0017) and cognitive breakdowns (8.8% vs 3.2%, p<0.0018) — a trade-off between expressive engagement and conversational difficulty.
The authors offer two candidate explanations for the caregiver-cohabiting group's shorter, faster responses: greater familiarity with regular social interaction, or interruptions from caregivers present during sessions. These alternatives are not disentangled, which matters because the second explanation would be an artefact of the recording environment rather than a property of the participant. Practically, the finding argues for incorporating living situation into user profiling: solo-living users may benefit from robots that encourage extended storytelling, while cohabiting users may be better served by the robot acting as a conversational facilitator.
Limitations and open questions
The paper concedes that the sample of eight participants limits statistical power and generalisability, and several specific caveats bear on the results. The engagement-group comparison uses four participants per group with an arbitrary three-session threshold. Four recordings were lost to technical failures, and their exclusion is not analysed for bias. The fatigue interpretation of within-session decline is confounded with reduced engagement, and the caregiver-presence explanation for living-situation differences is unresolved. Whether first-session latency and hesitation genuinely predict disengagement — rather than merely correlating with it in this small cohort — requires validation in longer deployments with larger cohorts, which the authors identify as the necessary next step.
Conclusion
This study provides an empirical characterisation of conversational engagement in robot-delivered iCST, showing that prompt personalisation substantially increases response duration and autobiographical expression, that verbal output declines within sessions consistent with fatigue, that first-session latency and hesitation mark subsequent disengagement, and that living situation shapes both the quantity and difficulty of conversational participation. The findings translate into concrete design guidance — autobiographically anchored prompts, phase-based task sequencing, latency-triggered adaptation, and socially contextualised interaction strategies — while leaving open the validation of these effects at scale.