Papers
Topics
Authors
Recent
Search
2000 character limit reached

Factors Influencing Conversational Engagement in Robot-Delivered Individual Cognitive Stimulation Therapy (iCST) for Dementia in Home Settings

Published 9 Jul 2026 in cs.RO | (2607.07998v1)

Abstract: Social robots offer a promising means of supporting cognitive therapies for dementia care by guiding structured conversation and therapeutic activities. However, little is known about the conversational dynamics that emerge during robot-delivered cognitive stimulation therapy (CST) sessions. This study analysed the interaction patterns from robot-delivered individual CST (iCST) sessions conducted with people living with dementia in home settings. Our Co-STAR (Cognitive Stimulation Therapy by an Autonomous Robot) system was deployed in the homes of eight PwDs for one week, who completed 30-minute sessions. Conversational metrics, including words per turn, speech production rate, response duration, response latency, and self-referential language, were analysed to examine how conversational engagement is shaped by prompt personalisation, interaction phase, and participant characteristics. The findings highlight three key interactional properties of robot-delivered iCST. First, personalised prompts significantly increase response duration, self-referential language, and overall engagement compared to generic prompts. Second, conversational behaviour changes within sessions, with a reduction in the verbal output and autobiographical engagement observed during later interaction phases, which suggests cognitive fatigue. Third, first-session conversational metrics were associated with long-term participation, while living situation influenced conversational engagement patterns. These findings provide empirical insights into the factors that shape conversational engagement in robot-delivered iCST. They inform the design of adaptive conversational robots for dementia therapy.

Summary

  • The paper shows that personalised prompts increased response duration by 70%, self-referential language from 29.2% to 67.8%, and positive sentiment from 13.5% to 29.3% compared with generic prompts.
  • Conversational output declined during the second half of sessions, while higher first-session latency and hesitation were associated with participation in three or fewer sessions, supporting adaptive pacing and shorter activities.
  • People living alone gave longer, faster, and more self-referential responses but also showed more self-corrections and cognitive breakdowns, highlighting the need to tailor robot interactions to social context.

Study context and motivation

Cognitive Stimulation Therapy (CST) is among the best-evidenced non-pharmacological interventions for mild-to-moderate dementia, but its delivery at scale is constrained by the need for trained facilitators and structured group settings. Individual CST (iCST) was developed to address this, yet adherence remains a persistent problem in home-based delivery. This paper by Akinrintoyo and Salomons examines the fine-grained conversational dynamics that arise when a socially assistive robot (SAR) delivers iCST in participants' own homes, moving beyond the acceptance and usability outcomes that have dominated prior SAR-for-dementia work (2607.07998).

The study addresses one hypothesis — that personalised prompts elicit higher conversational engagement than generic prompts (H1) — and three research questions: whether conversational behaviour changes within a session (RQ1), whether first-session conversational metrics predict long-term participation (RQ2), and whether living situation shapes engagement (RQ3).

System, deployment, and analysis pipeline

The Co-STAR system (Cognitive Stimulation Therapy by an Autonomous Robot) is built on a Misty platform with a tablet for visual prompts, a microphone, and a local mini-computer running a private cloud architecture; all speech processing and storage occurred on-device with no third-party cloud involvement, and no video data were collected. All robot utterances were pre-generated before deployment, and an adaptive response-latency mechanism adjusted turn-taking timing to each user's speaking pace.

Eight people with clinical dementia diagnoses (mean age 79.67 years) received 30-minute sessions over one week of in-home deployment, completing between one and eight sessions each (31 sessions total; four recordings lost to technical issues, leaving 27 analysed). Five turn-based activities were used: Word Association, Popular Places, Object Categorisation, Famous Faces, and Common Sayings, interleaved with open-ended, reminiscence-based follow-up prompts.

Transcription combined WhisperX for time-aligned ASR with WhisperD, a dementia-specific recogniser that preserves filler words ("uh", "um") with diagnostic relevance (2607.07998). Speaker diarisation used a pyannote-based pipeline, with Co-STAR's prompts as ground truth. A locally deployed Llama-3 model classified conversational markers (hesitation, self-correction, cognitive breakdown, self-reference) and categorised prompts as personalised or generic; all LLM outputs were manually verified. Metrics spanned speech production (words per turn, response duration, speech production rate, response latency), fluency/repair markers, and engagement markers (self-reference rate, positive sentiment). Non-parametric tests (Mann–Whitney U, chi-squared, Kruskal–Wallis) were used throughout, appropriate given the cohort size.

Personalised versus generic prompts

Personalisation produced the strongest effects in the study. With roughly balanced turn counts across conditions (174 vs 185), responses to personalised prompts were 70% longer than responses to generic prompts (17.3 vs 10.07 seconds, p<0.01p < 0.01). Self-referential language more than doubled (67.8% vs 29.2%, p<0.001p < 0.001), and positive sentiment also rose sharply (29.3% vs 13.5%, p<0.001p < 0.001).

Two nuances qualify this picture. First, words per turn showed only a non-significant trend (24.29 vs 17.32, p=0.113p = 0.113), so the duration effect is not fully mirrored in word count. Second, speech production rate was lower for personalised prompts (1.92 vs 2.08 words/s, p<0.05p < 0.05) while response latency did not differ (p=0.167p = 0.167). The authors interpret the slower pace as reflective autobiographical recall rather than retrieval difficulty, supported by the absence of significant differences in hesitation, self-correction, or cognitive breakdown rates. This interpretation is plausible but rests on null results from a small sample; it cannot be ruled out that personalisation does impose some processing cost. The design implication is direct: robots should prioritise prompts anchored in a PwD's autobiographical experience, since engagement gains occur without increased initiation delay.

Within-session dynamics

Comparing the first and second 15-minute halves of sessions revealed consistent decline: words per turn fell 21% (18.83 vs 14.74, p<0.001p < 0.001), speech production rate dropped (2.36 vs 2.12 words/s, p<0.001p < 0.001), and self-reference rate declined (42.8% vs 32.5%, p<0.001p < 0.001). Response latency and all fluency/repair markers remained stable. Qualitatively, early-session narratives were elaborated and temporally structured, whereas late-session responses became sparse and procedural ("Yes, it was good").

The authors attribute this pattern to cognitive fatigue, though they acknowledge it may equally reflect reduced engagement; the stable latency and repair rates argue against increased retrieval difficulty specifically. Either way, the result supports phase-adaptive session design: high-cognitive-load activities such as autobiographical recall should be front-loaded, with lower-load content later, and shorter, more frequent sessions may outperform single 30-minute blocks.

First-session metrics as predictors of sustained participation

Splitting participants into low-engagement (≤3 sessions, N=4) and high-engagement (>3 sessions, N=4) groups yielded a distinct first-session profile for those who disengaged: higher response latency (2.70 vs 1.63 s, p<0.001p < 0.001), slower speech production rate (2.066 vs 2.270 words/s, p<0.001p < 0.0010), and higher hesitation rates (8.2% vs 3.7%, p<0.001p < 0.0011).

A notable and somewhat counterintuitive finding is that low-engagement participants expressed more positive sentiment (21.5% vs 14.7%, p<0.001p < 0.0012), indicating that affective expression does not align directly with sustained engagement — a caution against using sentiment alone as an engagement proxy. Self-reference rates did not differentiate groups, unlike in the personalisation and session-phase analyses, suggesting autobiographical framing reflects the interaction itself rather than disposition toward continued use. With four participants per group, these associations are preliminary; the paper frames them as evidence that real-time detection of elevated latency could trigger adaptive simplification of prompts, but predictive validity has not been established.

Influence of living situation

Participants living alone produced significantly longer responses (21.9 vs 15.5 WPT, p<0.001p < 0.0013; 16.8 vs 9.1 s per turn, p<0.001p < 0.0014), responded faster (latency 1.55 vs 2.24 s, p<0.001p < 0.0015), and used more self-referential language (47.9% vs 32.2%, p<0.001p < 0.0016) than those living with a family caregiver. However, this greater narrative engagement came with significantly more self-corrections (7.6% vs 2.3%, p<0.001p < 0.0017) and cognitive breakdowns (8.8% vs 3.2%, p<0.001p < 0.0018) — a trade-off between expressive engagement and conversational difficulty.

The authors offer two candidate explanations for the caregiver-cohabiting group's shorter, faster responses: greater familiarity with regular social interaction, or interruptions from caregivers present during sessions. These alternatives are not disentangled, which matters because the second explanation would be an artefact of the recording environment rather than a property of the participant. Practically, the finding argues for incorporating living situation into user profiling: solo-living users may benefit from robots that encourage extended storytelling, while cohabiting users may be better served by the robot acting as a conversational facilitator.

Limitations and open questions

The paper concedes that the sample of eight participants limits statistical power and generalisability, and several specific caveats bear on the results. The engagement-group comparison uses four participants per group with an arbitrary three-session threshold. Four recordings were lost to technical failures, and their exclusion is not analysed for bias. The fatigue interpretation of within-session decline is confounded with reduced engagement, and the caregiver-presence explanation for living-situation differences is unresolved. Whether first-session latency and hesitation genuinely predict disengagement — rather than merely correlating with it in this small cohort — requires validation in longer deployments with larger cohorts, which the authors identify as the necessary next step.

Conclusion

This study provides an empirical characterisation of conversational engagement in robot-delivered iCST, showing that prompt personalisation substantially increases response duration and autobiographical expression, that verbal output declines within sessions consistent with fatigue, that first-session latency and hesitation mark subsequent disengagement, and that living situation shapes both the quantity and difficulty of conversational participation. The findings translate into concrete design guidance — autobiographically anchored prompts, phase-based task sequencing, latency-triggered adaptation, and socially contextualised interaction strategies — while leaving open the validation of these effects at scale.

Paper to Video (Beta)

No one has generated a video about this paper yet.

Whiteboard

No one has generated a whiteboard explanation for this paper yet.

Open Problems

We haven't generated a list of open problems mentioned in this paper yet.

Tweets

Sign up for free to view the 1 tweet with 0 likes about this paper.