Papers
Topics
Authors
Recent
Search
2000 character limit reached

Every(bot) Makes Mistakes: Coding Big Five Personalities, Context, and Tone into an LLM Chatbot Recovery Code Framework

Published 6 May 2026 in cs.HC | (2605.05391v1)

Abstract: Despite careful design involving classifiers, parameters, and safeguarding, errors during human/AI interaction are not rare. Poor error recovery can disrupt interaction flow, damage user trust, and decrease user engagement. Whilst existing work has explored LLM recovery, tone, context, and personality as separate design dimensions, no existing work has combined these variables into a structured guidance framework. This paper presents a recovery code that maps four common LLM chatbot task contexts to associated personality traits (four Big Five personalities: Conscientiousness, Agreeableness, Openness, and Extraversion), tones, and three-stage recovery instructions. A recovery evaluation rubric was also designed, comprising three dimensions (Recovery quality, Tone alignment, and Appropriateness) and nine sub-dimensions. The methodology is exploratory, with no participants used. A between-subjects design was employed across two conditions: Condition A (baseline, uncoded), four separate Claude Sonnet 4.6 agents received no recovery code training; Condition B (coded), four separate Claude Sonnet 4.6 models were trained on the recovery code. Identical 'user' prompts and error scenarios were used across both conditions. Eight LLM evaluator agents assessed the recovery responses using the evaluation rubric, producing scores out of 5 for each sub-dimension. Results found a 27.8% average performance increase in coded recovery responses (76.7%) compared to baseline responses (48.9%). Condition B performed strongest in the appropriateness dimension (83.3%), with notable improvement in personality appropriateness (75% versus 50%) and providing explanation (60% versus 20%). These findings suggest that structured personality, context, and tone-informed recovery codes can be successfully learnt and applied by LLM chatbots to improve error recovery quality across varying contextual tasks.

Authors (3)

Summary

  • The paper introduces a recovery code framework that maps four chatbot contexts to Big Five traits, task-specific tones, and three-stage instructions for identifying errors, reassuring users, and continuing the conversation.
  • Claude Sonnet 4.6 agents using the framework scored 76.7% overall versus 48.9% for baseline agents, with especially large gains in recovery quality, appropriateness, and explanation-giving.
  • The exploratory study suggests structured personality conditioning can improve error recovery, but human evaluation, diverse models, realistic errors, and adversarial user testing are needed to confirm generalizability.

Overview and motivation

LLM chatbot errors—hallucinations, omissions, miscommunication, and contextual mismatches—are an unavoidable feature of generative systems, even those equipped with classifiers, RLHF-based alignment, and safeguarding parameters. Prior work has established that poor recovery from these errors carries concrete costs: blunt refusals degrade user satisfaction, missing explanations damage trust, and post-error silence can be interpreted by users as social withdrawal or abandonment. Existing research has examined tone, personality conditioning, context-dependence, and recovery style as separate design dimensions, but no prior framework combines all four into a single structured guidance protocol. This paper addresses that gap with two artifacts: a recovery code framework that maps four task contexts to Big Five personality traits, tones, and three-stage recovery instructions, and a recovery evaluation rubric spanning three dimensions and nine sub-dimensions. The central empirical claim is that Claude Sonnet 4.6 agents trained on this code produce recoveries scoring 76.7% on average versus 48.9% for untrained baseline agents—a 27.8 percentage-point improvement.

The recovery code framework

The framework maps four common LLM task contexts—derived from usage studies indicating writing, social support, creative work, and learning dominate chatbot interactions—to four of Goldberg's Big Five traits, four literature-grounded tones, and four recovery scripts. Each recovery script instantiates the three-stage structure recommended in human/robot interaction research (identify the error, reassure the user, continue), with each stage phrased to embody the mapped trait's descriptors:

Context Trait Tone Code
Correcting grammar Conscientiousness Polite {C1; C; T1; R1}
Emotional support Agreeableness Warm {C2; A; T2; R2}
Brainstorming Openness Conversational {C3; O; T3; R3}
Learning a concept Extraversion Engaging {C4; E; T4; R4}

Neuroticism is deliberately excluded: injecting anxiety- or instability-associated intonation into a recovery protocol would be counterproductive, particularly given evidence that "verbal leakage" (a neuroticism-associated tone) triggers guilt in users. The design rationale draws on findings that affective-plus-cognitive combined feedback outperforms either style alone, and that no single recovery style dominates across contexts—motivating a context-indexed rather than universal recovery prescription.

Methodology

The study is explicitly exploratory, with no human participants. A between-subjects design compared Condition A (baseline agents receiving only the user prompt followed by the error prompt "I don't think that is right. Please try again.") against Condition B (agents first trained on the recovery code via a system-prompt-style training document, then given identical prompts and errors). Four synthetic user prompts were generated by Microsoft Copilot with memory disabled and instructions not to use real user data. Eight agent interactions were conducted (four per condition), and eight separate evaluator agents—also Claude Sonnet 4.6 instances with memory disabled—scored each transcript on nine sub-dimensions using a 1–5 Likert scale (maximum 45 per transcript). Condition B evaluators received the recovery code for transparency; Condition A evaluators did not.

The rubric's three dimensions are Recovery quality (identifying the error, reassuring the user, providing explanation, continued conversation), Tone alignment (alignment with task, naturalness), and Appropriateness (contextual relevance, personality appropriateness, tone appropriateness).

Results

The headline result is consistent across every condition task and every sub-dimension:

Dimension Condition A Condition B Difference
Recovery quality 38.8% 70.0% +31.2pp
Tone alignment 65.0% 80.0% +15.0pp
Appropriateness 51.7% 83.3% +31.6pp
Overall 48.9% 76.7% +27.8pp

No Condition A transcript exceeded 25/45 (55.6%), while no Condition B transcript fell below 31/45 (68.9%). Several sub-level results stand out. Baseline agents scored 1/5 (20%) on providing explanation—confirming prior findings that LLMs do not spontaneously explain during recovery—whereas coded agents reached 3/5 (60%). Coded agents achieved 4.75/5 (95%) average contextual relevance, suggesting the code effectively anchored responses to task context. Personality appropriateness rose from 50% to 75%. The largest single-task gap occurred in brainstorming (C3), where the coded agent more than doubled the baseline score (38/45 vs 17/45); the authors attribute this to the natural fit between open-ended generative tasks and the Openness trait plus conversational tone.

An unintended finding strengthens the practical case: despite explicit instruction to apply codes only after error notification, all Condition B agents embodied their assigned personality and tone from the start of the interaction, and all correctly identified the code they used at the reflection prompt. This pre-emptive embodiment—consistent with known LLM difficulty suppressing exposed knowledge—suggests the framework could function as a general conversational system prompt rather than solely an error-recovery mechanism.

Limitations and open questions

The authors are candid about several constraints. First, subjectivity bias: with no participants and a single researcher orchestrating both conditions, validity and generalisability are limited, and LLM-generated prompts, agents, and evaluators only partially mitigate this. Second, LLM bias: Sonnet 4.6's built-in clinical-safety parameters may have independently improved performance on the emotional support task (the CB:C2 agent described its approach as "non-clinical"), meaning some measured improvement may reflect vendor safeguarding rather than the code itself—the framework remains untested on models with different parameter regimes. Third, error scenarios were artificial: the error prompt was injected immediately regardless of whether an actual error existed, which may have depressed baseline recovery effort. Fourth, only polite user prompts were tested, so robustness to impolite or degrading user tone—an established performance risk—is unknown. Fifth, the failure of negatively framed instructions ("only apply after an error") prevented within-condition before/after comparison in Condition B, a problem the authors link to ironic process theory and recommend addressing through positively framed prompts. Finally, the evaluation relies entirely on LLM-as-judge scoring without inter-rater validation against human judgments, and with one evaluator per transcript there is no measure of scoring reliability.

Conclusion

This paper contributes a compact, psychologically grounded recovery code framework and a matching evaluation rubric, demonstrating that an off-the-shelf LLM can learn and apply structured personality–context–tone recovery guidance, yielding a 27.8 percentage-point improvement over baseline across all dimensions and tasks. The strongest gains appear precisely where baseline behavior is weakest—explanation-giving and personality-appropriate responding—and the observed pre-emptive trait embodiment hints at broader utility beyond error handling. The exploratory, single-model, LLM-evaluated design means the effect sizes should be treated as indicative rather than definitive; validation with human participants, diverse model architectures, realistic error timing, and adversarial user tones remains the necessary next step.

Paper to Video (Beta)

No one has generated a video about this paper yet.

Whiteboard

No one has generated a whiteboard explanation for this paper yet.

Open Problems

We haven't generated a list of open problems mentioned in this paper yet.