- The paper demonstrates that personality moderates AI coaching outcomes: resilient users benefited most from a static handbook, overcontrolled users gained self-efficacy from theory-driven AI, and undercontrolled users showed no psychological gains.
- The experiment found that Trucey shifted 16 of 17 negotiation-framework elements consistently across personality clusters, moving participants from distributive tactics toward collaborative, interest-based strategies.
- The findings support adaptive design based on readiness, including autonomous reference tools for resilient users, incremental scaffolding for overcontrolled users, and emotional-regulation or human support before negotiation training for undercontrolled users.
Overview and motivation
Workplace negotiation—requesting a raise, a promotion, or time off—is psychologically fraught, and prior research on AI-based negotiation coaching has largely assumed that such tools are uniformly effective across users. The paper "Not My Truce: Personality Differences in AI-Mediated Workplace Negotiation" (2604.00464) challenges this assumption by asking whether personality traits moderate who benefits from which form of negotiation support. Drawing on the Big Five inventory and the ARC typology (resilient, overcontrolled, undercontrolled), the authors conduct a between-subjects experiment (N=267) comparing three conditions: Trucey, a theory-driven AI coach grounded in Brett's negotiation framework; ControlAI, a general-purpose conversational AI coach serving as an active control; and ControlTheory, a static handbook serving as a passive control. The study is guided by three research questions concerning perceived effectiveness (RQ1), linguistic engagement (RQ2), and adoption of negotiation frameworks (RQ3).
The central claim of the paper is that coaching effectiveness emerges from an interaction between system architecture and the user's personality configuration rather than being a fixed property of the tool. This claim is supported by a striking dissociation: Trucey shifted users' strategic language toward theory-aligned negotiation elements almost universally, yet produced psychological gains only for specific personality profiles—and in one case, the traditional handbook outperformed the interactive AI.
System design and experimental method
Trucey implements a five-stage workflow: scenario assignment (salary raise, promotion, or time-off request), personality calibration of the participant's real supervisor via BFI-10 items, theory-guided advice generation covering Brett's pillars (interests, alternatives, legitimacy), role-based rehearsal in which the LLM embodies the supervisor's Big Five profile, and iterative feedback integration every two turns. The prototype was built on GPT-4o-mini with a Streamlit interface, cosine-similarity matching to pre-defined supervisor profiles, and few-shot prompting spanning five difficulty levels from distributive tactics to integrative strategies.
Participants were recruited through Prolific (N=267: Trucey n=134, ControlAI n=66, ControlTheory n=67), with sample size determined a priori for an effect size of f2=0.15, power 0.90, and α=.05. Measures included occupational self-efficacy (OSS-6), psychological empowerment (PEU), negotiation preparedness, usability (UMUX), intervention appropriateness (IAM), and lexico-semantic features of conversation logs. Framework adoption was quantified via BERT sentence-embedding cosine similarity between user/AI language and Brett's framework elements.
Personality clustering used k-means on z-standardized Big Five scores, yielding three clusters consistent with the ARC taxonomy: C_0 "undercontrolled" (n=74; high neuroticism, low agreeableness/conscientiousness), C_1 "resilient" (n=110; high extraversion, agreeableness, conscientiousness, openness; low neuroticism), and C_2 "overcontrolled" (N=2670; low openness and extraversion). The silhouette score of 0.197 is modest, though the authors note it is comparable to validated Big Five clustering studies and reflect the fuzzy, gradient-like nature of ARC types. A further design caveat acknowledged explicitly: the three conditions differ not only in interactivity but also in information structure, personalization, and scenario framing, so comparisons should be read as modality-level contrasts rather than isolations of a single treatment factor.
Perceived effectiveness: personality as moderator (RQ1)
Pooled across conditions, the clusters did not differ in change magnitudes on any psychological outcome (all N=2671), but they differed sharply in system perception: UMUX (N=2672, N=2673) and IAM (N=2674, N=2675), with resilient users rating all systems highest and undercontrolled users lowest. Within-cluster analyses revealed condition-specific effects:
| Cluster |
Key finding |
Effect |
| Undercontrolled (C_0) |
No significant condition effects on any outcome |
All N=2676 |
| Resilient (C_1) |
ControlTheory exceeded Trucey on empowerment, meaning, self-determination |
e.g., self-determination N=2677 0.385 vs. −0.246, N=2678 |
| Overcontrolled (C_2) |
Trucey exceeded ControlAI on self-efficacy |
0.200 vs. −0.266, N=2679 |
The resilient cluster's gain in self-determination under the static handbook (n=1340, n=1341) was the largest single change in the dataset and emerged exclusively under ControlTheory. The interpretation offered is that for users with strong internal resources, interactive coaching functions as an autonomy constraint, whereas reference material lets them leverage existing self-regulation—a result that directly contradicts the intuition that more interactivity is always better.
The overcontrolled cluster exhibited a notable perception–outcome dissociation: they rated Trucey and ControlAI nearly identically on usability (77.2 vs. 71.7) and appropriateness (4.01 vs. 3.95) despite deriving significantly greater self-efficacy benefit from Trucey. The authors attribute this to the cautious external evaluation style of low-openness individuals, meaning that subjective ratings would have masked a real benefit—an important caution for evaluation studies relying solely on perceived usefulness.
The undercontrolled cluster's null results across all conditions are framed as a "fundamental unreadiness" for self-directed digital interventions: high neuroticism heightens threat appraisal while low conscientiousness limits behavioral follow-through. The implication stated plainly is that no amount of theoretical grounding in the coaching content compensated for this mismatch.
Linguistic engagement (RQ2)
Analysis of nine lexico-semantic measures showed divergent adaptation patterns. Undercontrolled users displayed linguistic stasis—no significant differences between Trucey and ControlAI on any measure—mirroring their null psychological outcomes. Resilient users interacting with Trucey became more efficient and accessible: fewer words per sentence (n=1342) and per response (n=1343), higher readability (n=1344, n=1345), and lower formality (n=1346, n=1347), suggesting fluent internalization of structured guidance and treatment of the AI as a collaborative partner.
Overcontrolled users showed what the authors term a "fluency tax": without increasing verbosity, their outputs became significantly harder to read (n=1348, n=1349) and syntactically more complex (n=660, n=661). Interpreted through cognitive load theory, encountering an unfamiliar structured framework consumed the resources needed for clear expression. This linguistic strain marker is proposed as a real-time signal adaptive systems could monitor to adjust scaffolding.
Framework adoption (RQ3)
Using BERT embeddings against six categories of Brett's framework (17 elements total), the most robust finding emerged: 16 of 17 elements shifted significantly in the same direction across all three clusters, with the sole exception of holistic outcome analysis. Trucey pushed all users away from distributive approaches—basic strategy (n=662), limited interest exploration (n=663), and authority-based power dynamics (n=664, all largest in C_1)—toward structured, collaborative, evaluative framings. No element reversed sign across clusters.
This uniformity yields the paper's clearest functional division: personality determines the internalization of and reaction to coaching, but the framework's structure dictates the strategic output. Even overcontrolled users, whose shifts were smaller in magnitude, were anchored in the same collaborative logic. A secondary observation is that ControlAI, without specialized prompting, implicitly incorporated several of Brett's components by default, suggesting LLMs encode negotiation theory generically; targeted prompting steers outputs toward specific orientations (e.g., collaboration) beyond this default balance.
Design implications
The authors propose a "readiness floor"—a personality-driven threshold below which self-directed digital interventions may be ineffective—arguing that personalization should extend beyond stage-of-change tailoring to account for baseline psychological readiness. Three concrete interaction modes follow from the cluster findings: a fast-track mode for resilient users preserving autonomy through self-directed reference access; a scaffolded mode for overcontrolled users delivering guidance incrementally with visual scaffolds when linguistic strain is detected; and a pre-intervention mode for undercontrolled users involving readiness-building modules (emotional regulation, anxiety reduction) before any strategic content, with human handoff if strain persists.
An equity concern is raised prominently: undercontrolled profiles predict occupational attainment at effect sizes comparable to socioeconomic status and cognitive ability, so the workers who stand to benefit most from negotiation support are precisely those least served by self-directed digital interventions. Broad deployment without accounting for this mismatch risks exacerbating workplace inequality—empowering already well-resourced workers while leaving vulnerable users unsupported.
Limitations and open questions
The paper concedes several constraints on its claims. Participants were crowdworkers not facing imminent negotiations, limiting ecological validity; interactions were simulated rather than real negotiations with actual supervisors; outcomes were self-reported psychological constructs rather than behavioral or expert-evaluated performance, constraining claims about skill transfer; the cross-sectional between-subjects design prevents causal inference about readiness and rules out learning trajectories or delayed effects; the ControlTheory condition framed salary negotiation within career advancement and differed in information dosage, limiting direct condition comparisons; and the system operationalized a single negotiation framework within one conversational architecture, so generalization to other theories, AI paradigms, or formats remains untested. The "readiness floor" concept itself is offered as requiring empirical validation in other contexts. Open questions include whether psychological readiness translates into tangible workplace gains in field deployments, whether pre-intervention modules can enable previously unresponsive users to benefit downstream, and how hybrid AI–human support models perform relative to purely automated coaching.
Conclusion
This study demonstrates that AI-mediated negotiation coaching produces heterogeneous outcomes conditioned by personality configuration: resilient users gained most from a static handbook that preserved autonomy, overcontrolled users gained self-efficacy specifically from theory-driven AI despite not recognizing it in their ratings, and undercontrolled users responded minimally to every condition even as their strategic language shifted toward collaborative framings. The work reframes non-responsiveness as a design mismatch rather than user failure, arguing that equitable AI support requires systems capable of recognizing readiness limits and integrating preparatory or human scaffolding when self-directed intervention is insufficient.