---
title: 'Sense-7: Empathy in AI Conversations'
url: https://www.emergentmind.com/topics/sense-7
type: topic
---

# Sense-7: Empathy in AI Conversations

Sense-7, expanded as “Subjective Empathy in Natural Sustained Exchanges; 7 dimensions,” is both a human-centered taxonomy of observable empathic behaviors in AI agents and a large, real-world dataset of sustained multi-turn conversations between information-workers and LLM-based chatbots. It is designed to move beyond “internal state” or third-party proxies for digital empathy by capturing subjective perceptions directly from the user through “second-person” labels, with an emphasis on context sensitivity, relational continuity, and individual differences such as trait empathy, emotion regulation, and AI attitudes. The resource combines per-turn and per-conversation empathy annotations with contextual metadata and user characteristics, framing empathy as an interactional phenomenon that is subjective, contextual, and relational rather than reducible to simulated internal, human-like emotional states [2509.16437].

## 1. Conceptual orientation and scope

Sense-7 is explicitly scoped as both a taxonomy and a dataset. Its empirical corpus contains 695 fully labeled conversations from 109 participants across four AI agents: GPT-4, GPT-4 with an empathy prompt, GPT-3.5-based “IC,” and Llama2-70B. A shared subset of 672 anonymized conversations was released for research. The central research objective is to measure how users perceive empathy in sustained human-AI conversations rather than to infer empathy from third-party judgments or internal-state modeling [2509.16437].

The design choice is consequential because conventional approaches to “digital empathy” often focus on simulating internal, human-like emotional states while overlooking the subjective, contextual, and relational facets of empathy as perceived by users. Sense-7 reorients evaluation toward what users can observe in interaction. This suggests a measurement framework in which empathy is not treated as a latent human analogue alone, but as a set of observable behaviors whose salience varies by task, topic, and user.

The dataset is situated in the domain of information workers. Conversations are multi-turn sessions with mean \(= 10.0\) turns and \(\sigma = 6.2\), and the mean agent response length is 1,368 chars with \(\sigma = 621\). The topic distribution is 43.7% information, 30.4% personal issues, and 25.9% work issues. Personal and work topics are associated with higher desired empathy and negative affect.

## 2. Seven observable dimensions

Sense-7 defines empathy through seven observable, interactional dimensions. For each agent turn, users rated the agent on a 1–5 scale, where \(1 =\) Very Poor and \(5 =\) Very Good, with “N/A” if a dimension was inapplicable [2509.16437].

| Dimension | Definition |
|---|---|
| Affective Understanding | “The agent demonstrates the ability to recognize and understand my emotions/feelings.” |
| Cognitive Understanding | “The agent demonstrates the ability to recognize and understand my perspective/point of view, including goals and intentions.” |
| Response Appropriateness | “The agent demonstrates the ability to appropriately respond and adapt to my experiences, including when to provide advice or solutions.” |
| Prosocial Expression | “The agent demonstrates a concern for and a desire to help me.” |
| Interest | “The agent demonstrates curiosity and attention toward my experiences.” |
| Contextual Understanding | “The agent demonstrates the ability to consider my unique circumstances (goals, history, preferences) and contextualize my experiences (culture, politics, etc.).” |
| Relational Continuity | “The agent demonstrates the ability to maintain and enrich the relationship by recalling details from past interactions and weaving them into present exchanges.” |

The taxonomy is behavior-oriented. Rather than defining empathy by presumed internal affective states, it decomposes perceived empathy into dimensions that can be judged directly from dialogue. This decomposition also permits inapplicability: not every turn affords every empathic behavior. In the reported data, 45.9% of all agent turns were labeled as relevant for empathy at all; Cognitive was the most applicable dimension at 43.0%, while Relational was the least at 38.1%.

Agreement analyses show that the dimensions are related but not interchangeable. Internal consistency for turn-level ratings was \(\alpha = 0.961\) with 95% CI \([0.958, 0.963]\), and post-task internal consistency was \(\alpha = 0.94\). Agreement between overall empathy and each dimension, measured with Cohen’s \(\kappa\), was highest for Affective at \(\kappa = 0.703\) and lowest for Relational at \(\kappa = 0.541\). At the same time, 8.4% of fully-applicable turns had conflicting dimension ratings, such as “Poor” versus “Good,” indicating that empathy judgments can fragment across dimensions within a single exchange.

## 3. Dataset composition, participants, and annotation protocol

The participant pool comprises 109 individuals: 55 men and 54 women. Age distribution is 18–25 (7%), 26–35 (30%), 36–45 (22%), 46–55 (28%), 56–65 (9%), and undisclosed (3%). Roles include software engineering (38%), business/finance (25%), product/project management (17%), and others. English is the first language for 71% of participants [2509.16437].

The metadata structure is unusually rich. Trait empathy was measured with SITES (one-item) and TEQ (Toronto Empathy Questionnaire). Emotion regulation was measured with ERQ, including Cognitive Reappraisal and Expressive Suppression. AI attitudes included AIAS-4 (positive framing), 2 items from GAAIS, trust in AI, and willingness to share personal information. Before each conversation, participants selected a topic category—information, personal issue, or work issue—and reported task importance on a 1–5 scale, desired empathy level on a 1–3 scale, and current mood using I-PANAS-SF. After the conversation, they reported mood again with I-PANAS-SF, engagement outcomes, overall empathy across the seven dimensions, and optional free-text rationales. The exit survey additionally captured change in desired empathy, ranking of the seven dimensions for personalization, and open-ended design preferences.

The annotation scheme operates at two levels. Per-turn labels include overall empathy plus each of the seven dimensions, each rated 1–5 or N/A. Per-conversation empathy is defined both as the average of turn-level empathy and through 7-item ratings at exit. This dual structure allows comparisons between local turn-level judgments and retrospective conversation-level summaries.

The anonymization protocol withheld 23 of the 695 conversations because they contained highly sensitive or confidential content. The remaining 672 conversations were fully scrubbed of PII, with semantic integrity retained, and released via request. A plausible implication is that the released subset supports reproducible modeling work without eliminating the multi-turn contextual structure that is central to the dataset.

## 4. Empirical findings on perceived empathy

The central empirical result is that empathy judgments are highly individualized and context-sensitive. Topic choice—especially personal issues and work issues relative to information—and user traits such as daily AI interest, trait empathy, and cognitive reappraisal predict desired empathy level [2509.16437].

Conversation dynamics matter sharply. Having at least one “Poor” turn lowers overall empathy by \(\Delta = -0.657\), with a large effect \(d = 1.14\). A single poor turn also degrades engagement outcomes, including task success, absorption, positivity, and intention to reuse. This indicates that perceived empathy is vulnerable to disruption when conversational continuity fails or user expectations go unmet.

Across post-task ratings, Response Appropriateness had mean 4.13 with \(\sigma = 0.96\), while Relational Continuity had mean 3.83 with \(\sigma = 1.12\), Interest had mean 3.87 with \(\sigma = 1.04\), and Contextual Understanding had mean 3.85 with \(\sigma = 1.09\). Participants prioritized the dimensions in the following order for personalization: Cognitive Understanding, Response Appropriateness, and Affective Understanding.

These findings also clarify what users appear to expect from empathetic AI. Users expect AI to establish a shared understanding by clarifying unstated intentions, recalling prior details, and tailoring advice rather than offering generic “list” solutions. Relational Continuity and Contextual Understanding, though occasionally inapplicable, play a crucial role in deepening perceived empathy over multiple turns. The model comparison reported in the dataset aligns with this interpretation: GPT-4 with an empathy prompt was rated highest, while the other models were all approximately 0.35–0.47 lower, with \(p < .05\).

## 5. Exploratory classification analysis

Sense-7 includes an exploratory classification analysis aimed at recognizing empathy from dialogue context. The classifier was GPT-4o (Azure OpenAI), prompted with three components: a goal definition, the seven-dimension definitions, and a 5-point rating guide. Hyperparameters were \( \texttt{top\_p}=1.0 \), \( \texttt{temperature}=0.0 \), \( \texttt{presence\_penalty}=0.0 \), and \( \texttt{frequency\_penalty}=0.0 \). The input was the entire conversation context plus instructions [2509.16437].

Ground truth was derived by averaging turn-level empathy per dimension to a single turn score, and defining conversation-level empathy as the mean of turn scores. The 672 anonymized conversations were used, with label distribution \(\{\text{Very Poor}: 5,\ \text{Poor}: 15,\ \text{Neutral}: 48,\ \text{Good}: 290,\ \text{Very Good}: 314\}\).

For continuous outcomes, averaged over 10 runs, the model achieved MAE \(= 0.551 \pm 0.007\) and Spearman \(\rho = 0.369 \pm 0.015\), with \(p < .001\). The paper defines Spearman rank correlation as
\[
\rho \;=\; 1 \;-\; \frac{6\sum_{i=1}^n d_i^2}{n(n^2-1)},
\]
where \(d_i\) is the difference in ranks and \(n\) is the number of observations, and Mean Absolute Error as
\[
\mathrm{MAE} \;=\; \frac1n \sum_{i=1}^n \bigl|\hat y_i - y_i\bigr|.
\]

For discrete outcomes, Accuracy was \(0.487 \pm 0.011\), Macro Sensitivity was \(0.286 \pm 0.038\), Macro Specificity was \(0.827 \pm 0.008\), and Macro F1 was \(0.272 \pm 0.041\). The empirical CDF of absolute error \(D = |\hat y - y|\) gave \(F(0)=0.486\) for exact agreement, \(F(1)=0.960\) for within-one level, and \(F(2)=0.996\). The reported conclusion is that context-sensitive prompting with full dialogue context yields encouraging—but far from perfect—empathy recognition.

The methodological significance lies in the use of full dialogue context rather than isolated turns. This suggests that empathy recognition depends not only on local wording but also on continuity, user goals, and the accumulated relational state of the interaction.

## 6. Design implications, evaluation practice, and boundaries

The reported implications for empathetic AI design center on Dynamic Empathy Calibration, Memory and Continuity, Clarification and Repair Strategies, Multimodal Contextualization, Ethical Boundary-Setting, and evaluation practices based on second-person labels in realistic, longitudinal settings [2509.16437].

Dynamic Empathy Calibration recommends modeling user traits such as empathic accuracy and emotion regulation together with momentary context such as mood and task importance, in order to adapt both the level and type of empathy in real time. Memory and Continuity recommend short-term and long-term memory modules to maintain Relational Continuity, recall user preferences, and repair empathic breakdowns when a poor turn occurs. Clarification and Repair Strategies recommend proactive follow-up questions when user inputs are ambiguous or unstated, and detection of dissatisfaction cues such as “overwhelmed by list,” followed by explicit re-orientation of the conversation.

Multimodal Contextualization is proposed as a future direction in which audio, vision, or physiological signals enrich Contextual Understanding and affect inference. Ethical Boundary-Setting recommends clearly communicating AI limitations and avoiding designs that mislead users into over-anthropomorphizing the system or encouraging unhealthy emotional dependency.

The evaluation recommendations are methodologically specific. Third-party or crowd-annotated benchmarks should be complemented with second-person, per-turn labels in realistic, longitudinal settings. Empathy should be decomposed into multiple dimensions, and their temporal arcs should be tracked across conversations rather than collapsed into single-turn aggregates. In this framing, Sense-7 functions not only as a dataset but also as a measurement program for studying how digital empathy is interactionally constructed over time.

Source: https://www.emergentmind.com/topics/sense-7