Papers
Topics
Authors
Recent
Search
2000 character limit reached

Emotion In-Context Learning (EICL)

Updated 15 July 2026
  • Emotion In-Context Learning is a research family that conditions emotion-sensitive reasoning on retrieved contextual examples and auxiliary signals.
  • It applies across modalities—including medical consultations, fine-grained text and visual emotion recognition, and speech sentiment analysis—demonstrating measurable improvements.
  • EICL methods integrate dynamic soft labeling, prototype theory, and retrieval strategies to enable emotionally controlled generation and adaptive recognition.

Emotion In-Context Learning (EICL) denotes a set of in-context and retrieval-conditioned methods for emotion-sensitive reasoning, generation, recognition, and control. Recent papers use the acronym for substantially different mechanisms: an emotional control layer for medical consultation, emotionally similar demonstration retrieval for fine-grained text emotion recognition, context-specific explanation retrieval for visual emotion classification, few-shot personalization for speech emotion recognition, acoustically informed prompting for synthetic speech sentiment labels, and social-context-conditioned gesture generation for humanoid robots (Zuo et al., 22 Mar 2025, Ren et al., 8 Oct 2025, Seoh et al., 20 May 2025, Ihori et al., 10 Sep 2025, Huang et al., 10 Jun 2026, Huang et al., 2024). This suggests that EICL is best understood as a research family rather than a single canonical algorithm: its shared premise is that affective behavior is strongly context-dependent, and that this dependence can be exploited by conditioning inference on retrieved exemplars, natural-language explanations, auxiliary emotion models, speaker-specific support sets, or social-context analyses.

1. Conceptual scope and recurring formulations

Across the literature, EICL changes what counts as “context.” In some systems, context is the patient’s emotional state and consultation style; in others it is emotional proximity to prototype-like examples, cluster-specific visual explanations, acoustically similar speech segments, or a robot’s inferred social scene. The common operational move is to make emotion prediction or emotion-conditioned generation depend on explicit demonstrations or retrieved conditioning signals instead of relying only on a base model’s generic prior.

Setting EICL formulation Representative paper
Medical consultation Conditions generation on emotional and attribute information after TEIR retrieval (Zuo et al., 22 Mar 2025)
Fine-grained text emotion recognition Retrieves emotionally similar examples, uses dynamic soft labels, and applies exclusionary candidate control (Ren et al., 2024)
Fine-grained text emotion recognition Uses emotionally similar examples, prototype-based decision analysis, and a two-stage exclusion strategy (Ren et al., 8 Oct 2025)
Visual emotion understanding Retrieves cluster-specific natural-language explanations for emotion labels (Seoh et al., 20 May 2025)
Visual emotion understanding Generates emotion-aware contextual descriptions with a VLLM for downstream fusion and classification (Xenos et al., 2024)
Speech sentiment annotation Retrieves acoustically similar labeled demonstrations for multimodal prompting (Huang et al., 10 Jun 2026)
Personalized SER Conditions inference on a few labeled utterances from the target speaker (Ihori et al., 10 Sep 2025)
Humanoid robotics Infers social context and generates expressive gesture trajectories with optional language feedback (Huang et al., 2024)

The resulting landscape is technically heterogeneous. Some variants are training-free and operate entirely by prompting; some add auxiliary emotion models; some perform instruction tuning or meta-training; and some use the in-context step to control outputs that are not labels at all, such as medical responses or continuous robot motion sequences. The term therefore describes a mode of adaptation rather than a fixed architecture.

2. Prototype theory, emotional similarity, and fine-grained text recognition

In fine-grained emotion recognition, EICL emerges from a critique of standard ICL. The central claim is that semantically similar demonstrations are often emotionally inaccurate, and that this mismatch is especially damaging when labels are densely overlapping. "E-ICL: Enhancing Fine-Grained Emotion Recognition through the Lens of Prototype Theory" reinterprets prompt demonstrations as prototypes. A plug-and-play emotion auxiliary model, RoBERTa^{emo}_{large}, outputs an emotion vector V∈R768V \in \mathbb{R}^{768} and an emotion probability distribution P∈RNcP \in \mathbb{R}^{N_c}, retrieves the top k1=5k_1=5 examples by cosine similarity, constructs dynamic soft labels with α=0.2\alpha = 0.2, and divides candidate emotions into possible and impossible sets before exclusionary emotion prediction. The method is designed to avoid interference from irrelevant categories, requires no additional training of the LLM, and achieves superior emotion prediction performance on EDOS, Empathetic-Dialogues, EmpatheticIntent, and GoEmotions; the paper further reports that even when the emotion auxiliary model is lower than 10% of the LLMs, E-ICL can boost performance by over 4% on multiple datasets (Ren et al., 2024).

"Fine-Grained Emotion Recognition via In-Context Learning" sharpens the same line of argument by separating emotion reasoning from emotion decision-making. The paper uses prototype theory to argue that LLMs contain internal emotion prototypes HcjlH^l_{c_j}, and computes query-to-prototype similarity through

oh=1L∑l=1LHl⋅Hcjl.o_h = \frac{1}{L}\sum_{l=1}^{L} H^l \cdot H^l_{c_j}.

Its EICL method replaces semantically similar examples with emotionally similar examples retrieved by an auxiliary RoBERTaemoRoBERTa_{emo}, adds a dynamic soft-label strategy to represent mixed or ambiguous affect, and applies a two-stage exclusion strategy that first prioritizes a primary emotion set SpesS_{pes} and only then considers the secondary set SsesS_{ses}. Extensive experiments report that EICL significantly outperforms ICL on multiple datasets, and the paper’s hyperparameter analysis argues that moderate values of α\alpha, P∈RNcP \in \mathbb{R}^{N_c}0, and P∈RNcP \in \mathbb{R}^{N_c}1 work best because too much reliance on the auxiliary model or too many candidate labels can introduce noise (Ren et al., 8 Oct 2025).

A shared implication of these papers is that EICL in text emotion recognition is not merely few-shot prompting. It is a retrieval-and-decision pipeline in which the quality of the emotional neighborhood matters at least as much as topical similarity. This is why both papers present emotionally similar examples as more faithful prototypes than semantically similar ones.

3. Medical consultation and conversation-centered generation

In medical consultation, EICL is explicitly positioned as the generation-side complement to Terminology-Enhanced Information Retrieval (TEIR). TEIR focuses on what medical information to retrieve through terminology detection, terminology memory, and enhanced sentence generation; EICL focuses on how to generate the response after retrieval, especially in terms of sentiment, emotional tone, and attribute alignment. The framework uses a large unlabeled corpus and a Chinese medical consultation dataset of 803,564 records spanning 12 medical departments and more than 10 Q&A scenarios, with a 70% training / 30% testing split. The paper describes EICL as memorizing semantic and attribute information from unlabeled corpora, using P∈RNcP \in \mathbb{R}^{N_c}2, P∈RNcP \in \mathbb{R}^{N_c}3, and P∈RNcP \in \mathbb{R}^{N_c}4 as a probability-space view of initial emotional context, document-context influence, and iterative demonstration, and formulating suffix-tuning with fixed P∈RNcP \in \mathbb{R}^{N_c}5 and trainable P∈RNcP \in \mathbb{R}^{N_c}6 to maximize

P∈RNcP \in \mathbb{R}^{N_c}7

Its ablation study shows that patient satisfaction falls from 36.83% (221/600) to 15.33% (92/600) when EICL is removed, while ROUGE-L drops from 14.50 to 13.32 and Distinct-2 from 0.96 to 0.90; the paper interprets this pattern as evidence that EICL contributes emotional appropriateness, response quality, and lexical diversity rather than only token-level overlap (Zuo et al., 22 Mar 2025).

The same paper provides a concrete regeneration example. For the input “I have a fever today,” the naive answer “Drink more hot water.” is associated with Predicted sentiment: Negative, whereas the rewritten answer recommending fluids, juices, clear soups, or hot lemonade is associated with Predicted sentiment: Positive. The point is not sentiment classification alone, but emotion-controlled regeneration so that medically plausible content is also emotionally acceptable.

A related but distinct conversational line appears in emotion recognition in conversation (ERC). InitERC treats EICL as one-stage in-context instruction tuning. Its demonstration pool is

P∈RNcP \in \mathbb{R}^{N_c}8

where P∈RNcP \in \mathbb{R}^{N_c}9 is historical context, k1=5k_1=50 explicitly couples speaker identity with the current utterance, and k1=5k_1=51 is the emotion label. Relevant demonstrations are retrieved by

k1=5k_1=52

using Contriever-MS MARCO, with examples from the same dialogue excluded to avoid leakage. Training then optimizes

k1=5k_1=53

On IEMOCAP, MELD, and EmoryNLP, InitERC reaches weighted F1 scores of 92.65, 78.71, and 56.50, respectively, and its ablations report that zero-shot prompting is poor, in-context examples without fine-tuning collapse badly, and fine-tuning without in-context examples remains much worse than the full system (Ma et al., 16 Aug 2025).

These two strands share a common design logic. In medical dialogue, EICL conditions style and affect after medical retrieval; in ERC, it conditions emotion classification on demonstrations that jointly encode history, speaker, utterance, and label. In both cases, affective adequacy is treated as a context-conditioned property, not a post hoc surface embellishment.

4. Visual emotion understanding and explanation-based context injection

A major visual strand uses language as the intermediate carrier of emotional context. "VLLMs Provide Better Context for Emotion Understanding Through Common Sense Reasoning" studies emotion recognition in context rather than standard facial emotion recognition. Its two-stage pipeline first prompts LLaVa-1.5 with an image and the emotion class list to generate subject-specific, context-aware emotion descriptions, then feeds those descriptions and image features into a multimodal classifier built from a vision encoder, learnable query tokens, a Q-Former, and a fully connected classifier. The backbone is initialized with InstructBLIP weights from Salesforce/instructblip-flan-t5-xl without the LLM weights. The method reaches 93.08% accuracy on CAER-S, 26.66 mAP on BoLD, and 38.52 mAP on EMOTIC; on EMOTIC, the paper reports that visual-only ViT is around 29.19 mAP, text-only RoBERTa is around 34.25 mAP, simple concatenation fusion is 34.66 mAP, and zero-shot LLaVa is 16.98 mAP, which supports the claim that generated descriptions and visual features are complementary and that Q-Former fusion is materially better than direct zero-shot label prediction (Xenos et al., 2024).

EmoGist pushes the explanation idea toward a training-free EICL pipeline. Training images are embedded with mmE5 and stored in an hnswlib vector database; for each emotion label, the corresponding images are clustered with k-means, with k1=5k_1=54, and Qwen2.5-VL 72B generates cluster-specific natural-language explanations from 4 representative images. At test time, the system retrieves the closest cluster explanation by centroid similarity and feeds that explanation to a fast LVLM for classification. The ensemble variant, k1=5k_1=55, generates 3 explanations per cluster from 4 non-overlapping images and aggregates predictions by majority vote. On FI, k1=5k_1=56 reaches macro F1 scores of 47.719 with Qwen2.5-VL 7B, 47.055 with Aya Vision 8B, and 48.019 with InternVL2.5 8B-MPO; on Memotion, it reaches micro F1 scores of 72.106, 78.669, and 74.463 with the same backbones. The paper summarizes these gains as up to 8 points in macro F1 on FI and up to 13 points in micro F1 on Memotion, while also reporting that naive ICL variants such as Global Exp or ICL_all can be substantially worse than zero-shot (Seoh et al., 20 May 2025).

A plausible implication is that visual EICL often benefits when examples are compressed into task-relevant explanations rather than inserted as raw demonstrations. Both papers argue, in different ways, that context-dependent emotion definitions are more informative than generic label glosses.

5. Speech-centered EICL: retrieval-guided annotation and speaker personalization

In speech-based settings, EICL often appears as a mechanism for handling sparsity, subjectivity, and speaker variation without task-specific fine-tuning. "LLM-Based Synthetic Ground Truth Generation for Audio-Based Emotion Classification via In-Context Learning" frames EICL as retrieval-guided multimodal prompting for synthetic sentiment labeling in collaborative VR. Audio from two VR sessions is recorded in mono at 48 kHz with Meta Quest Pro headsets and transcribed with Whisper. For each segment, the system computes acoustic descriptors including pitch, loudness, intensity, and speaking rate in words/sec and syllables/sec, standardizes them, and retrieves demonstrations by Euclidean distance on the acoustic descriptor vector

k1=5k_1=57

The top-k1=5k_1=58 retrieved items, together with audio, transcripts, acoustic descriptors, and labels, are inserted into a prompt for Voxtral without task-specific fine-tuning. On 794 segment-level samples, acoustically informed retrieval-guided ICL improves macro-F1 from 0.30 to 0.49, recall from 0.30 to 0.59, and accuracy from 0.82 to 0.85 in the ground-truth-based retrieval comparison; the paper further reports consistent gains over wav2vec 2.0, NRC-VAD, and XLM-Roberta baselines, especially on the negative class, while noting that vanilla ICL collapses toward neutral predictions (Huang et al., 10 Jun 2026).

"Few-shot Personalization via In-Context Learning for Speech Emotion Recognition based on Speech-LLM" uses EICL for personalized SER. Its speech-LLM extends an instruction-tuned encoder-decoder LLM with a fixed speech encoder and a Q-Former that maps speech representations to a fixed number of query vectors. At inference, the model receives a common instruction, a target-speaker support set

k1=5k_1=59

and a new target utterance, and predicts the emotion text by conditioning on these few-shot examples. The common instruction used in experiments asks the model to choose from “Neutral, Surprise, Sadness, Joy, Fear, Disgust, Anger,” and the number of enrollment utterances varies from 0 to 7. On a new Japanese SER dataset of 800 native Japanese speakers, 50 utterances per speaker, and 7 emotions, the proposed method improves speaker-wise unweighted accuracy from 0.675 in the 0-shot speech LM to 0.698, 0.717, 0.732, 0.738, 0.747, 0.752, and 0.757 for 1-shot through 7-shot, respectively. The paper attributes the strongest general adaptability to meta-training under the TU+LUα=0.2\alpha = 0.20 regime, where target emotion may or may not appear in the support set and support labels may be same or different (Ihori et al., 10 Sep 2025).

These speech papers operationalize context at two different levels. One retrieves acoustically similar segments to stabilize synthetic annotation in noisy multi-user speech; the other uses a target speaker’s own labeled utterances as inference-time evidence of speaker-specific emotional style. In both cases, EICL replaces parameter-updating personalization or annotation with contextualized prompting.

6. Embodied EICL, runtime constraints, and the problem of LLM priors

The robotic literature extends EICL beyond label prediction into continuous control. "EMOTION: Expressive Motion Sequence Generation for Humanoid Robots with In-Context Learning" formalizes expressive gesture synthesis as a three-stage process. A first agent α=0.2\alpha = 0.21 takes image observation α=0.2\alpha = 0.22 and/or language instruction α=0.2\alpha = 0.23 and outputs a natural-language analysis and a gesture token,

α=0.2\alpha = 0.24

A second agent α=0.2\alpha = 0.25 maps the gesture token and a small demonstration set α=0.2\alpha = 0.26 to a motion sequence,

α=0.2\alpha = 0.27

with α=0.2\alpha = 0.28 and each state α=0.2\alpha = 0.29 represented by 22 real values corresponding to hand positions, orientations, and finger opening/closing for both hands. EMOTION++ adds iterative natural-language critique through

HcjlH^l_{c_j}0

with HcjlH^l_{c_j}1. The framework uses only two demonstrations, “idle” and “right-hand wave,” to generate 10 gestures across emblems, illustrators, affective displays, and regulators on the GR-1 humanoid robot. In an online within-subject study with HcjlH^l_{c_j}2, EMOTION is statistically comparable to human-oracle demonstrations overall, with naturalness HcjlH^l_{c_j}3 and understandability HcjlH^l_{c_j}4; EMOTION++ is significantly better than EMOTION overall, with naturalness HcjlH^l_{c_j}5 and understandability HcjlH^l_{c_j}6; and EMOTION++ receives significantly higher understandability than the human oracle overall, with HcjlH^l_{c_j}7. The paper also reports concrete design variables—hand position and orientation, motion pattern and temporal coordination, finger pose, balance between direct and subtle expressive motion, and user/context variability—and identifies practical limits: LLM-generated trajectories may not be collision-free or kinematically trackable, GPT-4o generation takes about 26.8 seconds on average for the initial sequence and 21.2 seconds for one feedback round, and robot finger dexterity constrains representable gestures (Huang et al., 2024).

Despite these successes, not all affective ICL behaves as genuine task adaptation. "The Strong Pull of Prior Knowledge in LLMs and Its Impact on Emotion Recognition" studies multilabel emotion recognition on SemEval 2018 Task 1 E-c and GoEmotions and argues that ICL in emotion recognition is often a struggle between prompt demonstrations and the model’s pre-existing task priors. Using temperature-0 inference, Jaccard Score, Micro F1, and Macro F1, the paper defines several task-prior proxies, compares regular ICL against the best-performing task prior across 5, 15, and 25 shots, and finds that LLMs frequently show little improvement over their priors, that prior pull generally increases with model size, and that larger models can become more ossified rather than more pliable. The strongest baseline finding is that LLMs underperform the BERT-based multi-label classifier Demux on both datasets, especially on GoEmotions, and that increasing demonstrations does not generally weaken prior pull (Chochlakis et al., 2024).

Taken together, these results define a central tension in EICL research. Retrieved or constructed context can materially improve fine-grained recognition, medical response quality, speech annotation, personalization, visual classification, and robot expressiveness, but the benefit depends on whether the contextual signal is emotionally faithful, domain-matched, and strong enough to overcome irrelevant categories, noisy auxiliary predictions, or pre-existing LLM priors. The field’s recurring design problem is therefore not simply how to add examples, but how to make those examples emotionally commensurate with the target decision or generation process.

Topic to Video (Beta)

No one has generated a video about this topic yet.

Whiteboard

No one has generated a whiteboard explanation for this topic yet.

Follow Topic

Get notified by email when new papers are published related to Emotion In-Context Learning (EICL).