Papers
Topics
Authors
Recent
Search
2000 character limit reached

VGAF-GEMS: Holistic Video Affect Profiling

Updated 19 July 2026
  • VGAF-GEMS is a multi-level video affect benchmark that integrates individual, group, and situational emotion annotations in real-world group videos.
  • It employs a combined manual and semi-automated annotation pipeline to achieve robust label reliability across diverse emotional dimensions.
  • The associated GEMS framework utilizes multimodal techniques, including a Swin-transformer and S3Attention, to effectively fuse visual and contextual cues for enhanced affect prediction.

Searching arXiv for the cited VGAF-GEMS paper and the original VGAF dataset paper for supporting context. VGAF-GEMS is a large-scale extension of the original VGAF dataset for multi-person video affect analysis. Whereas VGAF previously only included coarse group-level emotion annotations—valence: positive/neutral/negative—in 5-second videos, VGAF-GEMS introduces dense and multi-level annotations spanning individual, group, and situational emotion, together with contextual descriptions generated with a multimodal LLM. In the associated formulation, emotion comprehension is treated as prediction from fine-grained individual emotion to coarse grained group and event level emotion, and the benchmark is paired with GEMS, a multimodal Swin-transformer and S3Attention based framework for joint prediction across these levels (Kataria et al., 30 Jul 2025).

1. Corpus origin and task formulation

VGAF-GEMS extends VGAF by shifting the problem definition from coarse group affect recognition to holistic emotional profiling in dynamic, real-world video group settings. The extension adds frame-wise, for each detected individual, emotion labels; group-level emotion labels; situational annotations; and MLLM-generated contextual descriptions. This re-specification materially changes the benchmark from a single-label group task to a coupled multi-level prediction problem in which individual emotion, group emotion, and event-level understanding are modeled together (Kataria et al., 30 Jul 2025).

The dataset contains 4,183 videos, each 5 sec long, with fps 13–30. The standard split is 2,661 train, 766 validation, and 756 test. The stated aim is to predict basic discrete and continuous emotions, including valence and arousal, as well as individual, group and event level perceived emotions. A central implication of this design is that group affect is not treated as an isolated output; instead, it is embedded in a structured hierarchy linking local facial evidence, social interaction, and situational context (Kataria et al., 30 Jul 2025).

2. Annotation schema and representational scope

The annotation schema is explicitly multi-level. At the individual level, VGAF-GEMS provides categorical emotions with 9 classes and continuous valence and arousal. At the group level, it provides categorical group emotion with 9 classes and aggregate valence/arousal scores. At the situational level, it provides both a situation type and a situational emotion label. Context annotations consist of MLLM-generated descriptions for events, scene background, group interactions, and relationships (Kataria et al., 30 Jul 2025).

Level Label type Classes/Range
Individual Categorical Emotion 9 classes
Individual Valence (continuous) [−1,1][-1,1]
Individual Arousal (continuous) [−1,1][-1,1]
Group Categorical Emotion 9 classes
Group Valence (continuous) [−1,1][-1,1]
Group Arousal (continuous) [−1,1][-1,1]
Situational (Event) Situation Type 10 classes
Situational (Event) Situational Emotion 10 classes (multi-label)
Context MLLM-generated: event, background, etc. --

The 10 situation types are Quarrelling, Meeting, Funeral, Sports, Casual-gathering, Group-activities, Protest, Celebration-party, Fighting, and Show. The situational emotion space is also 10-way, with Active, Enthusiastic, Anxious, Angry, Attentive, Distressed, Sad, Relaxed, Rejected, and Happy. This label design couples basic emotion categories with continuous affect dimensions and social-scene semantics. A common misconception is to treat VGAF-GEMS as merely a refined group-emotion dataset; the annotation schema shows that it is instead a dense benchmark for joint individual, group, and event-level affect modeling (Kataria et al., 30 Jul 2025).

3. Annotation pipeline and reliability

The curation pipeline combines automation and manual verification. The paper describes a manual and semi-automatic annotation pipeline using Emolysis toolkit automation with human-in-the-loop corrections for improved label accuracy. Multiple annotators are used, with a third annotator for tie-break in cases requiring consensus. This design is intended to support fine-grained labels that are difficult to obtain reliably from fully automatic procedures alone (Kataria et al., 30 Jul 2025).

Reported inter-annotator agreement, using Cohen's Kappa, is 0.63 for individual emotion, 0.70 for valence, 0.65 for arousal, 0.8 for group emotion, 0.76 for situation, and 0.72 for situational emotion. These values are presented as evidence that label reliability remains strong despite the complexity of the task. The stronger agreement at group and situation level than at individual emotion level is consistent with the greater ambiguity typically associated with frame-wise person-specific affect in unconstrained social video; this suggests that VGAF-GEMS is challenging not only for models but also for annotation (Kataria et al., 30 Jul 2025).

The inclusion of MLLM-generated contextual descriptions, using VideoGPT, is also structurally important. These descriptions cover narrative and context understanding rather than only visible facial cues, and they form a bridge between visual evidence and situational labels. In effect, the benchmark formalizes contextual supervision as part of the dataset rather than leaving it entirely implicit (Kataria et al., 30 Jul 2025).

4. GEMS framework and multimodal modeling

The benchmark is paired with GEMS, which leverages a multimodal swin-transformer and S3Attention based architecture. The system pipeline is: input video to detected faces plus VideoGPT context; SWIN-B to per-face emotion embeddings; RoBERTa to context embedding; fusion plus S3Attention to prediction of individual, group, and situation-level emotions. The model is explicitly hierarchical and multimodal, with separate visual and contextual encoders followed by spatio-temporal fusion (Kataria et al., 30 Jul 2025).

At the visual stage, faces are detected frame-by-frame using MTCNN. Each face is then passed through a SWIN-B backbone, yielding emotional embeddings; the summary specifies that this backbone is pre-trained and frozen, and classifies 9 emotions. At the contextual stage, the entire video is processed using Video-ChatGPT (VideoGPT) with designed prompts to extract information about scene, interactions, events, and relationships, after which RoBERTa is used for text embedding (Kataria et al., 30 Jul 2025).

Fusion is performed by concatenating individual and contextual embeddings and passing them through an encoder block with auto-correlation layers for spatio-temporal modeling. S3Attention is decomposed into Smoothing, Sketching, and Sub-sequence attention: smoothing merges local/global information, reducing noise; matrix sketching performs key data selection; and low-rank approximation improves computational efficiency, with separate row and column attention mechanisms aggregating information. Output heads implement multi-task learning for individual, group, and situation-level predictions across both discrete and continuous targets. This architecture operationalizes the benchmark’s fine-to-coarse formulation rather than treating the three prediction levels as disjoint tasks (Kataria et al., 30 Jul 2025).

5. Benchmarking protocol and empirical results

The benchmarking setup includes both zero-shot evaluation and supervised learning. Zero-shot evaluation uses off-the-shelf Video-LLMs, specifically Video-ChatGPT and MiniCPM-V, to assess direct emotion understanding capabilities without training. Reported results are poor; for example, Video-ChatGPT attains group emotion acc 1.17% and situational emotion acc 15.77%. The paper interprets this as evidence of the complexity and novelty of VGAF-GEMS (Kataria et al., 30 Jul 2025).

For supervised learning, the comparison includes visual-only baselines, visual+context without S3Attention, and the full GEMS model with S3Attention. The metrics are Accuracy for discrete tasks and Mean Squared Error (MSE) and Concordance Correlation Coefficient (CCC) for continuous tasks. A headline result is the reported improvement in group emotion accuracy from 9.46% without S3Attention to 54.80% with the full GEMS model. The paper therefore attributes a substantial part of the performance gain to efficient and effective spatio-temporal fusion of multimodal information (Kataria et al., 30 Jul 2025).

The broader empirical conclusion is twofold. First, state-of-the-art generic video-LLMs perform very poorly on the benchmark. Second, even though GEMS outperforms other variants, the overall task remains challenging, with room for improvement. This is important for interpretation: VGAF-GEMS is not presented as a saturated benchmark but as a difficult testbed that exposes weaknesses in current multimodal affect models (Kataria et al., 30 Jul 2025).

6. Significance, applications, and nomenclature

The principal contribution claimed for VGAF-GEMS is that it is the first dataset to densely annotate individual, group, and situational emotions in dynamic, real-world video group settings. The benchmark thereby enables holistic emotional profiling, linking fine-grained individual facial expression to coarse social situation assessments. The stated applications include classroom engagement monitoring, social robotics, video summarization, surveillance, and affective computing research. A plausible implication is that the dataset may be especially useful wherever perceived affect depends jointly on local expression, collective behavior, and scene context (Kataria et al., 30 Jul 2025).

The nomenclature is potentially confusing because “GeMS” and “GEMS” are also used in unrelated research areas. In adaptive optics, GeMS denotes the Gemini Multi-Conjugate Adaptive Optics System associated with Wide Field Adaptive Optics, distortion correction, and atmospheric profiling (Bernard et al., 2016, Masciadri et al., 2016). In 3D reconstruction, GeMS denotes a Gaussian Splatting framework for extreme motion blur (Matta et al., 20 Aug 2025). In graph generation, GEMS denotes “Scene Expansion using Generative Models of Graphs” (Agarwal et al., 2022). VGAF-GEMS belongs instead to multimodal affective computing and specifically to group emotion profiling through multimodal situational understanding.

Within that domain, the dataset’s central significance lies in making context explicit. Rather than assuming that facial evidence alone is sufficient, the benchmark encodes scene background, interactions, and relationships as part of the prediction problem. This suggests a broader methodological shift in group affect analysis: robust performance may require models that integrate individual-level emotion cues with contextual and situational semantics, not merely stronger single-stream visual backbones.

Topic to Video (Beta)

No one has generated a video about this topic yet.

Whiteboard

No one has generated a whiteboard explanation for this topic yet.

Follow Topic

Get notified by email when new papers are published related to VGAF-GEMS.