Papers
Topics
Authors
Recent
Gemini 2.5 Flash
Gemini 2.5 Flash
110 tokens/sec
GPT-4o
56 tokens/sec
Gemini 2.5 Pro Pro
44 tokens/sec
o3 Pro
6 tokens/sec
GPT-4.1 Pro
47 tokens/sec
DeepSeek R1 via Azure Pro
28 tokens/sec
2000 character limit reached

Leveraging TCN and Transformer for effective visual-audio fusion in continuous emotion recognition (2303.08356v3)

Published 15 Mar 2023 in cs.CV

Abstract: Human emotion recognition plays an important role in human-computer interaction. In this paper, we present our approach to the Valence-Arousal (VA) Estimation Challenge, Expression (Expr) Classification Challenge, and Action Unit (AU) Detection Challenge of the 5th Workshop and Competition on Affective Behavior Analysis in-the-wild (ABAW). Specifically, we propose a novel multi-modal fusion model that leverages Temporal Convolutional Networks (TCN) and Transformer to enhance the performance of continuous emotion recognition. Our model aims to effectively integrate visual and audio information for improved accuracy in recognizing emotions. Our model outperforms the baseline and ranks 3 in the Expression Classification challenge.

User Edit Pencil Streamline Icon: https://streamlinehq.com
Authors (4)
  1. Weiwei Zhou (9 papers)
  2. Jiada Lu (3 papers)
  3. Zhaolong Xiong (1 paper)
  4. Weifeng Wang (5 papers)
Citations (21)

Summary

We haven't generated a summary for this paper yet.