Papers
Topics
Authors
Recent
Search
2000 character limit reached

FG-SGL: Fine-Grained Semantic Guidance Learning via Motion Process Decomposition for Micro-Gesture Recognition

Published 17 Mar 2026 in cs.CV | (2603.16269v1)

Abstract: Micro-gesture recognition (MGR) is challenging due to subtle inter-class variations. Existing methods rely on category-level supervision, which is insufficient for capturing subtle and localized motion differences. Thus, this paper proposes a Fine-Grained Semantic Guidance Learning (FG-SGL) framework that jointly integrates fine-grained and category-level semantics to guide vision--LLMs in perceiving local MG motions. FG-SA adopts fine-grained semantic cues to guide the learning of local motion features, while CP-A enhances the separability of MG features through category-level semantic guidance. To support fine-grained semantic guidance, this work constructs a fine-grained textual dataset with human annotations that describes the dynamic process of MGs in four refined semantic dimensions. Furthermore, a Multi-Level Contrastive Optimization strategy is designed to jointly optimize both modules in a coarse-to-fine pattern. Experiments show that FG-SGL achieves competitive performance, validating the effectiveness of fine-grained semantic guidance for MGR.

Summary

  • The paper introduces FG-SGL, a vision–language framework that decomposes each micro-gesture into initiator, receiver, direction, and motion type, then aligns these attributes with visual features.
  • FG-SGL achieves 78.13% top-1 accuracy on SMG, surpassing prior baselines, while reaching 62.58% on iMiGUE with RGB-only inference and outperforming CLIP-MG.
  • Ablations show that category prototype alignment contributes the largest gains, fine-grained alignment is especially valuable for similar classes, and staged optimization reduces conflicting gradients.

Overview and motivation

Micro-gestures (MGs) are subtle, involuntary body movements of short duration that reflect psychological states such as stress or cognitive load. Recognizing them from video is difficult for two reasons stated by the authors: motion amplitude is extremely small and easily obscured by appearance and background variation, and inter-class differences can be confined to pixel-level displacements (e.g., rubbing the eyes versus touching the eyebrow). The paper argues that existing MGR methods, which rely on category-level supervision alone, provide insufficient guidance for learning these localized motion cues. To address this, Wei et al. propose FG-SGL, a Fine-Grained Semantic Guidance Learning framework built on a pretrained vision–LLM (VLM) that injects structured, instance-aware textual semantics into visual representation learning (2603.16269).

Framework components: FG-Text, FG-SA, and CP-A

The framework rests on three components operating at different depths of the visual encoder. First, FG-Text is a human-annotated fine-grained textual dataset that decomposes each MG instance into four semantic attributes — initiator, receiver, direction, and motion type — composed into a natural-language description. Unlike prior MG benchmarks (SMG, iMiGUE, MA-52), which provide only coarse class labels, FG-Text supplies explicit supervision of internal action structure without altering the label space.

Second, Fine-Grained Semantic Alignment (FG-SA) aligns mid-level visual representations with instance-aware semantic embeddings via an InfoNCE-style contrastive loss, treating matched video–text pairs as positives and other in-batch descriptions as negatives. Third, Category Prototype Alignment (CP-A) applies the same contrastive objective at the category level, aligning high-level visual features with textual category prototypes to strengthen global inter-class separability.

Multi-Level Contrastive Optimization

A practical concern when combining contrastive objectives at multiple granularities is gradient conflict early in training, when representations lack structure. ML-CO addresses this with a two-stage schedule: classification loss plus CP-A first establish a discriminative category-level space; FG-SA is then introduced while both earlier objectives remain active. The total objective is a weighted sum L=Lcls+λfgLfg+λcpLcp\mathcal{L} = \mathcal{L}_{cls} + \lambda_{fg}\mathcal{L}_{fg} + \lambda_{cp}\mathcal{L}_{cp}. Training uses InternVL2.5-8B as backbone, with frozen vision weights, rank-8 LoRA adapters, a frozen text encoder, 8 frames at 448×448448\times448 resolution, bf16 precision, and AdamW with cosine scheduling.

Experimental results

On SMG, FG-SGL achieves 78.13% top-1 accuracy, outperforming all compared baselines including H2OFormer (75.44%, skeleton-based). On iMiGUE it reaches 62.58%, which trails several skeleton-based and RGB–skeleton methods (e.g., PL at 70.25%) but exceeds CLIP-based baselines such as CLIP-MG (61.82%), despite using only RGB at inference time. A summary of representative results:

Method Modality SMG iMiGUE
TSM RGB 65.41 –
JSSEL Skeleton 68.03 64.12
CVTCL RGB 65.08 66.12
PL RGB + Skeleton – 70.25
H2OFormer Skeleton 75.44 70.00
CLIP-MG RGB + Skeleton – 61.82
FG-SGL RGB 78.13 62.58

Ablations attribute consistent gains to each component. Removing FG-SA costs 1.62 points on SMG and 3.29 on iMiGUE; removing CP-A costs 4.91 and 5.50 points respectively, indicating category-level alignment is the larger single contributor but fine-grained alignment matters more on datasets with higher inter-class similarity. Replacing FG-Text with class-level text drops accuracy to 76.55% and 59.36%, directly demonstrating that holistic category descriptions cannot substitute for instance-aware decomposition. Disabling the progressive ML-CO schedule costs 0.43 points on SMG and 1.86 on iMiGUE, supporting the claim that jointly optimizing both objectives from scratch produces unstable, conflicting gradients.

Limitations and open questions

Several caveats bear on these results. The performance gap on iMiGUE relative to skeleton-based methods shows that RGB-only inference does not fully recover pose information even with fine-grained semantic supervision; whether richer semantic priors can close this gap is left open. FG-Text depends entirely on manual annotation with cross-validation, raising scalability concerns that the authors acknowledge and defer to future work on "more scalable semantic representations." Additionally, no analysis is provided of how sensitive results are to annotation quality, the choice of the four semantic dimensions, or the hyperparameters λfg\lambda_{fg} and λcp\lambda_{cp}, nor of generalization beyond SMG and iMiGUE.

Conclusion

FG-SGL demonstrates that decomposing micro-gesture motions into structured, instance-level semantic attributes and aligning them with mid-level VLM representations — coordinated with category prototype alignment under a coarse-to-fine optimization scheme — improves recognition of visually similar gesture classes. Its strongest result, state-of-the-art RGB-only performance on SMG, supports the paper's central claim that fine-grained semantic guidance captures localized motion cues unavailable from category labels alone, though its advantage over skeleton-informed methods on iMiGUE remains unresolved.

Paper to Video (Beta)

No one has generated a video about this paper yet.

Whiteboard

No one has generated a whiteboard explanation for this paper yet.

Open Problems

We haven't generated a list of open problems mentioned in this paper yet.