- The paper introduces FG-SGL, a vision–language framework that decomposes each micro-gesture into initiator, receiver, direction, and motion type, then aligns these attributes with visual features.
- FG-SGL achieves 78.13% top-1 accuracy on SMG, surpassing prior baselines, while reaching 62.58% on iMiGUE with RGB-only inference and outperforming CLIP-MG.
- Ablations show that category prototype alignment contributes the largest gains, fine-grained alignment is especially valuable for similar classes, and staged optimization reduces conflicting gradients.
Overview and motivation
Micro-gestures (MGs) are subtle, involuntary body movements of short duration that reflect psychological states such as stress or cognitive load. Recognizing them from video is difficult for two reasons stated by the authors: motion amplitude is extremely small and easily obscured by appearance and background variation, and inter-class differences can be confined to pixel-level displacements (e.g., rubbing the eyes versus touching the eyebrow). The paper argues that existing MGR methods, which rely on category-level supervision alone, provide insufficient guidance for learning these localized motion cues. To address this, Wei et al. propose FG-SGL, a Fine-Grained Semantic Guidance Learning framework built on a pretrained vision–LLM (VLM) that injects structured, instance-aware textual semantics into visual representation learning (2603.16269).
Framework components: FG-Text, FG-SA, and CP-A
The framework rests on three components operating at different depths of the visual encoder. First, FG-Text is a human-annotated fine-grained textual dataset that decomposes each MG instance into four semantic attributes — initiator, receiver, direction, and motion type — composed into a natural-language description. Unlike prior MG benchmarks (SMG, iMiGUE, MA-52), which provide only coarse class labels, FG-Text supplies explicit supervision of internal action structure without altering the label space.
Second, Fine-Grained Semantic Alignment (FG-SA) aligns mid-level visual representations with instance-aware semantic embeddings via an InfoNCE-style contrastive loss, treating matched video–text pairs as positives and other in-batch descriptions as negatives. Third, Category Prototype Alignment (CP-A) applies the same contrastive objective at the category level, aligning high-level visual features with textual category prototypes to strengthen global inter-class separability.
Multi-Level Contrastive Optimization
A practical concern when combining contrastive objectives at multiple granularities is gradient conflict early in training, when representations lack structure. ML-CO addresses this with a two-stage schedule: classification loss plus CP-A first establish a discriminative category-level space; FG-SA is then introduced while both earlier objectives remain active. The total objective is a weighted sum L=Lcls+λfgLfg+λcpLcp. Training uses InternVL2.5-8B as backbone, with frozen vision weights, rank-8 LoRA adapters, a frozen text encoder, 8 frames at 448×448 resolution, bf16 precision, and AdamW with cosine scheduling.
Experimental results
On SMG, FG-SGL achieves 78.13% top-1 accuracy, outperforming all compared baselines including H2OFormer (75.44%, skeleton-based). On iMiGUE it reaches 62.58%, which trails several skeleton-based and RGB–skeleton methods (e.g., PL at 70.25%) but exceeds CLIP-based baselines such as CLIP-MG (61.82%), despite using only RGB at inference time. A summary of representative results:
| Method |
Modality |
SMG |
iMiGUE |
| TSM |
RGB |
65.41 |
– |
| JSSEL |
Skeleton |
68.03 |
64.12 |
| CVTCL |
RGB |
65.08 |
66.12 |
| PL |
RGB + Skeleton |
– |
70.25 |
| H2OFormer |
Skeleton |
75.44 |
70.00 |
| CLIP-MG |
RGB + Skeleton |
– |
61.82 |
| FG-SGL |
RGB |
78.13 |
62.58 |
Ablations attribute consistent gains to each component. Removing FG-SA costs 1.62 points on SMG and 3.29 on iMiGUE; removing CP-A costs 4.91 and 5.50 points respectively, indicating category-level alignment is the larger single contributor but fine-grained alignment matters more on datasets with higher inter-class similarity. Replacing FG-Text with class-level text drops accuracy to 76.55% and 59.36%, directly demonstrating that holistic category descriptions cannot substitute for instance-aware decomposition. Disabling the progressive ML-CO schedule costs 0.43 points on SMG and 1.86 on iMiGUE, supporting the claim that jointly optimizing both objectives from scratch produces unstable, conflicting gradients.
Limitations and open questions
Several caveats bear on these results. The performance gap on iMiGUE relative to skeleton-based methods shows that RGB-only inference does not fully recover pose information even with fine-grained semantic supervision; whether richer semantic priors can close this gap is left open. FG-Text depends entirely on manual annotation with cross-validation, raising scalability concerns that the authors acknowledge and defer to future work on "more scalable semantic representations." Additionally, no analysis is provided of how sensitive results are to annotation quality, the choice of the four semantic dimensions, or the hyperparameters λfg and λcp, nor of generalization beyond SMG and iMiGUE.
Conclusion
FG-SGL demonstrates that decomposing micro-gesture motions into structured, instance-level semantic attributes and aligning them with mid-level VLM representations — coordinated with category prototype alignment under a coarse-to-fine optimization scheme — improves recognition of visually similar gesture classes. Its strongest result, state-of-the-art RGB-only performance on SMG, supports the paper's central claim that fine-grained semantic guidance captures localized motion cues unavailable from category labels alone, though its advantage over skeleton-informed methods on iMiGUE remains unresolved.