SimLabel: Similarity-Weighted Semi-supervision for Multi-annotator Learning with Missing Labels

Published 13 Apr 2025 in cs.MM | (2504.09525v1)

Abstract: Multi-annotator learning has emerged as an important research direction for capturing diverse perspectives in subjective annotation tasks. Typically, due to the large scale of datasets, each annotator can only label a subset of samples, resulting in incomplete (or missing) annotations per annotator. Traditional methods generally skip model updates for missing parts, leading to inefficient data utilization. In this paper, we propose a novel similarity-weighted semi-supervised learning framework (SimLabel) that leverages inter-annotator similarities to generate weighted soft labels. This approach enables the prediction, utilization, and updating of missing parts rather than discarding them. Meanwhile, we introduce a confidence assessment mechanism combining maximum probability with entropy-based uncertainty metrics to prioritize high-confidence predictions to impute missing labels during training as an iterative refinement pipeline, which continuously improves both inter-annotator similarity estimates and individual model performance. For the comprehensive evaluation, we contribute a new video emotion recognition dataset, AMER2, which exhibits higher missing rates than its predecessor, AMER. Experimental results demonstrate our method's superior performance compared to existing approaches under missing-label conditions.