Quantify human annotation disagreement and correction rates

Determine the inter-annotator agreement for the BanglaTurn corpus and quantify how often human review changes the automatically generated labels.

Background

BanglaTurn labels were produced from diarization and Gemini 2.0 Flash Lite proposals and then checked by a single native Bangla-speaking annotator. Because no second annotator independently reviewed the clips, the reliability of the labeling process cannot be assessed through inter-annotator agreement. In addition, corrections were made directly in place without recording the original proposals, so the frequency of human changes to the automatic labels is unavailable.

References

This protocol has two gaps. Because one person checked every label, we cannot report inter-annotator agreement. Because corrections were made in place without a log, we cannot report how often the human check changed the automatic labels.

— BanglaTurn: A Benchmark and Whisper-Based Model for End-of-Turn Detection in Bangla Speech  (2609.29371 - Maruf, 24 Sep 2026) in Section 3.3, Annotation protocol

The labels rest on a single annotator's check of automatic proposals, with no inter-annotator agreement and no record of how many labels the check changed, so the label error rate is unknown.

— BanglaTurn: A Benchmark and Whisper-Based Model for End-of-Turn Detection in Bangla Speech  (2609.29371 - Maruf, 24 Sep 2026) in Section 6.3, Uncertainty and label quality; Section 8, Limitations