Determine the number of annotators needed for fair soft-label estimates

Determine how many human annotators are required to obtain a fair estimate of the soft-label distribution representing human judgment variation.

Background

The study uses datasets with varying numbers of human annotations, and the authors note that the soft-label distributions in three of the five datasets are sparse because they contain fewer than six annotators per instance. Since the quality and fairness of an estimated human judgment distribution depend on the amount of annotation evidence, the required annotator count remains unresolved.

References

It is unclear how many annotators are needed to get a fair estimate of the soft-labels.

Post-hoc Alignment of LLM-judges to Human Judgment Distribution  (2609.01073 - Steindl et al., 1 Sep 2026) in Section 7, Limitations