Resolve whether longer listening windows reduce listener disagreement

Determine whether extending the evaluated speech window from ten to approximately twenty seconds resolves listener disagreement about synthetic versus human speech, rather than merely shifting listeners toward synthetic judgments.

Background

The validation study primarily uses ten-second onset-aligned clips, many of which contain only a few seconds of actual speech. A second round presents longer, silence-reduced clips to examine whether the limited stimulus causes disagreement.

Both contested clips and agreed control clips moved toward synthetic judgments in the longer-window condition, and the study did not obtain enough repeated judgments per clip to establish whether disagreement was actually resolved. This leaves unresolved whether additional speech improves classification or simply changes listener response criteria.

References

The contested-minus-control difference is indistinguishable from zero at this sample size, and there is no evidence yet that disagreement on the contested clips was resolved. Answering that question would need at least three judgments per clip, which the round did not reach, and its judgments are not pooled with the precision estimate above (analysis/analyze_round2.py).

The Machines Are Calling: Measuring Automated and Synthetic Voices in Unwanted Inbound Calls  (2609.11137 - Shen et al., 10 Sep 2026) in Appendix A, subsection “Round two”