Determine whether accent differences reflect voices or detector bias

Determine whether the higher frequency of the detector’s generic American accent label among synthetic-labeled calls reflects genuine differences in the callers’ voices or artifacts and biases of the shared detector and accent-tagging system.

Background

The commercial service supplies both the synthetic-speech score and an accent label. Synthetic-labeled calls receive a generic American label substantially more often than human-labeled calls, but the two outputs come from the same system and the same audio.

Because the accent tag is not independently measured, the observed difference may arise from genuine properties of the voices, from detector features that jointly influence both outputs, or from channel and preprocessing effects. An independent accent-labeling or acoustic analysis would be needed to resolve this attribution.

References

Whatever the callers intend, the synthesized voices in this traffic carry a generic American tag far more often than the human-labeled ones do; whether that is the voices or the tagger is what Appendix A cannot tell.

The Machines Are Calling: Measuring Automated and Synthetic Voices in Unwanted Inbound Calls  (2609.11137 - Shen et al., 10 Sep 2026) in Section 4.8, subsection “What the scripts do, what the channel hides, and what a single call still shows”; Appendix A, subsection “A shared-instrument null”