Effect of human-response-length information on Turing-test performance

Determine the extent to which including the human respondent’s answer length in the model prompt affected the outcome of the Finnish-language Turing Test, by comparing experimental conditions with and without information about human response length.

Background

The model-generated role prompt instructed ChatGPT 5.2 to produce answers of approximately the same length as the human respondent’s answer. The authors note that this design choice may have facilitated imitation by preventing response length from becoming a distinguishing cue, while also reducing the possibility that participants would identify the model through an experimental artefact rather than domain-specific competence. The unresolved issue is whether controlling response length materially influenced the model’s successful performance in the Finnish Turing Test. The authors propose directly comparing conditions in which the model does and does not receive information about human response length.

References

The extent to which this design choice affected the outcome remains unclear. Future studies could examine this directly by comparing conditions with and without information about human response length.

Cultural Competence in Context: A Large Language Model Passes the Turing Test in Finland  (2609.18394 - Segersven et al., 16 Sep 2026) in Section 4, Discussion, paragraph beginning “One limitation in our study concerns the inclusion of the human respondent's answer length in the model prompt.”