Determine the causes of cross-model repetition in open-ended LLM outputs
Determine the specific causes of the high semantic similarity and verbatim overlap observed across different large language models when generating responses to open-ended queries, especially identifying whether shared pretraining data pipelines across regions or contamination from synthetic data are responsible, given that the exact causes are currently unclear due to proprietary training details.
References
Although the exact causes remain unclear due to proprietary training details, possible explanations include shared data pipelines across regions or contamination from synthetic data. We highlight the need for future work to rigorously investigate the sources of such cross-model repetition.
However, it remains unclear whether this homogeneity is a temporary byproduct of a still-developing technology or an inevitable—potentially compounding—feature of statistical LLMs.
Our aggregate analysis does not establish whether the observed creative similarity is uniform across all types of open-ended prompts, so a prompt-level decomposition may help identify the factors driving model convergence.