Scale and significance of semantic-duplicate contamination in LLM training corpora
Determine the scale (prevalence) and significance (impact on evaluation outcomes) of contamination of large language model training corpora by semantic duplicates of benchmark test items—i.e., training examples whose substantive meaning matches benchmark items despite low or no syntactic overlap—and quantify how such contamination affects benchmark validity and interpretation.
References
Research on these semantic duplicates has often focused on their robustness to standard, n-gram based ‘decontamination’ methods, but the scale and significance of the phenomenon remains a mostly open question.
Since DeCon operates by surface-level $n$-gram matching, this analysis bounds verbatim overlap rather than semantic or schema-level similarity, which is particularly relevant for MCP-related tool definitions. We leave a deeper semantic audit of such overlap to future work.