Determine impact of benchmark contamination on LLM performance claims
Determine whether the reported performance gains of best-performing large language models on widely used NLP benchmarks are attributable to training data contamination and memorization due to inclusion of benchmark data in the training corpus, given the lack of transparency about training datasets. Establish methods to detect and quantify such contamination and its effect on evaluation outcomes to ensure valid comparisons and conclusions.
References
The lack of training data transparency associated with some of the best-performing LLMs means that we cannot be certain whether some of the performance gains are due to the memorisation of benchmarks being in the training datasets.
We investigated how these behaviors operate at inference time, but do not yet fully understand how they arise during training. Having more transparent data mixtures in general would help the field further study whether these behaviors arise from benchmark-guided model selection algorithms, data leakage, memorization, or other mechanisms.
We cannot rule out this form of benchmark contamination, and future versions of BavGround should consider replacing AI-generated general questions with human-authored items or questions drawn from independently verified sources.