2000 character limit reached
Is the Best Better? Bayesian Statistical Model Comparison for Natural Language Processing (2010.03088v1)
Published 6 Oct 2020 in cs.CL, cs.LG, and stat.ME
Abstract: Recent work raises concerns about the use of standard splits to compare natural language processing models. We propose a Bayesian statistical model comparison technique which uses k-fold cross-validation across multiple data sets to estimate the likelihood that one model will outperform the other, or that the two will produce practically equivalent results. We use this technique to rank six English part-of-speech taggers across two data sets and three evaluation metrics.
Collections
Sign up for free to add this paper to one or more collections.
Paper Prompts
Sign up for free to create and run prompts on this paper using GPT-5.