Determine the proportion of book-coverage results attributable to short matches

Determine the exact share of short matches contributing to each book memorization coverage result, given that the distribution of match lengths is not reported.

Background

The paper critiques the bmc@5 metric because it aggregates matches of at least five consecutive tokens, whereas substantially longer spans are generally required to distinguish memorization from coincidental language-model output. The authors state that the absence of match-length distributions prevents disentangling how much of any reported coverage percentage is produced by such short, potentially spurious matches. Establishing this share would be necessary to assess whether the headline coverage results provide valid evidence of memorization.

References

But it's very hard to disentangle because the paper doesn't report the distribution of match lengths, so the exact share of short matches in any given coverage number is unknown.

Playing Whack-a-Mole with misconceptions about memorization, extraction, and copyright  (2609.09320 - Cooper, 8 Sep 2026) in Section 5.1, “Five ‘words’ is too short as a minimum for claiming memorization”