Determine the relevance of memorized model copies to infringement rulings

Determine whether, and to what extent, the fact that language models can be copies of memorized training data will affect courts’ rulings on copyright infringement.

Background

The author distinguishes the technical possibility that LLMs may contain copies of memorized training data from the separate legal question of how that fact will be treated in copyright litigation. The paper explicitly leaves unresolved whether, or how significantly, memorized copies in models will matter when courts decide infringement cases.

References

(Yes, models can be copies of the training data they've memorized. But it's complicated, and I don't know if or how much this will matter when it comes to courts making rulings on infringement.)

Playing Whack-a-Mole with misconceptions about memorization, extraction, and copyright  (2609.09320 - Cooper, 8 Sep 2026) in Section 7, “Some closing thoughts about the field”