Papers
Topics
Authors
Recent
Gemini 2.5 Flash
Gemini 2.5 Flash
97 tokens/sec
GPT-4o
53 tokens/sec
Gemini 2.5 Pro Pro
44 tokens/sec
o3 Pro
5 tokens/sec
GPT-4.1 Pro
47 tokens/sec
DeepSeek R1 via Azure Pro
28 tokens/sec
2000 character limit reached

Understanding the Detrimental Class-level Effects of Data Augmentation (2401.01764v1)

Published 7 Dec 2023 in cs.CV and cs.LG

Abstract: Data augmentation (DA) encodes invariance and provides implicit regularization critical to a model's performance in image classification tasks. However, while DA improves average accuracy, recent studies have shown that its impact can be highly class dependent: achieving optimal average accuracy comes at the cost of significantly hurting individual class accuracy by as much as 20% on ImageNet. There has been little progress in resolving class-level accuracy drops due to a limited understanding of these effects. In this work, we present a framework for understanding how DA interacts with class-level learning dynamics. Using higher-quality multi-label annotations on ImageNet, we systematically categorize the affected classes and find that the majority are inherently ambiguous, co-occur, or involve fine-grained distinctions, while DA controls the model's bias towards one of the closely related classes. While many of the previously reported performance drops are explained by multi-label annotations, our analysis of class confusions reveals other sources of accuracy degradation. We show that simple class-conditional augmentation strategies informed by our framework improve performance on the negatively affected classes.

User Edit Pencil Streamline Icon: https://streamlinehq.com
Authors (7)
  1. Polina Kirichenko (15 papers)
  2. Mark Ibrahim (36 papers)
  3. Randall Balestriero (91 papers)
  4. Diane Bouchacourt (32 papers)
  5. Ramakrishna Vedantam (19 papers)
  6. Hamed Firooz (27 papers)
  7. Andrew Gordon Wilson (133 papers)
Citations (8)

Summary

We haven't generated a summary for this paper yet.