Papers
Topics
Authors
Recent
Gemini 2.5 Flash
Gemini 2.5 Flash
110 tokens/sec
GPT-4o
56 tokens/sec
Gemini 2.5 Pro Pro
44 tokens/sec
o3 Pro
6 tokens/sec
GPT-4.1 Pro
47 tokens/sec
DeepSeek R1 via Azure Pro
28 tokens/sec
2000 character limit reached

Less is more: Selecting informative and diverse subsets with balancing constraints (2104.12835v2)

Published 26 Apr 2021 in cs.CV, cs.AI, and cs.LG

Abstract: Deep learning has yielded extraordinary results in vision and natural language processing, but this achievement comes at a cost. Most models require enormous resources during training, both in terms of computation and in human labeling effort. We show that we can identify informative and diverse subsets of data that lead to deep learning models with similar performance as the ones trained with the original dataset. Prior methods have exploited diversity and uncertainty in submodular objective functions for choosing subsets. In addition to these measures, we show that balancing constraints on predicted class labels and decision boundaries are beneficial. We propose a novel formulation of these constraints using matroids, an algebraic structure that generalizes linear independence in vector spaces, and present an efficient greedy algorithm with constant approximation guarantees. We outperform competing baselines on standard classification datasets such as CIFAR-10, CIFAR-100, ImageNet, as well as long-tailed datasets such as CIFAR-100-LT.

User Edit Pencil Streamline Icon: https://streamlinehq.com
Authors (6)
  1. Srikumar Ramalingam (40 papers)
  2. Daniel Glasner (7 papers)
  3. Kaushal Patel (5 papers)
  4. Raviteja Vemulapalli (29 papers)
  5. Sadeep Jayasumana (19 papers)
  6. Sanjiv Kumar (123 papers)
Citations (5)

Summary

We haven't generated a summary for this paper yet.