Papers
Topics
Authors
Recent
Search
2000 character limit reached

Generalization Error Bounds for Learning under Censored Feedback

Published 14 Apr 2024 in cs.LG and stat.ML | (2404.09247v3)

Abstract: Generalization error bounds from learning theory provide statistical guarantees on how well an algorithm will perform on previously unseen data. In this paper, we characterize the impacts of data non-IIDness due to censored feedback (a.k.a. selective labeling bias) on such bounds. Censored feedback is ubiquitous in many real-world online selection and classification tasks (e.g., hiring, lending, recommendation systems) where the true label of a data point is only revealed if a favorable decision is made (e.g., accepting a candidate, approving a loan, displaying an ad), and remains unknown otherwise. We first derive an extension of the well-known Dvoretzky-Kiefer-Wolfowitz (DKW) inequality, which characterizes the gap between empirical and theoretical data distribution CDFs learned from IID data, to problems with non-IID data due to censored feedback. We then use this CDF error bound to provide a bound on the generalization error guarantees of a classifier trained on such non-IID data. We show that existing generalization error bounds (which do not account for censored feedback) fail to correctly capture the model's generalization guarantees, verifying the need for our bounds. We further analyze the effectiveness of (pure and bounded) exploration techniques, proposed by recent literature as a way to alleviate censored feedback, on improving our error bounds. Together, our findings illustrate how a decision maker should account for the trade-off between strengthening the generalization guarantees of an algorithm and the costs incurred in data collection when future data availability is limited by censored feedback.

Definition Search Book Streamline Icon: https://streamlinehq.com
References (42)
  1. Threshold bandits, with and without censored feedback. Advances In Neural Information Processing Systems, 29, 2016.
  2. Learning From Data. AMLBook, 2012.
  3. Margin based active learning. In Learning Theory: 20th Annual Conference on Learning Theory, COLT 2007, San Diego, CA, USA; June 13-15, 2007. Proceedings 20, pp.  35–50. Springer, 2007.
  4. Rademacher and gaussian complexities: Risk bounds and structural results. Journal of Machine Learning Research, 3(Nov):463–482, 2002.
  5. Equal opportunity in online classification with partial feedback. Advances in Neural Information Processing Systems, 32, 2019.
  6. A theory of learning from different domains. Machine learning, 79:151–175, 2010.
  7. A dvoretzky–kiefer–wolfowitz type inequality for the kaplan–meier estimator. In Annales de l’Institut Henri Poincare (B) Probability and Statistics, volume 35, pp.  735–763. Elsevier, 1999.
  8. Stability and generalization. The Journal of Machine Learning Research, 2:499–526, 2002.
  9. A simple non-iid sampling approach for efficient training and better generalization. arXiv preprint arXiv:1811.09347, 2018.
  10. Algorithmic decision making and the cost of fairness. In Proceedings of the 23rd acm sigkdd international conference on knowledge discovery and data mining, pp.  797–806, 2017.
  11. Region-based active learning. In The 22nd International Conference on Artificial Intelligence and Statistics, pp.  2801–2809. PMLR, 2019.
  12. Adaptive region-based active learning. In International Conference on Machine Learning, pp.  2144–2153. PMLR, 2020.
  13. Accurate inference for adaptive linear models. In International Conference on Machine Learning, pp.  1194–1203. PMLR, 2018.
  14. A probabilistic theory of pattern recognition, volume 31. Springer Science & Business Media, 2013.
  15. UCI machine learning repository, 2017. URL http://archive.ics.uci.edu/ml.
  16. Runaway feedback loops in predictive policing. In Conference on fairness, accountability and transparency, pp.  160–171. PMLR, 2018.
  17. Yair Goldberg. Hoeffding-type and bernstein-type inequalities for right censored data. arXiv preprint arXiv:1903.01991, 2019.
  18. Support vector regression for right censored data. 2017.
  19. Active learning for skewed data sets. arXiv preprint arXiv:2005.11442, 2020.
  20. Fair decisions despite imperfect predictions. In International Conference on Artificial Intelligence and Statistics, pp.  277–287. PMLR, 2020.
  21. Uniform chernoff and dvoretzky-kiefer-wolfowitz-type inequalities for markov chains and related processes. Journal of Applied Probability, 51(4):1100–1113, 2014.
  22. Generalization bounds for non-stationary mixing processes. Machine Learning, 106(1):93–117, 2017.
  23. Partitioned active learning for heterogeneous systems. Journal of Computing and Information Science in Engineering, 23(4):041009, 2023.
  24. Pascal Massart. The tight constant in the dvoretzky-kiefer-wolfowitz inequality. The annals of Probability, pp.  1269–1283, 1990.
  25. Minimum complexity regression estimation with weakly dependent observations. IEEE Transactions on Information Theory, 42(6):2133–2145, 1996.
  26. Stability bounds for non-iid processes. Advances in Neural Information Processing Systems, 20, 2007.
  27. Rademacher complexity bounds for non-iid processes. Advances in Neural Information Processing Systems, 21, 2008.
  28. Michael Naaman. On the tight constant in the multivariate dvoretzky–kiefer–wolfowitz inequality. Statistics & Probability Letters, 173:109088, 2021.
  29. Why adaptively collected data have negative bias and how to correct for it. In International Conference on Artificial Intelligence and Statistics, pp.  1261–1269. PMLR, 2018.
  30. David Pollard. Convergence of stochastic processes. Springer Science & Business Media, 2012.
  31. Unintended selection: Persistent qualification rate disparities and interventions. Advances in Neural Information Processing Systems, 34:26053–26065, 2021.
  32. Online learning with markov sampling. Analysis and Applications, 7(01):87–113, 2009.
  33. Fast learning from non-iid observations. Advances in neural information processing systems, 22, 2009.
  34. Learning from dependent observations. Journal of Multivariate Analysis, 100(1):175–194, 2009.
  35. Personalized federated learning with clustered generalization. 2021.
  36. Cross-validation for geospatial data: Estimating generalization performance in geostatistical problems. Transactions on Machine Learning Research, 2023.
  37. Dennis Wei. Decision-making under selective labels: Optimal finite-domain policies and beyond. In International Conference on Machine Learning, pp.  11035–11046. PMLR, 2021.
  38. Adaptive data debiasing through bounded exploration. Advances in Neural Information Processing Systems, 35:1516–1528, 2022.
  39. Bin Yu. Rates of convergence for empirical processes of stationary mixing sequences. The Annals of Probability, pp.  94–116, 1994.
  40. Gray learning from non-iid data with out-of-distribution samples. arXiv preprint arXiv:2206.09375, 2022.
  41. A generalization theory based on independent and task-identically distributed assumption. arXiv preprint arXiv:1911.12603, 2019.
  42. The generalization performance of erm algorithm with strongly mixing observations. Machine learning, 75(3):275–295, 2009.

Summary

No one has generated a summary of this paper yet.

Paper to Video (Beta)

No one has generated a video about this paper yet.

Whiteboard

No one has generated a whiteboard explanation for this paper yet.

Open Problems

We haven't generated a list of open problems mentioned in this paper yet.

Continue Learning

We haven't generated follow-up questions for this paper yet.

Collections

Sign up for free to add this paper to one or more collections.

Tweets

Sign up for free to view the 4 tweets with 4 likes about this paper.