Robustness to contamination of the verified-clean set
Measure how SAGE’s poison-detection performance degrades when the set of examples verified as clean contains mislabeled poisoned examples, thereby evaluating robustness to contamination of the verified-clean reference set.
References
A poison incorrectly included in $G_{\text{clean}$ would provide an erroneous reference for the similarity-weighted prediction. The assumption is easier to satisfy at our scale, since certifying a few dozen examples is more tractable than certifying the hundreds or thousands used elsewhere, but our experiments do not evaluate contamination of $G_{\text{clean}$ at any scale. Measuring how performance degrades under such contamination is an important robustness experiment left for future work.
— SAGE: Similarity-Based Cleaning of Poisoned Training Data from Verified Examples
(2610.01788 - Han et al., 1 Oct 2026) in Section Discussion and Conclusion, subsection Limitations, paragraph “Assumption of correct verified set”