Robustness to contamination of the verified-clean set

Measure how SAGE’s poison-detection performance degrades when the set of examples verified as clean contains mislabeled poisoned examples, thereby evaluating robustness to contamination of the verified-clean reference set.

Background

SAGE assumes that examples identified as clean are genuinely clean, even though it relaxes the stronger purity assumption made by conventional base-set defenses. A poisoned example mistakenly placed in the verified-clean subset can act as an erroneous reference in the similarity-weighted poison prediction and may therefore impair detection. The paper notes that its experiments do not test this form of contamination at any scale, leaving the degradation in performance unresolved.

References

A poison incorrectly included in $G_{\text{clean}$ would provide an erroneous reference for the similarity-weighted prediction. The assumption is easier to satisfy at our scale, since certifying a few dozen examples is more tractable than certifying the hundreds or thousands used elsewhere, but our experiments do not evaluate contamination of $G_{\text{clean}$ at any scale. Measuring how performance degrades under such contamination is an important robustness experiment left for future work.

— SAGE: Similarity-Based Cleaning of Poisoned Training Data from Verified Examples  (2610.01788 - Han et al., 1 Oct 2026) in Section Discussion and Conclusion, subsection Limitations, paragraph “Assumption of correct verified set”