Determine whether reported gains reflect semantic understanding or noisy-supervision artifacts

Determine whether performance gains reported for machine-learning methods that predict indirect control-flow edges in stripped binaries reflect genuine semantic understanding or artifacts caused by noisy supervision.

Background

The paper emphasizes that machine-learning models for indirect control-flow prediction are highly sensitive to ground-truth quality. Existing datasets, including Callee and CupidCall, rely on single-source labels that may over-approximate or under-approximate the true set of dynamic targets, and they lack a clean standardized test protocol.

Because label errors can substantially affect benchmark results, a reported improvement in precision, recall, or F1 may arise either from better semantic modeling or from artifacts of noisy supervision. The paper addresses this concern operationally by constructing a clean test protocol, but the general question of how much reported progress reflects genuine semantic understanding remains unresolved in the stated discussion.

References

As a result, it is often unclear whether reported gains reflect real semantic understanding or artifacts of noisy supervision.

Long-Range Indirect Control-Flow Prediction in Stripped Binaries via Dual Virtual Hubs and Multi-Task Graph Learning  (2609.03280 - Liu et al., 3 Sep 2026) in Challenge 3, Section 1 (Introduction)