Construct repository-level vulnerability data at training scale

Construct repository-level software-vulnerability training data at scale while retaining build configuration, cross-function information flow, and labels that remain valid at repository granularity.

Background

The paper explains that function-level slicing makes automatic labeling more affordable but removes calling context, build configuration, and interprocedural flow. This omission prevents function-sliced datasets from representing vulnerabilities whose manifestation depends on interactions across functions or on environmental configuration.

Repository-level benchmarks demonstrate that context-preserving evaluation is feasible, but they remain limited to hundreds of instances and are primarily used for evaluation rather than training. The unresolved problem is therefore to produce repository-level data at training scale without losing reliable vulnerability labels when moving beyond isolated functions.

References

The open problem is repository-level data at training scale: mining that retains build configuration and cross-function flow, and a label that survives the transition.

The Data Problem in Software Vulnerability Analysis: Artifacts, Quality, and Consumption  (2609.01503 - Nong et al., 1 Sep 2026) in Section 6, “Future Directions,” subsection “Preserve context: move mining up the granularity ladder”