Evaluate CGMIA on larger language models

Determine whether CGMIA remains effective for larger code-generation language models with substantially more parameters than the smaller models used in the study.

Background

The experiments use two NVIDIA 4090 GPUs and consequently restrict target and shadow models to smaller-scale models that can be fine-tuned within the available hardware budget. Although CGMIA performs effectively in these experiments, its scalability and detection effectiveness for substantially larger LLMs are not established. Resolving this issue requires evaluating the method on larger models to determine whether its membership signals and classifier performance persist at scale.

References

Consequently, the effectiveness of CGMIA on larger models with substantially more parameters remains uncertain and requires further investigation.

Keep Evaluation Fair: Detecting Data Leakage in Code Generation Benchmarks via Membership Inference Attacks  (2609.09865 - Zhao et al., 9 Sep 2026) in Section 10, Threats of validity, item (1)