Transferability of Stage-1 MLLM Validation to the Full Audit Distribution
Determine whether the Cohen’s kappa validation measured for Gemini 2.5 Flash under the eight-frame-plus-text E3 configuration on the 300-video Stage-1 reference sample transfers to the full Stage-2 video distribution.
References
First, the Stage-1 $\kappa = 0.42$ was measured on the 300-video reference sample, and whether it transfers to the full Stage-2 distribution is an open question this design cannot answer.
— Auditing Exposure to Harmful Content on TikTok using Multimodal Language Models: A Cross-National, Age-Stratified Study
(2608.17583 - Saffari et al., 18 Aug 2026) in Discussion and Conclusion, paragraph “MLLM auditing is feasible at this scale, conditional on three structural caveats”
Second, at scale E2 (native video) reports consistently higher harm rates than primary E3 (eight frames plus text) ($+3$--$12$ pp), and which modality is closer to human truth on the Stage-2 population is not answerable without a Stage-2 annotation pass.
— Auditing Exposure to Harmful Content on TikTok using Multimodal Language Models: A Cross-National, Age-Stratified Study
(2608.17583 - Saffari et al., 18 Aug 2026) in Discussion and Conclusion, paragraph “MLLM auditing is feasible at this scale, conditional on three structural caveats”