Determine the source of Skyra’s temporal-grid sensitivity

Determine whether the large difference in Skyra-SFT and Skyra-RL fake-recall rates between exact-five-second clips and clips with other final timestamps arises from duration-dependent visual content or from the temporal labels themselves, despite preserving the same sampled visual frames.

Background

RA-Bench evaluates the fine-tuned MLLM detectors Skyra-SFT and Skyra-RL using either official timestamps or frame-index labels attached to the same 16 sampled video frames. The authors observe that clips ending at exactly five seconds receive substantially higher fake-recall rates than clips with other final timestamps when official timestamps are shown.

Replacing absolute timestamps with frame indices largely eliminates this exact-five-second gap, while leaving the visual frames and their order unchanged. This indicates strong dependence on the temporal-label representation, but the authors do not establish whether the effect is caused by visual or duration-related properties correlated with the displayed timestamps, or by the labels themselves. Resolving this distinction is important for determining whether Skyra’s apparent detection performance reflects content-based forensic evidence or a protocol-specific prior.

References

These comparisons reveal a strong association between the displayed temporal grid and Skyra's predictions, but they do not determine whether the difference arises from duration-dependent visual content or from the temporal labels themselves.

Can We Defend Against AI-Generated Video Attacks on Real-World Crisis Events? A Systematic Evaluation of Detectors, Generators and Social Dissemination  (2608.14391 - Liang et al., 14 Aug 2026) in Section 4.1.3, paragraph beginning “The fixed-duration control reveals an additional anomaly”