Causal isolation of gradient starvation in Stage 3 forgetting

Determine the causal contribution of scale-conditioned gradient imbalance to Stage 3 forgetting in YOLOMG by rerunning Stage 3 with controlled batch-size and learning-rate settings that match Stage 2, thereby isolating gradient imbalance from concurrent optimization differences.

Background

The paper interprets the rapid collapse of large-target detection during CST Anti-UAV fine-tuning as evidence of a scale-conditioned gradient imbalance: CST contains no large targets, while small-target gradients dominate the shared detection heads. However, Stage 2 and Stage 3 use different batch sizes and learning rates, so the observed forgetting may also reflect optimization settings rather than scale shift alone.

The authors state that a controlled Stage 3 rerun is needed to separate these explanations. Establishing this causal attribution would clarify whether the proposed gradient-starvation mechanism, rather than a confounding change in optimization, is responsible for the approximately 18-fold increase in forgetting.

References

The probe supports the scale-conditioned account as a candidate mechanism; isolating its causal weight from the concurrent batch-size and learning-rate differences (Section~\ref{sec:discussion}) would require a controlled Stage~3 rerun and is left to future work.

— Catastrophic Forgetting in Sequential Thermal Anti-UAV Detection: The Role of Scale-Conditioned Gradient Imbalance  (2610.08315 - Nguyen et al., 6 Oct 2026) in Section 3.2, subsection “Direct Gradient Probe at the Stage 2→3 Boundary”; also discussed in Section 4.4, “Limitations”