Sensitivity of CANOPY to fault-quarantine treatment
Determine the effect of scoring every quarantined episode as a failure rather than removing it from the group, thereby quantifying the selection bias introduced by the CANOPY fault-quarantine rule.
References
Second, we did not log a per-step quarantine count, so we cannot report an exclusion rate or the sensitivity experiment a reader should reasonably want: retraining with every quarantined episode scored as a failure, which would upper-bound the selection effect. This is a gap in our instrumentation, and it is the second experiment we would add after the compute-matched group-size run of Appendix~\ref{app:cost}.
— Explore More, Drift Less: Outcome-Only Reinforcement Learning Can Suffice for Long-Horizon Interactive Agents
(2609.01245 - Pu et al., 1 Sep 2026) in Technical Appendix, Section 'Environment Reliability and Fault Quarantine', paragraph 'Three caveats for reproduction'