Sensitivity of CANOPY to fault-quarantine treatment

Determine the effect of scoring every quarantined episode as a failure rather than removing it from the group, thereby quantifying the selection bias introduced by the CANOPY fault-quarantine rule.

Background

The CANOPY training procedure excludes episodes attributed to exogenous serving-layer faults before reward scoring and advantage computation. This avoids treating infrastructure failures as behavioral failures, but it may also remove trajectories that were misclassified by the serving layer.

The paper reports that no per-step quarantine count was logged and therefore does not provide the exclusion rate or the proposed sensitivity analysis. Retraining with quarantined episodes scored as failures would provide an upper bound on the effect of the selection rule.

References

Second, we did not log a per-step quarantine count, so we cannot report an exclusion rate or the sensitivity experiment a reader should reasonably want: retraining with every quarantined episode scored as a failure, which would upper-bound the selection effect. This is a gap in our instrumentation, and it is the second experiment we would add after the compute-matched group-size run of Appendix~\ref{app:cost}.

Explore More, Drift Less: Outcome-Only Reinforcement Learning Can Suffice for Long-Horizon Interactive Agents  (2609.01245 - Pu et al., 1 Sep 2026) in Technical Appendix, Section 'Environment Reliability and Fault Quarantine', paragraph 'Three caveats for reproduction'