Explain TraceFlow’s deficits on Counting and Occlusion tasks

Determine the causal explanation for why TraceFlow’s bounded, progress-aligned guidance underperforms the unfrozen baseline on the RoboMemArena Counting and Occlusion suites, and establish an online switch criterion that prevents these deficits.

Background

TraceFlow improves performance on ordering and transfer tasks by guiding a frozen flow-matching action expert toward retrieved successful action windows and away from retrieved failed windows. However, the reported gains are task-specific: performance on the RoboMemArena Counting and Occlusion suites remains below the frozen baseline under every tested guidance cap and TraceBank scope.

The paper’s retrospective audit associates part of the Occlusion degradation with persistent subtask-selection errors, while Counting already loses performance before the identified persistent-error boundary. Because this analysis does not establish causation or provide a mechanism for disabling guidance when it is harmful, the authors explicitly leave both the causal explanation and an online mitigation criterion unresolved.

References

Thus Upper error alone does not explain both suites; this retrospective association supplies no causal explanation or online switch criterion. These deficits remain unresolved.

— TraceFlow: Guiding Frozen Flow-Matching Robot Policies with Success and Failure Traces  (2609.20646 - Zhang et al., 17 Sep 2026) in Section 'Limitations and Conclusion'