Quantitative comparison with point-cloud policies in moving eye-in-hand settings

Establish a reliable quantitative comparison between ObstaDiff and point-cloud imitation-learning policies, particularly DP3, when deployed with a continuously moving eye-in-hand camera in dense foliage.

Background

ObstaDiff is evaluated using image-aligned RGB-D observations from a moving eye-in-hand camera. The paper does not report DP3 as a quantitative baseline because the authors encountered unstable point-cloud reconstruction under camera motion, foliage occlusion, missing depth, registration, cropping, and sampling. Consequently, the relative performance of ObstaDiff and point-cloud policies in this deployment regime remains unresolved.

References

Our evaluation is also confined to a single indoor greenhouse testbed with three obstacle plant species and one target crop, and the comparison against point-cloud policies remains open.

ObstaDiff: Generalizable Diffusion Policy Learning via Obstacle-aware Representations  (2609.10918 - Wang et al., 10 Sep 2026) in Section 5, Limitations