Generalization to larger subject populations and diverse real-world scenes

Investigate the generalization ability of identity-aware human-object interaction motion captioning methods, including ID-HOINet, to larger subject populations and more diverse real-world scenes beyond the reorganized BEHAVE and InterCap datasets.

Background

The study evaluates ID-HOINet on reorganized BEHAVE and InterCap data containing only 18 subjects and a relatively limited range of objects, actions, and capture environments. Because of this restricted data coverage, it is unresolved whether the method generalizes to substantially larger subject populations and more varied real-world settings. The question concerns the robustness and applicability of identity-aware human-object interaction motion captioning beyond the controlled multi-view data used in the experiments.

References

The generalization ability of the proposed method to larger subject populations and more diverse real-world scenes therefore remains to be investigated.

Identity-Aware Human-Object Interaction Motion Captioning  (2608.20690 - Wang et al., 21 Aug 2026) in Section Limitations