Integrate large-capacity models into real-time robot control frameworks

Identify principled ways to seamlessly incorporate large-capacity models, such as foundation models, into robotic control frameworks for real-time inference, including the design and evaluation of hierarchical learning or slow–fast control schemes that ensure responsiveness and reliability.

Background

The survey notes that foundational and large-capacity models are increasingly used in embodied AI, but integrating them into real-time control remains challenging due to latency, responsiveness, and reliability constraints.

As an explicit open problem, the authors call for principled integration strategies—potentially via hierarchical learning or slow–fast control—to enable seamless, real-time use of large models within control loops for embodied agents.

References

Open research problems and viable approaches include: 1) mixing different proportions of prior data distribution when fine-tuning on the latest data to alleviate catastrophic forgetting , 2) developing efficient prototypes from prior distributions or curricula for task inference in learning new tasks, 3) improving training stability and sample efficiency of online learning algorithms, 4) identifying principled ways to seamlessly incorporate large-capacity models into control frameworks, potentially through hierarchical learning or slow-fast control, for real-time inference.

Aligning Cyber Space with Physical World: A Comprehensive Survey on Embodied AI  (2407.06886 - Liu et al., 2024) in Section 8, Challenges and Future Directions – Continual Learning

Open problems remain: unified context representations, validation of subjective norms, and latency of LLM/MPC inference.

Context-Aware Intelligent Vehicles  (2609.00682 - Liu et al., 1 Sep 2026) in Section 4, subsection “Context-Aware Planning and Control,” Takeaway

The existing version also exhibits limitations. First, the hierarchical slow--fast architecture introduces additional computational and memory overhead, motivating future research on model compression and asynchronous inference. Second, the current model does not explicitly capture multimodal futures or predictive uncertainty, which may limit its performance in ambiguous and rare driving scenarios. We leave solving them as future works.

Drive-HWM: Hierarchical World Models for Dynamic-Latent Guided Autonomous Driving  (2609.03572 - Fan et al., 3 Sep 2026) in Conclusion