Enabling general-purpose multimodal interaction at scale
Determine effective approaches to enable general-purpose multimodal interaction within large-scale models trained on interleaved vision–language data, so that a single system can flexibly process and generate interleaved inputs and outputs across tasks.
References
In particular, it remains unclear how to effectively learn long videos interleaved with text, how to enable general-purpose multimodal interaction, and how to efficiently predict tens of thousands of visual tokens, which pose stringent demands on pre-training, post-training, and inference, respectively.
This design presents several open design considerations, including whether the front cerebellum should perform additional reasoning based on the transcribed query before passing it to the subsequent stage, and whether a trainable back brain could directly receive the user's complete multimodal input stream in real time.