Discovering intrinsically well-aligned cross-modal corpora for pretraining
Identify and construct large-scale datasets that provide intrinsically well-aligned cross-modal supervision (for example, instructional videos with tightly aligned visual content and spoken words) to enable effective Transformer-based multimodal pretraining without incurring prohibitive labeling and alignment costs.
References
How to look for more corpora that intrinsically have well-aligned cross-modal supervision, such as instructional videos, is still an open problem.
We evaluate video-text, audio-text, and image-text settings separately, and do not train joint audio-visual models. This leaves open how well LAION-BVD supports learning unified representations across modalities, particularly for tasks that require tight synchronization between visual and auditory signals.