QR-CLIP: Introducing Explicit Open-World Knowledge for Location and Time Reasoning
Abstract: Daily images may convey abstract meanings that require us to memorize and infer profound information from them. To encourage such human-like reasoning, in this work, we teach machines to predict where and when it was taken rather than performing basic tasks like traditional segmentation or classification. Inspired by Horn's QR theory, we designed a novel QR-CLIP model consisting of two components: 1) the Quantity module first retrospects more open-world knowledge as the candidate language inputs; 2) the Relevance module carefully estimates vision and language cues and infers the location and time. Experiments show our QR-CLIP's effectiveness, and it outperforms the previous SOTA on each task by an average of about 10% and 130% relative lift in terms of location and time reasoning. This study lays a technical foundation for location and time reasoning and suggests that effectively introducing open-world knowledge is one of the panaceas for the tasks.
- J. A. Crowder and S. Friess, “Artificial psychology: The psychology of ai,” People, vol. 2, no. 3, pp. 4–5, 2012.
- D. Hassabis, D. Kumaran, C. Summerfield, and M. Botvinick, “Neuroscience-inspired artificial intelligence,” Neuron, vol. 95, no. 2, pp. 245–258, 2017.
- J. Wirtz, P. G. Patterson, W. H. Kunz, T. Gruber, V. N. Lu, S. Paluch, and A. Martins, “Brave new world: service robots in the frontline,” Journal of Service Management, vol. 29, no. 5, pp. 907–931, 2018.
- P. Salovey and J. D. Mayer, “Emotional intelligence,” Imagination, Cognition and Personality, vol. 9, no. 3, pp. 185–211, 1990.
- J. Schmidhuber, “On learning to think: Algorithmic information theory for novel combinations of reinforcement learning controllers and recurrent neural world models,” arXiv preprint arXiv:1511.09249, 2015.
- S. Harrer, P. Shah, B. Antony, and J. Hu, “Artificial intelligence for clinical trial design,” Trends in pharmacological sciences, vol. 40, no. 8, pp. 577–591, 2019.
- X. Fu, B. Zhou, I. Chandratreya, C. Vondrick, and D. Roth, “There’s a time and place for reasoning beyond the image,” in ACL, 2022.
- Z. Wang, X. Shan, X. Zhang, and J. Yang, “N24news: A new dataset for multimodal news classification,” in LREC, 2022.
- M. Yang, L. Jiao, F. Liu, B. Hou, S. Yang, Y. Zhang, and J. Wang, “Coarse-to-fine contrastive self-supervised feature learning for land-cover classification in sar images with limited labeled data,” IEEE TIP, vol. 31, pp. 6502–6516, 2022.
- H. Sun, X. Zheng, and X. Lu, “A supervised segmentation network for hyperspectral image classification,” IEEE TIP, vol. 30, pp. 2810–2825, 2021.
- Y. Pei, Y. Huang, Q. Zou, X. Zhang, and S. Wang, “Effects of image degradation and degradation removal to cnn-based image classification,” IEEE TPAMI, vol. 43, no. 4, pp. 1239–1253, 2019.
- M. Shoeybi, M. Patwary, R. Puri, P. LeGresley, J. Casper, and B. Catanzaro, “Megatron-lm: Training multi-billion parameter language models using model parallelism,” arXiv preprint arXiv:1909.08053, 2019.
- A. R. Fabbri, I. Li, T. She, S. Li, and D. Radev, “Multi-news: A large-scale multi-document summarization dataset and abstractive hierarchical model,” in ACL, 2019.
- H. Jiang and Y. Mu, “Joint video summarization and moment localization by cross-task sample transfer,” in CVPR, 2022, pp. 16 388–16 398.
- C. Ma, W. E. Zhang, M. Guo, H. Wang, and Q. Z. Sheng, “Multi-document summarization via deep learning techniques: A survey,” ACM Computing Surveys, vol. 55, no. 5, pp. 1–37, 2022.
- C. Conforti, J. Berndt, M. T. Pilehvar, C. Giannitsarou, F. Toxvaerd, and N. Collier, “Stander: An expert-annotated dataset for news stance detection and evidence retrieval,” in ACL, 2020.
- R. Zuo, X. Deng, K. Chen, Z. Zhang, Y.-K. Lai, F. Liu, C. Ma, H. Wang, Y.-J. Liu, and H. Wang, “Fine-grained video retrieval with scene sketches,” IEEE TIP, 2023.
- S. Reed, K. Zolna, E. Parisotto, S. G. Colmenarejo, A. Novikov, G. Barth-maron, M. Giménez, Y. Sulsky, J. Kay, J. T. Springenberg et al., “A generalist agent,” TMLR.
- R. Schwartz, J. Dodge, N. A. Smith, and O. Etzioni, “Green ai,” Communications of the ACM, vol. 63, no. 12, pp. 54–63, 2020.
- E. Strubell, A. Ganesh, and A. McCallum, “Energy and policy considerations for deep learning in nlp,” arXiv preprint arXiv:1906.02243, 2019.
- J. Rasley, S. Rajbhandari, O. Ruwase, and Y. He, “Deepspeed: System optimizations enable training deep learning models with over 100 billion parameters,” in SIGKDD, 2020.
- M. Ott, S. Edunov, A. Baevski, A. Fan, S. Gross, N. Ng, D. Grangier, and M. Auli, “fairseq: A fast, extensible toolkit for sequence modeling,” in ACL, 2019.
- Y. Huang, Y. Cheng, A. Bapna, O. Firat, D. Chen, M. Chen, H. Lee, J. Ngiam, Q. V. Le, Y. Wu et al., “Gpipe: Efficient training of giant neural networks using pipeline parallelism,” in NeurIPS, 2019.
- T. Brown, B. Mann, N. Ryder, M. Subbiah, J. D. Kaplan, P. Dhariwal, A. Neelakantan, P. Shyam, G. Sastry, A. Askell et al., “Language models are few-shot learners,” in NeurIPS, 2020.
- A. Radford, J. W. Kim, C. Hallacy, A. Ramesh, G. Goh, S. Agarwal, G. Sastry, A. Askell, P. Mishkin, J. Clark et al., “Learning transferable visual models from natural language supervision,” in ICML, 2021.
- L. Ouyang, J. Wu, X. Jiang, D. Almeida, C. L. Wainwright, P. Mishkin, C. Zhang, S. Agarwal, K. Slama, A. Ray et al., “Training language models to follow instructions with human feedback,” arXiv preprint arXiv:2203.02155, 2022.
- A. Nichol, H. Jun, P. Dhariwal, P. Mishkin, M. Chen et al., “Point-e: A system for generating 3d point clouds from complex prompts,” arXiv preprint arXiv:2212.08751, 2022.
- A. Radford, J. W. Kim, T. Xu, G. Brockman, C. McLeavey, I. Sutskever et al., “Robust speech recognition via large-scale weak supervision,” arXiv preprint arXiv:2212.04356, 2022.
- P. Dhariwal, H. Jun, C. Payne, J. W. Kim, A. Radford, I. Sutskever et al., “Jukebox: A generative model for music,” arXiv preprint arXiv:2005.00341, 2020.
- J. W. Rae, S. Borgeaud, T. Cai, K. Millican, J. Hoffmann, F. Song, J. Aslanides, S. Henderson, R. Ring, S. Young et al., “Scaling language models: Methods, analysis & insights from training gopher,” arXiv preprint arXiv:2112.11446, 2021.
- S. Borgeaud, A. Mensch, J. Hoffmann, T. Cai, E. Rutherford, K. Millican, G. B. Van Den Driessche, J.-B. Lespiau, B. Damoc, A. Clark et al., “Improving language models by retrieving from trillions of tokens,” in ICML, 2022.
- R. Anil, A. M. Dai, O. Firat, M. Johnson, D. Lepikhin, A. Passos, S. Shakeri, E. Taropa, P. Bailey, Z. Chen et al., “Palm 2: A state-of-the-art language model with improved multilingual, reasoning and coding capabilities,” arXiv preprint arXiv:2305.10403, 2023.
- K. He, X. Zhang, S. Ren, and J. Sun, “Deep residual learning for image recognition,” in CVPR, 2016.
- Z. Liu, Y. Lin, Y. Cao, H. Hu, Y. Wei, Z. Zhang, S. Lin, and B. Guo, “Swin transformer: Hierarchical vision transformer using shifted windows,” in ICCV, 2021.
- S. Min, X. Lyu, A. Holtzman, M. Artetxe, M. Lewis, H. Hajishirzi, and L. Zettlemoyer, “Rethinking the role of demonstrations: What makes in-context learning work?” arXiv preprint arXiv:2202.12837, 2022.
- J. D. M.-W. C. Kenton and L. K. Toutanova, “Bert: Pre-training of deep bidirectional transformers for language understanding,” in NACCL, 2019.
- A. Dosovitskiy, L. Beyer, A. Kolesnikov, D. Weissenborn, X. Zhai, T. Unterthiner, M. Dehghani, M. Minderer, G. Heigold, S. Gelly et al., “An image is worth 16x16 words: Transformers for image recognition at scale,” in ICLR, 2020.
- S. Zhang, Y. Liang, M. Gong, D. Jiang, and N. Duan, “Multi-view document representation learning for open-domain dense retrieval,” in ACL, 2022.
- T. Chen, S. Kornblith, M. Norouzi, and G. Hinton, “A simple framework for contrastive learning of visual representations,” in ICML, 2020.
- K. He, H. Fan, Y. Wu, S. Xie, and R. Girshick, “Momentum contrast for unsupervised visual representation learning,” in CVPR, 2020.
- W. Xia, T. Wang, Q. Gao, M. Yang, and X. Gao, “Graph embedding contrastive multi-modal representation learning for clustering,” IEEE TIP, vol. 32, pp. 1170–1183, 2023.
- R. Bommasani, D. A. Hudson, E. Adeli, R. Altman, S. Arora, S. von Arx, M. S. Bernstein, J. Bohg, A. Bosselut, E. Brunskill et al., “On the opportunities and risks of foundation models,” arXiv preprint arXiv:2108.07258, 2021.
- S. Arora, A. Narayan, M. F. Chen, L. J. Orr, N. Guha, K. Bhatia, I. Chami, F. Sala, and C. Ré, “Ask me anything: A simple strategy for prompting language models,” arXiv preprint arXiv:2210.02441, 2022.
- C. Zhou, Q. Li, C. Li, J. Yu, Y. Liu, G. Wang, K. Zhang, C. Ji, Q. Yan, L. He et al., “A comprehensive survey on pretrained foundation models: A history from bert to chatgpt,” arXiv preprint arXiv:2302.09419, 2023.
- M. Zhuge, D. Gao, D.-P. Fan, L. Jin, B. Chen, H. Zhou, M. Qiu, and L. Shao, “Kaleido-bert: Vision-language pre-training on fashion domain,” in CVPR, 2021.
- R. J. Chen, C. Chen, Y. Li, T. Y. Chen, A. D. Trister, R. G. Krishnan, and F. Mahmood, “Scaling vision transformers to gigapixel images via hierarchical self-supervised learning,” in CVPR, 2022.
- C. Luo, L. Jin, and J. Chen, “Siman: exploring self-supervised representation learning of scene text via similarity-aware normalization,” in CVPR, 2022.
- K. Sirotkin, P. Carballeira, and M. Escudero-Viñolo, “A study on the distribution of social biases in self-supervised learning visual models,” in CVPR, 2022.
- S. Paul, A. Norkin, and A. C. Bovik, “Self-supervised learning of perceptually optimized block motion estimates for video compression,” IEEE TIP, 2022.
- J. Xia, M. Zhuge, T. Geng, S. Fan, Y. Wei, Z. He, and F. Zheng, “Skating-mixer: multimodal mlp for scoring figure skating,” in AAAI, 2023.
- G.-P. Ji, M. Zhuge, D. Gao, D.-P. Fan, C. Sakaridis, and L. V. Gool, “Masked vision-language transformer in fashion,” Machine Intelligence Research, 2023.
- F. Wang, T. Kong, R. Zhang, H. Liu, and H. Li, “Self-supervised learning by estimating twin class distribution,” IEEE TIP, 2023.
- A. Chowdhery, S. Narang, J. Devlin, M. Bosma, G. Mishra, A. Roberts, P. Barham, H. W. Chung, C. Sutton, S. Gehrmann et al., “Palm: Scaling language modeling with pathways,” arXiv preprint arXiv:2204.02311, 2022.
- J.-B. Alayrac, J. Donahue, P. Luc, A. Miech, I. Barr, Y. Hasson, K. Lenc, A. Mensch, K. Millican, M. Reynolds et al., “Flamingo: a visual language model for few-shot learning,” in NeurIPS, 2022.
- A. Ramesh, M. Pavlov, G. Goh, S. Gray, C. Voss, A. Radford, M. Chen, and I. Sutskever, “Zero-shot text-to-image generation,” in ICML, 2021.
- R. Rombach, A. Blattmann, D. Lorenz, P. Esser, and B. Ommer, “High-resolution image synthesis with latent diffusion models,” in CVPR, 2022.
- M. Lewis, Y. Liu, N. Goyal, M. Ghazvininejad, A. Mohamed, O. Levy, V. Stoyanov, and L. Zettlemoyer, “Bart: Denoising sequence-to-sequence pre-training for natural language generation, translation, and comprehension,” in ACL, 2020.
- A. Wang, A. Singh, J. Michael, F. Hill, O. Levy, and S. R. Bowman, “Glue: A multi-task benchmark and analysis platform for natural language understanding,” arXiv preprint arXiv:1804.07461, 2018.
- Z. Yang, Z. Lin, P. Kang, J. Lv, Q. Li, and W. Liu, “Learning shared semantic space with correlation alignment for cross-modal event retrieval,” ACM TOMM, vol. 16, no. 1, pp. 1–22, 2020.
- G. Tahmasebzadeh, E. Kacupaj, E. Müller-Budack, S. Hakimov, J. Lehmann, and R. Ewerth, “Geowine: Geolocation based wiki, image, news and event retrieval,” in SIGIR, 2021.
- A. Lewkowycz, A. J. Andreassen, D. Dohan, E. Dyer, H. Michalewski, V. V. Ramasesh, A. Slone, C. Anil, I. Schlag, T. Gutman-Solo et al., “Solving quantitative reasoning problems with language models,” in NeurIPS, 2022.
- J. Degrave, F. Felici, J. Buchli, M. Neunert, B. Tracey, F. Carpanese, T. Ewalds, R. Hafner, A. Abdolmaleki, D. de Las Casas et al., “Magnetic control of tokamak plasmas through deep reinforcement learning,” Nature, vol. 602, no. 7897, pp. 414–419, 2022.
- L. Chen and S. Shang, “Region-based message exploration over spatio-temporal data streams,” in AAAI, 2019.
- B. Zhou, Q. Ning, D. Khashabi, and D. Roth, “Temporal common sense acquisition with minimal supervision,” in ACL, 2022.
- R. Han, X. Ren, and N. Peng, “Econet: Effective continual pretraining of language models for event temporal reasoning,” in EMNLP, 2021.
- Q. Ning, B. Zhou, H. Wu, H. Peng, C. Fan, and M. Gardner, “A meta-framework for spatiotemporal quantity extraction from text,” in ACL, 2022.
- J. Zhang and V. L. Patel, “Distributed cognition, representation, and affordance,” PragCog, vol. 14, no. 2, pp. 333–341, 2006.
- K. Srinivasan, K. Raman, J. Chen, M. Bendersky, and M. Najork, “Wit: Wikipedia-based image text dataset for multimodal multilingual machine learning,” in SIGIR, 2021.
- J. Deng, W. Dong, R. Socher, L.-J. Li, K. Li, and L. Fei-Fei, “Imagenet: A large-scale hierarchical image database,” in CVPR, 2009.
- P. Gatti, A. S. Penamakuri, R. Teotia, A. Mishra, S. Sengupta, and R. Ramnani, “Cofar: Commonsense and factual reasoning in image search,” in AACL-IJCNLP, 2022.
- A. Radford, J. Wu, R. Child, D. Luan, D. Amodei, I. Sutskever et al., “Language models are unsupervised multitask learners,” OpenAI blog, vol. 1, no. 8, p. 9, 2019.
- Y. Su, T. Lan, Y. Liu, F. Liu, D. Yogatama, Y. Wang, L. Kong, and N. Collier, “Language models can see: plugging visual controls in text generation,” arXiv preprint arXiv:2205.02655, 2022.
- Y. Su, T. Lan, Y. Wang, D. Yogatama, L. Kong, and N. Collier, “A contrastive framework for neural text generation,” in NeurIPS, 2022.
Paper Prompts
Sign up for free to create and run prompts on this paper using GPT-5.
Top Community Prompts
Collections
Sign up for free to add this paper to one or more collections.