DualFocus: Integrating Macro and Micro Perspectives in Multi-modal Large Language Models
Abstract: We present DualFocus, a novel framework for integrating macro and micro perspectives within multi-modal LLMs (MLLMs) to enhance vision-language task performance. Current MLLMs typically singularly focus on inputs at a predefined resolution, resulting in deficiencies in detailed questions involving local regions. We introduced a DualFocus mechanism where the model concentrates on the image from a macro perspective, responses to the question, and identifies suitable sub-regions to zoom in for subsequent micro perspective analysis. Via the integration of answers from both macro and micro perspectives, the model is adept at addressing tasks that encompass global, detailed, and combined considerations. To endows the DualFocus mechanism in MLLMs, we curated a tailored dataset derived from the Visual Genome (VG) and adapted it to align with the training regimen of DualFocus. Through comparative studies across different model sizes and benchmarks, we demonstrate DualFocus's superiority in balancing detailed examination with holistic insight, significantly reducing hallucination instances in MLLMs and improving their performance in various vision-language tasks.
- Qwen-vl: A frontier large vision-language model with versatile abilities. arXiv.org, 2023.
- Baichuan. Baichuan 2: Open large-scale language models. arXiv.org, 2023. URL https://arxiv.org/abs/2309.10305.
- Language models are few-shot learners. Advances in Neural Information Processing Systems (NeurIPS), 33:1877–1901, 2020.
- Shikra: Unleashing multimodal llm’s referential dialogue magic. arXiv.org, 2023a.
- Sharegpt4v: Improving large multi-modal models with better captions. arXiv.org, 2023b.
- Vicuna: An open-source chatbot impressing gpt-4 with 90%* chatgpt quality, March 2023. URL https://lmsys.org/blog/2023-03-30-vicuna/.
- Palm: Scaling language modeling with pathways. arXiv.org, 2022.
- Instructblip: Towards general-purpose vision-language models with instruction tuning, 2023.
- Bert: Pre-training of deep bidirectional transformers for language understanding. arXiv.org, 2018.
- Internlm-xcomposer2: Mastering free-form text-image composition and comprehension in vision-language large model. arXiv.org, 2024.
- Planting a seed of vision in large language model.
- Wanjuan: A comprehensive multimodal dataset for advancing english and chinese large models. arXiv.org, abs/2308.10755, 2023. URL https://api.semanticscholar.org/CorpusID:261049100.
- Cogagent: A visual language model for gui agents. arXiv.org, 2023.
- LoRA: Low-rank adaptation of large language models. In International Conference on Learning Representations, 2022. URL https://openreview.net/forum?id=nZeVKeeFYf9.
- Bliva: A simple multimodal llm for better handling of text-rich visual questions. ArXiv, abs/2308.09936, 2023. URL https://api.semanticscholar.org/CorpusID:261049015.
- Opera: Alleviating hallucination in multi-modal large language models via over-trust penalty and retrospection-allocation. arXiv.org, 2023.
- Gqa: A new dataset for real-world visual reasoning and compositional question answering. Conference on Computer Vision and Pattern Recognition (CVPR), 2019.
- Jelinek, F. Statistical methods for speech recognition. MIT press, 1998.
- Visual genome: Connecting language and vision using crowdsourced dense image annotations. IJCV, 2017.
- Otterhd: A high-resolution multi-modality model. Arxiv, 2023a.
- Otter: A multi-modal model with in-context instruction tuning. arXiv.org, 2023b.
- Blip: Bootstrapping language-image pre-training for unified vision-language understanding and generation. In Proceedings of the International Conference on Machine learning (ICML), pp. 12888–12900. PMLR, 2022.
- Evaluating object hallucination in large vision-language models. arXiv.org, 2023c.
- Monkey: Image resolution and text label are important things for large multi-modal models. Arxiv, 2023d.
- Improved baselines with visual instruction tuning. arXiv preprint arXiv:2310.03744, 2023a.
- Visual instruction tuning. arXiv.org, 2023b.
- Mmbench: Is your multi-modal model an all-around player? arXiv:2307.06281, 2023c.
- OpenAI. Chatgpt. https://openai.com/blog/chatgpt, 2022.
- OpenAI. Gpt-4 technical report, 2023.
- Training language models to follow instructions with human feedback. Advances in Neural Information Processing Systems (NeurIPS), 35:27730–27744, 2022.
- Kosmos-2: Grounding multimodal large language models to the world. arXiv.org, 2023.
- Gpt4point: A unified framework for point-language understanding and generation, 2023a.
- Gemini vs gpt-4v: A preliminary comparison and combination of vision-language models through qualitative cases, 2023b.
- Qwen. Introducing qwen-7b: Open foundation and human-aligned models (of the state-of-the-arts), 2023.
- Learning transferable visual models from natural language supervision. In Proceedings of the International Conference on Machine learning (ICML), pp. 8748–8763. PMLR, 2021.
- Exploring the limits of transfer learning with a unified text-to-text transformer. Journal of Machine Learning Research (JMLR), 21(1):5485–5551, 2020.
- Towards vqa models that can read. In Proceedings of the IEEE/CVF conference on computer vision and pattern recognition, pp. 8317–8326, 2019.
- Alpha-CLIP: A clip model focusing on wherever you want. arXiv.org, 2023.
- Team, I. Internlm: A multilingual language model with progressively enhanced capabilities. https://github.com/InternLM/InternLM, 2023.
- Llama: Open and efficient foundation language models. arXiv.org, 2023a.
- Llama 2: Open foundation and fine-tuned chat models, 2023b.
- V3det: Vast vocabulary visual detection dataset. In Proceedings of the IEEE/CVF International Conference on Computer Vision (ICCV), October 2023.
- Chain-of-thought prompting elicits reasoning in large language models. Advances in Neural Information Processing Systems (NIPS), 35:24824–24837, 2022.
- mplug-owl: Modularization empowers large language models with multimodality. arXiv.org, 2023.
- Internlm-xcomposer: A vision-language large model for advanced text-image comprehension and composition. arXiv.org, 2023.
- Mllm-dataengine: An iterative refinement approach for mllm. arXiv.org, 2023.
- Minigpt-4: Enhancing vision-language understanding with advanced large language models. arXiv.org, 2023.
Paper Prompts
Sign up for free to create and run prompts on this paper using GPT-5.
Top Community Prompts
Collections
Sign up for free to add this paper to one or more collections.