Papers

Topics

Authors

Recent

View all

Detailed Answer

Quick Answer

Concise responses based on abstracts only

Detailed Answer

Well-researched responses based on abstracts and relevant paper content.

Custom Instructions Pro

Preferences or requirements that you'd like Emergent Mind to consider when generating responses

Gemini 2.5 Flash

Gemini 2.5 Flash 86 tok/s

Gemini 2.5 Pro 45 tok/s Pro

GPT-5 Medium 23 tok/s Pro

GPT-5 High 25 tok/s Pro

GPT-4o 111 tok/s Pro

Kimi K2 178 tok/s Pro

GPT OSS 120B 452 tok/s Pro

Claude Sonnet 4 37 tok/s Pro

2000 character limit reached

EXGRA-MED: Extended Context Graph Alignment for Medical Vision- Language Models (2410.02615v3)

Published 3 Oct 2024 in cs.LG

Abstract: State-of-the-art medical multi-modal LLMs (med-MLLMs), such as LLAVA-MED and BIOMEDGPT, primarily depend on scaling model size and data volume, with training driven largely by autoregressive objectives. However, we reveal that this approach can lead to weak vision-language alignment, making these models overly dependent on costly instruction-following data. To address this, we introduce EXGRA-MED, a novel multi-graph alignment framework that jointly aligns images, instruction responses, and extended captions in the latent space, advancing semantic grounding and cross-modal coherence. To scale to large LLMs (e.g., LLaMa-7B), we develop an efficient end-to-end training scheme using black-box gradient estimation, enabling fast and scalable optimization. Empirically, EXGRA-MED matches LLAVA-MED's performance using just 10% of pre-training data, achieving a 20.13% gain on VQA-RAD and approaching full-data performance. It also outperforms strong baselines like BIOMEDGPT and RADFM on visual chatbot and zero-shot classification tasks, demonstrating its promise for efficient, high-quality vision-language integration in medical AI.

Citations (1)

View on Semantic Scholar

Collections

Summary

We haven't generated a summary for this paper yet.

Summarize Now

Paper Prompts

Explore 10 Community Prompts

Follow-up Questions

We haven't generated follow-up questions for this paper yet.

Generate Now

EXGRA-MED: Extended Context Graph Alignment for Medical Vision- Language Models (2410.02615v3)

Collections

Summary

Paper Prompts

Follow-up Questions

Authors (13)

Don't miss out on important new AI/ML research

EXGRA-MED: Extended Context Graph Alignment for Medical Vision- Language Models (2410.02615v3)

Collections

Summary

Paper Prompts

Follow-up Questions

Related Papers

Authors (13)

Don't miss out on important new AI/ML research