Extent of LVLMs’ capability to meet diverse clinical demands
Determine the extent to which large vision-language models (including general-purpose models such as DeepSeek-VL, GPT-4V/GPT-4o, Claude3-Opus, Gemini, and Qwen-VL, as well as medical-specific models such as MedDr, LLaVA-Med, Med-Flamingo, RadFM, and Qilin-Med-VL-Chat) can accommodate the diverse demands encountered in real-world clinical scenarios across modalities, tasks, departments, and perceptual granularities.
References
However, it remains unclear to what extent these LVLMs can accommodate the diverse demands in real clinical scenarios.
A preregistered, multisite field study comparing unaided-first vs. immediate guidance will test whether this design scales across scanners, protocols, and training populations.
The obvious extension to this study is to evaluate the same protocol on the top-tier flagship models, which we leave to future work.
Our benchmarking of state-of-the-art 3D medical VLMs shows that existing models remain close to chance on fine-grained anatomical recognition and struggle to couple CT-derived morphology with PET-derived metabolic signals, identifying joint structural-metabolic reasoning as a critical open problem for the pattern recognition community.