Multimodal Large Language Models in Healthcare: Architectures, Datasets, Applications, Evaluation Challenges, and Future Directions
Keywords:
Multimodal large language models, Medical artificial intelligence, Healthcare diagnostics, Clinical decision support, Clinical validationAbstract
Healthcare data are inherently multimodal, combining clinical notes, laboratory results, medical images, physiological signals, genomic information, audio, video, and patient-generated data. Traditional unimodal artificial intelligence systems and text-only large language models are useful for selected tasks, but they are limited when clinical reasoning requires simultaneous interpretation of several data sources. Multimodal large language models (MLLMs) extend language-model reasoning by integrating non-textual modalities through modality-specific encoders, alignment modules, fusion strategies, and generative decoders. This structured narrative review synthesizes the architecture, datasets, clinical applications, evaluation methods, limitations, and future directions of MLLMs in healthcare, with emphasis on practical readiness for clinical decision support rather than technical performance alone. Literature was searched across Scopus, Web of Science, PubMed, IEEE Xplore, ACM Digital Library, Google Scholar, arXiv, and medRxiv using terms related to multimodal large language models, medical vision-language models, healthcare foundation models, radiology report generation, medical visual question answering, speech-text reasoning, biosignal-text learning, and clinical AI evaluation. Studies were narratively synthesized by model architecture, modality, dataset type, clinical domain, evaluation metric, and deployment challenge. Current evidence suggests that healthcare MLLMs are most mature in radiology and image-text reasoning, with growing applications in pathology, cardiology, oncology, mental health, emergency medicine, surgery, dermatology, ophthalmology, rehabilitation, and rare disease support. Medical-domain adaptation through biomedical image-text pretraining, instruction tuning, clinical terminology grounding, and human evaluation improves task performance compared with general-domain models in several benchmarks. However, performance remains uneven across modalities and clinical settings. Key barriers include scarce paired multimodal datasets, hallucination, weak external validation, demographic and device bias, privacy constraints, high computational cost, limited interpretability, and the absence of standardized clinically meaningful evaluation protocols. MLLMs represent a promising direction for context-aware healthcare AI, but they are not yet ready for autonomous clinical deployment. Future research should prioritize clinically grounded benchmarks, prospective validation, transparent reporting, privacy-preserving learning, uncertainty estimation, bias auditing, and human-in-the-loop deployment pathways.
Downloads
Published
How to Cite
Issue
Section
License
Copyright (c) 2026 Salman Ashraf, Rafia TEHSEEN, Hira ASHRAF

All papers published in Applied Medical Informatics are licensed under a Creative Commons Attribution (CC BY 4.0) International License.