What is it about?

n this review, we examine recent pretrained multimodal deep learning models developed between 2020 and 2025, including CLIP, GPT-4V, GPT-4o, SigLIP, ViT, BLIP, Flamingo, PaLI, LLaVA, Florence-2, PaliGemma-2, Gemma-3, and Llama-4. The review analyzes 120 studies across 11 application domains, covering model architectures, modalities, datasets, applications, evaluation metrics, limitations, and future research directions. We also discuss key challenges facing multimodal AI, including data quality, computational cost, cross-modal alignment, scalability, robustness, and trustworthy AI. This work represents an important step in my research journey toward developing trustworthy and multimodal AI systems, particularly for applications in healthcare.

Featured Image

Why is it important?

I hope this review will be useful to researchers and students working in Multimodal AI, Deep Learning, Foundation Models, Computer Vision, NLP, and Healthcare AI.

Read the Original

This page is a summary of: A comprehensive review on recent pretrained multimodal deep learning models from architectures to future directions, Discover Informatics, October 2026, Springer Science + Business Media,
DOI: 10.1007/s44564-026-00022-1.
You can read the full text:

Read

Contributors

The following have contributed to this page