MRI
MRI India Journals Vol. 15 No. 2 (2026)

Multimodal Foundation Models: Architectures, Challenges, and Applications

Authors

  • Hayder Kareem Algabri Department of Cybersecurity Techniques Engineering , College of Engineering Techniques,University of Hilla.
  • Hayder Al-Ghanimi Department of Artificial Intelligence Techniques Engineering, Hilla University College, Babylon 51001, Iraq
  • K G Kharade Department of Computer Science, Shivaji University Kolhapur, Maharashtra, India

Keywords:

Multimodal Deep Learning Transformers Foundation Models Cross-Modal Representation Computer Vision Natural Language Processing Large Multimodal Models

Abstract

The fast progress of AI has significantly changed the research paradigm from single-modality systems that operate on independent modalities to multimodal systems that can handle multimodal information from text, image, sound, video, and sensor data. This transformation is primarily driven by Multimodal Foundation Models (MFMs) that leverage the large-scale pre-training on vast collections of data to achieve cross-modal understanding and generation skills. The current paper provides an overview of multimodal deep learning research in terms of its history ranging from early fusion approaches to recent Transformer-based multimodal networks. The existing models are classified based on their alignment strategy, namely contrastive, generative, and hybrid strategies along with mathematical formulations of some major objective functions like InfoNCE loss and cross-attention. Moreover, the paper surveys the state-of-the-art in vision-language tasks, embodied robotics, healthcare, and scientific discovery applications, evaluating the challenges of multimodal AI such as hallucinations, data biases, computation issues, and evaluation metrics like FID, CIDEr, and POPE. In conclusion, the current trends in multimodal AI are reviewed in terms of neuro-symbolic fusion, embodied AI, and edge computing.

Downloads

Published

2026-08-18

How to Cite

Algabri, H. K., Al-Ghanimi, H., & Kharade, K. G. (2026). Multimodal Foundation Models: Architectures, Challenges, and Applications. International Journal on Advanced Computer Theory and Engineering, 15(2), 142–156. Retrieved from https://journals.mriindia.com/index.php/ijacte/article/view/4012

Issue

Section

Articles

Similar Articles

<< < 24 25 26 27 28 29 30 31 32 33 > >> 

You may also start an advanced similarity search for this article.