Multimodal Foundation Models: Architectures, Challenges, and Applications
Keywords:
Abstract
The fast progress of AI has significantly changed the research paradigm from single-modality systems that operate on independent modalities to multimodal systems that can handle multimodal information from text, image, sound, video, and sensor data. This transformation is primarily driven by Multimodal Foundation Models (MFMs) that leverage the large-scale pre-training on vast collections of data to achieve cross-modal understanding and generation skills. The current paper provides an overview of multimodal deep learning research in terms of its history ranging from early fusion approaches to recent Transformer-based multimodal networks. The existing models are classified based on their alignment strategy, namely contrastive, generative, and hybrid strategies along with mathematical formulations of some major objective functions like InfoNCE loss and cross-attention. Moreover, the paper surveys the state-of-the-art in vision-language tasks, embodied robotics, healthcare, and scientific discovery applications, evaluating the challenges of multimodal AI such as hallucinations, data biases, computation issues, and evaluation metrics like FID, CIDEr, and POPE. In conclusion, the current trends in multimodal AI are reviewed in terms of neuro-symbolic fusion, embodied AI, and edge computing.
Downloads
Published
How to Cite
Issue
Section
License

This work is licensed under a Creative Commons Attribution-NoDerivatives 4.0 International License.