Foundation Models Beyond Language: A Survey of Vision, Biology, Robotics, and Multimodal Systems

Authors

  • Amit
  • Anshu
  • Riya
  • Shruti
  • Taruna

DOI:

https://doi.org/10.71143/yrb0wp14

Abstract

Foundation models — large neural networks pretrained on broad data and adapted to many downstream tasks — first rose to prominence in natural language processing, but the same underlying recipe of large-scale self-supervised pretraining followed by task-specific adaptation has since been extended well beyond text. This paper surveys foundation models developed for vision, biological sequence and structure data, robotics, and multimodal settings that jointly reason over combinations of these modalities. We introduce a taxonomy organized around modality and pretraining objective, describe representative architectures and self-supervised learning strategies used outside the language domain, and examine applications spanning medical imaging, autonomous robotics, genomics, drug discovery, and multimodal virtual assistants. We summarize benchmarks used to evaluate these systems, present a worked case study of a vision-language-action model for robotic manipulation, and discuss challenges including modality-specific data scarcity, the difficulty of defining meaningful scaling laws outside text, and the added complexity of evaluating embodied and multimodal competence. We conclude by outlining directions connecting foundation models across modalities with embodied learning, scientific discovery, and unified multimodal reasoning.

Downloads

Download data is not yet available.

Downloads

Published

26-12-2024

How to Cite

Amit, Anshu, Riya, Shruti, & Taruna. (2024). Foundation Models Beyond Language: A Survey of Vision, Biology, Robotics, and Multimodal Systems. International Journal of Research and Review in Applied Science, Humanities, and Technology, 1(2), 172-183. https://doi.org/10.71143/yrb0wp14