Foundation Models Beyond Language: A Survey of Vision, Biology, Robotics, and Multimodal Systems
DOI:
https://doi.org/10.71143/yrb0wp14Abstract
Foundation models — large neural networks pretrained on broad data and adapted to many downstream tasks — first rose to prominence in natural language processing, but the same underlying recipe of large-scale self-supervised pretraining followed by task-specific adaptation has since been extended well beyond text. This paper surveys foundation models developed for vision, biological sequence and structure data, robotics, and multimodal settings that jointly reason over combinations of these modalities. We introduce a taxonomy organized around modality and pretraining objective, describe representative architectures and self-supervised learning strategies used outside the language domain, and examine applications spanning medical imaging, autonomous robotics, genomics, drug discovery, and multimodal virtual assistants. We summarize benchmarks used to evaluate these systems, present a worked case study of a vision-language-action model for robotic manipulation, and discuss challenges including modality-specific data scarcity, the difficulty of defining meaningful scaling laws outside text, and the added complexity of evaluating embodied and multimodal competence. We conclude by outlining directions connecting foundation models across modalities with embodied learning, scientific discovery, and unified multimodal reasoning.
Downloads
Downloads
Published
Issue
Section
License

This work is licensed under a Creative Commons Attribution-NonCommercial 4.0 International License.







