Skip to main navigation Skip to search Skip to main content

BLIP-3: A Family of Open Large Multimodal Models

  • Le Xue
  • , Manli Shu
  • , Anas Awadalla
  • , Jun Wang
  • , An Yan
  • , Senthil Purushwalkam
  • , Honglu Zhou
  • , Viraj Prabhu
  • , Yutong Dai
  • , Michael S. Ryoo
  • , Shrikant Kendre
  • , Jieyu Zhang
  • , Shaoyen Tseng
  • , Gustavo A. Lujan-Moreno
  • , Matthew L. Olson
  • , Musashi Hinck
  • , David Cobbley
  • , Vasudev Lal
  • , Can Qin
  • , Shu Zhang
  • Chia Chih Chen, Ning Yu, Juntao Tan, Tulika Manoj Awalgaonkar, Shelby Heinecke, Huan Wang, Yejin Choi, Ludwig Schmidt, Zeyuan Chen, Silvio Savarese, Juan Carlos Niebles, Caiming Xiong, Ran Xu
  • Salesforce Ai Research
  • University of Washington
  • Intel Labs

Research output: Chapter in Book/Report/Conference proceedingConference contributionpeer-review

5 Scopus citations

Abstract

This paper introduces BLIP-3, an open framework for developing Large Multimodal Models (LMMs). The framework comprises meticulously curated datasets, a training recipe, model architectures, and a resulting suite of LMMs. We release 4B and 14B models, including both the pre-trained base model and the instruction fine-tuned ones. Our models undergo rigorous evaluation across a range of tasks, including both single and multi-image benchmarks. Our models demonstrate competitive performance among open-source LMMs with similar model sizes. Our resulting LMMs demonstrate competitive performance among open-source LMMs with similar model sizes, with the ability to comprehend interleaved image-text inputs. Our training code, models, and all datasets used in this work, including the three large-scale datasets we create and the preprocessed ones, will be open-sourced to better support the research community.

Original languageEnglish
Title of host publicationProceedings - 2025 IEEE/CVF International Conference on Computer Vision Workshops, ICCV-W 2025
PublisherInstitute of Electrical and Electronics Engineers Inc.
Pages6183-6194
Number of pages12
ISBN (Electronic)9798331589882
DOIs
StatePublished - 2025
Event2025 IEEE/CVF International Conference on Computer Vision Workshops, ICCV-W 2025 - Honolulu, United States
Duration: Oct 19 2025Oct 20 2025

Publication series

NameProceedings - 2025 IEEE/CVF International Conference on Computer Vision Workshops, ICCV-W 2025

Conference

Conference2025 IEEE/CVF International Conference on Computer Vision Workshops, ICCV-W 2025
Country/TerritoryUnited States
CityHonolulu
Period10/19/2510/20/25

Keywords

  • Image Understanding
  • LMMs
  • MLLMs
  • Multimodal foundation models

Fingerprint

Dive into the research topics of 'BLIP-3: A Family of Open Large Multimodal Models'. Together they form a unique fingerprint.

Cite this