TY - GEN
T1 - BLIP-3
T2 - 2025 IEEE/CVF International Conference on Computer Vision Workshops, ICCV-W 2025
AU - Xue, Le
AU - Shu, Manli
AU - Awadalla, Anas
AU - Wang, Jun
AU - Yan, An
AU - Purushwalkam, Senthil
AU - Zhou, Honglu
AU - Prabhu, Viraj
AU - Dai, Yutong
AU - Ryoo, Michael S.
AU - Kendre, Shrikant
AU - Zhang, Jieyu
AU - Tseng, Shaoyen
AU - Lujan-Moreno, Gustavo A.
AU - Olson, Matthew L.
AU - Hinck, Musashi
AU - Cobbley, David
AU - Lal, Vasudev
AU - Qin, Can
AU - Zhang, Shu
AU - Chen, Chia Chih
AU - Yu, Ning
AU - Tan, Juntao
AU - Awalgaonkar, Tulika Manoj
AU - Heinecke, Shelby
AU - Wang, Huan
AU - Choi, Yejin
AU - Schmidt, Ludwig
AU - Chen, Zeyuan
AU - Savarese, Silvio
AU - Niebles, Juan Carlos
AU - Xiong, Caiming
AU - Xu, Ran
N1 - Publisher Copyright:
© 2025 IEEE.
PY - 2025
Y1 - 2025
N2 - This paper introduces BLIP-3, an open framework for developing Large Multimodal Models (LMMs). The framework comprises meticulously curated datasets, a training recipe, model architectures, and a resulting suite of LMMs. We release 4B and 14B models, including both the pre-trained base model and the instruction fine-tuned ones. Our models undergo rigorous evaluation across a range of tasks, including both single and multi-image benchmarks. Our models demonstrate competitive performance among open-source LMMs with similar model sizes. Our resulting LMMs demonstrate competitive performance among open-source LMMs with similar model sizes, with the ability to comprehend interleaved image-text inputs. Our training code, models, and all datasets used in this work, including the three large-scale datasets we create and the preprocessed ones, will be open-sourced to better support the research community.
AB - This paper introduces BLIP-3, an open framework for developing Large Multimodal Models (LMMs). The framework comprises meticulously curated datasets, a training recipe, model architectures, and a resulting suite of LMMs. We release 4B and 14B models, including both the pre-trained base model and the instruction fine-tuned ones. Our models undergo rigorous evaluation across a range of tasks, including both single and multi-image benchmarks. Our models demonstrate competitive performance among open-source LMMs with similar model sizes. Our resulting LMMs demonstrate competitive performance among open-source LMMs with similar model sizes, with the ability to comprehend interleaved image-text inputs. Our training code, models, and all datasets used in this work, including the three large-scale datasets we create and the preprocessed ones, will be open-sourced to better support the research community.
KW - Image Understanding
KW - LMMs
KW - MLLMs
KW - Multimodal foundation models
UR - https://www.scopus.com/pages/publications/105035184030
U2 - 10.1109/ICCVW69036.2025.00644
DO - 10.1109/ICCVW69036.2025.00644
M3 - Conference contribution
AN - SCOPUS:105035184030
T3 - Proceedings - 2025 IEEE/CVF International Conference on Computer Vision Workshops, ICCV-W 2025
SP - 6183
EP - 6194
BT - Proceedings - 2025 IEEE/CVF International Conference on Computer Vision Workshops, ICCV-W 2025
PB - Institute of Electrical and Electronics Engineers Inc.
Y2 - 19 October 2025 through 20 October 2025
ER -