Skip to main navigation Skip to search Skip to main content

4D LangSplat: 4D Language Gaussian Splatting via Multimodal Large Language Models

  • Wanhua Li
  • , Renping Zhou
  • , Jiawei Zhou
  • , Yingwei Song
  • , Johannes Herter
  • , Minghan Qin
  • , Gao Huang
  • , Hanspeter Pfister
  • Harvard University
  • Tsinghua University
  • Brown University
  • Swiss Federal Institute of Technology Zurich

Research output: Contribution to journalConference articlepeer-review

12 Scopus citations

Abstract

Learning 4D language fields to enable time-sensitive, open-ended language queries in dynamic scenes is essential for many real-world applications. While LangSplat successfully grounds CLIP features into 3D Gaussian representations, achieving precision and efficiency in 3D static scenes, it lacks the ability to handle dynamic 4D fields as CLIP, designed for static image-text tasks, cannot capture temporal dynamics in videos. Real-world environments are inherently dynamic, with object semantics evolving over time. Building a precise 4D language field necessitates obtaining pixel-aligned, object-wise video features, which current vision models struggle to achieve. To address these challenges, we propose 4D LangSplat, which learns 4D language fields to handle time-agnostic or time-sensitive open-vocabulary queries in dynamic scenes efficiently. 4D LangSplat bypasses learning the language field from vision features and instead learns directly from text generated from object-wise video captions via Multimodal Large Language Models (MLLMs). Specifically, we propose a multimodal object-wise video prompting method, consisting of visual and text prompts that guide MLLMs to generate detailed, temporally consistent, high-quality captions for objects throughout a video. These captions are encoded using a Large Language Model into high-quality sentence embeddings, which then serve as pixel-aligned, object-specific feature supervision, facilitating open-vocabulary text queries through shared embedding spaces. Recognizing that objects in 4D scenes exhibit smooth transitions across states, we further propose a status deformable network to model these continuous changes over time effectively. Our results across multiple benchmarks demonstrate that 4D LangSplat attains precise and efficient results for both timesensitive and time-agnostic open-vocabulary queries.

Original languageEnglish
Pages (from-to)22001-22011
Number of pages11
JournalProceedings of the IEEE Computer Society Conference on Computer Vision and Pattern Recognition
DOIs
StatePublished - 2025
Event2025 IEEE/CVF Conference on Computer Vision and Pattern Recognition, CVPR 2025 - Nashville, United States
Duration: Jun 11 2025Jun 15 2025

Fingerprint

Dive into the research topics of '4D LangSplat: 4D Language Gaussian Splatting via Multimodal Large Language Models'. Together they form a unique fingerprint.

Cite this