Skip to main navigation Skip to search Skip to main content

MuseDance: A Diffusion-based Music-Driven Image Animation System

  • Zhikang Dong
  • , Weituo Hao
  • , Ju Chiang Wang
  • , Peng Zhang
  • , Pawel Polak
  • Stony Brook University
  • ByteDance Ltd.
  • Apple

Research output: Chapter in Book/Report/Conference proceedingConference contributionpeer-review

Abstract

Image animation is a rapidly developing area in multimodal research, with a focus on generating videos from reference images. While much of the work has emphasized generic video generation guided by text, music-driven dance image animation remains underexplored. In this paper, we introduce MuseDance, an end-to-end model that animates reference images using both music and text inputs. By integrating music as a conditioning modality, MuseDance generates personalized videos that not only adhere to textual descriptions but also synchronize character movements with the rhythm and dynamics of the music. Unlike existing methods, MuseDance eliminates the need for explicit motion guidance, such as pose sequences or depth maps, reducing the complexity of video generation while enhancing accessibility and flexibility. To support further research in this field, we present a new multimodal dataset comprising of 3,122 dance videos, each paired with the corresponding background music and text descriptions. Our approach leverages diffusion-based methods to achieve robust generalization, precise control, and temporal consistency, setting a new benchmark for the task of music-driven image animation. The dataset of this work is available at https://github.com/Dongzhikang/musedance.

Original languageEnglish
Title of host publicationProceedings - 2026 IEEE/CVF Winter Conference on Applications of Computer Vision, WACV 2026
PublisherInstitute of Electrical and Electronics Engineers Inc.
Pages3813-3824
Number of pages12
ISBN (Electronic)9798331555115
DOIs
StatePublished - 2026
Event2026 IEEE/CVF Winter Conference on Applications of Computer Vision, WACV 2026 - Tucson, United States
Duration: Mar 6 2026Mar 10 2026

Publication series

NameProceedings - 2026 IEEE/CVF Winter Conference on Applications of Computer Vision, WACV 2026

Conference

Conference2026 IEEE/CVF Winter Conference on Applications of Computer Vision, WACV 2026
Country/TerritoryUnited States
CityTucson
Period03/6/2603/10/26

Keywords

  • audio-visual learning
  • diffusion models
  • generative ai
  • multimodal learning

Fingerprint

Dive into the research topics of 'MuseDance: A Diffusion-based Music-Driven Image Animation System'. Together they form a unique fingerprint.

Cite this