TY - GEN
T1 - MuseDance
T2 - 2026 IEEE/CVF Winter Conference on Applications of Computer Vision, WACV 2026
AU - Dong, Zhikang
AU - Hao, Weituo
AU - Wang, Ju Chiang
AU - Zhang, Peng
AU - Polak, Pawel
N1 - Publisher Copyright:
© 2026 IEEE.
PY - 2026
Y1 - 2026
N2 - Image animation is a rapidly developing area in multimodal research, with a focus on generating videos from reference images. While much of the work has emphasized generic video generation guided by text, music-driven dance image animation remains underexplored. In this paper, we introduce MuseDance, an end-to-end model that animates reference images using both music and text inputs. By integrating music as a conditioning modality, MuseDance generates personalized videos that not only adhere to textual descriptions but also synchronize character movements with the rhythm and dynamics of the music. Unlike existing methods, MuseDance eliminates the need for explicit motion guidance, such as pose sequences or depth maps, reducing the complexity of video generation while enhancing accessibility and flexibility. To support further research in this field, we present a new multimodal dataset comprising of 3,122 dance videos, each paired with the corresponding background music and text descriptions. Our approach leverages diffusion-based methods to achieve robust generalization, precise control, and temporal consistency, setting a new benchmark for the task of music-driven image animation. The dataset of this work is available at https://github.com/Dongzhikang/musedance.
AB - Image animation is a rapidly developing area in multimodal research, with a focus on generating videos from reference images. While much of the work has emphasized generic video generation guided by text, music-driven dance image animation remains underexplored. In this paper, we introduce MuseDance, an end-to-end model that animates reference images using both music and text inputs. By integrating music as a conditioning modality, MuseDance generates personalized videos that not only adhere to textual descriptions but also synchronize character movements with the rhythm and dynamics of the music. Unlike existing methods, MuseDance eliminates the need for explicit motion guidance, such as pose sequences or depth maps, reducing the complexity of video generation while enhancing accessibility and flexibility. To support further research in this field, we present a new multimodal dataset comprising of 3,122 dance videos, each paired with the corresponding background music and text descriptions. Our approach leverages diffusion-based methods to achieve robust generalization, precise control, and temporal consistency, setting a new benchmark for the task of music-driven image animation. The dataset of this work is available at https://github.com/Dongzhikang/musedance.
KW - audio-visual learning
KW - diffusion models
KW - generative ai
KW - multimodal learning
UR - https://www.scopus.com/pages/publications/105041274857
U2 - 10.1109/WACV61042.2026.00372
DO - 10.1109/WACV61042.2026.00372
M3 - Conference contribution
AN - SCOPUS:105041274857
T3 - Proceedings - 2026 IEEE/CVF Winter Conference on Applications of Computer Vision, WACV 2026
SP - 3813
EP - 3824
BT - Proceedings - 2026 IEEE/CVF Winter Conference on Applications of Computer Vision, WACV 2026
PB - Institute of Electrical and Electronics Engineers Inc.
Y2 - 6 March 2026 through 10 March 2026
ER -