TY - GEN
T1 - Energy-based recurrent model for stochastic modeling of music
AU - Liu, Yingru
AU - Xie, Dongliang
AU - Wang, Xin
N1 - Publisher Copyright:
© 2019 IEEE.
PY - 2019/7
Y1 - 2019/7
N2 - The aim of this work is to more accurately model the stochastic process of music-related data, which is essential for many AI applications in musicology. When music is naturally represented as a sequence of vectorized frames, existing models generally cannot well capture the correlation of the elements inside each frame. We propose an energy-based model called Chain Graphical Recurrent Neural Network (CGRNN) to explore the correlation of elements for more accurate modeling of the dynamics of music. In CGRNN, a probabilistic substructure named Conditional spike and slab Restricted Boltzmann Machine (C-ssRBM) is defined to better model the conditional covariance and joint distribution of elements in a frame. Besides, CGRNN is capable of tracking the evolution of music and extracting sparse features with an efficient design of temporal transition. With the estimated stochastic process of music, we further implement CGRNN to generate melodious music automatically. Extensive empirical evaluations of multiple unsupervised learning tasks are conducted on symbolic MIDI and audio sounds to demonstrate the performance of our model.
AB - The aim of this work is to more accurately model the stochastic process of music-related data, which is essential for many AI applications in musicology. When music is naturally represented as a sequence of vectorized frames, existing models generally cannot well capture the correlation of the elements inside each frame. We propose an energy-based model called Chain Graphical Recurrent Neural Network (CGRNN) to explore the correlation of elements for more accurate modeling of the dynamics of music. In CGRNN, a probabilistic substructure named Conditional spike and slab Restricted Boltzmann Machine (C-ssRBM) is defined to better model the conditional covariance and joint distribution of elements in a frame. Besides, CGRNN is capable of tracking the evolution of music and extracting sparse features with an efficient design of temporal transition. With the estimated stochastic process of music, we further implement CGRNN to generate melodious music automatically. Extensive empirical evaluations of multiple unsupervised learning tasks are conducted on symbolic MIDI and audio sounds to demonstrate the performance of our model.
KW - Automated music generation
KW - Stochastic deep learning model
KW - Unsupervised learning
UR - https://www.scopus.com/pages/publications/85071036634
U2 - 10.1109/ICME.2019.00049
DO - 10.1109/ICME.2019.00049
M3 - Conference contribution
AN - SCOPUS:85071036634
T3 - Proceedings - IEEE International Conference on Multimedia and Expo
SP - 236
EP - 241
BT - Proceedings - 2019 IEEE International Conference on Multimedia and Expo, ICME 2019
PB - IEEE Computer Society
T2 - 2019 IEEE International Conference on Multimedia and Expo, ICME 2019
Y2 - 8 July 2019 through 12 July 2019
ER -