TY - GEN
T1 - Exploiting Deep Learning for Sentence-Level Lipreading
AU - Wu, Isabella
AU - Wang, Xin
N1 - Publisher Copyright:
© 2023 IEEE.
PY - 2023
Y1 - 2023
N2 - Lipreading, also called visual speech recognition, is an excellent technique to understand what a speaker says without audios. Based on deep learning, many studies have gained outstanding achievements in English lipreading. English is dominated by polysyllables, with the proportion of homonyms as low as 1%. Compared to English, Mandarin (the official language of Chinese) is dominated by monosyllables, with the ratio of homonyms as high as 72%. The high ratio of homonyms makes the lipreading in Mandarin much more challenging. However, little attention has been paid to Mandarin lipreading within the deep learning framework, especially at the sentence level. In this paper, we first introduce a dataset, named Mandarin-Lipreading, which is recorded in the controlled lab environment and is the largest dataset so far for sentence-level lipreading in Mandarin. We further investigate the modeling units and propose an end-to-end Mandarin lipreading system. The experimental results show that the proposed system achieves 14.77% CER on Mandarin-Lipreading and 8.7% CER for unseen speakers evaluation on the GRID corpus.
AB - Lipreading, also called visual speech recognition, is an excellent technique to understand what a speaker says without audios. Based on deep learning, many studies have gained outstanding achievements in English lipreading. English is dominated by polysyllables, with the proportion of homonyms as low as 1%. Compared to English, Mandarin (the official language of Chinese) is dominated by monosyllables, with the ratio of homonyms as high as 72%. The high ratio of homonyms makes the lipreading in Mandarin much more challenging. However, little attention has been paid to Mandarin lipreading within the deep learning framework, especially at the sentence level. In this paper, we first introduce a dataset, named Mandarin-Lipreading, which is recorded in the controlled lab environment and is the largest dataset so far for sentence-level lipreading in Mandarin. We further investigate the modeling units and propose an end-to-end Mandarin lipreading system. The experimental results show that the proposed system achieves 14.77% CER on Mandarin-Lipreading and 8.7% CER for unseen speakers evaluation on the GRID corpus.
KW - lipreading
KW - Mandarin
UR - https://www.scopus.com/pages/publications/85169589541
U2 - 10.1109/IJCNN54540.2023.10191888
DO - 10.1109/IJCNN54540.2023.10191888
M3 - Conference contribution
AN - SCOPUS:85169589541
T3 - Proceedings of the International Joint Conference on Neural Networks
BT - IJCNN 2023 - International Joint Conference on Neural Networks, Proceedings
PB - Institute of Electrical and Electronics Engineers Inc.
T2 - 2023 International Joint Conference on Neural Networks, IJCNN 2023
Y2 - 18 June 2023 through 23 June 2023
ER -