Skip to main navigation Skip to search Skip to main content

Does Speech Enhancement of Publicly Available Data Help Build Robust Speech Recognition Systems?

  • Stony Brook University
  • Nokia

Research output: Chapter in Book/Report/Conference proceedingConference contributionpeer-review

Abstract

Automatic speech recognition(ASR) systems play a key role in many commercial products including voice assistants. Typically, they require large amounts of high quality speech data for training which gives an undue advantage to large organizations which have tons of private data. We investigated if speech data obtained from publicly available sources can be further enhanced to train better speech recognition models. We begin with noisy/contaminated speech data, apply speech enhancement to produce 'cleaned' version and use both the versions to train the ASR model. We have found that using speech enhancement gives 9.5% better word error rate than training on just the original noisy data and 9% better than training on just the ground truth 'clean' data. It's performance is also comparable to the ideal case scenario when trained on noisy and it's ground truth 'clean' version.

Original languageEnglish
Title of host publicationAAAI 2020 - 34th AAAI Conference on Artificial Intelligence
PublisherAAAI Press
Pages13793-13794
Number of pages2
ISBN (Electronic)9781577358350
StatePublished - 2020
Event34th AAAI Conference on Artificial Intelligence, AAAI 2020 - New York, United States
Duration: Feb 7 2020Feb 12 2020

Publication series

NameAAAI 2020 - 34th AAAI Conference on Artificial Intelligence

Conference

Conference34th AAAI Conference on Artificial Intelligence, AAAI 2020
Country/TerritoryUnited States
CityNew York
Period02/7/2002/12/20

Fingerprint

Dive into the research topics of 'Does Speech Enhancement of Publicly Available Data Help Build Robust Speech Recognition Systems?'. Together they form a unique fingerprint.

Cite this