Mixed Bangla-English Spoken Digit Classification Using Convolutional Neural Network
Loading...
Date
Journal Title
Journal ISSN
Volume Title
Publisher
Abstract
In this era of the scientific revolution, speech recognition is an important field. People of the world are connecting by using technology. People are shifting from one country to another, sharing their culture and language. Speech recognition has made it easy by translating most of the languages into a readable format. Our world is moving forward through the era of the digital revolution. Still, there are rudimentary examples of research works on Bangla speech recognition with the advancement of automatic speech recognition (ASR). From a Bangladeshi perspective, we often feel the need of using mixed Bangla-English language in different use-cases, mostly in educational institutions and hospital environments. However, most research works focus on speech recognition in the English language, so we were motivated to develop a mixed Bangla-English language classifier to transcribe isolated mixed Bangla-English spoken digits. We have used an open-source dataset for English, and for Bangla, we created a dataset in a noisy environment by speakers of different ages, gender, and dialects. Finally, for the mixed dataset, we have used Mel Frequency Cepstral Coefficient (MFCC) for feature extraction and Convolutional Neural Network (CNN) classifier to train, test, and analyze data for two different experiments we found promising results.