Abstract
The relevance of the research is caused by the need of application of speech recognition technology for language teaching. The speech recognition is one of the most important tasks of the signal processing and pattern recognition fields. The speech recognition technology allows computers to understand human speech and it plays very important role in people’s lives. This technology can be used to help people in a variety way such as controlling smart homes and devices; using robots to perform job interviews; converting audio into text, etc. But there are not many applications of speech recognition technology in education, especially in English teaching. The main aim of the research is to propose an algorithm in which speech recognition technology is used English language teaching. Objects of researches are speech recognition technologies and frameworks, English spoken sounds system. Research results: The authors have proposed an algorithm based on speech recognition framework for English pronunciation learning. This proposed algorithm can be applied to another speech recognition framework and different languages. Besides the authors also demonstrated how to use the proposed algorithm for development English pronunciation practicing system based on iOS mobile app platform. The system also allows language learners can practice English pronunciation anywhere and anytime without any purchase.
Access this chapter
Tax calculation will be finalised at checkout
Purchases are for personal use only
References
Juang, B.H., Rabiner, L.R.: Automatic speech recognition–a brief history of the technology development (2015). https://web.ece.ucsb.edu/Faculty/Rabiner/ece259/Reprints/354_LALI-ASRHistory-final-10-8.pdf
Benesty, J., Sondhi, M.M., Huang, Y.: Springer Handbook of Speech Processing. Springer, Heidelberg (2008). https://doi.org/10.1007/978-3-540-49127-9
Jelinek, F.: Pioneering Speech Recognition (2015). https://www.ibm.com/ibm/history/ibm100/us/en/icons/speechreco/
Huang, X., Baker, J., Reddy, R.: A Historical perspective of speech recognition. Commun. ACM 57(1), 94–103 (2014)
Hanazawa, T., Hinton, G., Shikano, K., Lang, K.J.: Phoneme recognition using time-delay neural networks. IEEE Trans. Acoust. Speech Sig. Process. 37(3), 328–339 (1989)
Wu, J., Chan, C.: Isolated word recognition by neural network models with cross-correlation coefficients for speech dynamics. IEEE Trans. Pattern Anal. Mach. Intell. 15(11), 1174–1185 (1993)
Zahorian, S.A., Zimmer, A.M., Meng, F.: Vowel Classification for Computer based Visual Feedback for Speech Training for the Hearing Impaired, ICSLP, 2002
Hu, H., Zahorian, S.A.: Dimensionality reduction methods for HMM phonetic recognition. In: ICASSP (2010)
Sak, H., Senior, A., Rao, K., Beaufays, F., Schalkwyk, J.: Google voice search: faster and more accurate. Wayback Machine (2016)
Fernandez, S., Graves, A., Hinton, G.: Sequence labelling in structured domains with hierarchical recurrent neural networks. In: Proceedings of IJCAI (2007)
Graves, A., Mohamed, A., Schmidhuber, J.: Speech recognition with deep recurrent neural networks. In: ICASSP (2013)
Deng, L., Yu, D.: Deep Learning: Methods and Applications. Found. Trends Sig. Process. 7(3), 197–387 (2014)
Yu, D., Deng, L., Dahl, G.: Roles of pre-training and fine-tuning in context-dependent DBN-HMMs for real-world speech recognition. In: NIPS Workshop on Deep Learning and Unsupervised Feature Learning (2010)
Dahl, G.E., Yu, D., Deng, L., Acero, A.: Context-dependent pre-trained deep neural networks for large-vocabulary speech recognition. IEEE Trans. Audio Speech Sig. Process. 20(1), 30–42 (2012)
Deng, L., Li, J., Huang, J., Yao, K., Yu, D., Seide, F.: Recent advances in deep learning for speech research at microsoft. In: ICASSP (2013)
Jurafsky, D., James, H.M.: Speech and Language Processing: An Introduction to Natural Language Processing, Computational Linguistics, and Speech Recognition. Stanford University (2018)
Graves, A.: Towards end-to-end speech recognition with recurrent neural networks. In: ICML (2014)
Yannis, M.A., Brendan, S., Shimon, W.N., de Freitas, N.: LipNet: End-to-End Sentence-level Lipreading. Cornell University (2016)
Brendan, S., et al.: Large-Scale Visual Speech Recognition. Cornell University (2018)
National Center for Technology Innovation Speech Recognition for Learning (2010). http://www.ldonline.org/article/38655/
Follensbee, B., McCloskey-Dale, S.: Speech recognition in schools: an update from the field. In: Technology and Persons with Disabilities Conference (2018)
Forgrave, K.E.: Assistive technology: empowering students with disabilities. The Clearing House 7(3), 122–126 (2002)
Apple Inc: Speech framework (2010). https://developer.apple.com/documentation/speech
Author information
Authors and Affiliations
Corresponding author
Editor information
Editors and Affiliations
Rights and permissions
Copyright information
© 2019 ICST Institute for Computer Sciences, Social Informatics and Telecommunications Engineering
About this paper
Cite this paper
Phan, N.H., Bui, T.T.T., Spitsyn, V.G. (2019). Development English Pronunciation Practicing System Based on Speech Recognition. In: Vinh, P., Rakib, A. (eds) Context-Aware Systems and Applications, and Nature of Computation and Communication. ICCASA ICTCC 2019 2019. Lecture Notes of the Institute for Computer Sciences, Social Informatics and Telecommunications Engineering, vol 298. Springer, Cham. https://doi.org/10.1007/978-3-030-34365-1_12
Download citation
DOI: https://doi.org/10.1007/978-3-030-34365-1_12
Published:
Publisher Name: Springer, Cham
Print ISBN: 978-3-030-34364-4
Online ISBN: 978-3-030-34365-1
eBook Packages: Computer ScienceComputer Science (R0)