Development English Pronunciation Practicing System Based on Speech Recognition

Phan, Ngoc Hoang; Bui, Thi Thu Trang; Spitsyn, V. G.

doi:10.1007/978-3-030-34365-1_12

Ngoc Hoang Phan¹⁷,
Thi Thu Trang Bui¹⁷ &
V. G. Spitsyn¹⁸

Part of the book series: Lecture Notes of the Institute for Computer Sciences, Social Informatics and Telecommunications Engineering ((LNICST,volume 298))

Included in the following conference series:

316 Accesses

Abstract

The relevance of the research is caused by the need of application of speech recognition technology for language teaching. The speech recognition is one of the most important tasks of the signal processing and pattern recognition fields. The speech recognition technology allows computers to understand human speech and it plays very important role in people’s lives. This technology can be used to help people in a variety way such as controlling smart homes and devices; using robots to perform job interviews; converting audio into text, etc. But there are not many applications of speech recognition technology in education, especially in English teaching. The main aim of the research is to propose an algorithm in which speech recognition technology is used English language teaching. Objects of researches are speech recognition technologies and frameworks, English spoken sounds system. Research results: The authors have proposed an algorithm based on speech recognition framework for English pronunciation learning. This proposed algorithm can be applied to another speech recognition framework and different languages. Besides the authors also demonstrated how to use the proposed algorithm for development English pronunciation practicing system based on iOS mobile app platform. The system also allows language learners can practice English pronunciation anywhere and anytime without any purchase.

This is a preview of subscription content, log in via an institution to check access.

Access this chapter

Log in via an institution

Chapter: USD 29.95; Price excludes VAT (USA)

eBook: USD 39.99; Price excludes VAT (USA)

Softcover Book: USD 54.99; Price excludes VAT (USA)

Tax calculation will be finalised at checkout

Purchases are for personal use only

Institutional subscriptions

References

Juang, B.H., Rabiner, L.R.: Automatic speech recognition–a brief history of the technology development (2015). https://web.ece.ucsb.edu/Faculty/Rabiner/ece259/Reprints/354_LALI-ASRHistory-final-10-8.pdf
Benesty, J., Sondhi, M.M., Huang, Y.: Springer Handbook of Speech Processing. Springer, Heidelberg (2008). https://doi.org/10.1007/978-3-540-49127-9
Book Google Scholar
Jelinek, F.: Pioneering Speech Recognition (2015). https://www.ibm.com/ibm/history/ibm100/us/en/icons/speechreco/
Huang, X., Baker, J., Reddy, R.: A Historical perspective of speech recognition. Commun. ACM 57(1), 94–103 (2014)
Article Google Scholar
Hanazawa, T., Hinton, G., Shikano, K., Lang, K.J.: Phoneme recognition using time-delay neural networks. IEEE Trans. Acoust. Speech Sig. Process. 37(3), 328–339 (1989)
Article Google Scholar
Wu, J., Chan, C.: Isolated word recognition by neural network models with cross-correlation coefficients for speech dynamics. IEEE Trans. Pattern Anal. Mach. Intell. 15(11), 1174–1185 (1993)
Article Google Scholar
Zahorian, S.A., Zimmer, A.M., Meng, F.: Vowel Classification for Computer based Visual Feedback for Speech Training for the Hearing Impaired, ICSLP, 2002
Google Scholar
Hu, H., Zahorian, S.A.: Dimensionality reduction methods for HMM phonetic recognition. In: ICASSP (2010)
Google Scholar
Sak, H., Senior, A., Rao, K., Beaufays, F., Schalkwyk, J.: Google voice search: faster and more accurate. Wayback Machine (2016)
Google Scholar
Fernandez, S., Graves, A., Hinton, G.: Sequence labelling in structured domains with hierarchical recurrent neural networks. In: Proceedings of IJCAI (2007)
Google Scholar
Graves, A., Mohamed, A., Schmidhuber, J.: Speech recognition with deep recurrent neural networks. In: ICASSP (2013)
Google Scholar
Deng, L., Yu, D.: Deep Learning: Methods and Applications. Found. Trends Sig. Process. 7(3), 197–387 (2014)
Article MathSciNet Google Scholar
Yu, D., Deng, L., Dahl, G.: Roles of pre-training and fine-tuning in context-dependent DBN-HMMs for real-world speech recognition. In: NIPS Workshop on Deep Learning and Unsupervised Feature Learning (2010)
Google Scholar
Dahl, G.E., Yu, D., Deng, L., Acero, A.: Context-dependent pre-trained deep neural networks for large-vocabulary speech recognition. IEEE Trans. Audio Speech Sig. Process. 20(1), 30–42 (2012)
Article Google Scholar
Deng, L., Li, J., Huang, J., Yao, K., Yu, D., Seide, F.: Recent advances in deep learning for speech research at microsoft. In: ICASSP (2013)
Google Scholar
Jurafsky, D., James, H.M.: Speech and Language Processing: An Introduction to Natural Language Processing, Computational Linguistics, and Speech Recognition. Stanford University (2018)
Google Scholar
Graves, A.: Towards end-to-end speech recognition with recurrent neural networks. In: ICML (2014)
Google Scholar
Yannis, M.A., Brendan, S., Shimon, W.N., de Freitas, N.: LipNet: End-to-End Sentence-level Lipreading. Cornell University (2016)
Google Scholar
Brendan, S., et al.: Large-Scale Visual Speech Recognition. Cornell University (2018)
Google Scholar
National Center for Technology Innovation Speech Recognition for Learning (2010). http://www.ldonline.org/article/38655/
Follensbee, B., McCloskey-Dale, S.: Speech recognition in schools: an update from the field. In: Technology and Persons with Disabilities Conference (2018)
Google Scholar
Forgrave, K.E.: Assistive technology: empowering students with disabilities. The Clearing House 7(3), 122–126 (2002)
Article Google Scholar
Apple Inc: Speech framework (2010). https://developer.apple.com/documentation/speech

Download references

Author information

Authors and Affiliations

Ba Ria-Vung Tau University, 80, Truong Cong Dinh, Vung Tau, Ba Ria-Vung Tau, Vietnam
Ngoc Hoang Phan & Thi Thu Trang Bui
National Research Tomsk Polytechnic University, 30, Lenin Avenue, Tomsk, Russia
V. G. Spitsyn

Authors

Ngoc Hoang Phan
View author publications
You can also search for this author in PubMed Google Scholar
Thi Thu Trang Bui
View author publications
You can also search for this author in PubMed Google Scholar
V. G. Spitsyn
View author publications
You can also search for this author in PubMed Google Scholar

Corresponding author

Correspondence to Ngoc Hoang Phan .

Editor information

Editors and Affiliations

Nguyen Tat Thanh University, Ho Chi Minh City, Vietnam
Phan Cong Vinh
The University of the West of England, Bristol, UK
Abdur Rakib

Rights and permissions

Reprints and permissions

Copyright information

About this paper

Cite this paper

Phan, N.H., Bui, T.T.T., Spitsyn, V.G. (2019). Development English Pronunciation Practicing System Based on Speech Recognition. In: Vinh, P., Rakib, A. (eds) Context-Aware Systems and Applications, and Nature of Computation and Communication. ICCASA ICTCC 2019 2019. Lecture Notes of the Institute for Computer Sciences, Social Informatics and Telecommunications Engineering, vol 298. Springer, Cham. https://doi.org/10.1007/978-3-030-34365-1_12

Download citation

DOI: https://doi.org/10.1007/978-3-030-34365-1_12
Published: 01 November 2019
Publisher Name: Springer, Cham
Print ISBN: 978-3-030-34364-4
Online ISBN: 978-3-030-34365-1
eBook Packages: Computer ScienceComputer Science (R0)

Publish with us

Policies and ethics