An AI-powered English pronunciation training assistant using NVIDIA Riva and pygame

Authors

  • รวินทร์ ไชยสิทธิพร -

Abstract

Developing correct English pronunciation requires practice in order to build skill. The objective of this research is to apply artificial intelligence technology from NVIDIA Riva, which provides speech recognition and speech synthesis, as a tool for human–computer interaction. A large language model (LLM) is used to function as an intelligent agent that communicates through spoken dialogue, and pygame is employed to create a graphical interface in the form of an animated robot face. ROS2 serves as the middleware connecting the various system components. The system’s goal is to support English pronunciation training and to foster positive attitudes toward English learning.

The system is designed to run specifically on the Jetson Orin NX 16GB, a small, energy-efficient yet high-performance computer. This enables the system to operate offline, without relying on a network, which improves processing speed and enhances security—unlike typical cloud-based systems that often suffer from higher latency and lower security.

               Testing results show that the system performs well, with low response times and high processing accuracy from NVIDIA Riva. In conclusion, the techniques presented in this research are highly suitable for developing innovative conversational intelligent agents, offering significant potential benefits across many future applications.

References

Arunsirot, S. (2017). Implementing a Speech Analyzer Software to enhance English pronunciation competence of Thai students. วารสารศึกษาศาสตร์ มหาวิทยาลัยบูรพา, 28 (2). https://journal.lib.buu.ac.th/index.php/education2/article/view/4954

Arunsirot, S. (2017). Implementing a Speech Analyzer software to enhance English pronunciation competence of Thai students. วารสารศึกษาศาสตร์ มหาวิทยาลัยบูรพา, 28(2). https://journal.lib.buu.ac.th/index.php/education2/article/view/4954

กิตต์สุทธิ์, น., บุญมา, ค., เมืองประทับ, จ., นาไช, ว., & สุพนิธิ, ท. (2024, November 25). Improve English pronunciation at word level for Thai EFL learners in southern region using end-to-end automatic speech recognition. In Proceedings of the 2024: ICCE 2024: The 32nd International Conference on Computers in Education. https://doi.org/10.58459/icce.2024.4917

Jantaworn, P., Thiengburanathum, P., Denpaiboon, C., Wongthanavasu, S., & Keawthip, S. (2024). City to City of Learning in Thailand: From Research Synthesis by Large Language Models = เมืองสู่เมืองแห่งการเรียนรู้ของไทย : การสังเคราะห์งานวิจัยโดยแบบจำลองภาษาขนาดใหญ่. Asian Journal of Adult Education. https://so01.tci-thaijo.org/index.php/AJA/article/download/274059/178740/1114876

Huang, Y.-C., & Liao, L.-C. (2015). A study of text-to-speech (TTS) in children’s English learning. Teaching English With Technology, 15(1), 14-30. https://files.eric.ed.gov/fulltext/EJ1140575.pdf

Fitria, T. N. (2024). Utilizing Text-to-Speech (TTS) Technology in Creating Listening Materials for English Language Teaching (ELT). Journal of English Language and Culture, 15(1), 73–85. https://www.researchgate.net/publication/388142264_Utilizing_Text-to-Speech_TTS_Technology_in_Creating_Listening_Materials_for_English_Language_Teaching_ELT

NVIDIA Corporation. (2025, April 3). NVIDIA Riva user guide (Release 2.19.0). https://docs.nvidia.com/deeplearning/riva/user-guide/docs/index.html

Published

2026-04-30

How to Cite

[1]
ไชยสิทธิพร ร., “ An AI-powered English pronunciation training assistant using NVIDIA Riva and pygame ”, JSciTech, vol. 10, no. 1, Apr. 2026.