Design and Implementation of an Indonesian Speech Recognition Application Based on Mel Frequency Cepstral Coefficients and K-Nearest Neighbor Classification

Closed

Ali Nur Fathoni, Inaya Retno Putri, Hesti Khuzaimah Nurul Yusufiyah, Fendi Achmad, Nur Kholis, Yuli Sutoto Nugroho

2024 7th International Seminar on Research of Information Technology and Intelligent Systems: Advanced Intelligent Systems in Contemporary Society, ISRITI 2024 - Proceedings Conference paper Cited by 1 Quartile

Abstract

The Speech-to-Text application facilitates the recognition of human speech and its conversion into written form, ultimately simplifying the typing process. Voice data is analyzed by converting it into data parameters using the extraction feature. In the feature extraction stage, voice data is analyzed and converted into data parameters using the MFCC method. These parameters are then grouped and classified using the KNN algorithm, which categorizes the data based on the shortest distance between parameter values. This research is centered on the development of a machine learning model and its corresponding application designed for real-time conversion of speech to text. The dataset includes 800 samples, generated by 10 speakers who each repeated ten keywords eight times. The keywords correspond to the numbers zero through nine in Indonesian. The model achieved an accuracy rate of 85.90% in recognizing these words, with an optimal K value of 5. The findings of this study establish a robust basis for advancing the development of speech-to-text applications, with a specific emphasis on the Indonesian language. © 2024 IEEE.

Affiliations

Universitas Negeri Surabaya, Faculty of Engineering, Department of Electrical Engineering, Surabaya, Indonesia; Queen Mary University of London, School of Electronic Engineering and Computer Science, London, United Kingdom