Ali Nur Fathoni, Inaya Retno Putri, Hesti Khuzaimah Nurul Yusufiyah, Fendi Achmad, Nur Kholis, Yuli Sutoto Nugroho
The Speech-to-Text application facilitates the recognition of human speech and its conversion into written form, ultimately simplifying the typing process. Voice data is analyzed by converting it into data parameters using the extraction feature. In the feature extraction stage, voice data is analyzed and converted into data parameters using the MFCC method. These parameters are then grouped and classified using the KNN algorithm, which categorizes the data based on the shortest distance between parameter values. This research is centered on the development of a machine learning model and its corresponding application designed for real-time conversion of speech to text. The dataset includes 800 samples, generated by 10 speakers who each repeated ten keywords eight times. The keywords correspond to the numbers zero through nine in Indonesian. The model achieved an accuracy rate of 85.90% in recognizing these words, with an optimal K value of 5. The findings of this study establish a robust basis for advancing the development of speech-to-text applications, with a specific emphasis on the Indonesian language. © 2024 IEEE.
Universitas Negeri Surabaya, Faculty of Engineering, Department of Electrical Engineering, Surabaya, Indonesia; Queen Mary University of London, School of Electronic Engineering and Computer Science, London, United Kingdom