Anita Qoiriah, Dian Oktavia Putri, Yuni Yamasari, I.M. Suartana, Ricky E. Putra, Andi Iwan Nurhidayat
People's new habits in using social media, including spreading comments or hate speech that leads to cyberbullying behavior, is a concern because it has a serious impact on victims. Twitter has 400 million users with the majority being young users. Based on Pew Research Center and UNICEF surveys, the prevalence of cyberbullying among teenagers is quite high in both the United States and Indonesia, so a classification method that can classify cyberbullying well is needed to minimize the negative impact of cyberbullying and create a safer online space for users. On the Twitter platform, there is a 280 -character limit for uploading a tweet, which causes word variations that allow vocabulary mismatches in uploaded tweets. A method is needed to overcome word variations, such as FastText. The use of the FastText Pretrained Model as word embedding can help handle words that are not recognized (Out Of Vocabulary). In addition, to improve prediction accuracy, a classification prediction model built with ensemble learning, namely hybrid CNN-SVM, combines feature extraction by Convolutional Neural Network (CNN) with Support Vector Machine (SVM) classification capabilities. From the test results, it is found that at K = 20 CNN-SVM with Fast Text word embedding provides excellent results, namely accuracy reaching 9 1. 1 6 % in model formation and 7 4. 7 8 % in test data. Cyberbullying classification can be classified into 5 class labels Age, ethnicity, gender, religion, and not_cyberbullying.. © 2024 IEEE.
Universitas Negeri Surabaya, Department of Informatics Engineering, Surabaya, Indonesia