A Cross-Sampling Method for Hidden Structure Extraction to Improve Imbalanced Multiclass Classification Accuracy

Closed

Wiyli Yustanti, Nur Iriawan, Irhamah, I Kadek Dwi Nuryana, Aries Dwi Indriyanti

2023 2023 6th International Conference on Vocational Education and Electrical Engineering: Integrating Scalable Digital Connectivity, Intelligence Systems, and Green Technology for Education and Sustainable Community Development, ICVEE 2023 - Proceeding Conference paper Cited by 1 Quartile

Abstract

Class prediction problems in classification cases are often faced with irregular data conditions. This research contributes to providing a more structured hidden pattern extraction approach to data sets that already have class labels. The process of taking samples from data collection in this study is called the cross-sampling (CS) technique. The basic idea of this technique is to regroup the data using the appropriate clustering method. This study applied the Mini Batch K-Mean and Spectral algorithms to four public datasets with an imbalanced multiclass distribution. The grouping results are then assigned a label based on the original label reference using the pattern-matching concept on the first principal component through the PCA procedure. The labelling results are then used as external validation for the actual data labels to represent all classes. The validation results produce a confusion matrix used as the basis for the cross-sampling method. The dataset before cross-sampling and the results of cross-sampling were compared to the performance of the classification prediction accuracy using the F1-Score and Area Under Curve (AUC) measurements. The statistical hypothesis testing results show a significant difference in performance before and after the cross-sampling procedure. This difference is demonstrated by the accuracy of all the classification algorithms used, which increased significantly from the average performance value of 82.09% to 96.7%. © 2023 IEEE.

Affiliations

Universitas Negeri Surabaya, Department of Informatics Faculty of Engineering, Surabaya, Indonesia; Institut Teknologi Sepuluh Nopember, Faculty of Science and Data Analytics, Department of Statistics, Surabaya, Indonesia