Performance of Distance Metrics in SMOTE for Binary Imbalanced Classification
DOI:
https://doi.org/10.31294/co-science.v6i2.12543Keywords:
Binary Imbalanced Classification, Distance Metrics, SMOTE, MCC, G-MeanAbstract
Skewed class distribution continues to be one of the central obstacles in binary classification, since a learning model tends to lean toward the dominant class and consequently overlooks observations belonging to the under-represented class. The purpose of this research is to examine how the choice of distance measure inside SMOTE, specifically Euclidean, Manhattan, Chebyshev, and Hamming, affects predictive quality on imbalanced binary data. Ten publicly available binary datasets drawn from the KEEL repository, whose imbalance ratios span from 1.86 up to 15.80, were used in the experiment. Every dataset was preprocessed and partitioned into 80% for training and 20% for testing; oversampling with SMOTE was carried out on the training portion only, after which four learners, namely Naive Bayes, Decision Tree, Logistic Regression, and k-Nearest Neighbor, were assessed. Model quality was judged through the Matthews Correlation Coefficient (MCC) together with the G-Mean, as these two indicators describe imbalanced performance more faithfully than plain accuracy. The comparison revealed that pairing Euclidean-based SMOTE with Logistic Regression yielded the strongest average scores (MCC = 0.72; G-Mean = 0.79); Manhattan-based SMOTE reached its top MCC again with Logistic Regression (MCC = 0.68) and its top G-Mean with the Decision Tree (G-Mean = 0.79); Chebyshev-based SMOTE delivered the best overall combination together with the Decision Tree (MCC = 0.74; G-Mean = 0.84); and Hamming-based SMOTE performed best alongside Logistic Regression (MCC = 0.73; G-Mean = 0.81). Taken together, these outcomes suggest that the distance function chosen within SMOTE shapes the quality of the generated synthetic points and, in turn, the behavior of the trained classifier.
Downloads
References
Asniar, Maulidevi, N. U., & Surendro, K. (2022). SMOTE-LOF for noise identification in imbalanced data classification. Journal of King Saud University - Computer and Information Sciences, 34(6), 3413-3423. https://doi.org/10.1016/j.jksuci.2021.01.014
Carvalho, M., Pinho, A. J., & Bras, S. (2025). Resampling approaches to handle class imbalance: A review from a data perspective. Journal of Big Data, 12, Article 71. https://doi.org/10.1186/s40537-025-01119-4
Chandra, W., Suprihatin, B., & Resti, Y. (2023). Median-KNN Regressor-SMOTE-Tomek Links for handling missing and imbalanced data in air quality prediction. Symmetry, 15(4), Article 887. https://doi.org/10.3390/sym15040887
Chen, W., Yang, K., Yu, Z., Shi, Y., & Chen, C. L. P. (2024). A survey on imbalanced learning: Latest research, applications and future directions. Artificial Intelligence Review, 57, Article 137. https://doi.org/10.1007/s10462-024-10759-6
Dai, Q., Liu, J., & Zhao, J. L. (2023). Distance-based arranging oversampling technique for imbalanced data. Neural Computing and Applications, 35(2), 1323-1342. https://doi.org/10.1007/s00521-022-07828-8
Cruz Huayanay, A., Bazan, J. L., & Russo, C. M. (2025). Performance of evaluation metrics for classification in imbalanced data. Computational Statistics, 40(3), 1447-1473. https://doi.org/10.1007/s00180-024-01539-5
Elreedy, D., Atiya, A. F., & Kamalov, F. (2024). A theoretical distribution analysis of synthetic minority oversampling technique (SMOTE) for imbalanced learning. Machine Learning, 113(7), 4903-4923. https://doi.org/10.1007/s10994-022-06296-4
Feng, S., Keung, J., Zhang, P., Xiao, Y., & Zhang, M. (2022). The impact of the distance metric and measure on SMOTE-based techniques in software defect prediction. Information and Software Technology, 142, Article 106742. https://doi.org/10.1016/j.infsof.2021.106742
Hairani, H., Widiyaningtyas, T., & Prasetya, D. D. (2024). Addressing class imbalance of health data: A systematic literature review on modified Synthetic Minority Oversampling Technique (SMOTE) strategies. International Journal on Informatics Visualization, 8(3), 1310-1318. https://doi.org/10.62527/joiv.8.3.2283
Hemmatian, J., Hajizadeh, R., & Nazari, F. (2025). Addressing imbalanced data classification with Cluster-Based Reduced Noise SMOTE. PLoS ONE, 20(2), Article e0317396. https://doi.org/10.1371/journal.pone.0317396
Husain, G., Nasef, D., Jose, R., Mayer, J., Bekbolatova, M., Devine, T., & Toma, M. (2025). SMOTE vs. SMOTEENN: A study on the performance of resampling algorithms for addressing class imbalance in regression models. Algorithms, 18(1), Article 37. https://doi.org/10.3390/a18010037
Imani, M., Joudaki, M., Bagheri, A., & Arabnia, H. R. (2026). Why ROC-AUC is misleading for highly imbalanced data: In-depth evaluation of MCC, F2-score, H-measure, and AUC-based metrics across diverse classifiers. Technologies, 14(1), Article 54. https://doi.org/10.3390/technologies14010054
Jain, A., Dubey, A. K., Khan, S., Panwar, A., Alkhatib, M., & Alshahrani, A. M. (2025). A PSO weighted ensemble framework with SMOTE balancing for student dropout prediction in smart education systems. Scientific Reports, 15, Article 17463. https://doi.org/10.1038/s41598-025-97506-1
Li, Y., Yang, Y., Song, P., Duan, L., & Ren, R. (2025). An improved SMOTE algorithm for enhanced imbalanced data classification by expanding sample generation space. Scientific Reports, 15, Article 23521. https://doi.org/10.1038/s41598-025-09506-w
Lubis, A., Irawan, Y., Junadhi, J., & Defit, S. (2024). Leveraging K-Nearest Neighbors with SMOTE and boosting techniques for data imbalance and accuracy improvement. Journal of Applied Data Sciences, 5(4), 1625-1638. https://doi.org/10.47738/jads.v5i4.343
Lyu, J., Yang, J., Su, Z., & Zhu, Z. (2025). LD-SMOTE: A novel local density estimation-based oversampling method for imbalanced datasets. Symmetry, 17(2), Article 160. https://doi.org/10.3390/sym17020160
Maldonado, S., Vairetti, C., Fernandez, A., & Herrera, F. (2022). FW-SMOTE: A feature-weighted oversampling approach for imbalanced classification. Pattern Recognition, 124, Article 108511. https://doi.org/10.1016/j.patcog.2021.108511
Matharaarachchi, S., Domaratzki, M., & Muthukumarana, S. (2024). Enhancing SMOTE for imbalanced data with abnormal minority instances. Machine Learning with Applications, 18, Article 100597. https://doi.org/10.1016/j.mlwa.2024.100597
Misdram, M., Noersasongko, E., Purwanto, P., Muljono, M., & Pamuji, F. Y. (2023). Gaussian Based-SMOTE method for handling imbalanced small datasets. Jurnal Ilmiah Teknik Elektro Komputer dan Informatika, 9(4), 973-982. https://doi.org/10.26555/jiteki.v9i4.26881
Mukherjee, A., Qazani, M. R. C., Rana, B. M. J., Akter, S., Mohajerzadeh, A., Sathi, N. J., Ali, L. E., Khan, M. S., & Asadi, H. (2025). SMOTE-ENN resampling technique with Bayesian optimization for multi-class classification of dry bean varieties. Applied Soft Computing, 181, Article 113467. https://doi.org/10.1016/j.asoc.2025.113467
Nasaruddin, N., Masseran, N., Idris, W. M. R., & Ul-Saufie, A. Z. (2025). A SMOTE PCA HDBSCAN approach for enhancing water quality classification in imbalanced datasets. Scientific Reports, 15, Article 13059. https://doi.org/10.1038/s41598-025-97248-0
Pei, W., Xue, B., Zhang, M., & Shang, L. (2024). A survey on unbalanced classification: How can evolutionary computation help? IEEE Transactions on Evolutionary Computation, 28(2), 353-373. https://doi.org/10.1109/TEVC.2023.3257230
Rezvani, S., & Wang, X. (2023). A broad review on class imbalance learning techniques. Applied Soft Computing, 143, Article 110415. https://doi.org/10.1016/j.asoc.2023.110415
Swana, E. F., Doorsamy, W., & Bokoro, P. (2022). Tomek Link and SMOTE approaches for machine fault classification with an imbalanced dataset. Sensors, 22(9), Article 3246. https://doi.org/10.3390/s22093246
Taskiran, S. F., Turkoglu, B., Kaya, E., & Asuroglu, T. (2025). A comprehensive evaluation of oversampling techniques for enhancing text classification performance. Scientific Reports, 15, Article 21631. https://doi.org/10.1038/s41598-025-05791-7
Wainer, J. (2024). An empirical evaluation of imbalanced data strategies from a practitioner's point of view. Expert Systems with Applications, 256, Article 124863. https://doi.org/10.1016/j.eswa.2024.124863
Widiyaningtyas, T., Hairani, H., Prasetya, D. D., Pujianto, U., & Caesarendra, W. (2025). A modified SMOTE with noise filtering and Manhattan distance metric approach to address imbalanced health datasets. Engineering, Technology and Applied Science Research, 15(4), 25452-25459. https://doi.org/10.48084/etasr.11925
Wong, T. T., & Chung, P. C. (2025). A consistency analysis on four evaluation metrics for classifying imbalanced data. Knowledge and Information Systems, 67, 10639-10656. https://doi.org/10.1007/s10115-025-02544-w
Yulian Pamuji, F., Muslikh, A. R., Arief, R. M., & Muti, D. (2024). Komparasi metode Mean dan KNN Imputation dalam mengatasi missing value pada dataset kecil. Jurnal Informatika Polinema, 10(2), 257-264. https://doi.org/10.33795/jip.v10i2.5031
Zhang, Y., Deng, L., & Wei, B. (2024). Imbalanced data classification based on improved Random-SMOTE and feature standard deviation. Mathematics, 12(11), Article 1709. https://doi.org/10.3390/math12111709
Downloads
Published
Issue
Section
License
Copyright (c) 2026 Fandi Yulian Pamuji, Luthfi Indana, Mohammad Dwi Irfan Affandi

This work is licensed under a Creative Commons Attribution-ShareAlike 4.0 International License.


















