Performance of Distance Metrics in SMOTE for Binary Imbalanced Classification

Authors

DOI:

https://doi.org/10.31294/co-science.v6i2.12543

Keywords:

Binary Imbalanced Classification, Distance Metrics, SMOTE, MCC, G-Mean

Abstract

Skewed class distribution continues to be one of the central obstacles in binary classification, since a learning model tends to lean toward the dominant class and consequently overlooks observations belonging to the under-represented class. The purpose of this research is to examine how the choice of distance measure inside SMOTE, specifically Euclidean, Manhattan, Chebyshev, and Hamming, affects predictive quality on imbalanced binary data. Ten publicly available binary datasets drawn from the KEEL repository, whose imbalance ratios span from 1.86 up to 15.80, were used in the experiment. Every dataset was preprocessed and partitioned into 80% for training and 20% for testing; oversampling with SMOTE was carried out on the training portion only, after which four learners, namely Naive Bayes, Decision Tree, Logistic Regression, and k-Nearest Neighbor, were assessed. Model quality was judged through the Matthews Correlation Coefficient (MCC) together with the G-Mean, as these two indicators describe imbalanced performance more faithfully than plain accuracy. The comparison revealed that pairing Euclidean-based SMOTE with Logistic Regression yielded the strongest average scores (MCC = 0.72; G-Mean = 0.79); Manhattan-based SMOTE reached its top MCC again with Logistic Regression (MCC = 0.68) and its top G-Mean with the Decision Tree (G-Mean = 0.79); Chebyshev-based SMOTE delivered the best overall combination together with the Decision Tree (MCC = 0.74; G-Mean = 0.84); and Hamming-based SMOTE performed best alongside Logistic Regression (MCC = 0.73; G-Mean = 0.81). Taken together, these outcomes suggest that the distance function chosen within SMOTE shapes the quality of the generated synthetic points and, in turn, the behavior of the trained classifier.

Downloads

Download data is not yet available.

References

Asniar, Maulidevi, N. U., & Surendro, K. (2022). SMOTE-LOF for noise identification in imbalanced data classification. Journal of King Saud University - Computer and Information Sciences, 34(6), 3413-3423. https://doi.org/10.1016/j.jksuci.2021.01.014

Carvalho, M., Pinho, A. J., & Bras, S. (2025). Resampling approaches to handle class imbalance: A review from a data perspective. Journal of Big Data, 12, Article 71. https://doi.org/10.1186/s40537-025-01119-4

Chandra, W., Suprihatin, B., & Resti, Y. (2023). Median-KNN Regressor-SMOTE-Tomek Links for handling missing and imbalanced data in air quality prediction. Symmetry, 15(4), Article 887. https://doi.org/10.3390/sym15040887

Chen, W., Yang, K., Yu, Z., Shi, Y., & Chen, C. L. P. (2024). A survey on imbalanced learning: Latest research, applications and future directions. Artificial Intelligence Review, 57, Article 137. https://doi.org/10.1007/s10462-024-10759-6

Dai, Q., Liu, J., & Zhao, J. L. (2023). Distance-based arranging oversampling technique for imbalanced data. Neural Computing and Applications, 35(2), 1323-1342. https://doi.org/10.1007/s00521-022-07828-8

Cruz Huayanay, A., Bazan, J. L., & Russo, C. M. (2025). Performance of evaluation metrics for classification in imbalanced data. Computational Statistics, 40(3), 1447-1473. https://doi.org/10.1007/s00180-024-01539-5

Elreedy, D., Atiya, A. F., & Kamalov, F. (2024). A theoretical distribution analysis of synthetic minority oversampling technique (SMOTE) for imbalanced learning. Machine Learning, 113(7), 4903-4923. https://doi.org/10.1007/s10994-022-06296-4

Feng, S., Keung, J., Zhang, P., Xiao, Y., & Zhang, M. (2022). The impact of the distance metric and measure on SMOTE-based techniques in software defect prediction. Information and Software Technology, 142, Article 106742. https://doi.org/10.1016/j.infsof.2021.106742

Hairani, H., Widiyaningtyas, T., & Prasetya, D. D. (2024). Addressing class imbalance of health data: A systematic literature review on modified Synthetic Minority Oversampling Technique (SMOTE) strategies. International Journal on Informatics Visualization, 8(3), 1310-1318. https://doi.org/10.62527/joiv.8.3.2283

Hemmatian, J., Hajizadeh, R., & Nazari, F. (2025). Addressing imbalanced data classification with Cluster-Based Reduced Noise SMOTE. PLoS ONE, 20(2), Article e0317396. https://doi.org/10.1371/journal.pone.0317396

Husain, G., Nasef, D., Jose, R., Mayer, J., Bekbolatova, M., Devine, T., & Toma, M. (2025). SMOTE vs. SMOTEENN: A study on the performance of resampling algorithms for addressing class imbalance in regression models. Algorithms, 18(1), Article 37. https://doi.org/10.3390/a18010037

Imani, M., Joudaki, M., Bagheri, A., & Arabnia, H. R. (2026). Why ROC-AUC is misleading for highly imbalanced data: In-depth evaluation of MCC, F2-score, H-measure, and AUC-based metrics across diverse classifiers. Technologies, 14(1), Article 54. https://doi.org/10.3390/technologies14010054

Jain, A., Dubey, A. K., Khan, S., Panwar, A., Alkhatib, M., & Alshahrani, A. M. (2025). A PSO weighted ensemble framework with SMOTE balancing for student dropout prediction in smart education systems. Scientific Reports, 15, Article 17463. https://doi.org/10.1038/s41598-025-97506-1

Li, Y., Yang, Y., Song, P., Duan, L., & Ren, R. (2025). An improved SMOTE algorithm for enhanced imbalanced data classification by expanding sample generation space. Scientific Reports, 15, Article 23521. https://doi.org/10.1038/s41598-025-09506-w

Lubis, A., Irawan, Y., Junadhi, J., & Defit, S. (2024). Leveraging K-Nearest Neighbors with SMOTE and boosting techniques for data imbalance and accuracy improvement. Journal of Applied Data Sciences, 5(4), 1625-1638. https://doi.org/10.47738/jads.v5i4.343

Lyu, J., Yang, J., Su, Z., & Zhu, Z. (2025). LD-SMOTE: A novel local density estimation-based oversampling method for imbalanced datasets. Symmetry, 17(2), Article 160. https://doi.org/10.3390/sym17020160

Maldonado, S., Vairetti, C., Fernandez, A., & Herrera, F. (2022). FW-SMOTE: A feature-weighted oversampling approach for imbalanced classification. Pattern Recognition, 124, Article 108511. https://doi.org/10.1016/j.patcog.2021.108511

Matharaarachchi, S., Domaratzki, M., & Muthukumarana, S. (2024). Enhancing SMOTE for imbalanced data with abnormal minority instances. Machine Learning with Applications, 18, Article 100597. https://doi.org/10.1016/j.mlwa.2024.100597

Misdram, M., Noersasongko, E., Purwanto, P., Muljono, M., & Pamuji, F. Y. (2023). Gaussian Based-SMOTE method for handling imbalanced small datasets. Jurnal Ilmiah Teknik Elektro Komputer dan Informatika, 9(4), 973-982. https://doi.org/10.26555/jiteki.v9i4.26881

Mukherjee, A., Qazani, M. R. C., Rana, B. M. J., Akter, S., Mohajerzadeh, A., Sathi, N. J., Ali, L. E., Khan, M. S., & Asadi, H. (2025). SMOTE-ENN resampling technique with Bayesian optimization for multi-class classification of dry bean varieties. Applied Soft Computing, 181, Article 113467. https://doi.org/10.1016/j.asoc.2025.113467

Nasaruddin, N., Masseran, N., Idris, W. M. R., & Ul-Saufie, A. Z. (2025). A SMOTE PCA HDBSCAN approach for enhancing water quality classification in imbalanced datasets. Scientific Reports, 15, Article 13059. https://doi.org/10.1038/s41598-025-97248-0

Pei, W., Xue, B., Zhang, M., & Shang, L. (2024). A survey on unbalanced classification: How can evolutionary computation help? IEEE Transactions on Evolutionary Computation, 28(2), 353-373. https://doi.org/10.1109/TEVC.2023.3257230

Rezvani, S., & Wang, X. (2023). A broad review on class imbalance learning techniques. Applied Soft Computing, 143, Article 110415. https://doi.org/10.1016/j.asoc.2023.110415

Swana, E. F., Doorsamy, W., & Bokoro, P. (2022). Tomek Link and SMOTE approaches for machine fault classification with an imbalanced dataset. Sensors, 22(9), Article 3246. https://doi.org/10.3390/s22093246

Taskiran, S. F., Turkoglu, B., Kaya, E., & Asuroglu, T. (2025). A comprehensive evaluation of oversampling techniques for enhancing text classification performance. Scientific Reports, 15, Article 21631. https://doi.org/10.1038/s41598-025-05791-7

Wainer, J. (2024). An empirical evaluation of imbalanced data strategies from a practitioner's point of view. Expert Systems with Applications, 256, Article 124863. https://doi.org/10.1016/j.eswa.2024.124863

Widiyaningtyas, T., Hairani, H., Prasetya, D. D., Pujianto, U., & Caesarendra, W. (2025). A modified SMOTE with noise filtering and Manhattan distance metric approach to address imbalanced health datasets. Engineering, Technology and Applied Science Research, 15(4), 25452-25459. https://doi.org/10.48084/etasr.11925

Wong, T. T., & Chung, P. C. (2025). A consistency analysis on four evaluation metrics for classifying imbalanced data. Knowledge and Information Systems, 67, 10639-10656. https://doi.org/10.1007/s10115-025-02544-w

Yulian Pamuji, F., Muslikh, A. R., Arief, R. M., & Muti, D. (2024). Komparasi metode Mean dan KNN Imputation dalam mengatasi missing value pada dataset kecil. Jurnal Informatika Polinema, 10(2), 257-264. https://doi.org/10.33795/jip.v10i2.5031

Zhang, Y., Deng, L., & Wei, B. (2024). Imbalanced data classification based on improved Random-SMOTE and feature standard deviation. Mathematics, 12(11), Article 1709. https://doi.org/10.3390/math12111709

Cover Article

Downloads

Published

2026-07-22

How to Cite

Performance of Distance Metrics in SMOTE for Binary Imbalanced Classification. (2026). Computer Science (CO-SCIENCE), 6(2), 134-143. https://doi.org/10.31294/co-science.v6i2.12543

Most read articles by the same author(s)