inklap

Balanced DATA by DBSCAN and Weighted Arithmetic Mean to Improve Performance of Machine Learning Algorithms

Serkan GÜLDAL · Bitlis Eren Üniversitesi Fen Bilimleri Dergisi · 2021

Improvement of digital technology has caused the collected data sizes to increase at an accelerating rate. The increase in data size comes with new problems such as unbalanced data. If a dataset is unbalanced, the classes are not equally distributed. Therefore, classification of the data causes performance losses since the classification algorithms treat as the datasets are balanced. While the classification favors the majority class, the minority class is often misclassified. The majority of collected datasets, especially medical datasets, have an unbalanced distribution problem. To reduce the unbalance datasets, various studies have been performed in recent years. In general terms, these studies are undersampling, oversampling, or both to balance the datasets. In this study, an oversampling method is proposed employing distance and mean based resampling method to produce synthetic samples. For the resampling process, the distances between pairs are calculated by the Euclidean distance in the minority class. The calculated distances are considered in the sense of DBSCAN to obtain a sufficient amount of pairs. The new synthetic samples were formed between listed pairs by using the

📖 افتح في inklap 🔗 DOI 📮 اطلب بحثاً