A Comparative Study of Machine Learning Algorithms for Predictive Data Analytics
Keywords:
Abstract
Predictive data analytics converts historical observations into estimates of future or unknown outcomes, but the choice of learning algorithm strongly affects accuracy, robustness, interpretability, and computational cost. This paper presents a reproducible comparative methodology for six widely used supervised machine learning approaches: logistic regression, k-nearest neighbors (k-NN), support vector machines (SVM), random forests, gradient boosting, and a multilayer perceptron (MLP) neural network. The study uses the Wisconsin Diagnostic Breast Cancer benchmark, containing 569 observations and 30 numerical predictors derived from digitized cell-nuclei images. Data are stratified into 80% training and 20% testing partitions; scaling is applied to distance-, margin-, and gradient-based models, and model stability is examined through five-fold stratified cross-validation. Performance is assessed using accuracy, precision, recall, F1-score, ROC-AUC, and fit time. In the reproducible benchmark run, SVM and random forest reach the highest test accuracy (97.37%), while logistic regression obtains the highest ROC-AUC (0.996). The findings demonstrate that no single algorithm dominates every criterion: simpler models can be highly competitive on well-structured tabular data, whereas ensembles offer robustness and nonlinear representation at greater computational cost. The paper therefore emphasizes evaluation by multiple criteria rather than selection by accuracy alone
Downloads
Downloads
Published
How to Cite
Issue
Section
License

This work is licensed under a Creative Commons Attribution-NoDerivatives 4.0 International License.