A Comparative Study of Machine Learning Algorithms for Predictive Data Analytics
Keywords:
Abstract
Predictive data analytics increasingly relies on machine learning (ML) models whose performance varies with data structure, dimensionality, class distribution, and model assumptions. This methodology paper presents a controlled comparative framework for evaluating five widely used supervised learning algorithms: logistic regression, k-nearest neighbors (k-NN), support vector machine (SVM), random forest, and gradient boosting. Four public benchmark classification datasets representing binary and multiclass prediction tasks are evaluated using stratified five-fold cross-validation. Accuracy, macro-precision, macro-recall, macro-F1, and one-vs-rest ROC-AUC are used to reduce dependence on any single performance measure. Standardization is applied inside the cross-validation pipeline to distance- and margin-based models, while tree ensembles are trained on the original feature scales. The resulting benchmark shows that SVM achieves the highest mean macro-F1 (0.976), followed by k-NN (0.971), logistic regression (0.970), random forest (0.967), and gradient boosting (0.946). However, the differences are small and vary by dataset. Therefore, the study supports problem-specific model selection rather than assuming universal superiority of a single algorithm. The proposed framework is reproducible, transparent, and suitable for comparative predictive analytics studies in business, engineering, healthcare, and scientific data environments.
Downloads
Downloads
Published
How to Cite
Issue
Section
License

This work is licensed under a Creative Commons Attribution-NoDerivatives 4.0 International License.