ADVERTISEMENT

Home| Journals| Articles by Year| Audio Abstracts
 

Original Article



Machine learning approaches for diabetes prediction: A comparative analysis of classification models

Fatma Hilal Yagin.



Abstract
Download PDF Post

Early and accurate identification of individuals at risk for type 2 diabetes mellitus (T2DM) is a clinical priority given the global scale of the epidemic. Machine learning (ML) methods offer a data-driven complement to conventional screening approaches; however, systematic comparisons of multiple classifiers under identical experimental conditions remain valuable for guiding model selection in practice. This study evaluated four ML classification algorithms — Naïve Bayes (NB), Logistic Regression (LR), Random Forest (RF), and Bagged Classification and Regression Trees (CART) — for T2DM prediction using the Pima Indians Diabetes Database (n = 768). Physiologically implausible zero values were treated as missing and median-imputed within each training fold, class weighting was applied to address the class imbalance between diabetic (34.9%) and non-diabetic (65.1%) participants, and hyperparameters were selected via grid search. Models were evaluated using 10-fold stratified cross-validation and reported as mean ± 95% confidence interval. Bagged CART achieved the highest overall accuracy (0.77 ± 0.04), sensitivity (0.78 ± 0.06), and F1-score (0.70 ± 0.05), while Naïve Bayes attained the highest specificity (0.82 ± 0.06). All models achieved comparable discrimination (AUC-ROC 0.81 – 0.84). SHAP-based feature importance analysis of the Bagged CART model identified plasma glucose concentration as the dominant predictor, followed by body mass index and age. These findings support the potential value of ensemble tree-based methods as a component of non-invasive diabetes risk-screening tools and highlight the continued relevance of glycaemic variables and anthropometric measures in T2DM risk stratification. As this analysis is retrospective and confined to a single, demographically restricted cohort, these results should be regarded as hypothesis-generating rather than evidence of clinical readiness, pending external validation.

Key words: Type 2 diabetes mellitus, machine learning, classification, ensemble methods







Bibliomed Article Statistics

2
R
E
A
D
S

3
D
O
W
N
L
O
A
D
S
09
2026

Full-text options


Share this Article


Online Article Submission
• ejmanager.com




ejPort - eJManager.com
Author Tools
About BiblioMed
License Information
Terms & Conditions
Privacy Policy
Contact Us

The articles in Bibliomed are open access articles licensed under Creative Commons Attribution 4.0 International License (CC BY), which permits use, sharing, adaptation, distribution and reproduction in any medium or format, as long as you give appropriate credit to the original author(s) and the source, provide a link to the Creative Commons licence, and indicate if changes were made. To view a copy of this licence, visit http://creativecommons.org/licenses/by/4.0/.