Abstract
Objective
This scoping review aimed to explore how big data and machine learning techniques are currently applied to develop predictive models for the early identification and prediction of diabetes and cardiovascular diseases, while identifying methodological strengths, implementation gaps, and ethical considerations.
Methods
Following PRISMA-ScR guidelines and the Arksey and O’Malley framework, a comprehensive literature search was conducted across five major databases, covering studies published from 2013 to 2025. Studies were included if they used ML with big data sources to predict diabetes or cardiovascular diseases. Data were extracted, thematically synthesized, and quality appraised using a modified Mixed Methods Appraisal Tool (MMAT).
Results
Out of the screened studies, 66 studies were included in the final synthesis. Most employed supervised machine learning techniques such as decision trees and ensemble models, drawing on electronic health records, wearable sensors, and multi modal datasets. Deep learning approaches, while less common, showed promise in handling unstructured data. Key challenges included data heterogeneity, limited external validation, and underrepresentation of diverse populations. Ethical issues like algorithmic bias and lack of model interpretability were noted. Real-world implementation remained sparse, with few models integrated into clinical workflows.
Conclusion
Predictive modeling using big data and ML holds significant promise for early disease detection, yet translation into practice is hindered by methodological, infrastructural, and ethical challenges. Future efforts must prioritize transparency, inclusivity, and interdisciplinary collaboration to ensure responsible deployment in real-world healthcare settings.
1. Introduction
Chronic diseases such as diabetes and cardiovascular ailments have emerged as global health emergencies, not only due to their growing prevalence but also because of their long-term implications for individual well-being and national healthcare expenditures. 1 As societies contend with aging populations, sedentary lifestyles, and nutritional transitions, the incidence of these non-communicable conditions continues to rise at an alarming rate.2,3 Despite decades of advancements in medical interventions, a significant proportion of patients remain undiagnosed until complications have already progressed, narrowing the window for preventive care. 4 Traditional screening methods, while effective in structured clinical settings, often fall short in capturing the complex and dynamic nature of individual health trajectories. 5 As the burden of these diseases extends beyond clinical boundaries into economic and social spheres, there is an urgent need to reimagine early detection strategies through innovative, data-driven approaches. 6
In recent years, the convergence of health informatics, computational modeling, and biomedical data has created unprecedented opportunities for early disease prediction. 7 The exponential growth in digital health records, coupled with the widespread adoption of wearable devices and population-level health surveys, has generated massive datasets ripe for analysis. 8 These diverse and multilayered data sources, often referred to as big data, encompass not only large-scale datasets but also high-dimensional, longitudinal, and multi-source health data that require advanced computational approaches for analysis. 9 Such data offers the potential to uncover hidden associations, trends, and risk factors that remain elusive under conventional statistical frameworks. Alongside this data revolution, machine learning has emerged as a transformative analytical tool, enabling the construction of adaptive, predictive models capable of learning from patterns in historical and real-time data. 10 Unlike static rule-based systems, machine learning algorithms possess the ability to refine their predictive performance over time, making them particularly suited for identifying subtle precursors to chronic conditions such as diabetes and cardiovascular diseases. 11
The integration of big data analytics with machine learning methodologies holds promise not only for enhancing diagnostic accuracy but also for transforming clinical decision-making and health system planning. 12 Predictive models derived from machine learning can support a shift from reactive to proactive care by flagging high-risk individuals before symptoms emerge. 13 For clinicians, this means access to decision-support tools that complement clinical judgment with data-informed insights. 14 For patients, it opens the door to timely interventions and personalized management plans tailored to their unique risk profiles. 15 At the population level, such models can inform strategic resource allocation, guide public health policies, and address disparities by targeting vulnerable groups. 16 The move toward predictive and preventive care aligns with the broader goals of precision medicine, where individualized treatment pathways are informed by a comprehensive understanding of biological, behavioral, and environmental factors. 17
Despite the growing enthusiasm surrounding predictive modeling in healthcare, its translation from research to routine practice remains uneven. 18 While numerous studies have demonstrated the theoretical efficacy of machine learning models in predicting chronic diseases, challenges persist in terms of methodological consistency, clinical relevance, and ethical transparency.19–21 The complexity of integrating these tools into real-world settings, coupled with variability in data quality and population diversity, complicates their broader adoption. Moreover, critical considerations such as algorithmic fairness, model interpretability, and stakeholder engagement should be addressed to ensure responsible implementation. In this context, a comprehensive review of current evidence is vital to chart the landscape of existing approaches, identify methodological strengths and weaknesses, and illuminate paths forward. The present scoping review undertakes this task by systematically mapping the application of big data and machine learning in early detection models for diabetes and cardiovascular diseases, with the goal of informing future research, guiding policy, and enhancing clinical integration.
2. Methodology
2.1. Research design
This scoping review was designed in accordance with the methodological approach introduced by Arksey and O’Malley and subsequently enhanced by Levac and colleagues, 22 alongside adherence to the PRISMA-ScR reporting standards for scoping reviews. 23 The review aimed to comprehensively map existing literature examining the use of big data and machine learning techniques in the development of predictive models for the early detection of diabetes and cardiovascular diseases. A structured and systematic approach was adopted to promote methodological rigor while ensuring clarity, transparency, and reproducibility in the identification, selection, and synthesis of relevant evidence.
2.2. Identifying the research question
A well-defined research question was established to inform and direct each stage of the scoping review process. The question was structured using the Population–Concept–Context framework as recommended by the Joanna Briggs Institute. 24 The population of interest included individuals who were either at risk of, or living with, diabetes and cardiovascular diseases. The central concept focused on the application of big data and machine learning techniques for predictive modeling, while the context covered a wide range of healthcare settings, such as primary care, secondary care, and community-based services. Accordingly, the guiding research question was formulated as: What is the current landscape of predictive modeling using big data and machine learning for the early identification and prediction of diabetes and cardiovascular diseases?
2.3. Eligibility criteria
Inclusion and exclusion criteria were established as a priori to ensure consistency during the screening process. Studies were eligible for inclusion if they (1) involved predictive modeling for early detection of diabetes or cardiovascular conditions, (2) applied machine learning techniques, (3) utilized large-scale, longitudinal, high-dimensional, or multi-source health datasets (e.g., EHRs, wearable sensor data, population health databases, or genomic repositories), and (4) were published in peer-reviewed scientific sources, including journal articles, conference proceedings, and book chapters. Studies focusing on algorithm development, model validation, real-world implementation, or comparative performance analysis were all included. Both clinical and population-based research articles were eligible, including experimental, quasi-experimental, observational, and modeling studies. Exclusion criteria included editorials, letters to the editor, conference abstracts without full-text availability, dissertations, and non-English articles. Studies that applied traditional statistical methods without the use of machine learning algorithms were also excluded, unless explicitly compared against ML techniques.
2.4. Information sources and search strategy
An extensive literature search was undertaken across several major electronic databases to identify studies relevant to the review objectives. The databases searched were PubMed, IEEE Xplore, Scopus, Web of Science, and CINAHL, selected for their strong coverage of health sciences, biomedical research, and data-driven technologies. To enhance completeness, the reference lists of all included studies and pertinent review articles were also hand-searched to capture additional eligible publications not retrieved through database searching. Because the review focused on emerging machine learning methodologies, peer-reviewed conference proceedings indexed in databases such as IEEE Xplore were also considered eligible when they met all inclusion criteria. In addition to database searching, reference lists of included studies and relevant review articles were manually screened to identify potentially eligible studies. Records identified through hand-searching were incorporated into the screening process and are reflected in the overall study selection totals presented in the PRISMA flow diagram.
Search strategy.
2.5. Study selection process
All citations retrieved from the database searches were imported into reference management software to facilitate organization and screening. Automated filtering functions within the reference management software were used to identify duplicate records and clearly ineligible citations based on predefined criteria (e.g., document type and indexing information) prior to manual screening. Duplicate records were identified and removed prior to evaluation. Study selection was then carried out in two sequential stages. Initially, titles and abstracts were independently reviewed by two authors based on the predetermined inclusion and exclusion criteria. Articles deemed potentially relevant proceeded to the second stage, where full texts were examined in detail to determine final eligibility. Any discrepancies arising during either screening phase were addressed through discussion, and consensus was achieved with the involvement of a third reviewer when required. Screening decisions were conducted independently by two reviewers, and disagreements were resolved through discussion with a third reviewer when necessary. Because the review followed a consensus-based screening approach, formal inter-reviewer agreement statistics were not calculated. The study selection procedure was systematically recorded and illustrated using a PRISMA-ScR flow diagram to ensure transparency and reproducibility of the review process (Figure 1). PRISMA flow diagram.
2.6. Data extraction
A structured data extraction form was developed and pilot-tested before formal data charting. The form was designed to capture essential information such as authorship, publication year, study type, target population, type and source of big data, ML techniques used, disease prediction approach, performance metrics, validation methods, implementation context, and reported limitations or challenges. Additional fields included whether the study addressed ethical considerations or model interpretability. Data extraction was conducted independently by two reviewers to ensure accuracy and comprehensiveness. Periodic meetings were held to compare data entries and resolve discrepancies. Adjustments to the extraction template were made iteratively to accommodate new themes that emerged during the review process. All extracted data were organized in a spreadsheet for subsequent thematic analysis and synthesis.
2.7. Data synthesis and analysis
Given the heterogeneity in study designs, data sources, and ML techniques across the included studies, a narrative synthesis approach was adopted. Quantitative meta-analysis was not feasible due to the variability in outcome reporting and model performance metrics. Instead, the findings were synthesized thematically, focusing on trends in ML model development, data integration approaches, feature engineering techniques, and implementation outcomes. To synthesize this evidence, a conceptual, review-derived thematic model was developed (Figure 2) that maps the eight interrelated stages emerging from the reviewed literature: data collection from diverse sources, data preparation for feature engineering, model development using supervised and deep learning techniques, model evaluation for accuracy and reliability, clinical relevance in ensuring fairness across populations, transparency to support interpretability, real-world use through integration into clinical workflows, and ethical oversight as a cross-cutting governance layer. Each stage reflects a recurrent pattern identified across the 66 included studies, and the model is intended to illustrate a conceptual pipeline and related challenges rather than a validated implementation framework. Studies were grouped according to disease type (diabetes, CVD, or both), ML methodology (e.g., supervised learning, ensemble models, deep learning), and data type (structured, unstructured, multi-modal). Comparative table 2 were developed to illustrate variations in model performance, risk factor importance, and population settings. Furthermore, implementation barriers and ethical concerns such as bias, explainability, and patient consent were mapped and discussed qualitatively. Thematic model for early prediction of diabetes and cardiovascular diseases. (Caption: Conceptual, review-derived thematic model illustrating the stages commonly identified in predictive modeling studies for diabetes and cardiovascular diseases. The figure summarizes data collection, preparation, model development, evaluation, clinical relevance, transparency, real-world use, and ethical oversight). MMAT quality assessment.
2.8. Quality appraisal
A basic appraisal of methodological robustness was conducted to contextualize the findings and identify common strengths and weaknesses across the evidence base. The included studies were evaluated using the MMAT for studies employing machine learning in health settings. 91 Items assessed included clarity of research objectives, appropriateness of machine learning techniques, handling of missing data, validation procedures, and transparency in reporting results. Quality assessment was conducted independently by two reviewers. Any disagreements in scoring were resolved through discussion and consultation with a third reviewer when necessary. Studies were categorized as high, moderate, or low quality based on the number of appraisal criteria fulfilled. Studies meeting four or five criteria were classified as high quality, those meeting two or three criteria were classified as moderate quality, and those meeting fewer than two criteria were classified as low quality.
Studies rated as high quality generally demonstrated clearer methodological reporting, stronger validation procedures, and more comprehensive discussion of limitations and implementation considerations. In contrast, lower-quality studies frequently lack transparent validation strategies, detailed reporting of model development processes, or discussion of clinical applicability. Consistent with scoping review methodology, quality appraisal was conducted to provide context for interpreting the evidence rather than to exclude studies. Findings from the MMAT assessment informed interpretation of methodological strengths and limitations throughout the review and were considered when discussing the robustness, generalizability, and translational potential of the reported predictive models.
3. Results
3.1. Overview of included studies
Overview of included studies.
Performance metrics, validation methods, implementation context, and reported limitations or challenges.
3.2. Application of machine learning techniques in predictive modeling
The use of machine learning for the early identification of diabetes and cardiovascular diseases has progressed considerably, moving from simple predictive classifiers toward more sophisticated ensemble and deep learning models.25,26 According to the included studies, supervised learning techniques such as logistic regression, deciding trees, support vector machines, and random forests were most frequently applied.27,28 These approaches were commonly used to estimate disease risk or predict future onset by analyzing patterns within historical clinical and demographic data.29,30 Ensemble algorithms, particularly gradient boosting and extreme gradient boosting, were widely adopted because of their strong performance with complex, nonlinear relationships and their ability to manage class imbalance in health datasets.31,32
More recent research has increasingly explored deep learning methods, including convolutional neural networks and recurrent neural networks, due to their capacity to analyze unstructured and longitudinal data.33,34 Such models were applied to sources like clinical narratives, medical images, and time dependent physiological measurements, allowing for richer feature representation and, in some cases, higher predictive accuracy than traditional models relying solely on structured inputs.35,36 Despite these advantages, deep learning techniques were less commonly implemented, largely because they require substantial data volumes, advanced computational infrastructure, and pose challenges related to model transparency and interpretability.37,38 Overall, the choice of machine learning approach was shaped by the nature of the available data, the need for explainable outputs in clinical settings, and access to well annotated datasets.
3.3. Big data sources and data integration strategies
The predictive models analyzed in this review relied on a diverse array of big data sources, underscoring the growing reliance on data-driven approaches in healthcare. Electronic health records (EHRs) were the predominant data source, offering a rich mix of demographic, clinical, and laboratory data.39,40 Many studies leveraged longitudinal data from EHRs to capture temporal trends in patients’ health, enabling more accurate disease progression modeling.41,42 In addition to EHRs, several studies incorporated data from wearable sensors and mobile health applications, reflecting the increasing integration of continuous patient monitoring in predictive analytics.43,44
Some models were built using multi-source datasets that included claims data, genomic databases, population health surveys, and social determinants of health.45,46 These studies often used data fusion techniques to combine structured and unstructured information, enhancing the comprehensiveness of the predictive models.47,48 However, challenges related to data harmonization, missing values, and format inconsistencies were frequently reported. 49 To address these issues, various preprocessing techniques, including data imputation, normalization, and dimensionality reduction, were applied. 50 The integration of heterogeneous data types was found to improve model performance, particularly in identifying at-risk populations that traditional clinical indicators might overlook. 51
3.4. Feature selection and engineering practices
Feature selection and engineering emerged as critical steps in the development of accurate and generalizable predictive models. The reviewed studies applied a range of methods to identify relevant features from large datasets.52,53 Commonly used techniques included correlation-based filtering, recursive feature elimination, principal component analysis, and information gain 54 and 55. These methods were employed to reduce noise, enhance model performance, and mitigate the risk of overfitting. 56 Several studies emphasized the importance of domain knowledge in the selection of clinically meaningful features, such as HbA1c levels, blood pressure readings, cholesterol profiles, family history, and lifestyle behaviors.27,29,32
Moreover, temporal feature engineering was particularly prominent in studies using time-series data from EHRs and wearable devices. 57 Techniques such as sliding window analysis, lag feature construction, and sequence encoding were used to capture disease trajectory and detect subtle early warning signals.58,59 A subset of studies experimented with automated feature engineering using deep learning architectures or feature generation libraries. 60 These approaches allowed the discovery of latent patterns without manual intervention, but their results were often less interpretable to clinical stakeholders.61,62 Overall, successful feature engineering was a strong predictor of model performance, especially in real-world applications.
3.5. Model performance metrics and validation approaches
Across the included studies, model effectiveness was assessed using a broad set of indicators that captured both predictive accuracy and the capacity to generalize beyond the training data.33,37 Frequently reported measures included accuracy, sensitivity, specificity, precision, F1 score, and the area under the receiver operating characteristic curve (AUC ROC).40,79 Among these, AUC ROC emerged as the most applied metric because it provides a comprehensive summary of a model’s discrimination ability across different classification thresholds. 80 In investigations focused on early disease identification, greater emphasis was often placed on sensitivity so that individuals at elevated risk were not overlooked, even when this resulted in lower specificity.82,85
Approaches to model validation differed considerably across studies, largely influenced by methodological choices and data constraints. 63 Internal validation strategies, particularly k fold cross validation and bootstrap resampling, were widely adopted to examine model stability and reduce overfitting.64,65 Across the evidence base, internal validation approaches substantially outnumbered external validation methods, highlighting a persistent gap between model development and broader generalizability assessment. While internal validation provides important evidence of model consistency within a dataset, it offers limited insight into performance across different populations and healthcare settings.
In contrast, validation using external datasets was relatively uncommon. Only a small number of studies evaluated model performance on data drawn from different clinical settings or populations, limiting confidence in broader applicability.66,67 Studies that incorporated external validation generally provided stronger evidence of model transportability and clinical applicability than studies relying solely on internal validation.28,39,50 Those studies that incorporated both internal and external validation tended to present stronger evidence of model reliability and more frequently reported confidence intervals and tests of statistical significance.68,69 In addition, some studies extended performance assessment by using calibration plots and decision curve analysis, allowing evaluation of potential clinical usefulness rather than relying solely on discrimination-based metrics.70,71 Overall, the findings suggest that although predictive accuracy was frequently reported, rigorous validation beyond the development dataset remains limited, representing an important methodological challenge for future clinical implementation.
3.6. Disease focus and population diversity
The reviewed literature displayed varying degrees of focus between diabetes, cardiovascular diseases, and comorbid conditions. A significant portion of the studies focused on predicting type 2 diabetes due to its global prevalence and modifiable risk factors.25,34 Models for diabetes prediction frequently utilize metabolic indicators and behavioral data, including diet, exercise, and glucose levels.36–38 In contrast, studies targeting cardiovascular diseases often incorporated electrocardiogram signals, echocardiography reports, and biomarkers such as troponin and C-reactive protein 72 and 73. Some studies aimed to predict major adverse cardiovascular events (MACE) such as myocardial infarction or stroke within specified time horizons.74,75
While several studies addressed both diabetes and CVDs as interrelated conditions, fewer examined them in a truly integrated modeling framework.76–78 Population diversity across studies was uneven, with many models trained on data from single-country sources or institution-specific datasets. 79 As a result, there were concerns about geographic and demographic bias. Only a few studies specifically addressed underrepresented groups, such as ethnic minorities or rural populations.48,80,81 Those that did reported disparities in model performance, highlighting the importance of inclusive training datasets and subgroup analysis for equitable predictive healthcare. 82
3.7. Implementation challenges and real-world applications
Despite the promising performance of many models in research settings, real-world implementation remained limited. Several studies discussed challenges in integrating predictive models into clinical workflows, citing barriers such as lack of interoperability with existing health information systems, resistance from healthcare providers, and insufficient infrastructure for real-time analytics.43,83,84 Models that required manual data input or were computationally intensive were less likely to be deployed in resource-constrained environments. 85 Moreover, the absence of standardized protocols for model deployment, monitoring, and updating was a recurrent theme across studies. 86 Some implementation case studies demonstrated feasibility, particularly in hospital settings where EHR integration was more advanced.82,87 In such cases, predictive models were embedded into clinical decision support systems to alert providers of patients at elevated risk. 88 However, these examples remained exceptions rather than the norm. Most studies concluded with a call for pilot testing, stakeholder engagement, and iterative co-design with clinical teams to ensure alignment with user needs.89,90 The translation of predictive modeling research into practical applications continues to be a major area for development.29,36
3.8. Interpretability and explainability of models
Interpretability emerged as a recurring theme in the reviewed literature, particularly in relation to advanced algorithms such as neural networks and ensemble-based models.52,60 Concerns were frequently raised by clinicians and policy stakeholders about the opaque nature of these approaches and the challenges this creates for informed clinical judgment and policy formulation.41,45 In response, several studies applied post hoc explanation methods, including SHAP, LIME, and feature importance visualizations.59,74,81 These techniques enabled clearer insight into how individual predictors influenced model outputs and helped translate complex computational processes into more understandable explanations for end users.32,50,63
The literature consistently linked interpretability with trust, usability, and eventual adoption, especially in contexts involving high risk clinical decisions.26,81 Studies that incorporated intuitive visualization tools, simplified decision tree representations, or rule-based logic were more successful in engaging users without technical expertise.39,65,73 In some cases, hybrid modeling strategies were adopted, pairing transparent baseline models with more sophisticated algorithms to balance predictive performance with interpretability.49,57,76 Despite these advances, the use of explainability methods was inconsistent, and many studies offered limited reflection on the risks associated with black box decision making in healthcare.28,41,70 This shortcoming highlights the need for future research to prioritize inherently interpretable modeling approaches or systematically embedded strong explanation frameworks to support safe and responsible clinical use.55,62,79
3.9. Ethical considerations and algorithmic bias
Ethical dimensions of applying machine learning in predictive healthcare were acknowledged across many of the reviewed studies, although discussion was often brief and lacked critical depth.27,31,88 Commonly referenced concerns included data privacy, informed consent, fairness, and potential bias in algorithmic decision making.36,65,78 While studies relying on electronic health records or data from wearable technologies frequently noted approval from ethics committees or adherence to data protection regulations, detailed explanations of consent processes, data stewardship, and long-term governance arrangements were seldom provided.29,38,45,54,73,87
Algorithmic bias was highlighted as one of the most pressing ethical challenges. Several investigations observed uneven model performance across population subgroups defined by characteristics such as age, sex, or ethnic background.81,90 In some cases, models developed using data from urban clinical settings showed reduced accuracy when applied to rural or underserved populations.69,85 Despite recognition of these disparities, relatively few studies implemented concrete mitigation strategies, such as rebalancing training datasets, conducting subgroup specific calibration, or adopting fairness aware modeling techniques.50,71,86 The absence of routine bias audits and equity focused evaluations represents a significant gap in the existing evidence base.26,43,59 Addressing ethical considerations more systematically is essential to ensure that predictive models are credible, equitable, and suitable for wider deployment across diverse healthcare contexts. Importantly, ethical challenges were closely linked to modeling choices, as models developed from geographically restricted datasets or lacking explainability mechanisms were more vulnerable to concerns regarding fairness, accountability, and clinical trust.
4. Discussion
The exploration of predictive modeling using complex healthcare datasets and machine learning (ML) for the early diagnosis of diabetes and cardiovascular diseases (CVDs) reveals a dynamic yet fragmented research landscape. While several studies leveraged large-scale national or longitudinal databases, others relied on smaller institutional datasets, highlighting variability in the scale and scope of data sources across the evidence base. A striking observation is the divergence between the technical sophistication of the models and their real-world applicability. 92 Although the reviewed studies demonstrate significant progress in algorithm development and data integration, these advances largely represent multifactorial prediction models that combine multiple risk variables to estimate disease probability and often remain confined to research environments.93,94 The field appears to be driven largely by methodological innovation rather than end-user readiness. 3 This tension highlights a gap in translational focus, where model accuracy and performance metrics may overshadow critical concerns such as clinical usability, system interoperability, and healthcare provider engagement. 95 Moreover, a disproportionate emphasis on technical metrics without parallel consideration for contextual relevance restricts the models’ broader adoption across healthcare settings. 96
Another prominent theme that emerged relates to the variability in data quality and representation. Many models were developed using electronic health records and other structured data sources, but they often lacked external validation or were trained on homogenous populations. 97 This raises serious questions about equity, inclusivity, and the potential for algorithmic bias. 14 While some studies acknowledged performance disparities across demographic groups, comprehensive mitigation strategies were seldom incorporated.98,99 The absence of subgroup-specific analyses further reinforces the concern that predictive tools might inadvertently reinforce existing healthcare inequalities. 100 Equitable model development necessitates inclusive datasets, consistent subgroup evaluations, and the application of fairness-aware algorithms. 101 Furthermore, the underutilization of social determinants of health, which substantially influence chronic disease risk profiles, limits the models’ capacity to fully capture patient realities in varied social contexts and to progress toward true systems-based modeling that represents dynamic interactions, feedback loops, and multiscale biological, behavioral, social, and environmental mechanisms. 102
The review also underscores the critical yet often underdeveloped domain of interpretability and explainability in ML applications. In the clinical environment, the trustworthiness of a model is directly tied to its transparency. 103 While advanced models such as deep learning architectures offer predictive precision, their black-box nature creates a barrier for clinician acceptance. 20 The selective use of explainability tools like SHAP and LIME in some studies represents a step forward, but broader adoption remains inconsistent. 104 This lack of uniformity in interpretability practices weakens the potential for widespread clinical integration. 105 A future-oriented vision for predictive modeling should involve co-design with clinicians, employing models that provide actionable insights without compromising understandability. 106 Building hybrid systems that integrate interpretable baseline models with complex secondary layers could bridge this divide, enabling both accuracy and usability. 107
Furthermore, the review highlights a pressing need for stronger implementation of science in this domain. Despite promising technical outcomes, most studies lacked detailed discussions on how these models could be embedded into routine healthcare workflows. 108 Implementation challenges, including interoperability issues, provider resistance, data governance constraints, and infrastructure limitations, remain under-addressed. 109 Moreover, few studies discussed mechanisms for real-time model monitoring, adaptation to shifting data patterns, or post-deployment audits.110,111 The limited focus on these aspects suggests that the field is still in a proof-of-concept phase rather than being ready for operational scaling. 18 Advancing this research agenda requires a shift toward multidisciplinary collaboration that includes clinicians, data scientists, ethicists, health policy experts, and patients. 112 Such partnerships can help evaluate whether predictive models developed in research settings can be safely and effectively adapted to real-world healthcare environments. However, given the heterogeneity of study outcomes, limited external validation, and sparse implementation evidence, current findings should be interpreted as demonstrating potential rather than confirming readiness for large-scale clinical deployment. 113
5. Implications for practice and policy
The integration of predictive modeling into clinical practice offers a transformative shift from reactive to proactive healthcare, particularly in the early identification of diabetes and cardiovascular diseases. Healthcare professionals can leverage machine learning–driven decision-support tools to detect at-risk individuals before the manifestation of clinical symptoms, allowing timely interventions that can reduce disease progression and associated complications. Embedding predictive models within electronic health record systems can streamline workflows, support triage protocols, and tailor personalized care plans. However, to realize these benefits, clinical practitioners require appropriate training on model interpretability and ethical data usage. Health institutions should also prioritize interdisciplinary collaboration to ensure that technical solutions are aligned with practical care delivery realities.
From a policy perspective, the findings underscore the urgent need for regulatory frameworks that support the ethical deployment of machine learning technologies in public health. Policymakers should mandate fairness audits and subgroup performance analyses to prevent algorithmic bias that may deepen existing health disparities. Investment in data infrastructure, especially in low-resource and rural settings, is critical to ensure that predictive analytics serve diverse populations equitably. National health strategies should encourage the use of inclusive datasets, transparency in algorithm design, and real-time model evaluation. Additionally, policies that promote patient engagement, data privacy, and continuous monitoring of algorithmic outcomes will be essential to building trust and accountability. By aligning predictive modeling initiatives with robust governance and inclusive health goals, stakeholders can ensure that technological innovations translate into measurable improvements in population health outcomes.
6. Research gaps and future directions
The findings of this review reveal several gaps that warrant further investigation. First, the limited use of external validation across studies raises questions about the generalizability of current models. Future research should prioritize cross-institutional collaborations to access diverse datasets and enhance model robustness. Second, there is a need for standardized reporting guidelines for predictive modeling studies in healthcare, especially concerning data preprocessing, performance metrics, and ethical compliance. Adherence to such standards would improve reproducibility and comparability across studies.
Additionally, few studies explored the integration of social and environmental determinants of health into predictive models. Considering the multifactorial nature of diabetes and CVDs, future work should incorporate broader contextual variables to better reflect real-world risk factors. Although many reviewed models incorporated diverse clinical, demographic, and behavioral variables, most remained fundamentally predictive rather than systems oriented. Future research should explore system-level frameworks that integrate multiscale biological, behavioral, environmental, and social determinants while accounting for dynamic interactions and feedback mechanisms across the disease continuum. Such approaches may improve understanding of the complex network-based mechanisms underlying diabetes and cardiovascular diseases and support the development of more holistic and clinically meaningful prediction systems.
Longitudinal validation and post-deployment monitoring remain underdeveloped areas that could ensure model stability over time. Moreover, engaging patients and healthcare providers in the design and evaluation of predictive tools is essential for fostering trust, usability, and sustainability. These directions represent essential steps in bridging the gap between theoretical modeling and impactful healthcare solutions.
7. Conclusion
This scoping review highlights the growing interest in the application of big data and machine learning for the early identification and risk prediction of diabetes and cardiovascular diseases. Although many studies reported promising predictive performance, the evidence base remains heterogeneous, with substantial variation in disease outcomes, data sources, machine learning approaches, validation methods, and implementation settings. The reviewed literature demonstrates the potential of advanced analytical techniques, including ensemble learning and deep learning architectures, to leverage diverse data streams such as electronic health records, wearable sensor data, and population-level datasets for risk stratification and disease prediction.
Nevertheless, translation into routine clinical practice remains limited due to persistent challenges, including heterogeneous data sources, insufficient external validation, limited model transparency, and difficulties in embedding these tools within established healthcare systems. In addition, inadequate consideration of population diversity and fairness raises concerns about the potential reinforcement of existing health inequalities. The relatively small number of studies reporting external validation, prospective evaluation, or real-world implementation further limits confidence in the generalizability and clinical applicability of many existing models.
Closing the gap between technical advancement and practical use will require a stronger emphasis on collaboration across clinical, technical, and policy domains, alongside more transparent and standardized reporting practices. Future research should focus on real-world testing and validation, while ensuring that predictive models are interpretable, equitable, and aligned with clinical decision-making processes. Coordinated efforts among policymakers, healthcare organizations, and data scientists are essential to establish governance structures that protect privacy, address bias, and guide responsible implementation. Therefore, while current evidence suggests considerable potential for machine learning–based predictive modeling in diabetes and cardiovascular disease prevention, claims regarding clinical effectiveness and translational readiness should be interpreted cautiously until supported by stronger external validation, implementation research, and evidence from diverse healthcare settings. By integrating methodological innovation with ethical accountability and equity-oriented goals, predictive modeling may become a valuable component of future public health and clinical practice.
Supplemental material
Supplemental material - Predictive modeling for early detection and prediction of diabetes and cardiovascular diseases using big data and machine learning
Supplemental material for Predictive modeling for early detection and prediction of diabetes and cardiovascular diseases using big data and machine learning by Fahad Ahmed, Towsif Alam, Moustaq Karim Khan Rony, Afia Fairooz Tasnim, Mohammad Hossain, Durga Shahi, Arif Hosen, Adib Hossain, Mia Md Tofayel Gonee Manik in DIGITAL HEALTH.
Footnotes
Acknowledgements
The authors are deeply grateful to the Miyan Research Institute, International University of Business Agriculture and Technology, Dhaka, Bangladesh.
Author Contributions
Conceptualization: F.A., T.A., M.K.K.R., A.F.T. Methodology: M.K.K.R., T.A., A.H. Literature Search: F.A., A.F.T., M.H., D.S., A.H., M.M.T.G.M. Data Curation: F.A., T.A., A.F.T., M.H., D.S., A.H., M.M.T.G.M. Formal Analysis: M.K.K.R., A.H. Investigation: F.A., A.F.T., M.H., D.S., A.H. Writing – Original Draft: F.A., T.A., A.F.T. Writing – Review and Editing: F.A., M.K.K.R., M.H., D.S., A.H., M.M.T.G.M. Visualization: F.A., T.A., M.M.T.G.M. Supervision: M.K.K.R., F.A. Project Administration: M.K.K.R., T.A. Resources: M.H., D.S. Validation: A.H., F.A. All authors contributed substantially to the work, participated in drafting or critically revising the manuscript, and approved the final version.
Funding
The authors received no financial support for the research, authorship, and/or publication of this article.
Declaration of conflicting interests
The authors declared no potential conflicts of interest with respect to the research, authorship, and/or publication of this article.
Data Availability Statement
No new data were created or analyzed in this study.
Supplemental material
Supplemental material for this article is available online.
References
Supplementary Material
Please find the following supplemental material available below.
For Open Access articles published under a Creative Commons License, all supplemental material carries the same license as the article it is associated with.
For non-Open Access articles published, all supplemental material carries a non-exclusive license, and permission requests for re-use of supplemental material or any part of supplemental material shall be sent directly to the copyright owner as specified in the copyright notice associated with the article.
