Abstract
Background
With global aging, cognitive impairment, including Alzheimer's disease and related dementias, has become a critical public health challenge, driving the need for convenient screening tools to facilitate early intervention.
Objective
This study aimed to develop an efficient and noninvasive risk assessment model for identifying potential cognitive impairment in the elderly using machine learning algorithms based on comprehensive geriatric assessment (CGA).
Methods
We included 1410 participants aged 50 and older from geriatric clinics and community. Feature selection was performed on the CGA indicators using a combination of expert knowledge and machine learning. Logistic regression (LR), naive Bayes, support vector machines, neural networks, and random forests were comprehensively evaluated based on common classification performance metrics. The optimal machine learning algorithms and feature subset are used to construct the final prediction model. Shapley Additive exPlanations (SHAP) was used to explain the model.
Results
Thirteen noninvasive predictors were identified, including ‘Bathing’, ‘Age’, ‘Caregiver’, ‘Sleep duration’, ‘Homekeeping’, ‘Right Ear’, ‘Stand BWEO’, ‘Hobbies’, ‘Focusing Difficulty’, ‘UI effect’, and ‘Housework’. The LR model performed best on the test set, with an AUC of 0.877 and high accuracy (0.815), sensitivity (0.767), and specificity (0.827). The SHAP results illustrated the role of these key features in cognitive impairment, which is highly consistent with clinical knowledge.
Conclusions
This study identifies convenient, noninvasive predictors for screening-oriented prediction of cognitive impairment, develops an efficient machine learning model, and employs SHAP analysis for interpretation. This facilitates widespread screening, providing guidance for early detection and intervention in high-risk populations.
Introduction
Cognitive impairment, including Alzheimer's disease and related dementias, represents a critical global public health challenge as the population ages. It is characterized by a decline in memory, language, and problem-solving abilities, significantly impacting daily life and quality of life. 1 With an estimated 50 million people affected worldwide, the progressive nature of these conditions imposes a substantial economic burden on healthcare systems and societies.2,3
A key strategy to mitigate the impact of cognitive impairment is the development of efficient screening models for the early detection of high-risk individuals, thereby facilitating preventive interventions. Traditional approaches for risk identification have primarily relied on meta-analytic methods and conventional regression models. For instance, a landmark meta-analysis published in The Lancet identified 12 potential risk factors and established a life-course model of risk factors for dementia prediction. 4 However, these studies primarily focus on identifying risk factors without systematically raking their importance or quantifying integrated risk profiles.
Machine learning has emerged as a promising tool for disease prediction. Unlike traditional models, machine learning can not only compute specific risk values but also rank features by their relative importance. However, most existing machine learning models for cognitive impairment depend on data that are invasive, costly, or difficult to obtain, such as neuroimaging (CT, MRI), cerebrospinal fluid biomarkers, or genetic tests.5–7 Some studies have simplified the prediction methods, but still require invasive blood tests or rely on common clinical indicators and professional neuropsychiatric scale assessments, which limits their practicality for widespread self-screening in community settings.8,9
The comprehensiveness of candidate predictors is crucial to the performance of any predictive model. Many previous studies have been constrained by the scope of variables and are insufficient to cover all potential factors contributing to cognitive impairment. Comprehensive Geriatric Assessment (CGA) is a multidimensional evaluation method tailored for the elderly, with indicators that more thoroughly encompass medical, psychological, social, and functional aspects. There are no research utilizing CGA indicators as candidate predictors for the risk prediction of cognitive impairment. Given these challenges, this study aims to select a more comprehensive set of candidate predictors and develop a more accessible, simplified yet accurate model for predicting the risk of cognitive impairment. To improve prediction accuracy, we innovatively employed over 200 indicators from CGA as candidate predictors, ensuring a thorough coverage of predictive risk factors. Through clinical expertise, intelligent feature selection and machine learning algorithms, we identified 13 key features that are easily accessible and noninvasive. This significantly lowers the barrier for public self-screening, enhances screening accuracy and accessibility, and facilitates early detection and intervention in high-risk groups for cognitive impairment. Our approach represents a novel stride in the screening and prevention of dementia, offering an efficient and user-friendly tool for early identification of cognitive impairment risks. However, the predictive model constructed in this study is primarily designed to efficiently screen for individuals with potential cognitive impairment in the current state, rather than predicting future cognitive decline.
Methods
The overall modeling procedure of this research is shown in Figure 1. Using the open-source ML and data mining toolkit Orange v3.33 (https://orangedatamining.com), accompanied with Python programming language, we conducted the tasks of data cleaning, preprocessing, statistical analysis, and model training and validation.

The modeling and variable screening procedure.
Inclusion criteria
This study ultimately included 1410 participants aged 50 years and older. All participants were recruited between 2016 and 2019 from the Geriatrics Department of the First Affiliated Hospital of Chongqing Medical University and its surrounding communities. Each participant completed a comprehensive geriatric assessment, and all assessment data were complete. The study was approved by the hospital ethics committee, and informed consent was obtained from each participant before any data collection.
Cognitive function assessment
In the present study, cognitive impairment was defined as encompassing both mild cognitive impairment (MCI) and dementia patients. According to the Petersen standard, diagnosed as MCI: (1) Both subjective and objective examinations show mild cognitive impairment; (2) Daily living abilities are not affected; (3) Does not meet the diagnostic criteria for dementia; (4) Other diseases that may cause cognitive decline have been excluded; (5) Mini-Mental State Examination (MMSE) score: 21 ≤ score ≤ 26 indicates mild cognitive dysfunction.
Diagnosed as dementia according to the Diagnostic and Statistical Manual of Mental Disorders (DSM-V): (1) For patients with previously normal intelligence, subsequent acquired cognitive decline (memory, executive, language or visuospatial ability impairment) or mental behavioral abnormalities occur; (2) Affecting work ability or daily life; (3) And cannot be explained by delirium or other mental disorders; (4) MMSE score: illiterate ≤ 17 points, primary school education ≤ 20 points, secondary school education (including technical secondary school) ≤ 22 points, university education (including junior college) ≤ 23 points, defined as dementia.10–13
Exclusion criteria
Exclusion criteria are as follows: (1) altered consciousness and delirium, (2) pseudodementia attributable to depression and other mental diseases, (3) transient disturbances in consciousness and cognitive decline caused by medications, toxins, or other factors, (4) individuals who were unable to cooperate, unwilling to participate, or had incomplete clinical data.
Candidate variables
This study comprehensively collected candidate predictors related to cognitive impairment based on CGA, encompassing a total of 228 features across 19 evaluative aspects (Table 1). These aspects included basic demographic information, fundamental measurements, lifestyle and behavioral habits, life preferences, sleep status, daily living abilities, visual and auditory impairments, communication skill barriers, urinary incontinence, gastrointestinal issues, multiple comorbidities, physical health, swallowing function, cognitive function, the clock drawing test, and various neuropsychiatric scales such as the Geriatric Depression Scale-5 (GDS-5), Generalized Anxiety Disorder-7 (GAD-7), and the Patient Health Questionnaire-9 (PHQ-9). All these assessments were conducted by clinical professionals. The evaluation aspects and features number for 228 features are shown in Supplemental Table 1.
The questionnaire for data collection.
Variable screening
Based on the expert knowledge from our clinicians, we conducted an initial variable selection as follows. Firstly, we removed clinically meaningless variables (n = 4), including number, evaluation time, evaluation type and registered residence address. Subsequently, variables known to introduce bias during the data collection process (medical insurance region, ethnic group, hospitalization expenses, medical expenses and native place) were excluded (n = 5). Then, from the perspective of prediction task, we excluded variables that significantly indicated cognitive impairment and/or unsuitable as predictor factors, such as significant memory loss and communication disorders (n = 41). And finally, we removed variables that required professional assessment and not easily obtainable, such as Drinking tests, Grip strength, and total scores from various assessment scales (n = 18).
The data preprocessing was performed as follows: (1) excluded variables with a missing rate exceeding 50% (e.g. smoking cessation age) (n = 4); (2) certain variables were discretized or derived through arithmetic operations on other variables, yielding novel variables that are not directly available (e.g. wake up time period, pulse pressure difference, BMI2) (n = 57); (3) encoding of ordinal categorical variables in ascending order based on their influence relation; (4) performing one-hot encoding on nominal categorical variables.
To screen the remaining variables, ANOVA and Chi-squared test were used to conduct a significance test for the overall distribution differences between continuous and categorical variables respectively with the target variable. Considering non-significant (p ≥ 0.05), we removed 80 categorical variables but no numerical variables. Secondly, to reduce the impact of multicollinearity on model stability, the correlation between variables were profiled through Spearman correlation analysis. A Spearman correlation coefficient greater than 0.6 was used as a criterion for judgment, we eliminated variables with low correlation. This process removed 61 variables and finally retained 71 variables that had relatively independent effects on the target variable for the subsequent model development and validation procedures. The variable screening steps are detailed in Supplemental Table 2, and the list of 71 retained variables after screening is provided in Supplemental Table 3.
Model development and validation
This study aims to develop a dementia risk prediction model, which undertakes the classification task of discerning whether a participant belongs to Class 0 (Non-Cognitive Impairment) or Class 1 (Cognitive Impairment). We conducted the following steps under Python (v3.9) with the scikit-learn library (v1.0.2).
We divided the total 1410 participants into a training set (n = 1199) for developing the model using data collected from 2016 to the end of 2018, and a test set (n = 211) for evaluating the model using data collected in 2019. Before further analysis, we imputed missing values using the k-Nearest Neighbors (kNN) algorithm (with k = 3), and all continuous variable values were normalized to the range [0,1].
In order to determine the appropriate machine learning algorithm for developing the prediction model, we conducted 10-fold stratified cross-validation on the training set based on different machine learning algorithms including logistic regression (LR), naive Bayes (NB), support vector machines (SVM), neural networks (NN), and random forests (RF). The 10-fold stratified cross-validation strategy was to divide the training set into 10 folds (i.e. smaller subsets), each with the same ratio between Class 0 and 1 sample; and then select 9 folds as the internal training set and the remaining 1 fold as the internal validation set in turns. For each algorithm, we constructed corresponding prediction models and evaluated their performance using multiple metrics, including the area under the receiver operating characteristic curve (AUC), accuracy, sensitivity, specificity, positive predictive value (PPV), negative predictive value (NPV), F1 score and Brier score. The calibration performance of each model was further assessed using calibration plots to examine the agreement between predicted probabilities and observed outcomes. The 10-fold cross-validation results for all these metrics represented the overall performance of each machine learning algorithm, which were then used to determine the optimal model for subsequent analysis.
To build a parsimonious model, we applied feature selection approaches (combining features ranking and sequential forward selection) to obtain a smaller variable subset using the training set. In the first step, we used four feature importance ranking methods including Gain ratio, Gini, ReliefF and feature selection based on correlation and feedback (FCBF) to rank the variables in descending order. Next, we implemented the sequential forward selection method, adding variables one by one into the model training process based on their features ranking using the above optimal machine learning algorithm. Under the cross-validation scheme, we monitored the changes of the candidate models’ AUC as we gradually incorporated additional variables until the point at which the AUC demonstrated no further statistically significant increase. So, the optimal feature subset was confirmed.
In the end, we conducted the final modeling and optimization process, using the above optimal algorithm and optimal feature subset. Hyperparameters were tuned by grid search and cross-validation to maximize the model performance using the training set. Then, the final model was retrained on the complete training set by using the obtained optimal hyperparameters, and its validation was conducted using the test set.
Model interpretation
To further analyze the cognitive impairment risk prediction model at the variable level, we applied the SHapley Additive exPlanations (SHAP) method. 14 SHAP is a game theoretic approach that can explain the prediction output of the given machine learning model, based on the consistent and locally accurate attribution values. The SHAP feature importance and summary plot were used to draw the SHAP interpretation of importance and contribution to the model and interpret the model results by calculating the contribution of each feature to the predicted results. The SHAP value represents the contribution of one specific feature towards prediction performance: the larger the value, the higher the contribution. The SHAP dependency plot is employed to elucidate the impact of an individual feature on the model's output. Furthermore, the SHAP waterfall plot elucidates the model's influence on the prediction of a particular individual sample. We implemented the above analysis using the Shap library (v0.40.0).
Furthermore, based on the logistic regression model, we evaluated the contribution of each feature through feature decrement and increment analysis. We conducted feature decrement analysis which using an optimal subset model containing all 13 optimal features, and successively removing a single feature to monitor the decrease in AUC. On the other hand, we evaluated the contribution of each feature through feature increment analysis. We used a baseline model containing age and caregiver to build model, and successively adding a single feature from the remaining optimal features to observe the increase in AUC.
Results
Population characteristics
Table 2 shows the characteristics of the participants. Our study involved 625 males (44.33%) and 785 females (55.67%), with a median age of 78 years. According to the previous diagnostic criteria, 420 participants were diagnosed with cognitive impairment (29.8%), and 990 participants had normal cognition (70.2%). There was significant difference between the Normal cognition and cognitive impairment about age, gender, marital status (p < 0.05). In males, the risk of developing cognitive impairments is higher compared to females. With increasing age, the proportion of individuals experiencing cognitive impairments also gradually increases. Additionally, individuals who are widowed and living alone have a higher proportion of cognitive impairments.
Statistical characteristics of participants.
Chi-square test was used to compare the difference between normal and cognitive impairment participants, data is number of subjects (percentage).
**p < 0.05.
Selection of machine learning algorithm
Based on 71 input variables and 1 target variable, we built some preliminary models by applying five machine learning methods (SVM, NB, LR, NN, and RF) to the training set under the cross-validation scheme. Table 3 summarizes the performance metrics of all evaluated models. LR achieved the highest overall accuracy (0.8032) and the lowest Brier score (0.1407), reflecting both strong discrimination and excellent calibration. It also demonstrated the highest specificity (0.9197) with a competitive F1 score (0.6369). NB yielded the highest sensitivity (0.7294) and F1 score (0.6603), suggesting superior detection of positive cases but poorer calibration (Brier score = 0.2243) and lower specificity (0.7798). RF achieved comparable specificity to LR (0.9185) but markedly lower sensitivity (0.4589). SVM and NN exhibited moderate performance with sensitivity around 0.56 and Brier scores of 0.1664 and 0.2181, respectively.
The classification performance metrics of different machine learning methods (training set).
Figure 2 illustrates the model discrimination and calibration performance, where Figure 2(a) displays the merged ROC curves and Figure 2(b) presents the calibration plots. As shown, LR achieved the highest AUC (0.845), confirming its superior discrimination in identifying cognitive impairment. In the calibration plots, LR, SVM, and RF closely followed the ideal diagonal line, consistent with their lower Brier scores, while NB and NN deviated more notably. Specifically, NB tended to underpredict across most probability ranges, whereas NN exhibited instability in intermediate and high probability regions. Considering overall discrimination, calibration, and interpretability, LR demonstrated the most balanced performance. Therefore, we further optimized the LR model by tuning its parameters and selecting a simplified feature subset through various feature selection methods to enhance its classification accuracy and reliability.

(a) The ROC curves and (b) the calibration curves of different machine learning methods (training set).
Feature selection
Feature selection is an important preprocessing step for screening critical factors, which allows extract the prefect subset of features for building a optimum model. We employed a diversified approach to rank the importance of 71 input variables. Next, we used a forward feature selection algorithm to train the LR model through cross-validation and continuous addition of variables until the performance stabilized. Finally, we determine the best features based on the balance of quantity of features and model performance. As shown in Figure 3, the AUC of the model improved with the number of features but stabilized beyond 30 variables. FCBF outperformed other methods, particularly with smaller feature sets, suggesting it better captures potential risk factors for cognitive impairment. When the number of features less than 30, FCBF performed quite well and achieved the best AUC (0.848) with the top 13 variables, including ‘Bathing’, ‘Age’, ‘Right Ear’, ‘Hobbies_X’, ‘Focusing Difficulty’, ‘Sleep duration’, ‘Caregiver_C’, ‘Hobbies_W’, ‘Caregiver_A’, ‘Stand BWEO’, ‘Housework’, ‘UI effect’ and ‘Homekeeping’. A detailed description table including feature names, feature definitions, and feature values is provided in Supplemental Table 4.

The prediction performance of Top-N features based on different feature selection methods (training set).
Parameter optimization and model construction
Based on the 13 most important features, we conducted parameter tuning, e.g. regularization and learning rate, for the LR model by cross-validation on the training set. Using the optimal parameter setting, we retrain the LR model on the whole training set. The final prediction performance on the training set are shown in Supplemental Figure 1, with an AUC of 0.854 and an area under the precision recall curve (AUPRC) of 0.753. The AUC on the test set was 0.877 and the AUPRC was 0.681 (Figure 4(a) and (b)). In addition, the detailed classification performance of the optimized LR model demonstrated robust predictive ability, achieving an overall accuracy of 0.815, with a sensitivity of 0.767 and a specificity of 0.827. The PPV and NPV were 0.532 and 0.933, respectively, while the F1 score reached 0.629, indicating a balanced trade-off between precision and recall. The calibration plot revealed the LR model has a good predictive accuracy between the actual probability and the predicted probability (Figure 4(c)). The Brier score is roughly 0.137, which is between 0.1 and 0.25 and is usually considered good. The decision curve showed that the net benefit of the model decreases with the increase of the threshold probability, and the LR model performs better except for the larger threshold, which also indicates that it has good clinical utility (Figure 4(d)). As a comparison, we also trained another LR model using all 71 features (Supplemental Figure 2), achieving a slightly higher AUC (0.883) and lower AUPRC (0.667). These findings collectively demonstrate that the selected 13 features provided strong discriminative power while maintaining stable generalization performance.

The prediction performance of the logistic regression model with Top-13 features (test set).
Personalized interpretation
SHAP has the advantage of providing both a global explanation of feature importance and reflecting specific local explanation. 14 Therefore, this study employs SHAP to rank the importance of features in the LR model. Figure 5(a) illustrates the importance of the 13 key features, sorted in descending order from top to bottom. The SHAP summary plot (Figure 5(b)) further shows the positive and negative relationships between predictors and the target, where each point represents a single participant in the test set and its position on the horizontal axis displays the characteristic SHAP value of that participant. The color of the dots from blue to red indicates an increase in sample values from low to high. For binary variables, risks are lower when ‘Bathing’, ‘Caregiver_A’, ‘Stand BWEO’, and ‘Hobbies_W’ have a value of 1, whereas risks are higher for ‘Homekeeping’, ‘Right Ear’, ‘Hobbies_X’, and ‘Caregiver_C’ at the same value. The significance of different feature values in relation to the risk of cognitive impairment can be further understood by referring to the feature detail table (Supplemental Table 1). For instance, people who were self-caregivers, could do ‘Stand BWEO’ or had hobbies such as cooking or fishing were at lower risk of cognitive impairment, whereas those who were cared for by a spouse, were predominantly home-bound, had impaired hearing in the right ear or lacked hobbies were at higher risk. For multivalued variables like ‘Age’, ‘Sleep duration’, ‘Focusing Difficulty’, ‘UI effect’ and ‘Housework’, the risk of cognitive impairment increases with older age, longer sleep duration, poorer attention, and more severe urinary incontinence, while engaging in more housework appears to be protective.

SHAP values and feature importance. (a) Feature importance ranking based on SHAP. (b) Attributes of feature in SHAP. Each line represents a feature, and each dot represents a sample, with the color of the dot indicating the value of that feature variable. (c and d) The interpretation of model prediction results with two samples. (The values of each variable are normalized).
To demonstrate the interpretability of the model at the individual level, we present two typical cases using SHAP waterfall plots (Figure 5(c) and (d)). Figure 5(c) illustrates a patient with normal cognition, characterized by a low final prediction score (0.022) and a negative net SHAP value (−3.81). The blue bar indicates that this feature plays a role in reducing the risk of positive prediction at the current value, such as “Bathing” = 1 and “Caregiver_A” = 1 (“take a bath independently” and “mainly take care of by oneself”). Figure 5(d) illustrates a participant with cognitive impairment, characterized by a high final prediction score (0.90) and a positive net SHAP value (2.18). The red bar indicates that this feature plays a role in increasing the risk of positive prediction under the current value, such as “Housework” = 0 and “Right Ear” = 1 (“never do housework” and “abnormal hearing in the right ear”). The number in bold is the probability forecast value (f(x)), while the base value is the predicted value without providing input to the model. F(x) is the logarithmic ratio of each observation. The length of the arrows helps to visualize the extent to which the prediction is affected, and the effect is positively correlated with arrow length.
Furthermore, to evaluate the contribution of the final 13 features, we performed feature decrement and increment analyses using the logistic regression model. The results identified ‘Bathing’ as the most influential factor in the model: its exclusion caused the most substantial performance drop (Supplemental Table 5), whereas its inclusion yielded the largest AUC gain (Supplemental Table 6). Notably, although removing ‘Bathing’ significantly impaired the model, the AUC remained relatively high, suggesting that the model's performance depends on the relevance and stable combination of features rather than on any single predictor.
Discussion
This study innovatively utilized over 200 comprehensive and multi-dimensional assessment indicators from the CGA as candidate modeling factors. It encompasses comprehensive risk factors gathered through survey questionnaires, physical assessments, and neuropsychological evaluations, making the dataset not only targeted but also multifaceted, enhancing its predictive value significantly. This study compared logistic regression, naive Bayes, support vector machine, neural network and random forest algorithms for evaluation, and determined that logistic regression was the optimal model. In the final prediction model, 13 key features predicting cognitive impairment across various dimensions were highlighted, providing valuable insights into the complex interplay of various factors in cognitive health. Further SHAP analysis effectively demonstrated the positive or negative impact of these key features values on cognitive impairment prediction. These results are highly consistent with our existing clinical knowledge.
The results indicated ‘Bathing’ as a strong predictor of cognitive impairment, suggesting that individuals with bathing difficulties are at higher cognitive risk. This is likely because bathing is a complex activity of daily living (ADL) that integrates executive function, motor coordination, and sensory integration. 15 Growing evidence suggests that changes in the ability to perform ADL play a significant role in predicting cognitive decline.16–18 Therefore, difficulty with bathing may serve as a sensitive indicator of cognitive dysfunction or even early neurological impairment, reflecting compromised brain capacity for complex tasks. In addition, although the ‘Bathing’ feature had the greatest impact on model performance, the AUC remained high after its removal. This suggests reasonable redundancy among features, where the absence of a single feature can be compensated for by others. Such redundancy underscores the model's high robustness, reduces over-reliance on any single feature, thus enhances the risk of failure due to feature loss in practice and enhancing its clinical utility.
The feature ‘Housework’ involves not only the ability to perform ADL but also planning and decision-making abilities. It is a core component of higher-level executive functions. Impaired executive function is an early manifestation of various types of dementia, particularly vascular cognitive impairment and frontotemporal dementia. 19 Furthermore, reduced engagement in housework may also be related to apathy or a loss of social roles, which are psychosocial factors associated with cognitive decline. 20 This finding aligns with studies suggesting that engagement in housework may help reduce dementia risk. 21 In addition, it is well known that age is an immutable risk factor for cognitive impairment. The incidence of cognitive impairment increases with age, and our SHAP analysis and LR results fully also confirm this trend.
Notably, our study has identified several new sensitive predictors, including hearing abnormalities, urinary incontinence and weak balance. We found that hearing abnormalities in right ear were more predictive of cognitive dysfunction than hearing abnormalities in left ear. This finding has not been previously reported, but is consistent with existing research on the ‘right-ear advantage’.22,23 Urinary incontinence was identified as a potential marker of cognitive impairment, with studies indicating a connection between bladder control issues and the development of conditions like AD.24,25 Furthermore, difficulties in tasks like “standing with eyes open and feet apart” may indicate a higher risk of cognitive impairment, potentially reflecting deficits in planning abilities, motor coordination, and balance.26–28
Additionally, our model highlights the importance of key features like ‘Homekeeping’, ‘Hobbies’, ‘Sleep Duration’ and ‘Difficulty Focusing’. Among them, ‘Homekeeping’ reflects lower social engagement and depression tendency, which have been found to be associated with cognitive impairment. The model's demonstration of a positive correlation between cognitive function and engagement in hobbies aligns with the concept that hobbies contribute to cognitive reserve and dementia prevention. 29 Studies in Chinese cohorts have proven the role of sleep duration in cognitive performance and the association between attention and early symptoms of dementia, further emphasizing the multifaceted nature of cognitive health.30,31 These findings collectively underscore the complexity of cognitive impairment and the need for a multifaceted approach in its assessment and management. The model's ability to capture these diverse yet interconnected factors illustrate its potential as a valuable tool in early cognitive impairment detection and intervention planning.
This study presents a novel predictive model for cognitive impairment that outperforms traditional approaches. First, the model is constructed based on CGA, which addresses the issue of a narrow range of candidate risk factors in previous studies, ensuring that key risk factors are not overlooked. Second, the model incorporates 13 predictors that are easy to obtain, low-cost, and highly suitable for large-scale screening. Furthermore, applying SHAP analysis to the LR model not only improves predictive performance but also enhances interpretability by clarifying the importance and influence of each feature in predicting cognitive impairment. 32 The identification of novel sensitive predictors offers valuable insights for clinical validation of their significance in the early identification and intervention of cognitive impairment.
However, this study also has several limitations. Firstly, although this retrospective cross-sectional study found a strong correlation between the characteristics and the risk of diagnosis, it cannot determine the causal relationship between the characteristics and the risk of diagnosis through this strong correlation. It can only clarify the role of these characteristics in providing a hint for the definite diagnosis at present. Therefore, the model established in this study is more suitable as an auxiliary tool to support clinical screening and decision-making. In the future, we will further verify the prognostic predictive value of these indicators through longitudinal follow-up studies.
Additionally, the generalizability of our study is a key consideration, since all participants were recruited from a single center in Southwest China, representing a specific geographic and cultural context. While the combination of clinical and community samples was intended to improve representativeness, the applicability of our results to other populations or healthcare systems still awaits external validation. Therefore, the broad interpretation of our findings requires caution and must be confirmed through future multi-center research. Furthermore, while the CGA is comprehensive, the reliance on self-reported or subjectively measured data for a subset of predictors (e.g. certain lifestyle factors, mood) may introduce recall and reporting biases that could influence predictive accuracy. Future studies could benefit from incorporating objective measures (e.g. imaging data or biomarkers) to validate results. Consequently, we explicitly position this model not as a definitive clinical diagnostic tool, but as an efficient, noninvasive prescreening aid.
Finally, it is necessary to emphasize that although the model is designed for cross-sectional screening predictions, it is based on easily accessible and noninvasive indicators. It can efficiently fill the gap of preliminary screening tools for cognitive impairment in the community and provide a basis for subsequent referral and diagnosis. Its core value lies in “quickly identifying at-risk populations” rather than “predicting long-term prognosis”. Meanwhile, in our future research, we will employ a more detailed hierarchical sensitivity analysis to further verify the independent predictive contributions of each feature to cognitive impairment, in order to enhance the interpretability of the model.
Conclusion
In summary, this study found 13 important risk factors of cognitive impairment based on CGA and machine learning methods. We constructed an effective model for predicting cognitive impairment. In addition, while improving model reliability, we provided a personalized risk assessments for clinicians by combining machine learning model with SHAP analysis. While acknowledging its limitations, this model presents a promising tool for early detection and intervention in cognitive health, opening new possibilities for future research and clinical application.
Supplemental Material
sj-docx-1-alr-10.1177_25424823261415843 - Supplemental material for An efficient non-invasive model for predicting cognitive impairment based on comprehensive geriatric assessment: Machine learning and SHAP analysis
Supplemental material, sj-docx-1-alr-10.1177_25424823261415843 for An efficient non-invasive model for predicting cognitive impairment based on comprehensive geriatric assessment: Machine learning and SHAP analysis by Jia Zhang, Wenjie Li, Sha Wen, Xiangzhou Zhang, Yuan Gao, Yanhui Feng and Yang Lü in Journal of Alzheimer's Disease Reports
Footnotes
Acknowledgements
The authors thank the participants who took part in this study and the health staff who contributed to the data collection and maintenance.
Ethical considerations
The study was conducted in accordance with the Declaration of Helsinki, and approved by the Ethics Committee of the First Affiliated Hospital of Chongqing Medical University (protocol code: 2012-15; approval date: 18 July 2012).
Consent to participate
This retrospective study used fully anonymized data that do not contain any identifiable participant information. Hence, consent for participate in the study was not required and is not applicable. lt was approved by the institutional ethics board of The First Affiliated Hospital of Chongqing Medical University (approved on 18 July 2012, No.15).
Consent for publication
This retrospective study used fully anonymized data that do not contain any identifiable participant information. Hence, consent for publication in the study was not applicable.
Author contribution(s)
Funding
This work was supported by the grants from Chongqing Talent Plan (grant number cstc2022ycjh-bgzxm0184), Key Project of Technological Innovation and Application Development of Chongqing Science & Technology Bureau (grant number CSTC2021jscx-gksb-N0020) Science Innovation Programs Led by the Academicians in Chongqing under Project (grant number 2022YSZX-JSX0002CSTB) STI2030-Major Projects (grant number 2021ZD0201802), Program for Youth Innovation in Future Medicine, Chongqing Medical University (grant number W0166) and Chongqing Medical Scientific Research Project (Joint project of Chongqing Health Commission and Science and Technology Bureau) (grant number 2020MSXM027).
Declaration of conflicting interests
The authors declared no potential conflicts of interest with respect to the research, authorship, and/or publication of this article.
Data availability statement
The datasets generated and analyzed during the current study are not publicly available due to privacy concerns of data but are available from the corresponding author on reasonable request.
Supplemental material
Supplemental material for this article is available online.
References
Supplementary Material
Please find the following supplemental material available below.
For Open Access articles published under a Creative Commons License, all supplemental material carries the same license as the article it is associated with.
For non-Open Access articles published, all supplemental material carries a non-exclusive license, and permission requests for re-use of supplemental material or any part of supplemental material shall be sent directly to the copyright owner as specified in the copyright notice associated with the article.
