Abstract
In light of growing biomedical data, machine learning (ML) models offer tremendous potential for personalized prediction in medicine. However, the additional value provided by these computational tools should always be critically evaluated. Using the example of predicting walking ability after spinal cord injury (SCI), we highlight a popular scenario in which data-driven predictions are feasible but not clinically meaningful, as the task can be performed equally well by humans. We asked 11 human observers from diverse backgrounds (five researchers without clinical training but proven knowledge of SCI and the International Standards for Neurological Classification of SCI [ISNCSCI], and six neurologists experienced in SCI) to predict walking ability following SCI based on acute phase neurological status assessed by the ISNCSCI motor and sensory scores (≤40 days after injury [DAI]). Following an established clinical prediction rule, walking ability was defined by a binary label derived from the indoor walking ability subitem of the Spinal Cord Independence Measure. We compared the performance of human observers with extreme gradient boosting and logistic regression-based models, which represent popular approaches in clinical literature on SCI. Using 794 patients from the European Multicenter Study about SCI, we show that all approaches provide similar, excellent performance at population level (area under the receiver operating characteristic 0.93–0.95; accuracy 0.88–0.90). Importantly, predictions combined from multiple neurologists (accuracy: 0.89) were comparable with model-based predictions (accuracy: 0.88–0.90), whereas individual neurologists (accuracy: 0.79 [0.01]; mean [standard deviation]) were marginally outperformed by computational approaches (accuracy: 0.88–0.90), particularly for more heterogeneous incomplete injuries. Individual SCI researchers performed equally well compared with neurologists (accuracy: 0.78 [0.02]). Our results show that prediction of walking function following SCI, if described through a binary label, does not benefit from ML, as ensembles of clinical experts and researchers each achieve performance similar to a range of ML models and an established clinical prediction rule. This highlights two key considerations in clinical applications of data-driven prediction models in SCI: first, the importance of carefully choosing clinical outcome measures to target in a prediction task to achieve a true benefit, and second, the necessity of benchmarking human performance on specific tasks to determine whether meaningful differences are present.
Introduction
Data-driven models based on machine learning (ML) algorithms have achieved commendable success in a variety of health care applications, including process automation, clinician support systems, and personalized outcome predictions.1–3 The overarching aim of these approaches is to provide a connection between (possibly unstructured) input data and an output label even in the (initial) absence of an understanding of the mechanistic nature of this relationship. As such, data-driven models trained on large data collections are particularly powerful for prediction tasks that surpass the capacity of individual clinical experts. However, the abstraction of multifactorial or continuous outcomes in medicine to binary labels for use in prediction tasks,4–6 such as when time-to-event analyses are reduced from a regression problem to a binary question of whether the event occurs within a particular time window, or when a functional outcome recorded on an ordinal scale is converted to a binary label only distinguishing the best possible outcome from an impaired state, introduces pitfalls, 7 which are not always adequately addressed. In contrast to other areas of application for ML, 8 one aspect commonly overlooked in health care is a comparison between predictive models and human experts on the particular prediction task. While ML might yield promising performance, what counts in application is the improvement beyond human performance on the same question. This is often not evaluated, leaving the added value of the prediction model uncertain. Considering the challenges following initial model development in a research environment to achieve a practically useful implementation in routine clinical practice, 9 comprehensive evaluation, including a comparison with existing approaches, provides essential context to evaluate a model’s potential usefulness.
The complex pathophysiological processes following traumatic spinal cord injury (SCI) and the extensive heterogeneity in neurological and functional recovery of patients presenting similarly in acute clinical assessments represent such a setting. Considering the often life-long impairments of essential motor, sensory, and autonomic functions impacting patients’ quality of life and independence, it is key to set realistic expectations for recovery. 10 Personalized outcome prediction could anchor these expectations. In the acute phase after injury, the recovery of walking ability is of particular concern for affected individuals, 10 while in later stages other functional aspects such as bladder and bowel control gain importance. 11 Accordingly, prediction of walking outcome has been a commonly addressed problem, 12 and a range of statistical and machine-learning approaches have been investigated in recent years.13–16 The most widely referenced approach to walking prediction, based on a logistic regression, was published by van Middendorp et al. in 2011. 17 Importantly, this study binarizes a single subitem of the Spinal Cord Independence Measure (SCIM) to label individuals as “walkers” and “non-walkers.” 18 The promising results of this study (achieving outstanding performance) motivated further investigation of walking ability using predominantly linear or tree-based approaches and binary labels based on SCIM or specific walking tests (e.g., 6-minute-walk test).19–21 Despite the advantages and the growing popularity of walking function prediction models, the simplicity of many implementations yielding exceptional performance raises the question of whether a computational prediction model is required for this task. Pelletier-Roy et al. 22 suggested that the approach by van Middendorp et al. might not provide predictive performance significantly better than expert clinical judgment—in a limited set of patients. This reinforces the question of whether a binary classification task (walkers vs. non-walkers) based on a single assessment is too simplistic.
Our study provides a detailed benchmark of human performance in binarized prediction of walking ability to address this question (see Fig. 1). We hypothesize that experts perform equally well as computational models in predicting walking function 12 months post-injury indicating that a binary label insufficiently captures the complex process of neurological and functional recovery. In addition, we acknowledge that prediction models based solely on neurological classification may be insufficient to adequately address the multifaceted neurological and functional recovery after SCI. This comprehensive investigation using a large observational data source evaluated by multiple expert and nonexpert observers is facilitated through an interactive app. We score the performance of individuals and the aggregate of their judgments to quantify a potential benefit of computational prediction models of walking ability compared with human (clinical) expertise.

Overview of study. 794 adult patients enrolled in EMSCI (https://www.emsci.org/) met the inclusion criteria. Prediction of the patients’ walking ability was performed by clinicians whose answers were collected using an app (see Supplementary Fig. S1). These predictions were compared with the clinical prediction rule previously published by van Middendorp et al. (vM LR), a newly fitted all-feature LR model, and an XGBoost instance, which is a state-of-the-art tree-based model able to represent nonlinear feature interactions. Neurologists’, researchers’, and models’ predictions were compared using ROC–AUC, accuracy, sensitivity, specificity, and Venn diagrams of incorrect predictions. EMSCI, European Multicenter Study about Spinal Cord Injury; all-feature LR, logistic regression model developed in this study; vM LR, logistic regression model underlying clinical prediction rule as published by van Middendorp et al. (2011); XGBoost, extreme gradient boosting; ROC, receiver operating characteristic; AUC, area under the curve.
Methods
Study design and data source
Established in 2001, the observational European Multicenter Study about Spinal Cord Injury (EMSCI; https://www.emsci.org/) involves 29 trauma and rehabilitation centers throughout Europe and one in India. To date, this collaboration has amassed data on over 6,000 patients with traumatic spinal cord injuries. The study tracks the neurological status according to the International Standards for Neurological Classification of SCI (ISNCSCI) 23 and functional outcomes of patients with spinal injuries after acute spinal cord trauma or ischemia. Evaluations are systematically conducted at predetermined intervals during the first year post-injury, including the very acute phase (0–15 DAI), acute phase I (16–40 DAI), acute phase II (70–98 DAI), acute phase III (150–186 DAI), and the chronic phase (300–546 DAI). The EMSCI data included in this analysis were collected between 2001 and July 2020. Detailed information on the EMSCI study design, inclusion and exclusion criteria, and assessments is presented by Bourguignon et al. 24
Ethics approval
The EMSCI study is performed in accordance with the Declaration of Helsinki and was approved by all responsible institutional review boards. Throughout the study duration, EMSCI followed the ethics guidelines of the participating countries and implemented changes and new policies as required. Patients gave their written informed consent before being included in the EMSCI database.
Neurologists and SCI researchers
Six neurologists actively engaged in clinical practice at the Spinal Cord Injury Center of Balgrist University Hospital in Zurich, Switzerland, were included as expert observers and provided predictions for all patients using the app. Similarly, five researchers with past or current experience in SCI-related research provided predictions for all patients, allowing for comparison with nonexperts without clinical experience. These groups are referred to as clinicians and researchers, respectively. We collected the mean and standard deviation (SD) of the clinicians’ experience in neurology and SCI medicine measured in years, frequency of performing ISNCSCI exams, previous participation in ISNCSCI training courses, and familiarity with the prediction rule established by van Middendorp et al. Researchers’ experience was characterized based on the number of years working in SCI-related research, the number of publications and conferences with a focus on SCI contributed to or attended, and any ISNCSCI training completed. Three modes of scoring human observer performance were considered. (i) An ensemble aggregated the prediction probability by adding the assessors’ vote (0 = no walking ability or 1 = walking ability) and dividing it by the number of assessors, i.e., offering a probability of walking as rated by the group. Classifications (walker/non walker) were based on majority votes (walker label assigned if probability ≥0.5). (ii) Unanimous agreement rating required all observers in a group to agree on a prediction. A prediction was considered incorrect if the selection of any one of the observers did not match the actual walking ability in the chronic stage. (iii) The individual approach accounts for each observer independently. Note that unanimous agreement and individual predictions only provide binary classification but no associated probabilities.
Cohort definition: inclusion and exclusion criteria
To be included in our analysis, patients enrolled in the EMSCI had to meet the following criteria: available information on sex and age at injury; traumatic cause of SCI; complete ISNCSCI assessment (motor, pinprick, and light touch scores, anorectal exam) and injury severity according to the American Spinal Injury Association (ASIA) Impairment Scale (AIS) 23 within 40 DAI; and outcome assessment through SCIM subitem 12 (“indoor mobility”) 18 later than 150 DAI. We excluded patients younger than 18 years at time of injury, nontraumatic injuries, and patients with AIS grade E at baseline. Patients unable to comply with the instructions of the ISNCSCI exam do not have an assessment recorded in the EMSCI database. A large number of patients were excluded mainly due to missing or non-testable values in their SCIM assessment, resulting in a final cohort of 794 patients for analysis.
Prognostic variables
We chose patients’ age, sex, and the earliest available neurological examination scores as the prognostic variables. As older patients (>65 years) with SCI may have less capacity to translate neurological improvements into functional recovery compared with younger patients, 25 age was categorized into two groups following van Middendorp et al.: ≤65 years and >65 years. 17 The ISNCSCI exam comprises testing of motor scores graded on a six-point scale (0 = full paralysis and 5 = normal function), light touch sensory, and pinprick sensory scores (0 = absent, 1 = altered, and 2 = normal), as well as sacral sparing (voluntary anal contraction, VAC, [0 = absent and 1 = present] and deep anal pressure, DAP, [0 = absent and 1 = present]). 23 Presence/absence of DAP and VAC, while assessed as part of the ISNCSCI, were not included. Additional models including these features are described in Supplementary Methods S1.
Outcome assessment
The primary functional outcome was the ability to walk independently in the chronic phase 26 according to SCIM subitem 12 (ability to walk <10 m indoors) 18 assessed between 300 and 546 DAI as part of the EMSCI protocol. If this assessment was missing, we used the assessment between 150 and 186 DAI to balance data availability with the latest possible time point for the outcome assessment. 27 The mentioned SCIM subitem differentiates total assistance (0 points), wheelchair use (1 to 2 points), supervision required (3 points), walking with aids (4–7 points), and walking without aids (8 points). To classify patients as “non-walkers” or “independent walkers,” the following cut-off was applied: 17 patients with scores <4 were classified as “non-walkers” and those with scores ≥4 as “independent walkers.”
App implementation
The implementation of a mobile device app with a simple voting mechanism enabled human observers to easily evaluate the expected chronic phase walking ability of individuals with SCI. The app was developed using the Flutter framework (https://flutter.dev, version 3.13.6) and displays the information of one patient at a time with a voting option (“walker” vs. “non walker,” see Supplementary Fig. S1A). Users receive immediate feedback following their answers (see Supplementary Fig. S1B). Each patient profile consisted of the ISNCSCI motor and sensory scores from the earliest available assessment (see Cohort definition: inclusion and exclusion criteria and Prognostic variables), age at injury, and sex. The AIS grade was not provided. Observers were asked to complete at least one (of a total of eight) data subsets, each containing approximately 100 patients, in one session but had the opportunity to take breaks between subsets. Observers received a standardized introduction to the app using ten randomly selected patients, which were identical for all observers.
Implementation of machine learning models
Two ML model types representative of the current literature 12 were chosen to assess walking ability prediction: logistic regression and extreme gradient-boosted trees (XGBoost). 28 We train on the full set of input variables outlined in Prognostic variables to enable a direct comparison with human observers. The relevant models are referred to as all-feature LR and XGBoost, respectively. In addition, we used the clinical prediction rule as published by van Middendorp et al., referred to as vM LR, on the cohort defined in Cohort definition: inclusion and exclusion criteria. Details on model training and evaluation are reported in Supplementary Methods S2 and Supplementary Table S1. All ML models offer predicted probabilities on held-out test data in addition to binary labels.
Performance assessment, statistical analyses and feature importance
Performance was assessed by established evaluation metrics such as accuracy, precision, recall, and specificity. For all approaches offering probabilities of walking (ML models, clinician, or researcher ensemble), we further calculate the area under the receiver operator characteristic (AUROC). Statistical analysis of ROC curves was performed using the DeLong test 29 implemented in Python. 30 A 5% significance level was considered statistically significant, and the Bonferroni correction 31 was applied to account for multiple comparisons. Model calibration was assessed visually using calibration curves computed using scikit-learn. 32 Feature importance was calculated using the Shapley Additive Explanations (SHAP) method implemented using the shap library (version 0.45.1) in Python (see Supplementary Methods S3). 33
Data and code availability statement
The data used for this study, including pseudo-anonymized participant data and a data dictionary defining each field or variable within the dataset, can be made available on reasonable request to the corresponding author. Written proposals will be evaluated by the authors and EMSCI steering board, who will render a decision regarding suitability and appropriateness of the use of data. Approval of all authors and EMSCI steering board will be required and a data-sharing agreement must be signed before any data are shared. All code required for the analysis can be accessed on Gitlab (https://gitlab.ethz.ch/BMDSlab/publications/sci/sci-walking-app).
Results
Patient cohort summary
A total of 794 adult patients with complete ISNCSCI exams within 40 DAI and SCIM item 12 assessed later than 150 DAI (Fig. 2A) were included. Figure 2B summarizes key characteristics of this cohort in relation to the full EMSCI database.

Patient cohort summary.
Human observer cohort summary
The cohort of human observers comprised six clinicians and five researchers. Clinicians reported mean experience of 18.8 years in neurology (SD = 10.9) and 12.8 years specifically in SCI medicine (SD = 10.7). All clinicians reported conducting ISNCSCI examinations several times per week. Participation in ISNCSCI training as part of dedicated courses or at conferences was reported by five out of six clinicians (83.3%). Three out of six clinicians (50%) reported to be familiar with the clinical prediction rule by van Middendorp et al. Researchers reported a median of two years of experience working in SCI, having contributed to three articles and attended two conferences with a focus on SCI. No researcher reported having received formal training in conducting the ISNCSCI exam or having ever conducted the exam on patients.
Machine learning models show no significant performance gain over neurologists or SCI researchers
Table 1 summarizes the performance of the ML and human observer-based models for the patient population as a whole and patient subgroups stratified by acute phase AIS grades, demonstrating excellent performance. The performance of individual clinicians and researchers was comparable with the clinical prediction rule (vM LR) overall (Fig. 3A). Ensemble models of both groups of human assessors showed comparable performance with all statistical and ML prediction models and showed small improvements over these models for patients assessed as AIS B and AIS D in multiple metrics (Table 1).
Overview of Prediction Performance Overall and Stratified by AIS Grade
The best performance is highlighted in bold; multiple values are highlighted in case of a tie.
Rows for the individual neurologists and individual SCI researchers evaluation show mean and standard deviation across human raters.
Positive class prevalence is defined as the percentage of patients labeled as walkers in the respective cohorts.
The unanimous agreement model and individual neurologists’ and researchers’ predictions do not provide ROC–AUC values, as prediction probabilities were not available for this approach (see Supplementary Table S2).
The results for the unanimous agreement of researchers are shown in Supplementary Table S4.
Variability of performance for newly trained models was assessed across test sets generated as part of the five-fold cross validation: all-feature LR: accuracy = [0.87, 0.91], precision = [0.86, 0.90], recall = [0.81, 0.85], specificity = [0.92, 0.94], ROC–AUC = [0.94, 0.95]; XGBoost accuracy = [0.87, 0.89], precision = [0.86, 0.87], recall = [0.83, 0.87], specificity = [0.91, 0.91], ROC–AUC = [0.94, 0.95]; [25th, 75th] percentile of respective metric; see Supplementary Table S5.
AIS, American Spinal Injury Association (ASIA) Impairment Scale; ROC, receiver operating characteristic; AUC, area under the curve.

Comparison of prediction performance.
The Venn diagrams (Fig. 3B) visualize the number of incorrect predictions for each approach (neurologists and ML models) together with the incorrect predictions shared amongst approaches. We observe a noticeable increase in false predictions exclusively made by the unanimous agreement model for neurologists in comparison with the ensemble. Patients with an AIS C grade make up the largest percentage of the incorrect predictions across all approaches. However, the largest proportion of incorrect predictions exclusively made by XGBoost is made up of patients with an SCI classified as AIS B (41.2%).
The overall walking ability prediction of the all-feature LR, XGBoost, clinician ensemble, researcher ensemble, and vM LR showed only minor differences, achieving ROC–AUC values between 0.93 and 0.95 (Table 1, Fig. 3C). No statistically significant differences in ROC–AUCs (Fig. 3C) were observed (Supplementary Table S2). Considering alternative metrics (accuracy, specificity, and recall), different models performed best. Only the unanimous agreement models were outperformed across all metrics. Predictive performance of the newly implemented all-feature LR and XGBoost models improves only marginally with increasing cohort sizes (see Supplementary Table S3).
In addition to performance, prediction confidence is a relevant end-point, which we quantify for the entire cohort in Figure 3D and Supplementary Figure S2. Independent of the approach, correct predictions show a pronounced bimodal distribution with peaks at probabilities equal to 0 and 1, indicating high confidence. Probabilities for incorrect predictions were distributed more uniformly, emphasizing the uncertainty associated. Visual assessment of calibration indicates comparable results for all models, while ML models present greater variability (Supplementary Fig. S3). Finally, we evaluated feature importance (Fig. 3E). For the all-feature LR, the three most important features were age (binarized) and motor scores of the right and left quadriceps femoris (L3). These motor scores were also among the most important features of the XGBoost, for which, however, also the light touch score at L4 was of high importance. The inclusion of DAP and VAC as input features did not affect feature importance (Supplementary Fig. S4) or model performance (Supplementary Table S6).
Variation in performance by acute phase injury severity
An analysis of prediction performance stratified by injury severity as assessed by acute-phase AIS grade (within 40 DAI) reveals clear differences (Table 1). Considering the class imbalance present in the different severity grades (positive class prevalence [walking]: AIS A 7.1%, AIS B 32.1%, AIS C 59.4%, AIS D 93.0%), it is important to account for this upon evaluation. While accuracy is very high for all models for AIS A ([0.87, 0.96]) and AIS D ([0.88, 0.94]) evaluation metrics accounting for class imbalance emphasize that only AIS A patients are reliably predicted (ROC–AUC: [0.82, 0.86]). Here, in particular, negative outcomes are well predicted (specificity: [0.92, 0.99]), while results are less reliable for patients initially classified as AIS A who eventually walk (recall: [0.27, 0.62]). Performance for AIS D is notably worse overall (ROC–AUC: [0.65, 0.70]), stemming from a failure to correctly identify patients with AIS D grade assigned within 40 DAI who will not walk (specificity: [0, 0.23]). Performance in the case of moderate injury severities was much more variable across models, with accuracy for AIS B ranging from 0.58 (clinician unanimous agreement) to 0.86 (clinician ensemble) and accuracy for AIS C ranging from 0.32 (clinician unanimous agreement) to 0.74 (vM LR, XGBoost). While we also observe differences in recall and specificity for AIS B (recall: [0.26, 0.66]; specificity: [0.73, 0.97]), patients classified as AIS C showed more balanced performance (recall: [0.37, 0.82]; specificity: [0.37, 0.76]), implying that, despite the overall poorer performance, both walkers and non-walkers could be predicted equally well. Across all approaches the vM LR consistently achieved comparable or slightly superior ROC–AUC. The only exception can be seen for patients classified as AIS B for which the clinician ensemble performed best.
Figure 4A shows each clinician and researcher as an individual data point for accuracy, sensitivity, and specificity stratified by subgroups defined on acute-phase AIS grade. We observe that clinicians and researchers may be more prone to the assumption of ceiling and flooring effects, i.e., predicting AIS A patients generally as non-walkers (low recall) and AIS D patients as walkers (low specificity). Importantly, the predictive performance of the vM LR model generally falls in the range of individual human observer performance, independent of group. Figure 4B visualizes the ROC curves, stratified by injury severity, for the four models. We observe highly skewed distributions of prediction confidence for all approaches for AIS A and AIS D (Supplementary Fig. S5), while the distribution for AIS C (Fig. 4C) most closely resembles a uniform distribution for both correct and incorrect predictions, highlighting uncertainty in classifying these patients. Interestingly, no association between clinician experience, as evaluated by years of practice with a focus on SCI, and performance on the prediction task could be observed (Supplementary Fig. S6).

Comparison of model performance stratified by acute phase AIS grade.
A comparison of incorrect predictions at patient-level (Supplementary Fig. S7) revealed differences between injury severity groups, which mirror the differences in performance. Considering the clinician ensemble (Supplementary Fig. S7A), the subgroup with the highest absolute number of incorrect predictions was AIS C with 59 distinct incorrect predictions, followed by AIS B (28), AIS A (27), and AIS D (18). The subgroup with the largest common set of incorrect predictions for all approaches was AIS D (61%), followed by AIS A (52%), AIS B (39%), and the smallest common set for AIS C (31%). Our data collection also allowed us to assess the interrater variability in the predictions of clinicians (Supplementary Fig. S7B). This variability was highest for AIS C patients and lowest for AIS D patients.
An analysis of ML models retrained for specific acute phase AIS grades showed no relevant performance benefit (Supplementary Fig. S8, Supplementary Table S7) except for the XGBoost model retrained on patients with injury severity AIS B. We summarize feature importance for models trained on patient subsets of a specific AIS grade in Supplementary Figure S9. A combination of motor and sensory score features remains the most relevant, as determined based on SHAP analysis for all models trained on specific patient subsets, while age and sex appear to become relevant for less severe injuries (AIS C and D).
Discussion
In line with previous studies, our results demonstrate that a simplified binarized descriptor of walking ability (i.e., able to walk while not distinguishing different levels of walking performance) can be predicted outstandingly well from ISNCSCI scores (ROC–AUC > 0.93). 34 However, humans, even those without clinical training, achieved comparable results (no significant differences), confirming our initial hypothesis. While absolute performance varies depending on severity of the initial injury, equal performance between all predictive approaches was observed within each subgroup of the population. We thereby confirm the results of a study by Pelletier-Roy et al. based on a small cohort of less than 70 patients 22 in a much larger dataset comprising 794 patients and given modern prediction architectures based on tree ensembles. Noticeably, we expanded the analysis to include a group of human observers without clinical training but with educated knowledge of the ISNCSCI exam from research use. The observation of equal performance of experts (neurologists), nonexperts (SCI researchers), and data-driven models across studies raises two key questions: (i) Is a binary walking prediction task based on one dimension of ambulation sufficient to address the complex interplay of neurological recovery and rehabilitation of functional abilities?; (ii) Is the application of ML models to predictions of binarized walking ability relevant given performance comparable with experts with clinical experience and nonexperts?
Bolliger et al. 35 argue that the choice of measurement of walking ability in clinical trials is essential to determine relevant differences. Our results highlight that similar careful consideration should be applied to the definition of a label in a recovery prediction task, as the outcome in question possibly involves factors such as spasticity, 36 gait disturbances, 37 and the use of compensation strategies. 38 This is particularly relevant if insights on mechanistic underpinnings of recovery are to be derived from a prediction model.39–41 Popular outcome measures employed in walking prediction, such as SCIM subitems on mobility or the 6-min-walk-test 42 fail to capture mechanistic aspects, are prone to flooring and ceiling effects, and show varying patterns in longitudinal changes depending on the level of injury.35,43 Considering that the greatest benefit of ML in medicine will be derived from collaborative development and deployment demonstrating improvement of predictive performance through the integration of ML,44,45 our results indicate that using a binary label when predicting walking ability is most likely not presenting such a scenario.
Potential benefits of ML models, however, include the relative ease of systematically evaluating associations between predictive features and the outcome of interest and of obtaining prediction probabilities. While an analysis of feature importance cannot comprehensively account for confounding factors, it can still provide a starting point for hypothesis generation and a degree of insight into the prediction formation in case of known associations being represented. Our analysis of feature importance validates the previously reported importance of age and the motor score of the quadriceps muscle (L3). 17 Motor and sensory information of the segment S1 (gastrocnemius and soleus muscles) is rated less important, whereas the models determined sensory information from L3/L4 as more relevant. This may be due to the greater importance of proximal segments for functional walking with aids such as an ankle-foot orthosis, which may rely less on sensory input from distal segments. We suggest future research to include parameters such as aids used to more accurately reflect the importance of features. Although a detailed analysis of features considered to be important by clinical experts for their predictions is beyond the scope of this work, it should be noted that in our study expert predictions are only based on motor and sensory score information as recorded in the ISNCSCI exam as well as patient sex and age at time of injury. This is in contrast to the previous study by Pelletier-Roy et al., 22 who provided expert assessors with access not only to the ISNCSCI exam results but also to consultation notes from the initial spine surgery, information on medical history and comorbidities, and all available imaging reports. While EMSCI does not record this information, our findings suggest that clinical experts likely do not require access to this to perform as well as statistical and ML models. This limited value of detailed additional information on patient presentation further suggests that a binary prediction task presents an oversimplification. Our analysis of prediction probabilities for ML models and ensemble models based on human raters shows that calibration is more variable for the former (see Supplementary Fig. S3), while the latter present a relatively stronger shift toward very low or high probabilities within each subgroup of patients defined by injury severity (see Supplementary Fig. S5). An assessment of the confidence of individual human raters in their prediction was beyond the scope of this study but could be investigated in the future.
While we demonstrate no significant difference between expert performance, nonexpert performance, and model predictions overall, we revealed differences in prediction performance between acute-phase injury severity groups. Performance for sensory incomplete (AIS B) and motor incomplete (AIS C) subpopulations was noticeably worse for all prediction approaches. This implies that ISNCSCI, which is an assessment of body structure and function and not activity according to the International Classification of Functioning, Disability and Health, and demographic information alone may be insufficient to reliably predict walking ability. Interestingly, the clinician ensemble performed best for the AIS B subpopulation, which could potentially be explained by the relative uncertainty of model predictions as indicated by the near-uniform distribution of prediction probabilities (see Supplementary Fig. S5). This further highlights the need to carefully consider which prediction tasks should be addressed by computational approaches to add value beyond existing clinical expertise. One scenario is the prediction of walking ability for the AIS C subgroup. Here, all prediction algorithms performed consistently better than either clinicians or researchers. Human observers generally disagreed more frequently, emphasizing the difficulty of the task. Considering the potential for meaningful recovery in different severity groups, this is particularly concerning, as injuries of intermediate severity (AIS B and AIS C) have more realistic chances of achieving partial recovery with appropriate therapy.46–48 In addition, being able to identify the characteristics of the small number of very severely injured patients (AIS A) who recover walking function could be valuable to target therapy and improve understanding of recovery alike. While we were not able to identify individual features definitively distinguishing these two groups in this study (see Supplementary Table S8), additional data modalities such as electrophysiological measurements or fluid and imaging biomarkers could provide further insights in the future. Conversely, prediction of outcomes for AIS D was affected by ceiling effects in neurological recovery. Pelletier-Roy et al. 22 excluded AIS D patients from their analysis due to high positive class prevalence and sensitivity, again drawing attention to the impact a particular choice of prediction task and cohort to test it in can have on the relevance of outcomes. Our analysis, however, includes all injury severities to reflect the approach by van Middendorp et al. 17
While EMSCI is one of the largest and most comprehensive datasets available for SCI, the fact that functional outcomes are missing for a sizeable number of patients in the database limited the cohort included here. Further, the dichotomization of the SCIM subitem 12 to derive a label for the classification presents a limitation as it only describes one aspect of ambulation in a particular context. However, to enable a comparison with previously published approaches, this label was kept. A further limitation is that expert assessors were not provided with information on VAC and DAP. We account for this limitation by also excluding these features from the all-feature LR and XGBoost models, which does not impact model performance or feature importance (Supplementary Table S6 and Supplementary Fig. S4). While the results presented clearly indicate the equivalence of machine learning approaches and human raters on the task of predicting walking ability following SCI, future work could provide further evidence by approaching the comparison as a non-inferiority trial. 49 Finally, the inclusion of expert observers from a single center represents a limitation and potentially reduces the inter-rater variability, as all expert observers are familiar with the same standard operating procedures and received comparable training.
Conclusion
The present study compares the performance of highly experienced clinicians and researchers without clinical training but with experience in SCI with a published clinical prediction rule and machine learning algorithms on the task of predicting a binary label of walking ability in the chronic stage on a large dataset comprising early ISNCSCI exams after traumatic SCI. The results show a comparable performance of human experts and nonexperts, indicating a very limited benefit provided by machine learning models for this simple prediction task, which is widely addressed in the literature. Small differences are apparent for AIS B and AIS C injuries, with human experts performing better for the former, while a pronounced inter-rater variability might explain the worse performance for the latter. These results highlight the importance of comprehensive comparisons of performance of predictive models on clinical tasks, which include a representation of human capabilities. Our comparison indicates that careful consideration is required to define meaningful prediction tasks on clinical outcomes to realize the potential benefits from augmenting clinical expertise with machine learning.
Transparency, rigor and reproducibility summary
This study was not formally registered, as it is a retrospective analysis of data acquired as part of a registered observational study (https://clinicaltrials.gov/study/NCT01571531). The analysis plan was not formally preregistered. Neither statistical power and sample size calculations nor blinding of researchers performing the analysis were required for this study. Patient data were acquired as part of the ongoing European Multicenter Study about SCI. Predictions by human experts were acquired through a custom app between July 2023 and September 2023. All software required for data acquisition and analysis is open source and available free of charge. All code used for analysis is shared publicly through gitlab (https://gitlab.ethz.ch/BMDSlab/publications/sci/sci-walking-app). Key inclusion criteria, assessments, and the outcome measure are established standards in the field. Correction for multiple comparisons was performed according to the Bonferroni method. Anonymized data used in this study are available upon request to the corresponding author and in compliance with the European Union's General Data Protection Regulation (EU GDPR).
Authors’ Contributions
J.B.: Implemented the mobile app and machine learning models, performed the analysis and visualization of results, and drafted the article. L.P.L.: Implemented machine learning models, performed the analysis and visualization of results, reviewed all code used for analysis, and drafted the article. R.R., M.S., F.R., J.W., N.W., Y.B.K., R.A., D.M., H.S.C., T.L., and A.C.: Supported the data collection and clinical interpretation of the proposed findings. The EMSCI Study group collected data used in this analysis. S.B.: Designed the study, drafted the article, and provided continuous feedback. C.R.J.: Designed the study and provided data access and continuous feedback. All authors revised the article for intellectual content.
Footnotes
Acknowledgments
The authors would like to acknowledge the participating centers in the EMSCI network that were involved in patient care and collection of data necessary for this study. The EMSCI study group includes the following active members: Uniklinik Balgrist, Zentrum für Paraplegie, Zurich, Switzerland; Klinik für Paraplegiologie, Universitätsklinikum Heidelberg, Heidelberg, Germany; Klinikum Bayreuth, Bayreuth, Germany; UMC St. Radboud, Nijmegen, Netherlands; RKU—Universitäts‐ und Rehabilitationskliniken Ulm, Ulm, Germany; BG Kliniken Bergmannstrost]Halle, Halle, Germany; Institut Guttmann, Barcelona, Spain; SRH Klinikum Karlsbad‐Langensteinbach, Karlsbad, Germany; BG Unfallklinik Murnau, Murnau, Germany; Werner‐Wicker‐Klinik, Bad Wildungen, Germany; Motol Hospital, Prague, Czech Republic; Orthopädische Klinik Hessisch Lichtenau, Hessisch‐Lichtenau, Germany; Hospital Nacional de Parapléjicos, Toledo, Spain; Schweizer Paraplegiker‐Zentrum, Nottwil, Switzerland; Ospedale Fondazione Santa Lucia, Rome, Italy; Indian Spinal Injury Center, New Delhi, India; BG Unfallklinik Tübingen, Tübingen, Germany; Unfallkrankenhaus Berlin, Berlin, Germany; Klinik Bavaria, Kreischa, Germany; REHAB Basel, Basel, Switzerland; The Robert Jones and Agnes Hunt Orthopaedic Hospital, Oswestry, Great Britain; AUVA Wien, Vienna, Austria; Fondazione Salvatore Maugeri, Pavia, Italy. Special thanks are given to Prof. HJ Gerner (Heidelberg, Germany) as well as Prof. V Dietz and Prof. M Schwab (Zurich, Switzerland) for envisioning the value of the EMSCI network and laying the foundation for this research. Furthermore, we would like to thank Susanne Friedl, Maria Rasenack-De Vere Tyndall, Carl Zipser, and Luca Traini for their participation in our study, as well as Melina Giagiozis, Hugo Madge Leon, and Marc Bollinger. The authors acknowledge the use of an LLM to support the choice of appropriate vocabulary and improve the quality of the text. All text was written originally by the authors, and only parts of it were subsequently enhanced by the use of a LLM.
Author Disclosure Statement
The authors do not report any conflict of interest, financial, or otherwise.
Funding Information
This study was supported by the Swiss National Science Foundation (#PZ00P3_186101 and #IZLIZ3_200275, Jutzeler), Wings for Life Foundation (#2024_301, Brüningk, Jutzeler), and the International Foundation for Research in Paraplegia (#P192, Brüningk, Jutzeler). Sarah Brüningk is supported by the Botnar Research Centre for Child Health Postdoctoral Excellence Programme (#PEP-2021-1008). The funding sources of the study had no role in study design, data collection and analysis, decision to publish, or preparation of the article. The corresponding authors had full access to all the data in the study and had final responsibility for the decision to submit for publication.
Supplemental Material
References
Supplementary Material
Please find the following supplemental material available below.
For Open Access articles published under a Creative Commons License, all supplemental material carries the same license as the article it is associated with.
For non-Open Access articles published, all supplemental material carries a non-exclusive license, and permission requests for re-use of supplemental material or any part of supplemental material shall be sent directly to the copyright owner as specified in the copyright notice associated with the article.
