Abstract
Distracted driving is a major concern for road safety in the U.S. as it is associated with thousands of fatal motor vehicle crashes. Many distraction-involved crashes turn into fatal crashes when the vehicle departs the roadway (run-off-road [ROR]). Although studies focusing on both distracted driving and ROR crashes are ample, there has yet to be much research on ROR crashes involving distracted driving (RORIDD). This study aims to analyze the factors associated with RORIDD crashes using New Jersey crash data from 2015–2019. Various machine learning (ML) models—support vector machine, random forest, AdaBoost, Catboost, LightGBM, and XGBoost—were utilized to predict the injury severity. Accuracy, precision, and recall scores were utilized to evaluate model performance. An interpretable ML technique—Shapley additive explanation (SHAP) values—was employed to identify the most influential factors contributing to these crashes. The results showed that CatBoost performed relatively better than other models in predicting crash severity. The SHAP values demonstrated that alcohol involvement during daylight was more likely to result in severe injury crashes. In addition to examining RORIDD crashes, the study also provides a comparative perspective with general ROR crashes. This comparison helps highlight the unique role of distraction-related factors within the broader ROR crash context. Targeted safety countermeasures, such as education and enforcement about driving under the influence and buckling up, as well as improving the visibility of signs for wet roads and curves, could help mitigate the severity of RORIDD crashes. The findings of this study are expected to assist policymakers and practitioners in developing targeted countermeasures (e.g., variable speed, dashboard warning systems, and enforcement) to reduce RORIDD crashes and improve road safety in New Jersey.
Introduction
According to NHTSA, distracted driving is “any activity that diverts attention from driving, including talking on the phone or texting, eating, drinking, conversing with passengers, and fiddling with the stereo, entertainment, or navigation system—anything that diverts your attention away from the task of safe driving.” ( 1 ). Driving while distracted is a risky behavior that should be avoided behind the wheel. In the U.S. in 2020, 3,142 people died in motor vehicle crashes involving distracted driving, accounting for 8.1% of all fatalities ( 2 ). A higher prevalence of distraction has been reported recently because of the increased use of electronic devices, especially among young and newly licensed drivers ( 3 ). A recent survey in the U.S. revealed that almost 35% of the drivers who had received a license in the last 5 years used a mobile device to browse the internet, read or react to emails, or update social media while driving ( 3 ). According to NHTSA, 387 people died in motor vehicle crashes in the U.S. because of using a cell phone while driving in 2019 (13% of fatal crashes involving distractions) ( 4 ).
Run-off-road (ROR) crashes occur when vehicles experience departure from the designated roadway. Typically, ROR crashes involve single vehicles and are a major contributor to motor vehicle occupant fatalities and serious injuries. During the period 2016–2018 in the U.S., approximately 51% of all fatalities were ROR crashes, resulting in an average annual of 19,158 fatalities ( 5 ). Countermeasures to reduce the occurrence and severity of ROR crashes include vehicle-related features such as lane departure warnings (LDW) and infrastructure-based measures such as rumble strips. Substantial research has also explored potential causes of ROR crashes, particularly analyzing injury severity. New Jersey averaged 78 fatal traffic crashes annually resulting from ROR crashes involving distracted driving during 2017–2021 ( 6 ). Table 1 shows the fatal crashes from distracted driving, ROR crashes, and crashes that involve distracted driving and end up as ROR crashes in New Jersey from 2017 to 2021.
Statistics of Fatal Crashes Involving Distracted Driving and Run-off-Road Crashes in New Jersey from 2017 to 2021
Note: ROR = run-off-road; RORIDD = run-off-road involving distracted driving.
This study addresses the crucial need to identify factors influencing the severity of ROR crashes involving distracted driving (RORIDD). It contributes to the existing knowledge by conducting a thorough literature review, and summarizing important contributing factors and the methods used for interpretation. Specifically focusing on RORIDD crashes in New Jersey from 2015 to 2019, this study investigates site- and state-specific factors to gain a comprehensive understanding. We utilized machine learning (ML) models for crash severity prediction and a feature importance model to identify the top variables influencing the crash severity outcome. This study further evaluates the influence of these top contributing factors using Shapley additive explanation (SHAP) values and performs pairwise comparisons to assess their combined impact. Given that RORIDD crashes represent a subset of all ROR crashes, this study also incorporates a comparison with general ROR crashes. Presenting RORIDD findings alongside overall ROR outcomes reveals the influence of distraction-specific variables within the broader context of roadway departure events. The findings have practical implications for engineers, practitioners, and legislators in managing the occurrence of RORIDD crashes in New Jersey.
The paper is organized as follows. The significance of examining distracted driving and crashes involving ROR is underscored in the Introduction. The Literature Review section thoroughly examines existing literature, providing a concise overview of previous studies conducted on the relationship between ROR and distraction-related crashes involving drivers. The Methods section comprehensively describes the data, study design, and models employed in predicting crash severity. Following this, the Results and Discussion section presents and interprets the study’s findings. In conclusion, the final section of this study provides a comprehensive summary of the obtained outcomes, accompanied by a few recommendations.
Literature Review
Previous research on distracted driving has focused on identifying distraction behaviors and analyzing crashes involving distracted driving. Self-reported surveys, naturalistic studies, and roadside observational studies have been performed to identify distracted driving prevalence and find the patterns of distraction ( 7 – 11 ). Numerous researchers have previously focused on identifying the contributing factors to distracted driving crashes ( 12 , 13 ). These factors include crash attributes, environmental conditions, roadway properties, temporal variables, vehicular features, and driver attributes. Most of these factors also contribute to the severity and frequency of ROR crashes. Table 2 summarizes the findings on the factors which contribute to distracted driving crashes and ROR crashes.
Summary of Literature on the Factors Which Contribute to Distracted Driving (DD) and Run-off-Road (ROR) Crashes
Note: numbers in parentheses are reference citation numbers (see References section). NA = Not Availabe.
Environmental conditions such as adverse weather, dark-not-lit, and wet surface conditions increase the crash severity ( 13 – 16 , 18 – 20 ). In contrast, factors such as daylight, clear roads, and dry road surfaces increase the frequency of crashes ( 24 , 25 ). Concerning roadway features, higher speed limits, rural roadways, and undivided highways, crashes at the intersection increase the crash frequency ( 24 , 25 , 27 , 29 , 30 ). At the same time, the urban roads, annual average daily traffic (AADT), and the number of lanes decrease the severity of the crashes ( 12 , 14 – 17 , 22 , 26 ). The crashes during nighttime have higher crash severity, while the crashes during weekdays are less severe ( 15 , 17 , 22 ). Driver attributes such as speeding, involvement of alcohol/drugs, not wearing a seatbelt, fatigue or drowsiness, older drivers, and reckless driving increase the crash severity ( 12 , 13 , 15 – 17 , 22 , 26 , 32 ). However, factors such as young age and male gender have demonstrated mixed effects on crash severity ( 13 , 15 , 17 , 22 , 32 ). In previous studies, crashes involving a curve or horizontal curve and the angular type of collision showed an increased crash severity, while fixed-object crashes demonstrated mixed results ( 12 – 16 , 22 , 23 , 26 , 28 ). Certain factors, such as clear or urban roads, are found to contribute to a decrease in crash severity. However, because these areas experience higher traffic volumes, they naturally have more crashes overall, albeit with less severity. For instance, an urban setting experiences a higher frequency of crashes, with a lower proportion of severe crashes than rural roads ( 12 , 14 – 16 , 22 , 24 , 26 , 27 , 29 ). Drivers on clear roads experience more crashes, but the severity of crashes is less than on wet roads ( 17 , 23 – 25 ). It is important to note that some of the effects observed in Table 1 may not be entirely attributable to a specific condition or state, but rather to their comparison with another condition or state; for example, the higher severity of crashes occurring on weekends is relative to those on weekdays.
In addition to statistical models (regression, discrete choice, rule-based association, etc.), ML techniques have become increasingly popular for feature selection and incident severity modeling. Previous research has shown that utilizing ML over traditional statistical models has some advantages ( 36 , 37 ). In recent studies, ML models such as random forest (RF), support vector machine (SVM), neural networks, and boosting algorithms such as XGBoost have also been employed for crash severity prediction ( 38 – 43 ). ML techniques were once considered black-box approaches. However, these methods are now more transparent and interpretable because of recent advancements and expansions in ML model interpretation research. Researchers have employed interpretable ML to describe the impacts of the major variables in crash severity studies. Previous research has employed sensitivity analysis, permutation feature importance, and the partial dependence plot ( 40 , 42 , 44 ). In recent years, numerous studies have utilized SHAP to interpret the results of ML models ( 45 ). SHAP is a visualization metric that provides valuable insight into how various model features affect the model. Based on game theory, it was created in 1953 by Shapley ( 46 ). Consequently, in 2017, Lundberg and Lee developed a Python package to calculate the SHAP value for ML models.
This study contributes to traffic safety analysis concerning RORIDD crashes by providing a concise overview of existing practices and identifying gaps in crash prediction. It thoroughly compares six different ML approaches to assess their prediction performance. The study also employs two feature selection techniques to pinpoint the factors affecting RORIDD crash severity. SHAP dependence plots are utilized to interpret these crucial variables’ combined impact. The findings from explainable artificial intelligence (AI) are particularly valuable for transportation agencies, as they shed light on the causes and patterns of crash severity. This, in turn, enables the development of effective safety countermeasures to reduce RORIDD crashes in New Jersey. Additionally, the ML model outcomes guide future researchers in selecting an appropriate technique for crash severity analysis with their datasets.
Methods
Study Design
The study design includes the following major steps:
As illustrated in Figure 1, the study design provided a structured and systematic approach to conducting this investigation, ensuring a rigorous analysis of the crash data, and offering valuable insights to enhance road safety and prevent RORIDD incidents.

Study design of severity modeling for run-off-road crashes involving distracted driving in New Jersey (NJ).
Description of Database
From the Numetric crash query database, we used 5 years (2015–2019) of RORIDD crash data in New Jersey, which included 50,952 crashes ( 6 ). After data cleaning, a final dataset was achieved that consisted of 30,995 crashes (871 injury, 11,039 possible injury, and 19,085 no injury). While doing the query for the crash data, initial filters were based on “crash year” between 2015 and 2019; “distracted driving involved” = “yes,” and “run-off-road involved” = “yes.” It is worth noting that, while selecting the categories of distraction, it is important to consider whether these behaviors are counted as “distraction” by police officers when reporting a crash. Many factors are almost universal in all police crash reports, but distraction is not one of those ( 47 ). One reason behind this is a difference in jurisdiction, crash narratives, and distinct recognition of distraction in the crash reporting format. Therefore, the crash data used in this study specifically considers the context of New Jersey when selecting behaviors termed “distraction.” For instance, cellphone use, talking to passengers, eating/drinking, grooming, and tuning the radio are considered distractions by New Jersey police officers when archiving them as contributing factors to driving ( 48 , 49 ). It is worth mentioning that “drowsy driving” is a standalone category for driver-contributing factors in the crash reports in New Jersey ( 6 ). Therefore, this category was not considered a distraction in the crash database. Although drowsy driving is not classified as a form of distraction in the state of New Jersey, it is considered a major form of driver inattention and is considered as one of the risky driving behaviors by FHWA ( 50 ). To assess its influence as an independent factor, drowsy driving was included in the model. However, this inclusion does not imply a reclassification; rather, it is an effort to evaluate its impact specifically within the scope of this study, focusing only on cases classified as RORIDD events.
Variables were selected based on the identification in the literature review of contributing factors and engineering judgment. Fatal crashes merged with severe injuries are named “injury” crashes. Suspected minor and possible injuries are merged as “possible injury.”“No injury” crashes are “property damage only” crashes. Data cleaning was further performed, and 26 input variables were selected for further analysis. Variables were divided into five categories: temporal features (i.e., time of day, day of week, and season), driver behavior conditions (i.e., unsafe speeding, alcohol or drug intoxicated driver), roadway features (i.e., highway type, intersection, median type, road divided by…, temporary traffic control zone, area, total pedestrian and/or bicyclist involved, and work zone involved), environmental conditions (i.e., light conditions, environmental conditions, and surface conditions), and crash attributes (i.e., unrestrained occupant involved, single-vehicle involved, crash type, and curve-related).
Exploratory data analysis (EDA) was performed on the selected variables to examine their detailed characteristics. ML models’ performance relies on the quantity of the dataset. Conducting a comprehensive EDA can enhance quality control by eliminating redundant and incomplete information. The incomplete information was filtered out in two steps. First, the variables with more than 10,000 (20% of 50,000) missing values were removed from the dataset. AADT was one of the variables taken out of the dataset because of this. Second, the rows that had values coded as “not a number” or “unknown” were also taken out from the dataset. Through this process, a dataset of 50,952 crashes and 44 variables was trimmed to 30,995 crashes and 27 variables. Afterward, feature importance techniques were utilized to identify appropriate variables for analysis. Notably, all categorical and continuous variables were transformed into numeric values before being used as inputs for the ML models. Following variable importance and correlation tests, the final dataset was prepared, ensuring it was ready for further analysis. Figure 2 describes in detail the data cleaning and variable selection process, from the raw dataset to the final cleaned dataset.

Data preprocessing for crash severity modeling.
It is worth mentioning that the data cleaning process did not alter the distribution of injury severity. Initially, there were 50,957 crashes, with 1,314 injury, 18,789 possible injury, and 30,854 with no injury. The proportion of injury crashes was 2.6% in the total dataset compared with 2.8% in the cleaned dataset, indicating no major impact on crash severity distribution. Similarly, the data cleaning process did not change the proportion of user types involved. For instance, the initial dataset had 112 pedestrian-involved crashes (0.22% of total records), while the cleaned dataset had 60 (0.19% of cleaned records). Therefore, the trends in both crash severity and users involved did not experience any major shifts because of the data cleaning process.
A descriptive summary of input variables is provided in Table 3. According to the table, only 0.2% of the crashes involved pedestrians or bicyclists, and 28.3% of those crashes involved injury. Similarly, only 2.2% of the crashes involved unrestrained occupants, but 22.5% were severe (injury). Concerning roadway surfaces, the injury rate on dry roads (3.2%) was higher than on wet surfaces (1.5%). Most (79%) of the crashes were fixed-object crashes, while angular crashes were the most severe, with an injury rate of 4.8%. More than a quarter (27%) of the RORIDD crashes involved curves. Around 6.4% of crashes involving alcohol/drug use resulted in injuries, compared with 2.2% of injuries in crashes not involving alcohol or drug use. Older drivers’ involvement in crashes had a higher injury rate (4.0%) than the involvement of young drivers (1.8% injury). Table 3 reveals that crashes involving drowsy driving have a slightly lower reported injury rate (2.1%) than RORIDD crashes where drowsy driving was not a factor (2.9%). However, it is important to note that a higher percentage of possible injury crashes (46%) occurred in cases where drowsiness was reported, compared with crashes where drowsiness was not involved (34.4%). Similarly, while crashes involving reported cellphone use exhibit a lower injury rate (0.7%) compared with cellphones-not-involved crashes (2.9%), it is crucial to recognize that a higher proportion of no-injury crashes was observed in cases where cellphone use was not reported (61.7%) compared with crashes reporting cellphone use (54.5%). Given that cellphone use data is largely self-reported—except when confirmed through surveillance—the actual prevalence of cellphone-related crashes may be underreported. Additionally, the sample size for injury crashes involving cellphone use is relatively low, limiting the ability to draw strong conclusions. Despite these limitations, the reported data indicates an increased likelihood of crashes resulting in possible injury when the driver is distracted by drowsiness or cellphone use than when these distractions are not present, underscoring their role in RORIDD crash risk.
Distribution of Key Features
Feature Importance
The variable selection process can effectively address the dataset’s noise issue by identifying the redundant variables ( 51 ). RF, univariate selection, classification and regression tree (CART), discrete choice models, ExtraTree classifier, and XGboost are widely employed techniques for variable selection. The ML models assess the relative significance of the contributing variables by considering their influence on the crash severity prediction ( 52 , 53 ). Previous research has indicated that the CART approach, ExtraTree classifier, and XGboost are effective ranking algorithms ( 54 , 55 ). These three strategies are used in this study to determine the relative importance of variables connected to crash severity. Variable significance scores in these models highlight the contributions of the variables as key splitters of the regression tree in improving crash severity predictions ( 56 ). The CART model ranks variables based on the Gini index value, representing the DT’s impurity or entropy levels ( 57 ).
Similarly, an additional tree classifier computes variable evaluations using randomized trees and the Gini index ( 58 ). Classification and regression are performed by XGBoost using a gradient-boosting DT ( 59 ). Throughout the construction of each tree layer, the greedy algorithm maximizes the maximal gain of the target variable (crash severity). The method seeks to develop a tree by continually adding and separating trees. Each time a new tree is added, the algorithm learns a new function to match the residual from the previous prediction ( 60 ).
Description of Different ML Algorithms
Support Vector Machine (SVM)
Previous studies have detailed the SVM method in depth with regard to crash severity analysis ( 40 , 61 , 62 ). SVM can differentiate between separable and non-separable data ( 63 , 64 ). SVM aims to optimize the margin or separation between distinct crash severity groups. It makes certain faults simultaneously, which can be modified using a penalty parameter ( 65 ). SVM separates the dataset into various classes using hyperplanes, intending to maximize the distance of the data points from the hyperplanes. A kernel function is applied to perform the non-linear transformation of the linear kernel. The kernel function transforms the data into a higher-dimensional space. This study uses the radial basis function (RBF), a widely used kernel function for crash severity analysis (39). The equations for SVM are as follows:
Minimize
Subject to:
where
c = parameter to define decision boundary between severity classes,
P = regularization parameter,
ξ n = error parameter (margin violation),
b = intercept associated with decision boundaries,
ϕ(x n ) = the function of data transformation from X space into Z space, and
y n = crash severity for the nth observation.
The non-linear transformation of kernels and the RBF kernel are expressed as shown in Equations 4 and 5, respectively ( 66 ):
where
γ = kernel parameter.
The SVM model requires the determination of two parameters, P and γ, with RBF serving as the kernel function. The SVC package from the Scikit library is used to execute SVM on the Python platform ( 67 ). The computation time of SVM is one of its problems. The SVM grid search approach increases the computation time of the models by repeatedly iterating over different combinations of the hyperparameters to find the best one. One method to shorten this computation time is to optimize the variable size. Using a correlation matrix and feature selection to identify the key variables decreased the computation time in this study.
Random Forest (RF)
The RF classifier is tree-based. RF employs two distinct classification algorithms: random feature selection and bagging ( 68 ). Bagging generates each tree individually, whereas random feature gathering generates DTs rapidly. RF chooses the characteristics of the subgroups at random rather than using every feature in the DTs. The output of a fresh dataset is predicted by RF by averaging the outputs of distinct random bootstrap training data. As demonstrated by Equations 6 and 7, nodes inside DTs can be constructed using either the Gini index or the entropy index, which reflect the optimal data splicing with the least amount of distance between its branches ( 69 ):
where
S = the number of classes in the dataset, and
f i = the observed class’s relative frequency (or proportion).
Equation 8 provides a mathematical expression for the final RF result:
where
j = the number of severity classes,
i = the number of DT ranging from one to the ith DT, and
arg max = the max value or majority vote.
RF generates a fair prediction with a higher number of DTs in the algorithm. However, doing so could increase the computation time of the model. In general, the model gives a fair prediction with little hyperparameter tuning. Therefore, this algorithm is faster than SVM. The RandomForestClassifier module, part of Python's scikit-learn library, is used to implement the Random Forest (RF) model ( 67 ).
Boosting Methods
The boosting approaches increase prediction performance by combining numerous weak classifiers into a single strong classifier. The classification model used four boosting approaches: XGboost, LightGBM, AdaBoost, and Catboost. The gradient boosting approach compensates the losses at each iteration by regressing the gradient vector function ( 70 ). A gradient-boosting model changes the order of each DT, beginning with the weak DT, which serves as the foundation for the base DT. XGboost is a slow-boosting strategy that uses sequential model training to reduce misclassification errors at each iteration. Catboost is a boosting algorithm that takes numerical and category variables as input. It manages the variables during the training period, saving time on preprocessing. LightGBM is a boosting method that, leaf by leaf, constructs more accurate and sophisticated DTs.
XGBoost, commonly known as an “ensemble approach,” incorporates gradient-boosted DTs. Friedman first proposed this approach in 2002 ( 71 ). XGBoost has demonstrated promising results across various fields of research, including transportation safety ( 72 , 73 ). This algorithm comprises a series of DTs, each of which learns from and impacts its predecessor. The formulation of XGboost is described with equations in some previous studies ( 74 , 75 ). The XGBoost algorithm gives a fair prediction result on a clean dataset with less noise. Therefore, data preprocessing or cleaning is an extremely important step to achieve a good prediction accuracy of the model. Organizing the categories of the variables, eliminating redundant information, and removing variables with less information for severity prediction was performed in this study to achieve a good prediction result from the XGBoost model. XGBClassifier from the package xgboost, LGBMClassifier from the package lightgbm, and CatBoostClassifier from the package catboost were used in Python to run the boosting algorithms ( 76 – 78 ). The following is the formulation of XGboost using K tree functions:
where
k = the number of additive trees,
t = the number of iterations,
The objective function for minimizing the loss
where
T = the number of leaves,
N = the total number of crashes in the sample data.
Model Evaluation
In this study, we assessed the performance of severity prediction in accuracy, precision, and recall ( 54 ). “Accuracy” is the proportion of correctly predicted severity classes about the total number of crashes. “Precision” is the quotient obtained by dividing the projected number of crashes for a specific crash severity (such as injury, no injury, or possible injury) as determined by the ML algorithm’s total number of expected crashes within that category. “Recall,” conversely, is operationally defined as the ratio of accurately predicted crash severity instances to the overall number of crashes within the corresponding severity category.
SHAP Values
SHAP, created by Lundberg and Lee, could be used to analyze the output of ML models ( 46 ). SHAP, based on local explanations and game theory, presents a method for estimating the importance of each attribute ( 45 , 79 ). SHAP values are based on numerous axioms that help fairly assign each feature’s contribution. Lundberg et al. built on this model using “tree explainer” methods ( 80 ). This approach quickly assesses the risk components of a SHAP value at both global and local levels ( 81 ). Previous works have fully described SHAP interpretation. The SHAP package of Python is utilized to implement the SHAP interpretation on the dataset ( 74 , 82 , 83 , 84 ).
Let us consider a hypothetical model designed to forecast an output denoted as v (N), utilizing a group labeled as N (comprising n features). Following the marginal contribution of each feature, the SHAP framework allocates the influence of individual features (φ representing the contribution of feature i) on the model’s output, denoted as v(N) ( 85 ). The computation of SHAP values is guided by various axioms aimed at fairly distributing the impact of each feature, as expressed by the equation:
where
S = a subset of N.
Based on the subsequent additive feature attribution technique, a linear function of binary features g is defined as:
where
z' ∈ {0, 1}M = 1 when a feature is observed (otherwise, it equals 0), and
M = the number of input features ( 46 ).
Lundberg further improved the model by utilizing “tree explainer” methods to compute SHAP’s global and local values ( 80 , 81 ).
Results and Discussion
Exploratory Data Analysis (EDA)
A correlation analysis was conducted on the chosen variables. The correlation analysis reveals a strong positive relationship (correlation coefficient >0.7) between the time-light condition, environment-surface condition, highway type-median type, and median type-road divided by…. In contrast, all remaining variables exhibit correlation coefficients below 0.7, suggesting a lack of significant association among them. A feature importance analysis was conducted to ascertain the significance of input variables. Figure 3 depicts the ascending relative relevance of these variables, as determined by the two feature selection strategies. The most influential factors identified by both approaches are crash month (season), roadway system, crash hour (hour of the day), crash type, lighting conditions, and “road divided by….” The analysis identifies which associated variables must be omitted from the study to prevent redundancy in the ML model by identifying correlated variables. The variables “highway type,”“environmental condition,”“work-zone related,”“median type,” and “temporary traffic control zone” were excluded from the study based on the correlation matrix analysis and Figure 3. A total of 21 input variables were finally chosen to construct an ML model with “crash severity” as the dependent variable.

Feature importance of input variables: (a) RORIDD (Extratree), (b) ROR (Extratree), (c) RORIDD (XGBoost), and (d) ROR (XGBoost)
While the presence of cellphone use or drowsy driving did not emerge as dominant influences in the overall model predictions, their relative importance exhibited a slight increase in the context of RORIDD crashes compared with general ROR crashes, particularly in the ExtraTree classifier. This observation suggests that these distraction-related factors may have a more relevant role when restricted to RORIDD crashes. However, it is important to note that both variables remain in the lower portion of the ranked feature importance plots, indicating that their overall contribution to crash prediction is still limited compared with other key factors. Future analyses incorporating additional statistical metrics or alternative modeling approaches may help further clarify these patterns. This finding was consistently reflected across different feature importance techniques, highlighting its contribution to the accuracy and relevance of the results.
Model Performance
A classification model is constructed using ML techniques, as outlined in the previous section. The classification model’s codes are constructed utilizing a Python package, scikit-learn, an open-source library ( 86 ). Like some previous studies, this study has allocated 80% of the entire dataset for training classification models and the remaining 20% for evaluating the developed models ( 87 , 88 ). The training and test sets are partitioned randomly to ensure that the performance of the test set is robust when applied to an entirely unfamiliar dataset. To enhance the precision of the model evaluation, a K-fold cross-validation technique with a value of n = 5 was employed on each ML method. Moreover, consistency between the training and test sets is ensured when employing different ML approaches for training and testing a classification model. Significantly, the hyperparameters of the algorithms were adjusted, and the optimal combination of hyperparameters for each model was selected to predict crash severity.
Figure 4 illustrates the use of a confusion matrix as a simple tool to assess the performance of a classification model concerning the original class versus the expected class. The confusion matrix summarizes the accuracy of predictions, with diagonal elements representing correct predictions and non-diagonal elements indicating incorrect ones. In Figure 4, the overall accuracy of the test dataset for all ML algorithms applied ranges from 60% to 65%. Among them, RF exhibits the lowest accuracy (60%), while AdaBoost performs the best (65%). However, relying solely on overall accuracy may not fully capture the classification model’s performance. Therefore, a more comprehensive comparison is made using recall and precision values.

Performance evaluation of different machine learning algorithms: (a) support vector machine, (b) random forest, (c) XGBoost, (d) LightGMB, (e) CatBoost, and (f) AdaBoost.
Furthermore, it is important to identify the class type that is crucial for the classification model. This study considers crash severity related to injuries as the most critical type. Based on this consideration, both XGBoost and CatBoost demonstrate similar overall accuracy values (63%–64%) and recall values (5%) for the injury class. Concerning precision, XGBoost is outperformed by CatBoost by a very slight margin, achieving 26% precision for the injury class, compared with CatBoost’s precision of 30%. The model’s overall performance is more significantly influenced by a balanced combination of accuracy, precision, and recall rather than solely focusing on maximizing precision or recall values. Therefore, the CatBoost model performs relatively better in the severity prediction of RORIDD for the specific years chosen in New Jersey.
Interpretable ML (SHAP Values)
Using SHAP values, the effect of contributing factors on crash severity prediction was discussed further. The XGBoost algorithm has been widely employed in explainable AI to acquire SHAP values ( 85 , 89 ). In this study, we also looked at how factors affect the SHAP values used by the XGBoost models to predict how severe a crash will be. The global importance of the input factors was used to find the average absolute SHAP values for each feature in the dataset (Figure 5a). The more important a factor is to the mean SHAP value, the bigger its value. Figure 5a shows how important each input variable is for each of the three crash severity levels: injury, no injury, and possible injury. However, it is most important to know the impact of the variables on the crashes resulting in injury, which is shown in the global SHAP plot in Figure 5b. According to Figure 5b, “driving under the influence of alcohol or drugs” is the factor that most affects how severe a crash will be, followed by “surface condition,”“road type,” and “occupant not restrained.” Even though the methods for selecting variables shown in Figure 3 gave similar results, the impact of features on predicting severity for each class needed to be classified from the methods for selecting features. This extra information makes the model easier to understand.

(a) Global Shapley additive explanation (SHAP) importance values and (b) summary plots for crashes resulting in injury.
In Figure 5b, summary diagrams illustrate the impact of the input variables on predicting injury severity. Each point on the graph represents a SHAP value corresponding to a specific input variable instance. The y-axis represents the input variables in ascending significance order. From blue (low) to red (high), the color of each dot reflects the value of the corresponding input variable. Input variables such as “alcohol involved,”“unrestrained occupant involved,” and “older driver involved” demonstrate a positive SHAP value for crashes resulting in injury. This implies that higher values for these variables, such as “alcohol involved = yes,” result in greater SHAP values and a higher likelihood of crashes resulting in injury. Similar to “rural or urban,”“intersection related” and “young driver involved” produced a higher SHAP value on the negative side, indicating a negative impact (lower values indicate a lower likelihood of injury crash). According to previous research, the presence of alcohol or drugs, elderly drivers, and so forth, increases the severity of crashes ( 90 ). Similarly, a clear surface increases the severity of collisions caused by careless drivers ( 65 ). ROR crashes have been found in a previous study to increase crash severity for unrestrained drivers distracted by cellphone use ( 21 ). An interesting trend observed in the dataset is that reported cellphone use is associated with a lower likelihood of injury in RORIDD crashes. However, given the small sample size of injury crashes involving cellphone use (only five cases) and the likelihood of underreporting, this finding should be interpreted with caution. Similarly, the effect of drowsy driving on injury severity follows a similar trend, exhibiting a lower injury rate for drowsy driving involved in RORIDD crashes. As both of these cases are self-reported, unless proven by dashcam/surveillance cameras, further research with more comprehensive data is necessary to better understand these associations.
To gain deeper insights into variables associated with distraction, a further analysis of SHAP values of possible injury for RORIDD and ROR was performed. It was found that the proportion of crashes leading to possible injury was supported by the descriptive statistics. Figure 6 illustrates that distractions such as “drowsy driving” and “cellphone use” increase the chance of possible injury in RORIDD and general ROR crashes. Drowsy driving, in particular, has a greater impact on crashes resulting in possible injury for RORIDD crashes than general ROR crashes. Although distracted driving variables have a smaller effect on injury crashes overall, they do contribute to an increase in crashes that result in possible injury. This effect is more pronounced in RORIDD crashes than in general ROR crashes. A prior study in New Jersey indicated that cellphone use could lead to more crashes with potential injury, especially on wet roads, involving pedestrians or bicyclists, and on undivided highways ( 23 ).

Shapley additive explanation (SHAP) summary plot for “crashes resulting in possible injury” for: (a) run-off-road involving distracted driving (RORIDD) and (b) run-off-road (ROR) crashes.
Figure 7 presents the SHAP dependency analysis for crashes resulting in injury, showcasing the fluctuations of the SHAP value in response to different input variables. The SHAP values depicted in both Figures 5 and 7 offer unique perspectives on the influence of the features. Figures 5 and 6 depict the average or aggregated SHAP values, illustrating the primary impacts of variables. Conversely, Figure 7 presents the interaction effects of the variables, providing a more comprehensive understanding of the dispersion and variability of SHAP values concerning the input variables. The primary emphasis of Figure 7a pertains to the impact of drivers under the influence of alcohol or drugs across various lighting conditions. Notably, it reveals that daylight has a higher SHAP value for “alcohol or drug involved,” indicating that such drivers are more susceptible to injury during ROR incidents involving distraction in daylight (light_condition = 3). This observation aligns with findings from a previous study that also reported more severe crashes occurring in daylight because of the potential lack of caution during daytime driving ( 24 ).

SHAP dependence plots of the top contributing factors for injury severity of RORIDD crashes: (a) light condition with alcohol or drugged driver involved, (b) older driver (65+) involved with surface condition, (c) curve related with surface condition, and (d) crash type with surface condition
Furthermore, the study identifies that crashes involving older drivers (older driver involved = 1) show more crashes resulting in injury when the road surface is wet. This aligns with findings from numerous previous studies, consistently indicating that older drivers and wet road surfaces contribute to increased crash severity ( 90 ). The low friction of wet surfaces can be associated with vehicle loss of control, particularly when traveling at higher speeds ( 20 , 21 ). The involvement of the curve (curve involved = 1) results in a lower likelihood of injury crashes for roadway departure crashes when the surface condition is toward the wet side (surface_condition = 1, as shown in the bar chart of Figure 7(c)). Drivers usually follow low speeds at curves when the surface becomes wet, which might be the reason for the decreased likelihood of injury crashes. A previous study found wet surfaces to increase the likelihood of crashes resulting in possible injury ( 22 ). The likelihood of injury crashes is higher for fixed object RORIDD crashes (crash type = 1) when the surface condition is toward the wet side (surface_condition = 1, as shown in the bar chart of Figure 7(d)). In a previous study, fixed object crashes have been found to be a major contributing factor for severe ROR crashes ( 35 ).
A further analysis was conducted to examine the interaction between driver inattention factors, such as “cellphone use” and “drowsy driving,” on injury and crashes resulting in possible injury in RORIDD crashes. The concurrent presence of cellphone use and alcohol/drug involvement in ROR crashes was associated with a lower likelihood of reported injury in this dataset (Figure 8). However, given that previous studies have found that alcohol consumption and cellphone use individually contribute to higher crash severity, this unexpected statistical association warrants further investigation. It is important to note that this analysis is limited to ROR crashes, which may differ from prior studies that examined a broader range of crash types, including non-ROR and non-distraction-related crashes ( 22 ). Future research should explore potential confounding factors that may help explain this finding. When examining surface conditions, it was found that wet surfaces are associated with a higher likelihood of crashes resulting in possible injury (Figure 8). This finding aligns with previous research. Furthermore, drowsy driving crashes are more likely to result in crashes resulting in possible injury when they involve a single vehicle ( 90 ). This observation is consistent with previous studies on drowsy driving crashes under various environmental conditions ( 91 ). Notably, drowsy drivers during daylight hours (light_condition = 3) have a higher likelihood of severe injury crashes than those driving in other lighting conditions.

SHAP dependence plots for distracted driving behaviors in RORIDD crashes: (a) cellphone use in possible injury, (b) cellphone use in injury, (c) drowsy in possible injury, and (d) drowsy in injury.
Conclusions
This study aimed to utilize ML and explainable AI in predicting the injury severity of RORIDD crashes in New Jersey. We began by preparing an extensive database, aggregating 30,995 crash records classified into three levels of severity: injury, no injury, and crashes resulting in possible injury. EDA was conducted to identify correlations among the variables and five variables were removed based on feature importance ranking and correlation matrix. After this process, a set of 22 variables was selected to develop the ML model, with the target variable being crash severity. Subsequently, we employed six different ML algorithms to establish the classification model. Using a confusion matrix, the performance of the model was thoroughly analyzed, considering three crucial evaluation metrics: overall accuracy, precision, and recall. Among the six ML models, CatBoost emerged as a relatively better performer across all three evaluation metrics, indicating the superiority of boosting methods over the SVM and RF models. Furthermore, SHAP dependence plots were utilized to reveal that alcohol consumption combined with daylight conditions had a higher likelihood of leading to injury crashes. An extended analysis of general ROR crashes was conducted to identify the specific contribution of distracted driving factors to RORIDD crashes. The analysis revealed that distractions, such as cellphone use and drowsy or fatigued driving, contribute significantly to the likelihood of crashes resulting in possible injury. Specifically, cellphone use in wet conditions and drowsy driving in single-vehicle crashes were associated with an increased likelihood of moderate-severe crashes (crashes resulting in possible injury).
These findings hold significant implications for transportation safety professionals, engineers, and policymakers in devising appropriate countermeasures to prevent ROR crashes involving distracted driving. According to FHWA, human factors such as speeding, impaired driving, distracted driving, and failing to wear a seatbelt lead to more ROR crashes and increase their severity ( 92 ). Therefore, strategic safety measures targeting those behaviors could prevent ROR crashes or decrease the severity of ROR crashes. Some of the findings of this study point out the importance of corrective measures to address human factors. For instance, stringent law enforcement efforts could target alcohol-involved driving during daylight hours. Enforcing unbuckled drivers through educational awareness campaigns could help reduce the injury severity of RORIDD crashes. Special attention should also be given to enhancing safety measures in spot locations with wet roads, especially by rehabilitating the pavement surface ( 92 ). Moreover, improving road lighting conditions or installing retroreflective paint materials can contribute to safer driving experiences ( 93 ). According to FHWA, the installation and improvement of the visibility and location of roadside warning signs could also be advantageous in inclement weather conditions or at curves ( 94 ). Finally, the use of advanced features such as LDW and Lane Keep Assist could help drivers stay on track ( 95 ). LDW has been found particularly useful in preventing drowsy, fatigued drivers from falling asleep and departing lanes ( 96 ). Both drivers’ lane-keeping performance and propensity to get drowsy could be mitigated by the introduction of driver warning and feedback technology on the dashboard, which could contribute to fewer RORIDD crashes ( 97 ). The use of cell phones could be mitigated by the use of onboard warning technologies, while roadway departure in adverse weather conditions could be mitigated by making variable speed limits in critical areas ( 98 , 99 ). Overall, this study provides valuable insights that can guide the development of effective strategies to mitigate ROR crashes involving distracted driving, ultimately enhancing road safety and safeguarding the well-being of all road users.
This study acknowledges several limitations that should be considered. Firstly, it only analyzed crash data spanning a 5-year period, which might not fully capture long-term trends or changes in RORIDD incidents. Secondly, the presence of missing values in the dataset required their removal, leading to a reduction in the overall data used for analysis. Lastly, the merging of crash severity classes (i.e., fatal, major injury, and minor injury) was necessary because of their low frequency, potentially limiting the granularity of severity assessment. To address these limitations, future studies on RORIDD crashes could benefit from examining crash data over more extended periods to obtain a more comprehensive understanding of the phenomenon. Another limitation of this study is the inability to analyze specific types of distraction for each case. Because of the nature of crash reporting, the only distraction type consistently documented is cellphone use, which is often underreported because of confirmation bias. As a result, some cases where cellphone use contributed to RORIDD crashes may remain unidentified. Additionally, the impact of other distractions, such as eating, drinking, reaching for objects, talking to passengers, or tuning the radio, could not be assessed individually because of reporting limitations. Furthermore, considering drowsy driving as a factor leading to driver inattention introduces another challenge. Some cases where drowsy driving contributed to ROR crashes may have been classified under “no distraction” since drowsy driving is not categorized as a form of distraction in New Jersey crash records. Lastly, the use of “no cellphone” was limited in the injury class, since most distraction categories were perceived as cellphone use by people who were unaware of other types of distraction.
Footnotes
Author Contributions
The authors confirm contribution to the paper as follows: study conception and design: M. Jalayer, A. Hasan; data collection: A. Hasan, M. Jalayer; analysis and interpretation of results: A. Hasan, S. Das, M. Jalayer; draft manuscript preparation: A. Hasan, M. Jalayer, S. Das. All authors reviewed the results and approved the final version of the manuscript.
Declaration of Conflicting Interests
The authors declared the following potential conflicts of interest with respect to the research, authorship, and/or publication of this article: Two of the authors (Dr. Subasish Das and Dr. Mohammad Jalayer) serve on the Transportation Research Record’s editorial board.
Funding
The authors received no financial support for the research, authorship, and/or publication of this article.
