Abstract
Objective:
To develop and validate machine learning models to predict 5-year survival in oral cancer using a large population-based registry and to rank prognostic factors.
Methods:
We analyzed Surveillance, Epidemiology, and End Results (SEER) data from 1992 to 2020. After applying inclusion criteria, 39 904 patients with complete data on 21 variables were included from 53 611 cases (25.6% exclusion). Selection bias was assessed by comparing included and excluded patients. The outcome was binary 5-year survival. Four models were trained and evaluated using nested fivefold cross-validation: XGBoost, LASSO, Random Forest, and logistic regression. Performance was assessed using the Brier score, area under the receiver operating characteristic curve (AUC), sensitivity, and specificity. Calibration was evaluated using slope and intercept with 95% confidence intervals. Model explainability used permutation feature importance and SHAP values.
Results:
Random Forest showed the best discrimination (AUC 77.6%, accuracy 71.4%, Brier 0.186) and was selected for risk stratification. However, it overestimated risk in lower deciles. Logistic regression and LASSO showed better calibration, with slopes near 1.0, but slightly lower discrimination (AUCs 75.5% and 76.9%). SHAP analysis identified localized stage as the strongest protective factor (importance 100.0), followed by age (91.1) and chemotherapy (29.3). Excluded patients had more unstaged tumors (3.9% vs 1.7%, P < .001).
Conclusion:
Random Forest provides strong risk stratification, but miscalibration limits its use for absolute risk prediction. Logistic regression and LASSO may be preferable when accurate probabilities are needed. External validation is required before clinical use.
Introduction
Cancers of the lip and oral cavity remain among the most commonly diagnosed malignancies within head and neck cancers and contribute substantially to the global cancer burden. In the United States, an estimated 434 915 individuals were living with oral or oropharyngeal cancer in 2021, reflecting both improved survival and continued disease incidence. 1 In 2024, approximately 58 450 new cases and 12 230 deaths are expected. 2 Although tobacco control efforts previously contributed to declining incidence, recent data show a gradual increase of about 1% per year since 2009. 1 This shift likely reflects evolving risk factor profiles, including tobacco use, alcohol consumption, and human papillomavirus-related disease. 3 Commonly affected sites include the tongue, tonsils, gingiva, and floor of the mouth, which differ in biological behavior and clinical presentation.
Despite the accessibility of the oral cavity to direct examination, a large proportion of cancers are still diagnosed at advanced stages. Nearly 70% of cases present with regional or distant disease, which is associated with worse outcomes and more intensive treatment. 1 Improving early detection remains a major priority. The Healthy People 2030 initiative aims to increase early-stage diagnosis to at least 34.2%. 4 Survival outcomes vary substantially by stage. Data from the Surveillance, Epidemiology, and End Results (SEER) program indicate an overall 5-year relative survival of approximately 69%, but this ranges from about 86% for localized disease to 40% for distant disease.1,5,6 These differences highlight the importance of early detection while also reflecting the influence of tumor biology, treatment, access to care, and sociodemographic factors.7-10
In parallel, there has been growing interest in applying machine learning methods to improve outcome prediction in oncology. These approaches can capture complex, nonlinear relationships and interactions among variables and are well suited for large population-based datasets. Prior studies in oral cancer have applied machine learning to survival prediction, often focusing on specific subsites such as tongue squamous cell carcinoma.11,12 While informative, these studies are frequently limited by small sample sizes, single-institution data, or narrow clinical scope, which may reduce generalizability. 13
Another limitation in the literature is the limited emphasis on model interpretability and systematic evaluation of predictor importance. Although many studies report performance metrics such as accuracy or area under the receiver operating characteristic curve, fewer examine how individual variables contribute to predictions. This is critical for clinical translation, where understanding the relative influence of factors such as stage, age, treatment, and sociodemographic characteristics is essential. In addition, inconsistent validation strategies and limited use of robust methods such as nested cross-validation raise concerns about overfitting and reproducibility. 13
This study addresses these gaps using a large, population-based dataset and a rigorous analytical framework. We included all major oral cavity subsites and applied multiple machine learning algorithms to enable direct comparison of performance. Nested cross-validation was used to improve the reliability of performance estimates and reduce overfitting. 14 To enhance interpretability, we applied both permutation importance and SHAP (Shapley Additive Explanations) methods to quantify and rank predictor contributions. Although the prognostic relevance of factors such as age, stage, and treatment is well established,7,8,10 their relative importance when considered simultaneously within a population-based machine learning framework remains unclear. This distinction is important because clinical decisions typically involve multiple interacting factors rather than isolated variables.
Compared with previous oral cancer machine-learning studies, which were typically limited to single institutions, specific subsites (such as tongue cancer), or modest sample sizes,11-13 this study uses a large, population-based registry that includes all major oral cavity subsites. We additionally apply nested cross-validation to reduce overfitting and systematically rank prognostic factors using permutation importance and SHAP values within a unified modeling framework. We hypothesized that 5-year survival in oral cancer can be predicted with good accuracy using machine learning models trained on population-based data, and that stage at diagnosis would be the most influential predictor. The primary objective was to develop and internally validate predictive models using SEER data.15,16 The secondary objective was to rank the relative importance of clinical and demographic predictors to better understand their contribution to survival and inform future clinical and research applications.
Materials and Methods
Data Source and Study Population
Data for this study were obtained from the Surveillance, Epidemiology, and End Results (SEER) Research Plus incidence dataset, specifically the November 2022 submission, which includes cases diagnosed between 1992 and 2020 and was released in April 2023. 15 The SEER program collects cancer incidence and survival data from population-based registries that together cover approximately 48% of the United States population. The program is supported by the National Cancer Institute and is widely used for epidemiologic and outcomes research due to its standardized data collection and broad population coverage. 16
From this dataset, we identified 53 611 individuals diagnosed with oral cancer and extracted 33 variables for analysis, as detailed in Appendix 1. 17 This study was designed as a retrospective cohort analysis using registry-based data. Reporting follows the TRIPOD + AI (Transparent Reporting of a Multivariable Prediction Model for Individual Prognosis or Diagnosis using Artificial Intelligence) guidelines, 18 and the completed checklist is provided in Supplementary File 1.
Selection of Variables and Participants
Eligibility Criteria
Variable selection was conducted prior to participant inclusion to ensure a structured and reproducible modeling process. Of the 33 initially available variables, 21 predictors were selected based on clinical relevance to survival outcomes and a predefined missing data threshold of less than 30%. When multiple variables captured overlapping clinical information, such as alternative coding of the same characteristic, a single representative variable was retained to reduce redundancy and improve model stability. The full list of selected predictors is provided in Appendix 1.
Participants were eligible for inclusion if they had complete data across all 21 selected predictors. Individuals were excluded if key variables, including diagnosis confirmation or tumor grade recode, were recorded as unknown. To ensure that all included patients had sufficient follow-up time for outcome assessment, only those diagnosed on or before December 31, 2015, were retained. This allowed for a minimum of 5 years of follow-up through the end of the dataset on December 31, 2020, and reduced the risk of bias related to administrative censoring.
After applying these criteria, the final analytic cohort consisted of 39 904 participants, corresponding to an exclusion of 13 707 individuals or 25.6% of the initial sample. The primary contributors to missingness included variables related to time from diagnosis to treatment, median household income, and detailed surgical procedure coding.
To evaluate potential selection bias introduced by complete-case analysis, we compared included and excluded participants across key demographic and clinical characteristics, including age, sex, race or ethnicity, stage, grade, and chemotherapy use. These comparisons were conducted using chi-square tests, and the results are presented in Supplemental Table S1.
Outcome Variable
The primary outcome was 5-year survival status, defined as a binary variable. Patients who survived longer than 5 years from the time of diagnosis were classified as “Yes,” while those who died within 5 years were classified as “No.” Outcome status was determined using vital status information recorded in SEER, with follow-up available through December 31, 2020.
Predictors
All 21 predictors were measured at or near the time of diagnosis and during the initial treatment period. These variables were grouped into three broad categories. Sociodemographic factors included age at diagnosis, sex, race or ethnicity, marital status, and median household income. Clinicopathological variables included primary tumor site, tumor grade, laterality, diagnostic confirmation, histologic subtype, and combined stage. Treatment-related variables included receipt of surgery, radiation therapy, chemotherapy, treatment sequencing, and time from diagnosis to treatment.
Statistical Analysis
Exploratory Data Analysis
Initial exploratory analyses were conducted to describe the study population and assess distributions of key variables. Categorical variables were summarized using frequencies and percentages. Bivariate associations between each predictor and 5-year survival were evaluated using chi-square tests to provide an initial assessment of potential relationships. A significance threshold of P < .05 was used for these analyses.
Multivariable Predictive Model Development
We developed and compared 4 machine learning models to predict 5-year survival: logistic regression, Least Absolute Shrinkage and Selection Operator (LASSO), Random Forest, and eXtreme Gradient Boosting (XGBoost). These models were implemented using the MachineShop package.
Logistic regression was included as a baseline model due to its interpretability and widespread use in clinical research. LASSO was selected for its ability to perform regularization and automatic variable selection, which can improve model performance in the presence of correlated predictors. Random Forest was chosen for its ensemble-based structure and ability to capture nonlinear relationships and interactions without requiring explicit specification. XGBoost was included due to its strong predictive performance and computational efficiency in structured data settings.
For logistic regression and LASSO models, continuous variables, specifically age at diagnosis and time from diagnosis to treatment, were modeled as linear terms. For tree-based methods, these variables were retained in continuous form, allowing the algorithms to identify nonlinear patterns and interactions directly.
Data Preprocessing
Prior to model training, all predictors were processed using the recipe package. Continuous variables were standardized, and categorical variables were appropriately encoded. Although scaling is particularly important for penalized regression methods such as LASSO, it was applied uniformly across all models to ensure consistency in model comparison.
Missing Data
A complete-case analysis approach was used, excluding all participants with missing values in any of the selected predictors. This resulted in the exclusion of 25.6% of the initial cohort. To assess the potential impact of this approach, we compared included and excluded participants across key variables, as described above.
Sample Size Considerations
No formal sample size calculation was performed prior to analysis. Instead, all eligible participants meeting inclusion criteria were included. The final sample size of 39 904 participants, including 15 622 observed events, exceeds commonly recommended thresholds for prediction modeling studies and provides sufficient statistical power for model development and validation.
Model Fitting and Training
To reduce the risk of overfitting and obtain reliable estimates of model performance, we used nested fivefold cross-validation. 14 This approach involves 2 levels of resampling. The inner loop is used for hyperparameter tuning, while the outer loop is used to evaluate model performance on unseen data. This separation helps prevent optimistic bias in performance estimates. Each model was trained using the same set of 21 predictors. Logistic regression and LASSO models assumed linear relationships between predictors and the log-odds of the outcome, whereas Random Forest and XGBoost were able to capture more complex, nonlinear relationships and interactions. Final performance metrics were calculated by averaging results across all outer folds.
Model Calibration and Explainability
Model calibration was assessed to evaluate the agreement between predicted probabilities and observed outcomes. Calibration plots were generated by plotting observed event rates against predicted probabilities. In addition, calibration slope and intercept were calculated, with ideal values of 1 and 0, respectively. These results are reported with 95% confidence intervals in Supplemental Table S3.
To improve interpretability, we assessed global feature importance using permutation-based methods. For the Random Forest model, SHAP values were computed to provide a more detailed understanding of how individual predictors influenced model predictions. These values were used to examine both the magnitude and direction of predictor effects, and results are presented in Supplemental Table S2 and Figure 2.
Model Evaluation
Model performance was evaluated using multiple complementary metrics, including Brier score, area under the receiver operating characteristic curve (ROC AUC), accuracy, sensitivity, and specificity. These metrics were derived from the outer folds of the nested cross-validation framework to ensure unbiased estimates of generalization performance. Comparative performance across models is presented in Table 3. Where applicable, performance estimates are reported with 95% confidence intervals, as detailed in Supplemental Table S3.
Results
Descriptive Statistics
The final analytic sample included 39 904 patients, of whom 39.15% survived for at least 5 years following diagnosis. Most patients were between 55 and 69 years of age at diagnosis (approximately 53%). The cohort was predominantly male (60.48%), non-Hispanic White (77.78%), and married (43.43%). A substantial proportion of patients were from higher-income groups, with 34.97% residing in areas with a median household income of at least $75 000. In terms of tumor characteristics, the anterior tongue was the most common primary site (43.03%), followed by the gums (15.73%), while the palate, excluding the soft palate and uvula, was the least common site (4.01%). Most tumors were moderately differentiated (grade II, 53.88%) and classified as localized at diagnosis (56.84%). Nearly all cases (99.81%) were confirmed by positive histology. Regarding treatment, only 19.62% of patients received chemotherapy, whereas the majority initiated treatment within 2 months of diagnosis (91.83%). Descriptive details are provided in Table 1.
Demographic and Clinical Characteristics of the Study Cohort (N = 39 904).
Comparison of Included versus Excluded Patients
A comparison between included (n = 39 904) and excluded (n = 13 707) patients is presented in Supplemental Table S1. Excluded individuals differed in several key characteristics. They had a more than double the proportion of unstaged tumors (3.9% vs 1.7%, P < .001), a lower proportion of localized disease (52.5% vs 56.8%), and a higher proportion of regional disease (32.7% vs 30.9%). Differences were also observed across race and ethnicity, with a higher proportion of non-Hispanic Black patients in the excluded group (8.1% vs 5.4%). There was no statistically significant difference in tumor grade distribution between the 2 groups (P = .110). These findings indicate that the analytic cohort is more completely documented and presents with an earlier stage at diagnosis compared to excluded patients. This suggests a potential for modest overestimation of survival by the predictive models, a key limitation which is addressed in the discussion.
Bivariate Analysis
Bivariate associations between predictors and 5-year survival are summarized in Table 2. Age at diagnosis showed a strong association with survival (χ2 = 2372.15, P < .001). More than half of patients in most age groups survived at least 5 years, with the exception of those aged 80 to 84 years (43.21%) and those aged 85 years and older (31.34%), where survival rates were substantially lower. Sex, race or ethnicity, and primary tumor site were also significantly associated with 5-year survival. In contrast, laterality was not significantly associated with survival (χ2 = 7.69, P = .191). All reported chi-square statistics were verified to ensure accuracy following correction of a prior data reporting error.
Bivariate Associations Between Patient Characteristics and 5-Year Survival.
Prediction Model Performance
The performance of the machine learning models is presented in Table 3. Among the models evaluated, the Random Forest model demonstrated the best discrimination, with the highest area under the receiver operating characteristic curve (AUC = 77.6%) and the highest accuracy (71.4%). It also had the lowest Brier score (0.186).
Performance Metrics of Predictive Models.
The LASSO model achieved the highest sensitivity (82.0%), suggesting better identification of patients who did not survive 5 years, but this came at the cost of lower specificity (52.8%). Logistic regression performed comparably, with an AUC of 75.5%, supporting its continued relevance as a baseline model. The XGBoost model showed lower discrimination (AUC = 70.2%) and a higher Brier score (0.230), although its performance remained within the range reported in prior studies. Overall, model discrimination across all approaches ranged from 70.2% to 77.6%.
Calibration plots for all models are presented in Figure 1. Logistic regression and LASSO showed good agreement between predicted and observed risks across most deciles. In contrast, the Random Forest model substantially overestimated risk in the lowest predicted-risk deciles and underestimated risk in some mid-range deciles, indicating important miscalibration despite its favorable Brier score. The XGBoost model displayed intermediate calibration, with under-prediction at low risk and over-prediction at high risk. Quantitative calibration metrics, including the calibration slope and intercept for each model, are provided in Supplemental Table S3).

Calibration plots for all models. Points represent deciles of predicted risk (n ≈ 3990 per decile). Agreement between predicted probabilities (x-axis) and observed 5-year survival (y-axis) for the four machine learning models. Each point represents one decile of predicted risk (n = 3990 per decile). The dashed black line denotes perfect calibration (slope = 1, intercept = 0). Points below the line indicate overestimation of risk (predicted risk higher than observed), and points above the line indicate underestimation. Model colors: XGBoost (red), Random Forest (blue), LASSO (green), and logistic regression (purple). Quantitative calibration metrics (slope and intercept with 95% confidence intervals) for each model are shown in Supplemental Table S3. Logistic regression and LASSO showed the best overall calibration, whereas Random Forest substantially overestimated risk in the lowest predicted risk deciles.
Model Specification and Use
Random Forest was selected as the primary model for risk stratification and interpretability analyses because it achieved the highest discriminative performance. This model can be applied using the full set of 21 predictors listed in Appendix 1. For a given patient, the model generates a predicted probability of 5-year survival ranging from 0 to 1. However, because Random Forest was miscalibrated at low predicted risks, these probabilities are best interpreted for relative risk stratification (ranking patients by risk) rather than as exact absolute risks. For applications requiring accurate absolute probability estimates, logistic regression and LASSO currently provide more reliable calibrated probabilities. All models require recalibration and external validation before clinical use. Variable importance rankings, presented in Table 4, can be used to understand the relative contribution of each predictor to model performance.
Variable-Level Importance Rankings.
Model Interpretation
Permutation feature importance analysis identified stage at diagnosis as the most influential predictor, with a normalized importance score of 100.0. This was followed by age (91.1) and receipt of Radiation (66.2), indicating that both clinical severity and treatment-related factors play a central role in survival prediction.
In the Random Forest SHAP analysis (Figure 2), localized stage exerted the strongest protective effect, with consistently negative SHAP values indicating substantially lower predicted 5-year mortality among patients diagnosed with localized disease. Older age, particularly ⩾80 years, was associated with positive SHAP values, reflecting higher predicted mortality. Chemotherapy showed a bidirectional pattern, with both positive and negative SHAP values, consistent with confounding by indication: it is more often administered in advanced disease but may confer benefit in selected patients. Well-differentiated tumors, receipt of surgery, and being married were associated with lower predicted risk, whereas poorly differentiated grade, absence of radiation, and higher tumor sequence (non-first primaries) tended to increase predicted mortality. Time to treatment and specific primary sites showed SHAP values near zero, indicating minimal contribution to model predictions.

SHAP Beeswarm plot–Random Forest model. SHapley Additive exPlanations (SHAP) values for the 15 most influential predictors of 5 year survival in the Random Forest model, based on a random sample of 2000 patients from the full cohort (N = 39,904). Each point represents an individual patient. The position on the x axis indicates the SHAP value for that feature (impact on the log odds of 5 year mortality), with positive values increasing predicted mortality risk and negative values decreasing risk. Features are ordered on the y axis by mean absolute SHAP value (most important at the top). Point color represents the feature value (blue = low, red = high). The vertical dashed line at zero indicates no contribution to the prediction.
Discussion
Our analysis supported our primary hypothesis, demonstrating that cancer stage at diagnosis is the most powerful determinant of 5-year survival in oral cancer.1,5 This study achieved 2 principal goals. First, we developed and validated machine learning models to predict 5-year survival using a large, population-based cohort. Random Forest was selected as the primary model for risk stratification and interpretability analyses because it achieved the highest discrimination (AUC: 77.6%, accuracy: 71.4%). However, calibration plots revealed that Random Forest substantially overestimated risk in the lowest predicted-risk deciles, highlighting an important limitation. Second, our explainability analysis revealed a clear prognostic hierarchy: the strongest predictor of 5-year survival was localized cancer stage, which was more influential than any other variable, including age and treatment-related factors such as chemotherapy.
The Random Forest model achieved the strongest discriminative performance (AUC 77.6%, accuracy 71.4%, Brier score 0.186). However, calibration plots showed that it markedly overestimated risk among patients it classified as low risk and slightly underestimated risk in some intermediate-risk groups. The XGBoost model showed lower performance (AUC 70.2%, Brier score 0.230) and also demonstrated non-ideal calibration, with under-prediction at low risk and over-prediction at high risk. By contrast, logistic regression and LASSO had slightly lower AUCs (75.5% and 76.9%, respectively) but were substantially better calibrated across the range of predicted probabilities, making them preferable for applications requiring accurate absolute risk estimates.
The modest specificity observed across models (52.8%-66.6%) indicates that they are more effective at identifying patients who will die within 5 years than those who will survive. This pattern may still be clinically useful, as it supports the identification of high-risk patients who may benefit from closer monitoring or more aggressive management. 19 The near-identical performance of the LASSO model (AUC 76.9%) suggests that a substantial portion of the predictive signal is captured by linear relationships. 8 At the same time, the superior performance of Random Forest indicates that nonlinear interactions between predictors contribute additional predictive value. Together, these findings suggest that oral cancer survival is driven by a combination of strong linear effects, particularly stage and age, along with more complex interactions captured by ensemble methods. 13
Calibration is critical when predicted probabilities are intended to guide clinical decisions or risk communication. In our study, the simpler regression-based models (logistic regression and LASSO) produced more reliable absolute risk estimates than the tree-based models, particularly at the lower end of the risk spectrum. This suggests that, although Random Forest may be preferable for ranking patients according to risk, its raw predicted probabilities should be interpreted with caution or subjected to post-hoc calibration before being used as clinical decision thresholds.
Our finding that localized stage is the dominant prognostic factor has important implications. It reinforces the idea that shifting the stage distribution toward earlier disease through improved detection may have a greater impact on population-level survival than incremental advances in treatment for advanced cancer.1,5 This aligns with the goals of initiatives such as Healthy People 2030.
Age demonstrated a clear gradient, with increasing mortality risk observed in older age groups. Patients aged 85 years and older were more likely to have positive SHAP values, indicating a higher predicted probability of death. 20 This likely reflects both biological factors, such as reduced physiological reserve, and clinical factors, including differences in treatment intensity. Chemotherapy showed a bidirectional pattern, with approximately half of patients exhibiting positive SHAP values and half showing negative values. This likely reflects confounding by indication, where chemotherapy is more commonly used in advanced disease but may still provide benefit in selected patients. This type of pattern is not easily captured using traditional regression approaches and highlights the value of model interpretability methods.
Our bivariate findings were consistent with these results, demonstrating lower survival among older patients and those with more advanced disease. Survival patterns for favorable prognostic factors, such as localized stage and well-differentiated histology, were also consistent with prior studies.7,8 An unexpected finding was that undifferentiated (Grade IV) tumors showed slightly higher survival than poorly differentiated (Grade III) tumors. This contradicts the expected gradient of decreasing survival with increasing tumor grade. This finding may be due to the relatively small number of Grade IV cases, leading to unstable estimates. It may also reflect differences in treatment intensity that are not captured in registry data. This result should be interpreted with caution and warrants further investigation. In contrast to some previous studies that emphasized metastasis or tumor grade as dominant predictors, our analysis identified age as a more influential factor after stage. This may reflect the broader and more heterogeneous population included in this study. The least important predictors were site sequence, laterality, and time from diagnosis to treatment.
This study has several strengths, including the use of a large, population-based dataset and the application of multiple machine learning models within a consistent analytical framework.15,16 The use of nested cross-validation reduces the risk of overfitting and provides more reliable estimates of model performance. 14 In addition, the use of permutation importance and SHAP values enhances interpretability and supports clinically meaningful insights.
Several limitations should be considered. First, complete-case analysis resulted in the exclusion of 25.6% of the initial cohort, introducing potential selection bias. Excluded patients had more advanced disease and a higher proportion of unstaged tumors, indicating that the analytic cohort represents a healthier and better-documented subset. As a result, survival predictions may be overly optimistic, and generalizability may be limited. Second, although nested cross-validation provides strong internal validation, external validation in independent datasets is required before clinical implementation. 13 Third, the SEER database lacks detailed clinical variables such as performance status, tobacco and alcohol use, HPV status, and surgical margin data, which may improve predictive performance if available. Fourth, the study protocol was not prospectively registered, which may introduce reporting bias.
Despite these limitations, the findings have important implications. The dominant role of localized stage reinforces the need for improved early detection strategies, including routine screening in dental and primary care settings.3,4 At the same time, the moderate performance of the models suggests that registry-based variables alone may not be sufficient for precise individualized prediction. Future work should focus on integrating additional data sources, including molecular and genomic information, to improve predictive accuracy and clinical utility. 21
Conclusion
In conclusion, this study demonstrates that population-based cancer registry data can be used to develop interpretable machine learning models for predicting 5-year survival in oral cancer. Random Forest was selected as the primary model for risk stratification and explainability due to its superior discrimination, with SHAP analysis confirming localized stage as the dominant prognostic factor. However, due to miscalibration at low predicted risks, its probabilities require recalibration before use as absolute risk estimates. For applications requiring accurate absolute probabilities, logistic regression or LASSO are currently preferable. All models require external validation prior to clinical implementation.
Supplemental Material
sj-docx-1-cix-10.1177_11769351261442875 – Supplemental material for A Machine Learning Approach to Prognostication in Oral Cancer: Analysis of the Surveillance, Epidemiology, and End Results Database
Supplemental material, sj-docx-1-cix-10.1177_11769351261442875 for A Machine Learning Approach to Prognostication in Oral Cancer: Analysis of the Surveillance, Epidemiology, and End Results Database by Philips Okeagu, Saranya Murukutla, Afamdi Iwuchukwu and Chukwuebuka Ogwo in Cancer Informatics
Supplemental Material
sj-pdf-1-cix-10.1177_11769351261442875 – Supplemental material for A Machine Learning Approach to Prognostication in Oral Cancer: Analysis of the Surveillance, Epidemiology, and End Results Database
Supplemental material, sj-pdf-1-cix-10.1177_11769351261442875 for A Machine Learning Approach to Prognostication in Oral Cancer: Analysis of the Surveillance, Epidemiology, and End Results Database by Philips Okeagu, Saranya Murukutla, Afamdi Iwuchukwu and Chukwuebuka Ogwo in Cancer Informatics
Footnotes
Appendix
Predictor Variables Selected for Model Development.
| Variable name | Description/categories | Type |
|---|---|---|
| Sociodemographic | ||
| Age at diagnosis | Continuous (y) | Continuous |
| Sex | Male, female | Categorical |
| Race/Ethnicity | Non-Hispanic White, Non-Hispanic Black, Hispanic, Other | Categorical |
| Marital status | Married, single, divorced/separated, widowed | Categorical |
| Median household income | Categorical | Categorical |
| Clinicopathological | ||
| Primary site | Tongue, gums, floor of mouth, etc. | Categorical |
| Grades recode | Grade I (well diff.) to IV (undifferentiated) | Categorical |
| Laterality | Left, right, midline, paired | Categorical |
| Diagnosis confirmation | Positive histology, etc. | Categorical |
| Histology recode | Broad groupings (eg, squamous cell carcinoma) | Categorical |
| Combined stage | Localized, regional, distant | Categorical |
| Treatment-related | ||
| Treatment surgery primary site | Yes, No | Categorical |
| Radiation | Yes, No | Categorical |
| Chemotherapy | Yes, No | Categorical |
| Treatment surgery radiation sequence | eg, surgery before radiation, etc. | Categorical |
| Systemic surgery sequence | Sequence of systemic therapy and surgery | Categorical |
| Months from diagnosis to treatment | Continuous (months) | Continuous |
| Cancer history & identification | ||
| Sequence number | Order of multiple primary cancers | Continuous/Categorical |
| Record number | Unique identifier | Identifier |
| Site sequence | Information on multiple tumors | Categorical |
| Outcome | ||
| 5-Y survival | Deceased at ⩽5 y, Alive at > 5 y | Binary |
Acknowledgements
Ethical Considerations
Not applicable. This study used de-identified data from the publicly available Surveillance, Epidemiology, and End Results (SEER) database. The SEER database contains anonymized patient information collected by the National Cancer Institute and is exempt from Institutional Review Board approval. Access to the data was granted with permission from SEER following standard application procedures.
Consent to Participate
No patient consent was required as the data are publicly available and de-identified.
Consent for Publication
Not applicable. This manuscript does not contain any individual person’s data in any form (including individual details, images, or videos). All data presented are aggregated and anonymized.
Author Contributions
Philips Okeagu: Conceptualization, Methodology, Formal analysis, Writing – Original Draft, Writing – Review & Editing. Saranya Murukutla: Conceptualization, Methodology, Formal analysis, Writing – Original Draft. Afamdi Iwuchukwu: Data Curation, Formal analysis, Writing – Original Draft, Visualization. Chukwuebuka Ogwo: Conceptualization, Methodology, Formal analysis, Writing – Original Draft, Writing – Review & Editing, Supervision. All authors reviewed and approved the final manuscript and agree to be accountable for all aspects of the work.
Funding
The authors received no financial support for the research, authorship, and/or publication of this article.
Declaration of conflicting interests
The authors declared no potential conflicts of interest with respect to the research, authorship, and/or publication of this article.
Data Availability Statement
All data used in this study were obtained from the Surveillance, Epidemiology, and End Results (SEER) database.
15
The specific dataset used was the SEER Research Plus incidence data from the November 2022 submission (covering 1992-2020), released in April 2023. Data are available upon request from SEER (
) following standard application procedures. The analytical code used for model development and evaluation is available from the corresponding author upon request.
Supplemental Material
Supplemental material for this article is available online.
References
Supplementary Material
Please find the following supplemental material available below.
For Open Access articles published under a Creative Commons License, all supplemental material carries the same license as the article it is associated with.
For non-Open Access articles published, all supplemental material carries a non-exclusive license, and permission requests for re-use of supplemental material or any part of supplemental material shall be sent directly to the copyright owner as specified in the copyright notice associated with the article.
