Abstract
Objectives
Hyperlipidemic acute pancreatitis progresses rapidly to severe acute pancreatitis. Early identification of disease severity is critical for improving outcomes. This study aimed to investigate the risk factors associated with severe acute pancreatitis and to develop and validate a novel predictive model to support clinical decision making.
Methods
This retrospective cohort study included 502 patients with hyperlipidemic acute pancreatitis. A total of 502 patients with hyperlipidemic acute pancreatitis were retrospectively enrolled and randomly assigned to a training set (n = 351) and a validation set (n = 151) in a 7:3 ratio. Least absolute shrinkage and selection operator regression and multivariate logistic regression were used for model development. Model performance was comprehensively evaluated using the receiver operating characteristic curve, calibration curves, the Hosmer–Lemeshow test, Brier score, calibration slope, calibration-in-the-large, and decision curve analysis.
Results
Multivariate logistic regression confirmed the bedside index for severity in acute pancreatitis score (odds ratio = 7.042, 95% confidence interval: 3.850 to 14.145, p < 0.001) and metabolic score for insulin resistance (odds ratio = 1.053, 95% confidence interval: 1.023 to 1.087, p < 0.001) as independent risk factors for severe acute pancreatitis. The resulting predictive model demonstrated excellent discriminative ability in both the training set (area under the curve = 0.904, 95% confidence interval: 0.852 to 0.955) and the validation set (area under the curve = 0.885, 95% confidence interval: 0.812 to 0.958). In the training set, the area under the curve of the predictive model was significantly higher than those of the individual indicators, including metabolic score for insulin resistance, triglyceride-glucose index, triglyceride-glucose body mass index, triglycerides/high-density lipoprotein cholesterol, and the bedside index for severity in acute pancreatitis score (all p < 0.05). In the validation set, the model yielded only a minimal improvement in area under the curve over the bedside index for severity in acute pancreatitis score alone (difference = 0.020), which was not statistically significant (p = 0.326). The calibration slope and calibration-in-the-large were 1.00 and 0.00 in the training set and 0.82 and −0.54 in the validation set, respectively. Calibration curves, the Hosmer–Lemeshow test, and Brier scores collectively indicated good model fit and high predictive accuracy. Furthermore, decision curve analysis showed that the combined model provided superior net clinical benefit across a wide range of threshold probabilities.
Conclusion
The combined model incorporating the bedside index for severity in acute pancreatitis score and metabolic score for insulin resistance serves as a preliminary risk stratification tool for the early identification of severe acute pancreatitis in patients with hyperlipidemic acute pancreatitis.
Keywords
Background
Acute pancreatitis (AP) is a common gastrointestinal disorder characterized by both local and systemic inflammatory responses, with a highly variable clinical course. Mild AP is usually self-limiting; however, approximately 20% of patients progress to moderately severe or severe AP, with a mortality rate ranging from 20% to 40%. 1 A meta-analysis by Iannuzzi et al.,2 including 44 studies published between 1961 and 2016, demonstrated a gradual global increase in the incidence of AP, with an average annual percent change of 3.07%. The major etiologies of AP include biliary disease, hyperlipidemia, and alcohol consumption, 1 although their distribution varies significantly across regions. 3 In China, with changes in dietary patterns and the rising prevalence of metabolic syndrome (MS), the incidence of hyperlipidemia-induced acute pancreatitis (HLAP) has increased steadily, becoming the second leading cause of AP.4–7 Compared with AP of other etiologies, HLAP is more likely to progress to severe acute pancreatitis (SAP), leading to higher rates of complications and mortality. 5 Therefore, early assessment of disease severity in patients with HLAP is of great clinical importance.
MS is a cluster of metabolic disorders characterized by dyslipidemia, obesity, hyperglycemia, and hypertension, all of which are closely associated with an increased risk of adverse clinical outcomes in AP. 2 Insulin resistance (IR) is the central pathophysiological feature of MS, and accumulating evidence suggests that IR is strongly correlated with the severity and prognosis of HLAP.8–10 The metabolic score for insulin resistance (METS-IR) is a novel surrogate marker that integrates anthropometric parameters with glucose and lipid metabolism indicators. It is simple, cost-effective, and has shown good performance in assessing IR. 11 In recent years, METS-IR has been increasingly applied in metabolic diseases and has demonstrated promising predictive value12–14; however, its role in evaluating the severity of HLAP remains to be fully elucidated.
Currently, several scoring systems are available for assessing the severity of AP, such as the Ranson score and the acute physiology and chronic health evaluation II (APACHE II) score. However, their clinical application in emergency settings is limited because of the large number of variables required, computational complexity, and time consumption. The bedside index for severity in acute pancreatitis (BISAP) score, in contrast, is simple, rapid, and can be completed within 24 h of admission. It has been validated as an effective tool for predicting disease severity in patients with AP.15,16 HLAP is closely associated with IR and lipid metabolism disorders. 2 However, these metabolic factors are not incorporated into the BISAP score, which may limit its predictive performance in patients with HLAP.
Therefore, this study aimed to retrospectively evaluate the predictive value of METS-IR and the BISAP score, both individually and in combination, for SAP. Furthermore, a nomogram prediction model was constructed to facilitate early identification of patients with SAP and enable timely intervention, ultimately improving clinical outcomes.
Methods
Study population and grouping
We retrospectively analyzed data from consecutive hospitalized patients diagnosed with HLAP at Tianyou Hospital, Wuhan University of Science and Technology, between January 2020 and December 2024 as well as consecutive patients with HLAP admitted to the General Hospital of Central Theater Command of the People's Liberation Army between January 2019 and December 2022. There was no overlap between patients from the two centers included in this study. All cases were cross-verified through unique identification numbers (hospitalization numbers) in the electronic medical record system to ensure the independence of the study participants, and all patient data were deidentified to ensure anonymity. No prospective sample size calculation was conducted in this study; however, a post hoc assessment based on the 10 events per variable (EPV) criterion verified the robustness of the data for multivariate modeling. The final model incorporated two independent predictive factors (METS-IR and BISAP score). Based on the 10 EPV criterion, the minimum number of SAP outcome events required was calculated to be 20 cases, and the estimated minimum total sample size was approximately 193 cases according to the actual SAP incidence of 10.36%. A total of 502 patients with HLAP were enrolled in this study and divided into a training set (n = 351) and a validation set (n = 151) in a 7:3 ratio. A total of 52 SAP events were observed in the entire cohort (34 in the training set and 18 in the validation set), with an EPV of 17 for the training set and 26 for the entire cohort. Both values were significantly higher than the general threshold of 10 EPV, indicating that the sample size was sufficient for reliable parameter estimation and fully met the requirements for constructing a multivariate model.
This inclusion criteria were as follows: (a) fulfillment of the diagnostic criteria for AP 17 and (b) triglyceride (TG) level ≥11.3 mmol/L or TG level between 5.65 and 11.3 mmol/L accompanied by lipemic serum. The exclusion criteria were as follows: (a) concomitant AP attributable to other etiologies; (b) history of chronic pancreatitis; (c) presence of active malignancy; (d) pregnancy or lactation; (e) time from symptom onset to hospital admission exceeding 72 h; and (f) incomplete clinical data.
According to the Revised Atlanta Classification of 2012, 18 patients were categorized as having mild acute pancreatitis (MAP), moderately severe acute pancreatitis (MSAP), or SAP. MAP was defined as no organ failure; MSAP as transient organ failure resolving within 48 h and/or local or systemic complications; and SAP as organ failure persisting for more than 48 h.
A flowchart of the study is shown in Figure 1. This study was conducted in accordance with the Declaration of Helsinki (1975), as revised in 2024, and was approved by the Ethics Committees of Tianyou Hospital, Wuhan University of Science and Technology (Approval Number: LL2025-07-22-01) and the General Hospital of Central Theater Command (Approval Number: (2024)090-01). The requirement for written informed consent was waived because the retrospective nature of the study. The reporting of this study conforms to the Strengthening the Reporting of Observational Studies in Epidemiology (STROBE) guidelines. 19

A flowchart of the study.
Data collection
Clinical data were retrospectively collected from all enrolled patients with HLAP, including sex; age; body mass index (BMI); history of hypertension, diabetes mellitus, and fatty liver disease; and laboratory parameters on admission, including C-reactive protein(CRP), neutrophil count (NEUT), lymphocyte count (LYM), red cell distribution width (RDW), platelet count (PLT), albumin (ALB), total bilirubin (TBil), direct bilirubin (DBil), alanine aminotransferase (ALT), aspartate aminotransferase (AST), blood urea nitrogen (BUN), creatinine (Cr), fasting plasma glucose (FPG), potassium (K+), calcium (Ca2+), TG, high-density lipoprotein cholesterol (HDL-C), low-density lipoprotein cholesterol (LDL-C), activated partial thromboplastin time (APTT), prothrombin time (PT), and D-dimer.
In addition, the following composite indices were calculated for each patient: triglyceride-glucose (TyG) index, triglyceride-glucose body mass index (TyG-BMI), triglyceride/high-density lipoprotein cholesterol ratio (TG/HDL-C), METS-IR, and BISAP score.
Calculation formulas
The following formulas were used to calculate the metabolic indices:
BMI = Weight (kg)/[Height (m)]2
TyG = Ln[FPG(mg/dL) × TG (mg/dL)/2] 8
TyG-BMI = TyG × BMI (kg/m2) 9
TG/HDL-C = TG(mg/dL) ÷ HDL-C(mg/dL) 10
METS-IR = Ln[2 × FPG(mg/dL) + TG(mg/dL)] × BMI(kg/m2) ÷ Ln[HDL-C(mg/dL)] 11
BISAP score. The BISAP comprises five binary components, each assigned 1 point: 20 (a) BUN >8.9 mmol/L (>25 mg/dL); (b) impaired mental status; (c)p of systemic inflammatory response syndrome (SIRS); (d) age >60 years; and (e) pleural effusion on imaging. The total BISAP score ranges from 0 to 5.
Statistical analysis
Statistical analyses were performed using Statistical Package for the Social Sciences (SPSS) version 27.0 (IBM Corp.; Armonk, NY, USA) and R software version 4.5.0 (R Foundation for Statistical Computing; Vienna, Austria). Variables with a missing rate of <10% were imputed using multiple imputation. The imputation model included all study variables; predictive mean matching (PMM) was used for continuous variables, and logistic regression imputation was used for categorical variables. Five imputed datasets were generated, and statistical analyses were performed using the pooling rules for multiple imputation. Variables with a missing rate of >10% were excluded. The missing rate for each variable is detailed in Table S1. The entire study cohort was randomly divided into a training set and a validation set in a 7:3 ratio using a computer-generated random number sequence in SPSS 27.0, and the randomization process was verified to ensure no selection bias between the two sets. Continuous variables conforming to a normal distribution were expressed as mean ± SD (x̄ ± s) and compared among groups using one-way analysis of variance (ANOVA). Skewed continuous variables were reported as the median with interquartile range (M (Q1 and Q3)) and compared using the Kruskal–Wallis H test. Categorical variables were presented as frequencies (percentages) and analyzed using the chi-square (χ2) test. Variables were screened using least absolute shrinkage and selection operator (LASSO) regression, and 10-fold cross-validation was applied to determine the optimal λ value. The variables retained after screening were included in a multivariable logistic regression analysis to identify independent risk factors for SAP. A p value <0.05 was considered statistically significant. Multicollinearity among key predictive variables was assessed using the variance inflation factor (VIF), with VIF <10 indicating no significant multicollinearity. Based on the independent risk factors without multicollinearity, a nomogram prediction model was constructed. The model's discriminative ability was evaluated using the receiver operating characteristic (ROC) curve, and its calibration was assessed using a calibration curve. Calibration curves were constructed and bias-corrected with 1000 bootstrap resamples to minimize small sample bias and improve the reliability of the results. Model calibration was comprehensively evaluated using the Hosmer–Lemeshow test, Brier score, calibration slope, and calibration-in-the-large. Clinical decision curve analysis (DCA) was performed to assess the clinical utility of the predictive model. A complete-case sensitivity analysis was performed to validate the robustness of the results (Figures S1 to S3 and Tables S2 and S3).
Results
Baseline and clinical characteristics of patients in the training set
Table 1 presents the baseline characteristics of the 351 patients in the training set, stratified into three severity groups according to the Revised Atlanta Classification: MAP (n = 181), MSAP (n = 136), and SAP (n = 34). Significant intergroup differences (all p < 0.05) were observed in NEUT, DBil, FPG, TG, HDL-C, D-dimer, BMI, TyG index, TyG-BMI, TG/HDL-C ratio, METS-IR, and BISAP score. Notably, patients in the SAP group exhibited significantly higher levels of DBil, FPG, TG, D-dimer, BMI, TyG index, TyG-BMI, TG/HDL-C, METS-IR, and BISAP scores than those in the MAP and MSAP groups, whereas HDL-C levels were significantly lower in the SAP group.
Baseline characteristics of the MAP, MSAP, and SAP groups in the training set.
ALB: albumin; ALT: alanine aminotransferase; APTT: activated partial thromboplastin time; AST: aspartate aminotransferase; BISAP: bedside index for severity in acute pancreatitis; BMI: body mass index; BUN: blood urea nitrogen; Ca2+: calcium; Cr: creatinine; CRP: C-reactive protein; DBil: direct bilirubin; FPG: fasting plasma glucose; HDL-C: high-density lipoprotein cholesterol; K+: potassium; LDL-C: low-density lipoprotein cholesterol; LYM: lymphocyte count; MAP: mild acute pancreatitis; METS-IR: metabolic score for insulin resistance; MSAP: moderately severe acute pancreatitis; NEUT: neutrophil count; PLT: platelet count; PT: prothrombin time; RDW: red cell distribution width; SAP: severe acute pancreatitis; TBil: total bilirubin; TG: triglycerides; TG/HDL-C: triglyceride/high-density lipoprotein cholesterol ratio; TyG: triglyceride-glucose; TyG-BMI: triglyceride-glucose body mass index.
Variable selection
Variable selection was performed using LASSO regression, with the optimal λ value determined by 10-fold cross-validation. A total of seven variables with nonzero coefficients were identified: METS-IR, diabetes mellitus, BISAP score, CRP, PLT, BUN, and TG (Figure 2).

Selection of predictive factors using LASSO regression. (a) Coefficient profile plot of the LASSO regression model, from which seven predictive factors were ultimately retained; (b) determination the optimal λ value using 10-fold cross-validation in the LASSO model.
Construction of the nomogram model based on logistic regression
The seven predictors retained by LASSO regression were included in a multivariable logistic regression analysis. The results demonstrated that the BISAP score (odds ratio (OR) = 7.042, 95% confidence interval (CI): 3.850 to 14.145, p < 0.001) and METS-IR (OR = 1.053, 95% CI: 1.023 to 1.087, p < 0.001) were independent risk factors for SAP (Table 2). Multicollinearity testing showed that the VIF values for both METS-IR and BISAP score were 1.093, which was below the critical threshold of 10, indicating no significant multicollinearity between the two variables. Therefore, both variables were jointly incorporated into the model construction. Based on these independent risk factors, a novel predictive model was established using the following formula:
Multivariable logistic regression analysis.
BISAP: bedside index for severity in acute pancreatitis; BUN: blood urea nitrogen; CI: confidence interval; CRP: C-reactive protein; METS-IR: metabolic score for insulin resistance; OR: odds ratio; PLT: platelet count; SE: standard error; TG: triglycerides.
logit(P) = −7.293 + 0.051 × METS-IR + 1.952 × BISAP. A corresponding nomogram was plotted according to the regression results (Figure 3). The total score was calculated by summing the points corresponding to each variable, thereby predicting the individual probability of SAP occurrence.

Nomogram for predicting the risk of SAP. To use the nomogram, draw a vertical line from the specific value of each variable to the “Points” axis to determine the corresponding score. Sum the scores for all variables to obtain the total points, then locate this total on the “Total Points” axis and draw a vertical line down to the “Risk of SAP” axis to estimate the predicted probability.
Discriminatory performance of the prediction model
The predictive performance of the model and individual indicators was evaluated using ROC curves. In the training set, the model achieved an area under the curve (AUC) of 0.904 (95% CI: 0.852 to 0.955), with a sensitivity of 78.4% and a specificity of 91.1%. In the validation set, the model's AUC was 0.885 (95% CI: 0.812 to 0.958), with a sensitivity of 80.0% and a specificity of 85.9%. Compared with the individual predictive indicators, the model showed higher AUCs than the TyG index, TyG-BMI, TG/HDL-C, METS-IR, and BISAP score in the training set (Figure 4 and Table 3). Further comparison of AUC differences was performed using the DeLong test (Table 4). The AUC differences between the model and each individual indicator in the training set were statistically significant (all p < 0.05), with a difference of 0.034 compared with the BISAP score (95% CI: 0.006 to 0.062, p = 0.018) and 0.161 compared with METS-IR (95% CI: 0.085 to 0.237, p < 0.001), indicating that the model had significantly superior predictive performance compared with the individual indicators. In the validation set, the combined model yielded an AUC improvement of 0.020 over the BISAP score alone, and the difference was not statistically significant (Z = 0.982, p = 0.326), suggesting no significant incremental benefit over the BISAP score. Additionally, the differences in AUC between METS-IR and other IR-related indicators, including the TyG index, TyG-BMI, and TG/HDL-C, were not statistically significant (p > 0.05), suggesting comparable discriminative ability among the individual IR indicators, whereas the combined model achieved additional predictive gain when integrated with the BISAP score.

ROC curves for different indicators predicting SAP. (a) ROC curves for IR-related markers in the training set; (b) ROC curves for the predictive model versus individual indicators in the training set; (c) ROC curves for the predictive model versus individual indicators in the validation set. In the training set, METS-IR demonstrated a higher AUC than the TyG index, TyG-BMI, and TG/HDL-C (AUCs: 0.743 vs. 0.691, 0.663, and 0.682, respectively). The predictive model achieved an AUC of 0.904 (95% CI: 0.852 to 0.955) in the training set and 0.885 (95% CI: 0.812 to 0.958) in the validation set, outperforming all individual indicators in both sets.
Predictive performance of each indicator for SAP.
AUC: area under the curve; BISAP: bedside index for severity in acute pancreatitis; CI: confidence interval; METS-IR: metabolic score for insulin resistance; SAP: severe acute pancreatitis; TG/HDL-C: triglyceride/high-density lipoprotein cholesterol ratio; TyG: triglyceride-Glucose; TyG-BMI: triglyceride-glucose body mass index.
Comparison of AUC differences among different indicators.
AUC: area under the curve; BISAP: bedside index for severity in acute pancreatitis; CI: confidence interval; METS-IR: metabolic score for insulin resistance; TG/HDL-C: triglyceride/high-density lipoprotein cholesterol ratio; TyG: triglyceride-glucose; TyG-BMI: triglyceride-glucose body mass index.
Clinical utility and calibration of the predictive model
The predictive model was evaluated for both clinical utility and calibration. DCA demonstrated that within probability thresholds of 10%–80% in the training set and 10%–58% in the validation set, the clinical net benefit of the model for predicting SAP exceeded that of the “treat-all” and “treat-none” strategies, indicating favorable clinical applicability (Figure 5(a) and (b)). To further validate the superiority of the combined model, DCA was performed to compare the combined model with the BISAP score and METS-IR (Figure S4). Within the clinically relevant threshold ranges in both the training and validation sets, the combined model achieved greater net benefit than the two individual indicators, confirming its superior clinical decision-making value.

DCA and calibration of the predictive model for SAP. (a) DCA for the training set; (b) DCA for the validation set; (c) calibration curve for the training set; (d) calibration curve for the validation set. When the probability threshold ranged from 10% to 80% in the training set and from 10% to 58% in the validation set, the model's clinical net benefit for predicting SAP was superior to both the “intervention for all” and “intervention for none” strategies. The calibration curves were corrected using 1000 bootstrap resampling iterations, and the Hosmer–Lemeshow test indicated good model fit in both the training set (p = 0.143) and the validation set (p = 0.559).
Calibration curves were bias-corrected using 1000 bootstrap resamples. The model exhibited good calibration in both the training and validation sets, with the corrected curves closely fitting the ideal line (Figure 5(c) and (d)). The mean absolute error (MAE) was 0.013 in the training set and 0.021 in the validation set, indicating high consistency between the predicted and observed probabilities. The Hosmer–Lemeshow test indicated no significant deviation from ideal fit (training set: χ2 = 12.191, p = 0.143; validation set: χ2 = 6.791, p = 0.559). The Brier score further confirmed minimal deviation between the predicted probabilities and actual observations, demonstrating excellent calibration (Brier score = 0.064 in the training set and 0.070 in the validation set). The model showed perfect calibration in the training set (calibration slope: 1.00; calibration-in-the-large: 0.00) and good overall calibration in the validation set (calibration slope: 0.82; calibration-in-the-large: −0.54).
Discussion
To the best of our knowledge, this study is the first to investigate the association between METS-IR—a novel biomarker of IR—and the severity of HLAP. Moreover, we innovatively integrated METS-IR with the conventional BISAP score to develop a combined predictive model for SAP. Our findings demonstrate that both METS-IR and the BISAP score are independent risk factors for progression to SAP in patients with HLAP. The predictive model incorporating these two variables exhibited excellent discriminative performance, with AUCs of 0.904 in the training set and 0.885 in the validation set—significantly outperforming either predictor used alone. Furthermore, the model's reliability was rigorously validated through calibration plots, the Hosmer–Lemeshow goodness-of-fit test, Brier scores, and DCA, confirming its strong calibration accuracy and substantial clinical utility. These results provide a novel, efficient tool for early risk stratification of HLAP severity. Furthermore, they offer compelling clinical evidence for the role of IR in pathophysiological progression of HLAP.
METS-IR, proposed in 2018 as a novel tool for assessing IR, was found by Bello-Chavolla et al.11 in a prospective cohort study to maintain high consistency with the “gold standard” for IR assessment, the hyperinsulinemic-euglycemic clamp (HEC) technique (AUC: 0.84, 95% CI: 0.78 to 0.90). In recent years, the reliability of METS-IR has gained increasing recognition, demonstrating high predictive value for metabolic diseases such as hypertension, diabetes, and metabolic dysfunction–associated steatotic liver disease (MASLD). Zeng et al.12 observed that participants in the highest METS-IR quartile had a 2.89-fold increased prevalence of hypertension than those in the lowest quartile, highlighting its utility in hypertension assessment. A 16-year prospective cohort study of 5438 adults in Korea found that METS-IR could predict the incidence of MAFLD, with AUCs of 0.824 (95% CI: 0.814 to 0.834) and 0.831 (95% CI: 0.821 to 0.842), confirming its efficacy as an IR marker. 13 More recently, Zhang et al. 14 reported a positive correlation between METS-IR and MS components in middle-aged and older Chinese adults, noting that a METS-IR value exceeding 32.89 indicated an elevated risk of MS (AUC: 0.713). Despite its widespread clinical application, research on the association between METS-IR and HLAP remains limited. Our study identified METS-IR as an independent risk factor for SAP (OR = 1.053, 95% CI: 1.023 to 1.087, p < 0.001), demonstrating high predictive value for SAP in both the training (AUC: 0.743) and validation (AUC: 0.719) sets. The underlying mechanism is closely related to IR-mediated lipotoxic injury and the amplification of inflammatory cascades, which together drive the progression of HLAP toward severe disease. Under conditions of IR, enhanced lipolysis in visceral adipose tissue releases a large amount of free fatty acids (FFAs) into the circulation. Unbound FFAs exhibit strong cytotoxicity, directly damaging pancreatic acinar cells and vascular endothelial cells, leading to ischemia and acidosis, and also activating trypsinogen, which further exacerbates FFA toxicity and aggravates pancreatic injury. 21 Additionally, FFAs can induce intracellular calcium overload, mitochondrial dysfunction, and oxidative stress, thereby amplifying the inflammatory response and promoting cellular injury.22,23 Furthermore, IR is intrinsically a state of chronic low-grade inflammation, characterized by elevated levels of pro-inflammatory mediators such as nuclear factor kappa-light-chain-enhancer of activated B cells (NF-κB), tumor necrosis factor-alpha (TNF-α), and interleukin 6 (IL-6). These inflammatory mediators can exacerbate the inflammatory response in AP through both immune and nonimmune pathways, intensifying damage to the pancreas and other organs and contributing to SIRS, multiple organ dysfunction, and local complications.24,25
This study further compared the predictive performance of METS-IR with traditional IR-related indicators, including the TyG index, TyG-BMI, and TG/HDL-C, in the training set. Although the AUC of METS-IR was generally higher than that of these indicators, DeLong's test indicated that the differences were not statistically significant, suggesting that the discriminative abilities of various IR-related markers for predicting HLAP severity are generally comparable. This finding may be partly attributable to overlap in the components of these indices, as they all incorporate metabolic parameters such as blood glucose and triglycerides, resulting in similar predictive information. Nevertheless, METS-IR integrates multiple metabolic dimensions, including FPG, TG, BMI, and HDL-C, which theoretically allows for a more comprehensive reflection of insulin resistance and may offer advantages in clinical interpretability. In addition, the limited sample size may have affected the ability to detect significant differences, highlighting the need for future large-scale, multicenter studies to further validate these findings.
The BISAP score is a widely used clinical tool for assessing the severity of AP. It consists of five parameters: BUN level, mental status, SIRS, age, and the presence of pleural effusion, offering the advantages of simplicity, rapidity, and ease of bedside implementation. A retrospective study of 463 patients with AP demonstrated that BISAP exhibits strong predictive performance for SAP (AUC = 0.895; 95% CI: 0.862 to 0.929). 15 Notably, in a set specifically focused on HLAP, BISAP outperformed other established severity scores, including Ranson (AUC = 0.825), APACHE II (AUC = 0.807), and modified computed tomography severity index (MCTSI) (AUC = 0.831), with an AUC of 0.852 for SAP prediction. 16 Consistent with these findings, our study identified the BISAP score as an independent predictor of SAP (OR = 7.042; 95% CI: 3.850 to 14.145; p < 0.001), demonstrating robust discriminatory ability in both the training (AUC = 0.870) and validation (AUC = 0.865) sets, further supporting its stability as an early risk identification tool. Notably, when the BISAP score was combined with METS-IR, the predictive performance of the model was further improved. In the training set, the combined model achieved an AUC of 0.904 (95% CI: 0.852 to 0.955), with a sensitivity of 78.40% and a specificity of 91.10%. In the validation set, the AUC was 0.885 (95% CI: 0.812 to 0.958), with a sensitivity of 80.00% and a specificity of 85.90%. Further DeLong testing showed that the combined model was significantly superior to the BISAP score, METS-IR, and other IR-related indices in the training set (all p < 0.05). These findings suggest that integrating metabolism-related information with traditional clinical scoring can provide additional predictive gain and achieve higher identification accuracy. This combined model is expected to provide a more precise basis for early risk stratification and intervention decision making in patients with HLAP in clinical practice. Although the combined model was significantly superior to the BISAP score in the training set, the improvement in the validation set was modest (AUC difference = 0.020) and not statistically significant (p = 0.326). We speculate that this may be due to the limited sample size affecting the ability to detect the difference, and future large-sample studies are needed for further validation.
The combined model developed in this study demonstrated robust discriminative performance, good calibration, and strong clinical utility in both the training and validation sets, indicating its stable and reliable predictive capability. The model requires only simple scoring calculations and routine laboratory tests, enabling rapid risk stratification and making it an efficient and cost-effective assessment tool. It holds significant value for predicting the severity of HLAP, facilitating early health guidance and intervention for high-risk patients, and may contribute reducing the incidence of SAP. However, several limitations should be acknowledged. First, this model included only patients with HLAP. In real-world emergency settings, the etiology of AP often cannot be confirmed immediately upon admission; therefore, the applicability of this model to AP caused by non-HLAP etiologies requires further validation. Moreover, the combined model showed only a minimal improvement in the validation set, which may have been affected by the relatively small sample size. Thus, its incremental clinical value needs to be verified in larger samples. Second, internal validation was performed solely through a 7:3 random split, lacking an independent external validation cohort. Furthermore, all participants were recruited from two tertiary hospitals in Wuhan. This geographic concentration may introduce selection bias, as dietary patterns and the prevalence of metabolic disease in Wuhan may differ from those in other regions. Therefore, the extrapolation and generalizability of the model require further validation. Third, because of the retrospective design, certain targeted treatments and lifestyle factors were not adjusted for in the model, which might have slightly influenced the results. Fourth, only baseline METS-IR levels at admission were collected; the absence of dynamic longitudinal data limits our ability to establish a causal link between changes in METS-IR and the progression of HLAP to severe disease. Finally, the model was constructed using traditional LASSO and logistic regression methods without exploring optimization through machine learning algorithms, nor did it incorporate imaging features or genetic factors, which may limit further improvements in predictive accuracy. Given these limitations, this model should currently be regarded as only an exploratory tool and is not yet ready for widespread clinical use as a standardized prediction tool nationwide. Future research will involve multicenter, prospective studies to construct independent external validation sets comprising patients with HLAP from diverse regions and healthcare settings to optimize model universality. We also plan to incorporate dynamic METS-IR measurements and adjust for potential confounders to develop dynamic prediction models. Additionally, machine learning algorithms and the integration of imaging and genetic markers will be explored to further enhance model performance and biological plausibility.
Conclusion
This study demonstrates that both METS-IR and the BISAP score are effective predictors of disease severity in patients with HLAP. The predictive model integrating these two markers exhibits excellent discriminative performance and strong clinical utility. Owing to its simplicity, low cost, and reliance on routinely available clinical data, this combined model represents a promising risk stratification tool for early severity assessment in patients with HLAP under real-world conditions. Monitoring METS-IR and the BISAP score may facilitate timely assessment of disease severity, enabling earlier identification and intervention for high-risk patients with HLAP, thereby potentially improving clinical outcomes.
Supplemental Material
sj-docx-1-imr-10.1177_03000605261467013 - Supplemental material for Development and validation of a predictive model for the severity of hyperlipidemic acute pancreatitis based on the bedside index for severity in acute pancreatitis and metabolic score for insulin resistance: A retrospective cohort study
Supplemental material, sj-docx-1-imr-10.1177_03000605261467013 for Development and validation of a predictive model for the severity of hyperlipidemic acute pancreatitis based on the bedside index for severity in acute pancreatitis and metabolic score for insulin resistance: A retrospective cohort study by Longhui Kou, Huan Li, Wei Li, Xin Cheng, Yi Li, Bolun Zhang, Weitian Xu, Hui Long and Qingming Wu in Journal of International Medical Research
Supplemental Material
sj-docx-2-imr-10.1177_03000605261467013 - Supplemental material for Development and validation of a predictive model for the severity of hyperlipidemic acute pancreatitis based on the bedside index for severity in acute pancreatitis and metabolic score for insulin resistance: A retrospective cohort study
Supplemental material, sj-docx-2-imr-10.1177_03000605261467013 for Development and validation of a predictive model for the severity of hyperlipidemic acute pancreatitis based on the bedside index for severity in acute pancreatitis and metabolic score for insulin resistance: A retrospective cohort study by Longhui Kou, Huan Li, Wei Li, Xin Cheng, Yi Li, Bolun Zhang, Weitian Xu, Hui Long and Qingming Wu in Journal of International Medical Research
Footnotes
Acknowledgments
All authors sincerely thank all the physicians and nurses in the Department of Gastroenterology at Tianyou Hospital, Wuhan, and the Department of Gastroenterology at the General Hospital of Central Theater Command for their support and assistance in the collection of case data.
Ethics approval and consent to participate
This study was approved by the Ethics Committees of Tianyou Hospital, Wuhan University of Science and Technology (approval No. LL2025-07-22-01) and the General Hospital of Central Theater Command (approval No. (2024)090-01), and the requirement for informed consent was waived due to the retrospective nature of the study.
Consent for publication
Not applicable.
Author contributions
Longhui Kou and Huan Li : Conceptualization, Data curation, Formal analysis, and Writing–original draft; Wei Li, Xin Cheng, Yi Li, Bolun Zhang, and Weitian Xu: Data curation and Conceptualization. Hui Long and Qingming Wu: Conceptualization, Supervision, and Writing–review & editing. All authors read and approved the final manuscript.
Funding
The authors received no financial support for the research, authorship, and/or publication of this article.
Declaration of conflicting interests
The authors declare no potential conflicts of interest with respect to the research, authorship, and/or publication of this article.
Availability of data and material
The datasets used and analyzed during the current study are available from the corresponding author.
Supplemental material
Supplemental material for this article is available online.
References
Supplementary Material
Please find the following supplemental material available below.
For Open Access articles published under a Creative Commons License, all supplemental material carries the same license as the article it is associated with.
For non-Open Access articles published, all supplemental material carries a non-exclusive license, and permission requests for re-use of supplemental material or any part of supplemental material shall be sent directly to the copyright owner as specified in the copyright notice associated with the article.
