Abstract
Background
Large language models are increasingly applied in medical education and clinical decision support. However, comparative evaluations of their performance on pediatric content—spanning multiple subspecialties and question complexities—remain limited. This study sought to assess the performance of two advanced large language models, DeepSeek-R1 and GPT-4o, using a comprehensive set of Pediatrician Licensing Exam practice questions.
Methods
We administered 280 expert-validated pediatric questions covering 11 subspecialties and three levels of clinical complexity (A1: direct, A2: simple cases, A3: complex cases) to DeepSeek-R1 and GPT-4o. Each model was tested twice, with a 4-week interval. Performance metrics, including per-run accuracy, consistent accuracy (correct in both runs), aggregate accuracy (correct in either run), and inter-run consistency, were compared using chi-square tests and Bonferroni correction.
Results
DeepSeek-R1 outperformed GPT-4o in both runs (1st run: 90.7% vs. 81.8%, P=0.003; 2nd run: 86.8% vs. 79.3%, P = 0.02). It also showed higher aggregate accuracy (92.5% vs. 85.7%, P=0.01) and consistent accuracy (85.0% vs. 75.4%, P = 0.01). Performance declined with increasing question complexity for both models, with DeepSeek-R1 demonstrating numerically higher accuracy across all question types. For consistent accuracy, DeepSeek-R1 achieved 89.2% vs. 79.2% on A1 questions (P=0.05), 83.8% vs. 71.6% on A2 questions (P=0.11), and 80.2% vs. 73.3% on A3 questions (P=0.36). At the subspecialty level, DeepSeek-R1 showed more uniform performance, achieving >90% consistent accuracy in four subspecialties: Neonatology (94.7%), Infectious Diseases (93.8%), Nephrology (91.9%), and Cardiovascular diseases (90%). In comparison, GPT-4o reached >90% consistent accuracy in three subspecialties—Genetics (100%), Infectious Diseases (100%), and Cardiovascular diseases (90%)—with greater variability across other domains.
Conclusion
DeepSeek-R1 demonstrated higher accuracy than GPT-4o on pediatric questions. The observed variability across different runs warrants attention, particularly in contexts that require highly consistent and reproducible performance.
Introduction
Large language models (LLMs) have recently gained significant attention across a broad range of domains, including medicine.1–3 Among the most widely adopted LLMs in the medical field are those built on the generative pre-trained transformer (GPT) architecture, particularly the GPT-4 family. These models have demonstrated strong performance on a variety of medical benchmarks,4–6 including United States Medical Licensing Examination (USMLE)-style questions and clinical decision support tasks. 7–15 However, GPT-based models are fundamentally general-purpose systems. While they excel in knowledge breadth, they may fall short in domain-specific reasoning and context-sensitive interpretation that are often critical in clinical medicine. Concerns regarding their accuracy, interpretability, and reliability persist, particularly in high-stakes environments such as diagnostics and medical education.16–19
Newer models such as DeepSeek-R1 have been developed with enhanced reasoning capabilities and greater emphasis on transparency and structured problem-solving.20–23 Preliminary evidence suggests that DeepSeek-R1 may outperform ChatGPT models on tasks requiring precise factual recall and clinical knowledge retrieval, particularly in USMLE Step 1 and Step 2 contexts. Additionally, in more complex, case-based reasoning scenarios typical of Step 3, DeepSeek has shown a tendency to remain closer to the correct answer trajectory. 24 These findings point to a potential advantage in both foundational knowledge and clinical application.
Nonetheless, direct, head-to-head evaluations of DeepSeek and GPT-4-based models remain limited, especially in subspecialized fields such as pediatrics. Given the complexity, nuance, and age-specific variability inherent in pediatric medicine, a rigorous comparison in this domain is warranted to assess whether newer models like DeepSeek offer meaningful improvements in accuracy, consistency, and reasoning fidelity.
To address this gap, the present study conducts a comparative evaluation of DeepSeek-R1 and GPT-4o using a set of 280 pediatric licensing exam practice questions. These questions span 11 pediatric subspecialties and are categorized into three levels of clinical complexity. We examine multiple performance metrics—including per-run accuracy, aggregate accuracy, consistent accuracy, and response repeatability—to benchmark the capabilities of these state-of-the-art models in high-stakes pediatric knowledge tasks.
Methods
Pediatric questions tested
To evaluate the performance of DeepSeek and ChatGPT, we tested both models using 280 single-answer multiple-choice questions for preparing Pediatrician Licensing Exam in our hospital. These questions spanned 11 subspecialties: Hematology (n=18), Cardiovascular diseases (n=20), Digestion (n=21), Genetics (n=7), Healthcare (n=45), Rheumatology/Immune diseases (n=11), Infectious Diseases (n=16), Nephrology (n=37), Neurology (n=24), Neonatology (n=38), and Respiratory diseases (n=43). Additionally, questions were categorized into three clinical types: A1: Directly asked questions (n=120), A2: Simple clinical cases (n=74), and A3: Relatively complex clinical cases (n=86).
Study design
To accurately assess model performance and minimize potential bias, each question was processed twice by both models, with a 4-week interval between runs. Notably, all questions were administered in isolated sessions to minimize contextual influence from prior prompts.4,5 Specifically, each question was entered independently in a new chat session (i.e., without preceding conversation history), ensuring that model responses were not affected by earlier queries. Both models were prompted exclusively with the predefined questions, without additional guidance, follow-up prompts, or contextual priming. This approach was chosen to standardize input conditions and reduce variability related to prompt engineering or conversational memory. However, it should be highlighted that the open use of artificial intelligence, particularly in unsecured or uncontrolled environments, constitutes a data breach, as it can expose sensitive information without proper safeguards. To prevent unauthorized access or leakage, the ideal artificial intelligence system for handling protected data should be closed-loop.25,26
Model responses were recorded and compared against expert-validated answers. Performance was assessed using four metrics: • Per-Run Accuracy: The proportion of correct responses in each individual run, reflecting immediate performance. • Consistent Accuracy: The proportion of questions answered correctly in both runs, measuring correct response stability. • Aggregate Accuracy: The proportion of questions answered correctly in either run, representing overall capability. • Consistency: The proportion of questions with the same answers in both runs, measuring response repeatability.
Statistical analysis
Accuracy rates are reported as frequencies and percentages. Categorical comparisons used chi-square tests, with statistical significance defined as P ≤ 0.05. Additional pairwise comparisons across question types and subspecialties with Bonferroni correction for multiple testing were conducted. Analyses were performed using JMP Pro, version 16 (SAS Institute Inc., Cary, NC).
Results
Performance of large language models on pediatric questions
Accuracy of large language models in answering pediatric questions.
aP=0.52 for comparison of the 1st run versus the 2nd run.
bP=0.18 for comparison of the 1st run versus the 2nd run.
cP=0.52 for comparison of consistent versus aggregate accuracy.
d

Similarly, DeepSeek-R1 showed superior in both aggregate accuracy (92.5% vs. 85.7%, P = 0.01) and consistent accuracy (85% vs 75.4%, P = 0.01) to GPT-4o (Figure 1). Notably, DeepSeek-R1’s consistent accuracy was significantly lower than its aggregate accuracy (85.0% vs. 92.5%; P = 0.007). Although GPT-4o showed the same trend (75.4% vs. 85.7%), this difference was not statistically significant (P = 0.52) (Table 1). In addition, both models showed comparable inter-run consistency (91.4% vs. 87.9%, P = 0.21) (Figure 1).
Performance of large language models across question types
a
b
c
d
e

Performance of large language models across subspecialties
aP values between the 1st and 2nd run.
bP values for consistent accuracy between DeepSeek-R1 and GPT-4o.
cP values for aggregate accuracy between DeepSeek-R1 and GPT-4o.
dP values across subspecialities.

In the case of GPT-4o, both aggregate and consistent accuracy rates differed significantly across subspecialties (P = 0.02 and P = 0.03, respectively). GPT-4o achieved >90% consistent accuracy in four subspecialties: genetics (100%), infectious diseases (100%), cardiovascular diseases (90%), and nephrology (73%). However, consistent accuracy was below 80% in all other subspecialties. Notably, performance was particularly low (<70%) in healthcare (68.9%), digestion (66.7%), neurology (66.7%), and rheumatology/immune diseases (54.6%) (Figure 3). Pairwise comparisons across subspecialties showed that consistent accuracy was higher in infectious diseases (100%) than in neonatology (76.3%), nephrology (73%), hematology (72.2%), healthcare (68.9%), and neurology (66.7%) (P = 0.045, P = 0.02, P = 0.04, P = 0.01, P = 0.02, respectively). Consistent accuracy in respiratory disease (79.1%) was also higher than in rheumatology/immune disease (54.6%) (P = 0.006). However, none of these differences remained statistically significant after Bonferroni correction for multiple comparisons (adjusted significance threshold P < 0.001) (Table S2).
When comparing the two models directly (Figure 3), DeepSeek-R1 achieved significantly higher consistent accuracy than GPT-4o in neonatology (94.7% vs. 76.3%, P = 0.046). DeepSeek-R1 also outperformed GPT-4o in nephrology, with a higher consistent accuracy (91.9% vs. 73%, P = 0.06) and a significantly higher aggregate accuracy (100% vs. 78.4%, P = 0.005), although the difference in consistent accuracy was near-significant.
Discussion
In this study, the primary aim of comparing two different LLMs was to provide a head-to-head evaluation of their performance within a clinically relevant pediatric questions, where accuracy and reproducibility are critical. Given the rapid evolution of LLMs and their increasing integration into medical education and clinical decision support, it is essential to understand not only their overall performance but also how they differ across domains, levels of complexity, and repeated use. The ultimate goal of this comparison is to inform clinicians, educators, and researchers about the relative strengths and limitations of contemporary models, particularly with respect to reliability (e.g., consistent accuracy) and robustness across subspecialties. Such insights are necessary for guiding the appropriate selection, implementation, and future development of LLMs in high-stakes medical settings. Overall, our findings suggest that DeepSeek-R1 demonstrated higher performance than GPT-4o on this set of pediatric questions, particularly in consistent accuracy, which may reflect more stable outputs across repeated queries. Differences between models were observed across question types and subspecialties, with DeepSeek-R1 showing higher performance on directly asked (A1) questions and in subspecialties such as neonatology and nephrology.
Recent studies have reported that DeepSeek demonstrates performance that is comparable to, and in some cases higher than ChatGPT-based models across various medical domains. For example, GPT-4o and DeepSeek-R1 achieved similar accuracy on the Polish specialty examination in infectious diseases (73.9% vs 71.4%). 27 In a study using 100 clinicopathologic cases from the New England Journal of Medicine, DeepSeek-R1 and GPT-4 showed no statistically significant difference in accuracy (35% vs. 39%; P = 0.63), although overall performance remained low. 28 In contrast, DeepSeek-R1 achieved higher accuracy than GPT-4o on the Chinese National Medical Licensing Examination (92.0% vs. 87.2%; P < 0.05). 29 Similarly, in American College of Gastroenterology self-assessments, a search-augmented DeepSeek model scored 81.5%, and its base R1 model scored 77.1%, both outperforming GPT-4 (62.4%; P < 0.001). 30 DeepSeek has also demonstrated higher accuracy than ChatGPT in prostate cancer radiotherapy questions across both English and Chinese contexts. 31
However, while prior studies have primarily focused on accuracy, fewer have evaluated repeatability, which is an important consideration in clinical and educational contexts. In our study, we examined consistent accuracy, defined as the proportion of questions answered correctly in both independent runs. Although both models exhibited high aggregate accuracy, DeepSeek-R1 showed a smaller gap between aggregate and consistent accuracy (92.5% vs. 85.0%, Δ = 7.5%) compared with GPT-4o (85.7% vs. 75.4%, Δ = 10.3%), suggesting more stable performance across repeated quires. Similarly, in a study using item-analyzed multiple-choice questions on blood physiology, DeepSeek demonstrated slightly higher reliability (93%) than ChatGPT (90%). 32 These findings highlight that aggregate accuracy alone may overestimate performance if correct responses are not consistently reproduced. For application such as examination preparation or other settings requiring dependable outputs, reproducibility remains an important consideration, and our results suggest that DeepSeek-R1 may offer advantages in this aspect within the context of this dataset.
Although pairwise comparisons across question types did not show statistically significant differences, both models demonstrated decreased performance with increasing question complexity, with A3-type items—representing more complex clinical scenarios—yielding the lowest scores. DeepSeek-R1 showed numerically higher performance across all question types, with a statistically significant difference observed only in consistent accuracy for A1 questions. These findings are consistent with results from the Chinese National Medical Licensing Examination, where DeepSeek-R1 achieved higher accuracy than GPT-4o in low-difficulty questions (95.9% vs. 92.0%; P < 0.05), while no significant differences were observed in high-difficulty items. 29 In our study, the difference in A3 aggregate accuracy approached significance, suggesting a possible trend that warrants further investigation.
Subspecialty-level analysis revealed differences in performance patterns. DeepSeek-R1 showed more uniform performance across subspecialties, with no significant intra-group variability, whereas GPT-4o demonstrated greater variation, with statistically significant differences across domains. For instance, GPT-4o achieved high accuracy in genetics and infectious diseases but showed lower performance in subspecialties such as rheumatology/immune, neurology, digestion, and healthcare systems, where DeepSeek-R1 maintained more consistent accuracy. These patterns may suggest differences in how the models process and apply information; however, such interpretations remain speculative and were not directly evaluated in this study. Notably, differences in consistent accuracy identified in pairwise comparisons across several subspecialties did not remain statistically significant after Bonferroni correction, indicating that these variations should be interpreted with caution and warrant further investigation. Although DeepSeek-R1 demonstrated advantages in certain metrics in our dataset, these findings should be interpreted within the context of this study and not generalized broadly. Recent studies comparing large language models across medical domains report variable and context-dependent performance. In pediatric ophthalmology, GPT-4 and DeepSeek-R1 demonstrated comparable results, suggesting that differences between models may be domain-specific. 33 In gastroenterology board-style assessments, ChatGPT-o1 achieved the highest accuracy (84%), followed by GPT-4o (82%) and DeepSeek-R1 (79%), with all models showing high consistency and outperforming both the passing threshold and practicing clinicians. 34 In contrast, a guideline-based cardiovascular study found higher accuracy for ChatGPT models compared to DeepSeek, with moderate inter-rater reliability. 35 Similarly, on the Chinese National Medical Licensing Examination, DeepSeek-R1 (91%) performed among the top models alongside Doubao-1.5-pro (92%), while GPT-4o showed lower accuracy (83%). 36 In hypertension case vignettes, GPT-4 achieved the highest accuracy and safety among LLMs but remained below expert-level performance. 37 The existing evidence suggests that LLM performance varies by clinical domain and task type, with no single model consistently outperforming others across all settings.
It is important to clarify that the aim of this study is not to support the use of LLMs for direct clinical diagnosis or independent medical decision-making. Rather, the objective is to evaluate and compare the performance of two contemporary LLMs using standardized, exam-style pediatric questions within an educational and benchmarking framework. The dataset consists exclusively of de-identified multiple-choice questions and does not include real patient data, clinical records, or personal health information. The application of LLMs in real-world clinical settings involves important ethical, legal, and data privacy considerations. 38 Accordingly, these models are not intended to replace clinician judgment or function as standalone diagnostic tools. Their potential use is more appropriately considered in supportive or educational contexts, where considerations of safety, accountability, and data protection can be more effectively addressed.
This work provides several practical insights. First, by evaluating LLM performance on pediatric licensing-style questions across multiple subspecialties and levels of complexity, the study offers a structured benchmark of how these models perform in a clinically relevant knowledge domain. This allows readers, particularly clinicians and educators, to better understand where current models perform reliably and where limitations persist. Second, the inclusion of repeat testing and metrics such as consistent accuracy provides information on response reproducibility, which is not captured by single-run accuracy alone. This is particularly relevant for readers considering the use of LLMs in educational or supportive contexts, where stability of outputs is important. Third, the subspecialty- and complexity-level analyses highlight that model performance is not uniform, offering readers a more nuanced understanding of strengths and weaknesses across different types of clinical reasoning tasks. Overall, rather than promoting direct clinical application, the study aims to inform readers about the current capabilities and limitations of LLMs in pediatric knowledge assessment, thereby supporting more informed, cautious, and context-appropriate use of these tools.
However, several limitations should be considered when interpreting the results. First, the question set was derived from a single institution’s pediatric licensing preparation materials. Although these questions were expert-validated and covered a broad range of subspecialties and complexity levels, they may not fully represent the diversity of pediatric clinical scenarios across different regions or healthcare systems. Second, while a 4-week interval between model runs was used to assess consistency, both DeepSeek-R1 and GPT-4o may undergo unobservable backend updates over time, introducing a potential confounder in evaluating true intra-model repeatability. Third, although questions were categorized into three predefined complexity levels, these classifications may not fully capture the spectrum of clinical difficulty encountered in clinical practice. In addition, the relatively small number of questions in certain subspecialties—particularly genetics and rheumatology/immune—limits statistical power and may increase variability in performance estimates within those categories. Lastly, we did not assess hallucination rates, output validity, or explainability. Although accuracy and consistency were the primary focus, these metrics alone do not fully characterize model performance or suitability for real-world applications. Despite these limitations, this study provides a direct comparison of two contemporary LLMs within a pediatric context. Given the increasing interest in AI applications in medical education and decision support, these findings may offer useful insights for clinicians, educators, and researchers.
Conclusions
In this comparative evaluation of two advanced LLMs using pediatric licensing-style questions, DeepSeek-R1 demonstrated higher overall accuracy and greater consistency across repeated runs. Although GPT-4o performed competitively in certain domains, greater variability was observed across some subspecialties. These findings suggest that performance may differ depending on context, and that consistency remains an important consideration in evaluating model reliability. Within the scope of this dataset, DeepSeek-R1 showed more stable performance across question types and subspecialties; however, these results should be interpreted cautiously and not generalized beyond this setting. Future studies should extend beyond accuracy metrics to assess longitudinal performance, incorporate more diverse and clinically representative datasets, and further evaluate factors such as output stability, validity, and transparency to better inform the potential role of LLMs in medical education and clinical support contexts.
Supplemental material
Supplemental material - Accuracy and stability of DeepSeek-R1 and GPT-4o on pediatric questions
Supplemental material for Accuracy and stability of DeepSeek-R1 and GPT-4o on pediatric questions by Qibo Hu, Lin Lei, Xuechun Wang, Fangshu Liu, Guanghua Che in DIGITAL HEALTH
Footnotes
Acknowledgments
Large language models (DeepSeek-R1 and GPT-4o) were used only as study subjects to generate responses to the pediatric questions evaluated in this research. No AI tools were used in the preparation of the manuscript itself, including writing, data analysis, figure generation, or reference management.
Ethical consideration
This study does not require Ethics Committee or Institutional Review Board approval because it does not involve human or animal subjects, nor does it include patient information or identifiable personal data. Consequently, participant consent was waived for the same reasons.
Author Contributions
Funding
The authors received no financial support for the research, authorship, and/or publication of this article.
Declaration of conflicting interests
The authors declared no potential conflicts of interest with respect to the research, authorship, and/or publication of this article.
Data Availability Statement
The original contributions presented in the study are included in the article, further inquiries can be directed to the corresponding author.
Guarantor
GC.
Supplemental material
Supplemental material for this article is available online.
References
Supplementary Material
Please find the following supplemental material available below.
For Open Access articles published under a Creative Commons License, all supplemental material carries the same license as the article it is associated with.
For non-Open Access articles published, all supplemental material carries a non-exclusive license, and permission requests for re-use of supplemental material or any part of supplemental material shall be sent directly to the copyright owner as specified in the copyright notice associated with the article.
