Abstract
Background
Depression is a leading global cause of disability. The rapid emergence of generative artificial intelligence (GenAI), particularly large language models (LLMs) like ChatGPT, offers new opportunities for digital psychiatry. However, the Web of Science Core Collection (WoSCC) -indexed research landscape of GenAI in depression has not yet been systematically mapped.
Objective
This study aimed to systematically evaluate the WoSCC-indexed research landscape, hotspots, and emerging trends of GenAI in depression through bibliometric analysis.
Methods
A bibliometric analysis was conducted on publications from the WoSCC (January 2023–July 2025). Additionally, PubMed was searched to identify relevant clinical and translational studies for contextual interpretation. Analyses utilized CiteSpace, VOSviewer, and Bibliometrix.
Results
We identified 115 publications, with publication output increasing markedly from 2023 to mid-2025. The United States and China led in volume, with Harvard University as a key contributor. International collaboration involved 39 countries but remained regionally concentrated. Co-citation analysis revealed 10 clusters, including depression management, deep learning, and NLP. Keyword bursts highlighted trends in “large language models,” “ChatGPT,” and “digital health.” Top-cited works focused on conversational agents and LLM evaluation.
Conclusion
This study provides one of the first bibliometric analyses of WoSCC-indexed research on GenAI in depression, highlighting increasing scholarly attention to conversational agents, large language models, and digital mental health applications. Future research should prioritize clinical validation, safety, and interdisciplinary collaboration to strengthen the evidence base for responsible implementation.
Keywords
Introduction
Depression is one of the most prevalent mental health disorders worldwide and a leading cause of disability and years of productive life lost across all age groups.1,2 It contributes substantially to years lived with disability and is associated with high comorbidity, premature mortality, and reduced quality of life. Global epidemiological studies have shown that more than 280 million people are affected by depression, and its burden has continued to rise over the past three decades. 3 The COVID-19 pandemic further exacerbated this trend, leading to a sharp increase in the prevalence of depressive and anxiety disorders in many regions. 4 This aligns with global research priorities emphasizing urgent action in mental health science during the pandemic. 5 These findings underscore the urgent need for innovative approaches to improve early detection, monitoring, and treatment of depression.
Prior to the emergence of large language models (LLMs), artificial intelligence in mental health primarily relied on “discriminative” paradigms. These earlier approaches utilized conventional machine learning (ML) and natural language processing (NLP) to analyze static data—such as social media posts or voice biomarkers—for the passive detection of depressive symptoms.6,7 While effective for screening, these traditional models faced significant limitations in scalability, user engagement, and the ability to foster a therapeutic alliance.8–11
To address these limitations, recent exploratory studies have begun to test the capabilities of LLMs in psychiatric contexts. For instance, Levkovich and Elyoseph (2023) demonstrated that ChatGPT could generate treatment recommendations for depression that were largely consistent with clinical guidelines, offering a scalable alternative for initial psychoeducation. 12 Similarly, Liu et al. (2025) evaluated the psychometric properties of ChatGPT-4 in administering depression screening tools (such as the PHQ-9), finding high agreement with valid instruments. 13 These early applications suggest that unlike rigid rule-based chatbots, LLMs can dynamically adapt to patient inputs, potentially bridging the gap between screening and intervention. Crucially, it is necessary to distinguish the unique contribution of generative AI (GenAI) from earlier AI paradigms within this domain. As highlighted in a recent bibliometric analysis by Ren et al. (2025), the prevailing AI landscape in depression has focused heavily on detection and diagnosis using deep learning and multimodal data. 14 Traditional “discriminative” AI approaches merely focus on “classifying” whether a patient is depressed based on pattern recognition.6,7 In contrast, GenAI, particularly LLMs, utilizes transformer architectures to “generate” human-like, empathetic responses, capable of handling complex medical queries.15,16 In the broader literature, this technological development has been discussed as an important extension of earlier AI applications in psychiatry. 17 However, such characterizations are interpretive and should not be regarded as direct bibliometric findings.
However, this rapid expansion of interest has outpaced a systematic analysis of the field itself, and the WoSCC-indexed research landscape specifically focusing on GenAI systems (e.g., large language models and conversational agents) in depression has not yet been systematically mapped. As a result, the intellectual structure, key research frontiers, dominant clinical applications, and emerging geographical hubs within the indexed literature remain unclear. This temporal concentration is notable, as the public release of ChatGPT in late 2022 was followed by the first wave of WoSCC-indexed publications specifically addressing GenAI applications in depression in 2023. Therefore, this study aimed to conduct a bibliometric analysis of WoSCC-indexed literature on GenAI in depression from 2023 to 2025, summarizing publication trends, geographical and institutional contributions, collaborative networks, co-citation structures, and thematic hotspots. By mapping this indexed body of literature, we sought to provide a structured overview of current research activity, identify knowledge gaps, and inform future evidence-based development and governance of GenAI applications in depression.
Materials and methods
Data source and literature search strategy
The bibliometric data used in this study were retrieved from the Web of Science Core Collection (WoSCC, https://www.webofscience.com/), which is a comprehensive bibliographic database with real-time updates covering journals, books, and proceedings. 18 The primary bibliometric dataset was derived from the WoSCC, given its structured citation indexing and compatibility with network analysis tools. However, bibliometric indicators inherently favor publications that have accumulated citation impact, which may underrepresent very recent clinical trials or qualitative feasibility studies. To complement this limitation, a supplementary search was conducted in PubMed to identify newly published clinical trials, feasibility studies, and early translational research related to GenAI in depression (detailed search strategies are provided in Supplementary Table 1). These PubMed-identified studies were used solely for contextual interpretation in the Discussion section and were not merged into the WoSCC dataset or included in quantitative bibliometric network analyses.
In this study, “Generative Artificial Intelligence (GenAI)” refers to artificial intelligence systems designed to generate human-like text or multimodal outputs based on large-scale pretrained models, primarily LLMs built on transformer architectures. In contrast to traditional discriminative machine learning models that focus on classification or prediction tasks, GenAI systems are capable of producing interactive, context-aware responses. For clarity and consistency, the term “GenAI” is used throughout this manuscript to denote generative artificial intelligence systems, while “LLMs” specifically refers to large language model architectures that constitute a major subset of GenAI technologies. A static bibliometric analysis was conducted on publications related to GenAI and depression. All data were retrieved on 23 July 2025 to ensure consistency. The systematic search strategy was developed based on previous reviews of GenAI 19 and expert consultation, and was agreed upon by all authors. The following search formula was applied:
TS = ((“generative artificial intelligence” OR “AIGC” OR “generative AI” OR “AI Generated Content” OR “ChatGPT” OR “GPT-3.5″ OR “GPT-4″ OR “Generative pretrained transformer” OR “large language model” OR “OpenAI” OR “Google Bard” OR “Microsoft Copilot” OR “Google Gemini” OR “LLaMA” OR “Bing Chat” OR “Anthropic Claude” OR “Deepseek”) AND (“depression*” OR “depressive*” OR “depressed*” OR “melancholia*”)). The search was limited to “Article” and “Review” types published in English between January 2023 and July 2025. Conference abstracts, editorial materials, letters, meeting reports, and early access items without complete bibliographic metadata were excluded to ensure consistency in citation analysis. Preprints and gray literature were not included because WoSCC primarily indexes peer-reviewed publications, and citation-based bibliometric methods rely on standardized indexing records. Non-English publications were excluded to maintain consistency in keyword normalization and network analysis, as bibliometric software may inconsistently process multilingual metadata. To avoid inconsistencies due to database updates, literature retrieval, extraction, and downloading were all performed by the same investigator on the same day. The final WoSCC bibliometric dataset was determined using the pre-specified database filters, including publication period, language, document type, and availability of complete bibliographic metadata. After these automated filters were applied, two reviewers independently checked the retained records to confirm that they met the predefined eligibility criteria. This verification step did not lead to any additional exclusions. Although this study is a bibliometric analysis rather than a systematic review, the literature identification process was reported with reference to PRISMA 2020 principles to enhance transparency, and the accompanying checklist was adapted for reporting transparency rather than full systematic-review compliance (Supplementary Checklist 1). Any uncertainties during record verification were resolved through discussion and, when necessary, consultation with a third author. From the exported records, we documented authors, affiliations, titles, keywords, and cited references. All WoSCC records with cited references were saved as tag-delimited plain text files for subsequent analysis. The inclusion of platform- and model-specific terms (e.g., ChatGPT, GPT-4, Claude) was intended to enhance sensitivity during early-stage field development, where terminology often remains closely tied to specific publicly released models. However, this strategy may also have preferentially captured studies that explicitly referenced well-known commercial platforms, while underrepresenting studies using open-source, fine-tuned, or domain-specific models that did not mention such brand names directly. Broader conceptual terms such as “generative artificial intelligence” and “large language model” were therefore also included to partially mitigate this bias, although complete coverage cannot be assumed.
Data analysis and graph acquisition
Bibliometric analyses were performed using CiteSpace and VOSviewer (version 1.6.20). Additional network visualization was conducted with Bibliometrix (version 4.1.4) in RStudio. To complement the bibliometric findings, recent clinical and translational studies identified through PubMed were summarized descriptively for contextual interpretation only and were not included in the quantitative bibliometric analyses. International collaboration was visualized using ArcGIS (ArcMap 10.8). Data tables and basic charts were generated with Microsoft Excel and GraphPad Prism. The bibliographic network was constructed on the basis of coauthorship, co-occurrence, and citation analyses. Full counting was applied, without fractional adjustment for multi-authored papers. VOSviewer was used to generate keyword co-occurrence networks with clustering algorithms. CiteSpace was further employed to identify highly cited references and keywords with strong citation bursts over time.
In CiteSpace, the time slicing was set from 2023 to 2025 (1-year per slice), with node types including reference and keyword. The selection criteria were based on the top 50 most cited or occurring items per slice, and pruning was performed using the Pathfinder algorithm to simplify network structure. In VOSviewer, full counting was applied for keyword co-occurrence analysis. A minimum occurrence threshold of five was selected a priori as a pragmatic balance between retaining recurrent and representative keywords and maintaining network interpretability in a relatively small dataset. This threshold yielded 20 qualifying keywords, which allowed visualization of the main thematic structure while reducing noise from infrequently occurring terms. Network visualization parameters followed default normalization settings. These tools and configurations were selected to balance sensitivity and interpretability in a relatively small and rapidly evolving dataset: CiteSpace was used for temporal and co-citation analyses, including burst detection and dual-map overlay visualization, whereas VOSviewer was used for co-authorship and keyword co-occurrence mapping because of its strengths in network clustering and visual clarity.
Statistical analysis
This study primarily employed descriptive bibliometric statistics rather than inferential statistical testing. Publication counts, frequencies, percentages, annual growth trends, citation data, co-authorship networks, co-citation networks, keyword co-occurrence, clustering, and burst detection were analyzed to characterize the knowledge structure of WoSCC-indexed research on GenAI in depression. Analyses were conducted using CiteSpace, VOSviewer, Bibliometrix, Microsoft Excel, and GraphPad Prism. No formal hypothesis-testing procedures were performed.
Results
Overview of publication status in GenAI and depression
Bibliometric and citation analyses provide quantitative approaches for evaluating research activity and knowledge structures within a scientific field. Web-based data help researchers understand their interactions and reasonably predict future research behavior. As shown in Figure 1, we performed a comprehensive bibliometric analysis of GenAI and depression on the basis of a schematic representation of the workflow. Identification and selection of WoSCC-indexed studies on GenAI and depression.
The main information of the data.
*Calculated using available records through July 23, 2025; the 2025 data represent a partial year and should be interpreted cautiously.
Annual publication trends
To explore publication trends in this field, the annual and cumulative numbers of publications related to GenAI and depression were analyzed for the period from January 2023 to July 23, 2025. As shown in Figure 2(a), publication output increased markedly during the study period, rising from 11 papers in 2023 to 48 papers in 2024 and 56 papers by July 23, 2025, for a total of 115 publications. Because the 2025 count represents only a partial year rather than a complete annual output, direct comparison with the full-year counts for 2023 and 2024 should be interpreted cautiously. Nevertheless, the fact that the number of publications recorded by July 2025 had already exceeded the full-year output of 2024 suggests rapidly increasing scholarly attention following the public release of large language models in late 2022. Publication trends in WoSCC-indexed literature on GenAI and depression (2023–mid-2025). (A) Annual and cumulative publications. (B) The most productive countries/regions. (C) Top 10 most productive institutions and journals.
Regarding the national distribution (Figure 2(b)), the United States ranked first with 34 publications, followed by China (23 publications) and England (13 publications). Together, these three countries accounted for 70 publications, which exceeded the combined output of all other countries (45 publications, including South Korea, Israel, and Germany). This finding highlights the dominant role of the United States, China, and England in research on GenAI in depression.
At the institutional level (Figure 2(c)), Harvard University led the field with 17 publications, followed by Beijing Normal University and the Icahn School of Medicine at Mount Sinai (5 publications each). Institutional names were standardized based on the primary affiliation field in WoSCC, and the same aggregation rule was applied consistently across the dataset: subordinate units, medical schools, and affiliated hospitals were grouped under their parent university or institution when they clearly belonged to the same organizational system. Thus, this normalization approach was applied not only to Harvard University and its affiliated entities, but also to comparable institutional structures. In terms of journal distribution, Frontiers in Psychiatry and JMIR Mental Health published the largest number of papers (n = 9 each), followed by the Journal of Medical Internet Research (n = 8), IEEE Access (n = 5), and the Journal of Affective Disorders (n = 5).
Analysis of partnerships and cooperation among nations
Using ArcGIS, a world map was generated to visualize the distribution of publications across 39 countries that contributed to research on GenAI in depression (Figure 3(a)). The United States was the most productive country with 34 publications, followed by China (23 publications), the United Kingdom (13 publications), South Korea (9 publications), and Israel (8 publications). Publications from the remaining countries were fewer than eight. Geographical distribution and international collaboration. (a) World map of publication distribution generated by ArcGIS. (b) Country collaboration network by VOSviewer; node size reflects publication volume, line thickness reflects collaboration strength. (c) Collaborative network map with clusters indicated by colors. (d) Temporal overlay of collaborations, with colors indicating the average publication year of each country’s records. The scale begins at approximately 2024.0 because substantive cross-country collaboration activity was mainly observed from 2024 onward.
International cooperation networks were further analyzed using VOSviewer and Scimago Graphica (Figure 3(a) and 3(c)). In Figure 3(b), each node represents a country, with node size proportional to publication volume, link thickness reflecting the strength of collaboration, and colors indicating clusters of cooperation. A total of 39 countries engaged in international collaboration, with the United States occupying the most central and dominant position.
The cooperation frequency analysis (Figure 3(c); Supplementary Table 2) revealed that the strongest collaborative ties were between the United States and the United Kingdom (3 collaborations), as well as between the United Kingdom and Israel (3 collaborations). Additional frequent collaborations included the United States with China, India, Lebanon, and Turkey (2 collaborations each), the United Kingdom with India (2 collaborations), and India with Norway (2 collaborations). These findings highlight the central role of the United States in global collaborative research in this field. Furthermore, the temporal overlay visualization (Figure 3(d)) illustrates the evolution of international collaboration, where colors indicate the average publication year of each country’s records included in the collaboration network. Although the overall study period began in 2023, the color scale in Figure 3(d) starts at approximately 2024.0 because substantive cross-country collaboration activity in the retained dataset was mainly concentrated from 2024 onward. Earlier nodes, such as the United States and the United Kingdom, formed the initial collaboration network, whereas more recent nodes, such as China and South Korea, highlight the rapid expansion and increasing participation of Asian countries in 2024 and 2025.
Supplementary visualizations further supported these observations. In Supplementary Figure 1, the United States appeared as the largest and most connected node, repeatedly confirming its central position in the international research network. In contrast, during the past six months, research activity has increasingly shifted toward China and South Korea, reflecting the growing recognition of the clinical value of GenAI. This trend further emphasizes the leading academic influence of the United States while highlighting the emerging contributions of Asian countries. Additionally, the contributions of individual researchers are visualized in Supplementary Figure 2, where Levkovich I and Wang X emerged as the most relevant authors in this emerging field, each contributing three publications.
The evolution of research disciplines
To visualize the citation relationships among journals and the thematic distribution of research, a dual-map overlay of journals was generated using CiteSpace (Figure 4). In this map, the left side represents the citing journals and the right side represents the cited journals. Each point corresponds to a journal, and the curved citation lines trace the knowledge flow between different disciplines. Journal dual-map overlay and reference co-citation clusters. (a) Dual-map overlay showing citation trajectories from citing journals on the left to cited journals on the right. Category labels are reported as generated by CiteSpace. (b) Co-citation network of references generated by CiteSpace, with 10 major clusters labeled by LLR algorithm.
As shown in Figure 4(a), the thickest citation trajectory extended from journals categorized by CiteSpace as Psychology/Education/Health to journals categorized as Psychology/Education/Social. This pattern suggests that GenAI-related depression research is primarily situated within psychology, education, health, and social science domains. In addition, a secondary path from Psychology/Education/Health journals to Health/Nursing/Medicine journals suggests cross-disciplinary exchange between psychological, health-related, nursing, and medical research. Taken together, these trajectories support an interpretation of the field as interdisciplinary, while also underscoring that dual-map patterns are partly shaped by journal classification conventions.
Analysis of cocited references
The 115 included publications cited a total of 6,083 references, which were used to construct the co-citation network. Frequently co-cited references can be regarded as the intellectual base of the field. By analyzing the co-citation network, the background and knowledge foundations of research on GenAI in depression were identified. CiteSpace, using the log-likelihood ratio (LLR) algorithm, detected ten major clusters (Figure 4(b)). Each cluster was automatically labeled with keywords extracted from the cited references.
The ten largest clusters were: “effectiveness feasibility” (Cluster #0, size = 52), “mental health” (Cluster #1, size = 41), “depression management” (Cluster #2, size = 37), “deep learning” (Cluster #3, size = 36), “natural language processing” (Cluster #4, size = 36), “validated questionnaire” (Cluster #5, size = 35), “social media platform” (Cluster #6, size = 25), “early detection” (Cluster #7, size = 20), “providing mental health assessment” (Cluster #8, size = 20), and “side chain” (Cluster #9, size = 20). Because CiteSpace cluster labels generated by the log-likelihood ratio (LLR) algorithm are assigned automatically from cited-reference terms, not all labels are equally interpretable. In the present network, Cluster #9 (“side chain”) appeared to be an anomalous automated label that did not map clearly onto the core topic of GenAI in depression and was therefore not used as a basis for substantive thematic interpretation.
Excluding the anomalous label of Cluster #9 from higher-level interpretation, the remaining co-citation clusters suggested several broadly interpretable thematic areas. Clusters #0 and #1 reflected the overall scope of the field, emphasizing feasibility and mental health. Clusters #2, #7, and #8 were related to application-oriented themes in depression care, including management, early detection, and assessment. Clusters #3 and #4 represented core methodological and technical foundations, particularly deep learning and natural language processing. Clusters #5 and #6 highlighted measurement approaches and data sources, including validated questionnaires and social media platforms. Overall, these interpretable clusters suggest that the indexed literature is structured around methodological foundations, application-related themes, and measurement or data-related domains.
Analysis of top-cited articles
The top 10 papers with the highest number of citations.
Most of the highly cited articles were published between 2023 and 2024. The top 10 articles received a total of 363 citations, ranging from 18 to 86. The analysis revealed that the study published in NPJ Digital Medicine (2023) evaluated AI-based conversational agents for the promotion of mental health and well-being and received the highest number of citations (86). This article systematically examined the effectiveness of GenAI and multimodal systems delivered through mobile platforms, highlighting their role in reducing depression and psychological distress.
The second most cited article was published in Journal of Consumer Psychology (2023, 45 citations), which investigated the limitations of GenAI chatbots in crisis detection and raised important safety concerns. Another influential work, appearing in Translational Psychiatry (2023, 41 citations), analyzed applications of natural language processing in mental health research.
Furthermore, six of the ten top-cited papers focused on the application of ChatGPT and related GenAI tools in mental health care, covering aspects such as user dependency, clinical evaluation, adoption factors, comparative assessments with human experts, empathy in patient–AI interactions, and explainable AI frameworks for depression and suicide detection. These included studies published in International Journal of Educational Technology in Higher Education (2024, 37 citations), Annals of Biomedical Engineering (2024, 34 citations), Frontiers in Psychiatry (2024, 32 citations), Telemedicine and e-Health (2024, 27 citations), JMIR Mental Health (2024, 22 citations), Journal of Neurology (2024, 21 citations), and Cognitive Systems Research (2024, 18 citations).
Overall, the highly cited publications mainly emphasized three aspects: the effectiveness and safety of AI-driven conversational agents, the clinical and behavioral implications of AI adoption, and the technological development of GenAI in mental health. These articles constitute the cornerstone of research in this area and collectively define the current research landscape.
Frequency and clustering analysis of keywords
Keywords used in each scientific publication represent its essence, and keyword co-occurrence accurately identifies the most active research areas in the field. In this study, only keywords with a frequency of ≥5 occurrences were included. This threshold was selected a priori to balance sensitivity and interpretability in a relatively small dataset of 115 publications. A total of 20 keywords met this threshold, allowing the co-occurrence network to focus on recurrent themes while reducing noise from isolated or infrequently used terms. For example, singular and plural forms (e.g., “large language model” and “large language models”) were merged; case variations (e.g., “ChatGPT” and “chatgpt”) were unified; and closely related expressions were consolidated to avoid artificial inflation of keyword frequency. The top 20 terms ranked by frequency are shown in Figure 5(a). Among these keywords, “Depression” was the most frequent, appearing 55 times, followed by “Artificial Intelligence” (n = 54) and “Mental Health” (n = 32). “Generative Artificial Intelligence” appeared 15 times, ranking 10th among the high-frequency terms. Notably, “Chatbot” (n = 18) and “ChatGPT” (n = 18) also showed high frequencies. However, this prominence should be interpreted cautiously, because the inclusion of model-specific terms in the search strategy may have increased the retrieval of studies explicitly referencing well-known platforms such as ChatGPT. Comprehensive details are provided in Supplementary Table 3. Keyword analysis. (a) Top 20 most frequent keywords. (b) VOSviewer keyword co-occurrence network with seven thematic clusters (colors). (c) Citation overlay visualization; yellow nodes indicate higher average citations. (d) Temporal overlay visualization; colors from blue (ca. 2023) to yellow (ca. 2025) indicate topical evolution.
These keywords were grouped into seven clusters, with closely related terms clustered together (Figure 5(b)). Cluster 1, shown in red, was mainly related to the clinical features and diagnosis of mental disorders, including “major depressive disorder,” “anxiety disorders,” “disorders,” and “biomarkers.” Cluster 2, shown in green, focused on digital mental health technologies and detection methods, including “depression detection,” “digital mental health,” “help-seeking,” and “conversational agent.” Cluster 3, shown in blue, was closely related to the implementation and evaluation of mental health interventions, including “mental health interventions,” “care,” “efficacy,” and “medication.” Cluster 4, shown in yellow, emphasized specific applications of AI technology and system construction, represented by “computational modeling,” “data models,” “electronic health records,” and “biological system modeling.” Cluster 5, shown in light blue, highlighted machine learning and deep learning core technologies, including “machine learning,” “deep learning,” “accuracy,” and “determinant.” Cluster 6, shown in purple, focused on ChatGPT and its application scenarios, including “chatgpt,” “chatbot,” “help,” and “explainable artificial intelligence.” Cluster 7, shown in orange, emphasized the technological development of large language models, including “gpt-4,” “life,” “exercise,” and “biased.” To reveal the hotspots and trends of GenAI in depression research, a visual knowledge map was formed on the basis of keyword co-occurrence and clustering analysis.
Figure 5(c) displays the citation overlap co-occurrence network of keywords in the fields of mental health and artificial intelligence, whereas Figure 5(d) shows the temporal overlap co-occurrence network. Each node corresponds to a specific keyword, with node size reflecting frequency and the distance between nodes representing the strength of associations. In Figure 5(c), a color gradient visualizes citation intensity: the least-cited keywords appear in blue, whereas the most-cited keywords appear in yellow. The most frequently cited keywords included “artificial intelligence,” “depression,” “machine learning,” and “mental health.” In terms of temporal distribution (Figure 5(d)), earlier keywords appeared in blue (e.g., “mental health,” “major depressive disorder,” “health care”), whereas recent keywords appeared in yellow (e.g., “chatbot,” “chatgpt,” “digital health,” “conversational agent”). By analyzing the temporal evolution of keywords, the focus of research has shifted from foundational mental health theories and diagnostic methods to concrete applications of AI technology and the development of digital intervention programs.
Further analysis of research frontiers
The keyword burst analysis is shown in Figure 6(a). Through an integrated examination of the fifteen strongest citation burst terms that have persisted for a minimum of 1 year, we found that the keyword “large language models” (2023–2024) has garnered the most recent attention. Keywords such as “digital health” (2023–2024), “care” (2023–2024), and “disorders” (2023–2024) have been prominently utilized in recent studies. This pattern indicates a growing research interest in applying advanced AI technologies to mental health care, suggesting that future studies will continue to focus on these emerging topics. To further establish connections between research trends and hot spots in GenAI for mental health, a word cloud (Figure 6(b)) of frequently used keywords was generated, highlighting “artificial intelligence,” “depression,” “mental health,” and “machine learning” as central themes. Furthermore, the journal citation network (Figure 6(c)) visualizes the most influential sources in the field. The overlay colors, representing average citation impact, confirm that journals like Frontiers in Psychiatry and JMIR Mental Health (yellow nodes) are not only productive (as shown in 3.2) but also serve as central hubs with high citation impact, linking various research clusters. These findings demonstrate that current research on GenAI applications has primarily focused on mental health disorders, particularly depression detection and digital intervention solutions. Furthermore, the analysis reveals a shifting research focus from traditional mental health concepts toward AI-driven approaches utilizing large language models and conversational agents. Research frontiers and influential journals. (a) Top 15 keywords with strongest citation bursts (CiteSpace). Red bars indicate burst duration. (b) Keyword cloud. (c) VOSviewer journal citation network; color (blue to yellow) indicates average citation impact.
Discussion
The discussion below distinguishes between findings directly supported by bibliometric analyses and narrative contextualization informed by the broader literature and the supplementary PubMed review.
Rapid growth of GenAI in depression research
The publication data showed a marked increase beginning in 2023, temporally aligning with the public release of ChatGPT and similar large language models. Within the indexed literature, this pattern indicates rapidly increasing scholarly attention to GenAI applications in depression research. However, these bibliometric trends reflect growth in publication activity rather than direct evidence of clinical implementation or practice change.
Global collaboration and regional disparities
Our analysis demonstrates a distinct geographic concentration, with the United States and China accounting for the majority of the research output. While leading institutions such as Harvard University have formed highly connected internal collaboration clusters, cross-regional cooperation remains limited. Most partnerships emerge within national borders rather than spanning East and West (Figure 3). This concentration mirrors broader patterns in global health science, where research activity is often skewed toward high-income nations.20,21 Such geographic imbalance raises concerns about the generalizability of GenAI models, as cultural context is a critical determinant in the manifestation and treatment of depression. Although international co-authorship is associated with increased research visibility and citation impact,22,23 our results show that genuine cross-regional collaboration in this emerging field remains modest. Future stakeholders should prioritize inclusive multicenter studies to ensure that GenAI tools are validated across diverse cultural populations.
Beyond research productivity disparities, geographic concentration has implications for model generalizability. Many widely used GenAI systems, including large language models, are predominantly trained on English-language corpora and Western-centric datasets. As cultural context significantly shapes the expression, idioms, and stigma of depression, models optimized on Western linguistic patterns may fail to accurately interpret culturally specific symptom narratives in low- and middle-income countries. This raises concerns about diagnostic sensitivity, linguistic bias, and potential misclassification when such systems are deployed globally. Therefore, future development of GenAI tools in depression care should emphasize culturally adaptive training datasets, multilingual validation studies, and context-aware fine-tuning to enhance equitable applicability. Such regional concentration also highlights the need for cross-cultural validation to ensure that GenAI tools remain clinically relevant across diverse health care systems and sociocultural contexts.
Evolution of intellectual structure and thematic frontiers
Based on the co-citation clusters and keyword co-occurrence patterns, the field can be broadly organized into three dimensions: foundational technologies (e.g., NLP and deep learning), clinical application themes (e.g., depression management and early detection), and measurement-related topics (e.g., validated questionnaires). Keyword evolution showed a change in thematic emphasis from earlier prediction-oriented terms, such as “machine learning,” toward more recent high-frequency terms including “conversational agents,” “ChatGPT,” and “digital health.” These observations indicate changing thematic attention within the indexed literature rather than a bibliometrically proven shift in clinical paradigms.
Comparison with previous related studies
The present study should also be interpreted in relation to prior work in adjacent domains. A recent bibliometric analysis by Ren et al. focused primarily on artificial intelligence applications in depression detection and diagnosis, emphasizing conventional AI approaches such as machine learning, deep learning, and multimodal predictive modeling rather than GenAI-specific systems. 14 By contrast, the current study specifically maps WoSCC-indexed literature on GenAI in depression, with particular attention to large language models, conversational agents, and digital mental health applications that have emerged rapidly since late 2022. In addition, recent systematic reviews have examined large language models and GenAI in broader mental health or psychiatric contexts, highlighting both their clinical promise and their limitations in safety, accuracy, and evidence maturity.17,24–26 Taken together, these earlier studies provide important contextual grounding, whereas the present analysis contributes a field-level bibliometric mapping of how GenAI-related depression research has been structured, cited, and thematically organized within the indexed literature.
Bridging the gap: From bibliometric volume to clinical evidence
Selected recent clinical trials and feasibility studies relevant to AI-/chatbot-supported mental health applications.
Abbreviations: HADS, Hospital Anxiety and Depression Scale.
GenAI comhpared with earlier AI paradigms
The following comparison is provided as narrative context from the broader literature rather than as a direct output of the bibliometric mapping. Traditional discriminative AI primarily focuses on pattern recognition, such as classifying whether a patient may have depression on the basis of static inputs including social media text or voice biomarkers.6,7 While useful for screening, these models are generally limited in their capacity for real-time interaction. 16 By contrast, GenAI, particularly large language models (LLMs), can generate context-sensitive text responses and therefore has been discussed in the broader literature as enabling more interactive use cases in mental health settings. 17
Recent reviews and empirical studies outside the bibliometric indicators have discussed this potential in relation to screening, psychoeducation, and clinician support. 17 For example, ChatGPT has been reported to generate depression-related recommendations that are in some cases broadly aligned with clinical guidance, 12 and ChatGPT-4–adapted questionnaires have shown moderate-to-strong agreement with validated instruments in a cross-sectional study. 13 These studies provide contextual interpretation for why recent indexed publications increasingly discuss conversational agents and digital mental health applications, but they should not be interpreted as direct bibliometric findings from the present analysis. 29
Modalities of GenAI in depression care
While our bibliometric findings identify increasing attention to conversational agents, digital health, and large language models, they do not directly generate clinical taxonomies. To contextualize these thematic trends within the broader literature, we present an interpretive synthesis of how GenAI applications have been described across the clinical pathway of depression care. First, in detection and screening, GenAI systems have been explored for active conversational screening (e.g., LLM-assisted questionnaire delivery) 28 as well as passive identification of depressive signals from digital biomarkers derived from text, voice, or typing behavior.30,31 Second, in diagnostic support contexts, such systems have been evaluated for assisting clinicians in summarizing patient information and retrieving relevant clinical knowledge, although these applications remain investigational and require clinical oversight. 32 Third, in treatment and intervention settings, GenAI tools have been studied as platforms for delivering psychoeducation and structured digital therapeutic content (e.g., CBT-informed modules), and are often conceptualized as adjunctive support tools rather than replacements for clinician-led therapy. 33 Fourth, for monitoring and relapse prevention, emerging studies describe the potential of dialogue-based systems to support symptom tracking and flag high-risk language patterns, including signals related to suicidal ideation; however, robust prospective validation remains limited. 34 Finally, in research and development contexts, generative models have also been explored for hypothesis generation, drug discovery pipelines, and clinical trial design optimization. 35 Across these domains, common themes discussed in the literature include personalization, human-in-the-loop oversight, and the need for robust ethical and privacy governance frameworks.34 Importantly, this framework represents an integrative interpretation of the current literature rather than a structure statistically derived from bibliometric clustering. 36
Challenges in clinical implementation: Safety and accuracy
Beyond bibliometric limitations, the clinical deployment of GenAI in depression faces substantial hurdles. Safety remains the primary concern. Unlike passive detection models, generative agents can produce “hallucinations” or unsafe advice. A 2025 content analysis of GenAI responses to suicide inquiries found that while models have improved, they still occasionally fail to provide appropriate crisis resources or may offer ambiguous responses that could be dangerous for high-risk users. 37 Beyond safety, the unsupervised deployment of generative AI in mental health raises regulatory, ethical, and governance challenges. Clear guidelines, oversight mechanisms, and accountability frameworks are needed to ensure responsible use, protect patient privacy, and mitigate potential harm. Furthermore, diagnostic accuracy and cultural competence remain insufficient. A recent systematic review concluded that while GenAI models show promise in psychoeducation, their ability to accurately diagnose depression across diverse cultural contexts is weak, and longitudinal validation data are scarce. 25
While clinicians anticipate opportunities, surveys indicate significant hesitation regarding accountability and ethical risks. 38 Future research must therefore prioritize the development of “safety guardrails” specifically designed for psychiatric contexts before unsupervised use can be recommended.26,39
Implications of research trends for clinical translation
Our bibliometric findings suggest increasing scholarly attention to screening-related, monitoring-related, and intervention-oriented applications of GenAI in depression research. These patterns may help identify areas of emerging interest that could inform future clinical translation and research prioritization. However, bibliometric evidence alone cannot establish clinical function, implementation readiness, or real-world effectiveness. At present, the indexed literature more appropriately suggests potential adjunctive directions for psychoeducation, symptom monitoring, and supportive digital mental health applications, all of which still require rigorous prospective validation, clinical oversight, and robust governance to address risks related to safety, bias, privacy, and accountability.29,40,41
Future research agenda
Based on current bibliometric trends and emerging evidence, future studies should focus on rigorous clinical validation through large-scale RCTs, development of safety guardrails and ethical governance frameworks for unsupervised LLM deployment, integration of human-in-the-loop models to maintain accountability and quality of interventions, and creation of culturally adaptive and multilingual systems to ensure equitable applicability across diverse populations.
Limitations of the bibliometric analysis
Several methodological limitations of this study should be acknowledged. First, the quantitative visualization and network construction were primarily based on WoSCC to ensure metadata consistency and compatibility with bibliometric software. However, reliance on a single database may introduce coverage bias. In addition, only English-language publications were included, which may have excluded relevant studies published in other languages and may have further limited the representativeness of the indexed literature. Although WoSCC is widely used in bibliometric research due to its structured citation indexing, other databases such as Scopus or Embase may include additional regional journals, conference proceedings, or early online publications not indexed in WoSCC. Inclusion of multiple databases could potentially alter publication counts, collaboration networks, and citation patterns. Future studies integrating multi-database sources with standardized deduplication procedures may provide a more comprehensive mapping of this evolving field. Second, although our search strategy incorporated both general GenAI terminology and the names of widely used large language models (e.g., ChatGPT, GPT-4, LLaMA, Claude), this approach may have introduced a systematic retrieval bias toward studies that explicitly referenced well-known commercial platforms. As a result, studies using open-source, fine-tuned, or domain-specific models without brand-name mentions may have been underrepresented. This limitation has interpretive implications: the prominence of terms such as “ChatGPT” in the keyword analysis may partly reflect the structure of the search strategy rather than the true distribution or dominance of specific GenAI platforms in the broader field. Future bibliometric studies using expanded conceptual terminology and broader model descriptors may provide a more balanced representation of the evolving literature.
Third, the 2025 publication count covered only the period from January to July 23, 2025, rather than the full calendar year. Therefore, annual growth estimates and year-to-year comparisons involving 2025 should be interpreted cautiously. Bibliometric indicators, such as publication counts and citation frequencies, are also inherently lagging indicators of scientific progress. This limitation is particularly relevant in the context of GenAI, where the literature is evolving rapidly and new models, applications, preprints, conference papers, and early online publications may emerge faster than they can be indexed and cited. 24 Fourth, citation-based indicators may introduce structural bias. Publications from high-income countries, English-language journals, or early entrants in an emerging field are more likely to accumulate citations, potentially amplifying visibility disparities. As a result, bibliometric metrics may underrepresent contributions from emerging research communities or recently published studies that have not yet accrued citation impact.
Conclusion
This bibliometric study mapped the WoSCC-indexed research landscape of GenAI in depression from 2023 to mid-2025. The analysis identified rapid growth in publication output and increasing scholarly attention to large language models, conversational agents, and digital mental health applications within the indexed literature. While GenAI may hold potential to complement psychiatric care, bibliometric findings alone cannot establish clinical effectiveness or implementation readiness. The findings therefore highlight the need for rigorous prospective validation, broader database coverage, and inclusive international collaboration to support the responsible and equitable development of this field.
Supplemental material
Supplemental material - Generative artificial intelligence in depression research: A bibliometric analysis of WoSCC-Indexed literature
Supplemental material for Generative artificial intelligence in depression research: A bibliometric analysis of WoSCC-Indexed literature by Hongfei Chen, Lin Chen, Jin Yang, Aifa Tang and Yafei Yang in Digital Health.
Supplemental material
Supplemental material - Generative artificial intelligence in depression research: A bibliometric analysis of WoSCC-Indexed literature
Supplemental material for Generative artificial intelligence in depression research: A bibliometric analysis of WoSCC-Indexed literature by Hongfei Chen, Lin Chen, Jin Yang, Aifa Tang and Yafei Yang in Digital Health.
Supplemental material
Supplemental material - Generative artificial intelligence in depression research: A bibliometric analysis of WoSCC-Indexed literature
Supplemental material for Generative artificial intelligence in depression research: A bibliometric analysis of WoSCC-Indexed literature by Hongfei Chen, Lin Chen, Jin Yang, Aifa Tang and Yafei Yang in Digital Health.
Supplemental material
Supplemental material - Generative artificial intelligence in depression research: A bibliometric analysis of WoSCC-Indexed literature
Supplemental material for Generative artificial intelligence in depression research: A bibliometric analysis of WoSCC-Indexed literature by Hongfei Chen, Lin Chen, Jin Yang, Aifa Tang and Yafei Yang in Digital Health.
Supplemental material
Supplemental material - Generative artificial intelligence in depression research: A bibliometric analysis of WoSCC-Indexed literature
Supplemental material for Generative artificial intelligence in depression research: A bibliometric analysis of WoSCC-Indexed literature by Hongfei Chen, Lin Chen, Jin Yang, Aifa Tang and Yafei Yang in Digital Health.
Footnotes
Ethical considerations
Because this article does not contain any studies with human or animal subjects.
Author contributions
H.C.: writing – review and editing; writing – original draft; methodology; investigation; funding acquisition. L.C.: writing–review and editing; writing – original draft; methodology; investigation. J.Y.: writing–review and editing; writing – original draft; methodology; investigation. A.T.: writing – review and editing; writing – original draft; supervision; project administration; conceptualization. Y.Y.: writing – review and editing; writing – original draft; supervision; conceptualization; data curation.
Funding
The authors disclosed receipt of the following financial support for the research, authorship, and/or publication of this article: The present study was supported by the National Natural Science Foundation of China (grant no. 82300869), the 78th Batch of General Funding Projects of China Postdoctoral Science Foundation (2025M782291), The Sanming Project of Medicine in Shenzhen (No. SASM202201024), Health Commission of Sichuan Province Medical Science and Technology Program(grant no.24QNMP083), Chengdu Science and Technology Bureau Technology Innovation R & D Project(2024-YF05-00989-SN), Sichuan Medical (Youth Innovative) Scientific Research Projects (grant nos. Q23002, S22002, S22049 and Q20013), the Innovation Team Project of the Affiliated Hospital of Chengdu University (CDFYCX202207), the Chengdu Medical Research Projects (grant nos. 2022051, 2023128, 2022181,2023087 and 2023343), Young Talents Program of the Affiliated Hospital of Chengdu University and the Program of the Affiliated Hospital of Chengdu University (grant nos. 2020YZZ08, Y202204 and Y202207).
Declaration of conflicting interests
The authors declared no potential conflicts of interest with respect to the research, authorship, and/or publication of this article.
Data Availability Statement
Data will be made available upon reasonable request.
Disclosures
All authors declare that they have no competing financial interests to disclose.
Guarantor
Yafei Yang is the guarantor of this work.
Supplemental material
Supplemental material for this article is available online.
References
Supplementary Material
Please find the following supplemental material available below.
For Open Access articles published under a Creative Commons License, all supplemental material carries the same license as the article it is associated with.
For non-Open Access articles published, all supplemental material carries a non-exclusive license, and permission requests for re-use of supplemental material or any part of supplemental material shall be sent directly to the copyright owner as specified in the copyright notice associated with the article.
