Abstract
Background
Alzheimer's disease (AD) lacks effective early diagnostic tools. N-glycosylation dysregulation involving the enzyme MGAT3 is implicated in AD pathogenesis, with untapped clinical potential for biomarker development.
Objective
To identify and validate MGAT3-associated diagnostic biomarkers for AD using integrative multi-disciplinary approaches.
Methods
Bibliometric analysis of N-glycosylation-AD research; large-scale GEO dataset preprocessing, WGCNA-integrated bioinformatics, machine learning, and penalized regression in well-characterized independent training/validation sets.
Results
Bibliometric analysis revealed −1.94% annual growth, 21.76% international co-authorship, leading contributions from the U.S., Japan, and China; Journal of Biological Chemistry as the core journal; Taniguchi N and Kizuka Y as key authors; Lee JH (2010) and Zielinska DF (2010) as highly cited works; key keywords including “bisecting GlcNAc” and “MGAT3/GnT-III”; and thematic shifts toward amyloid pathology, MGAT3-driven glycosylation, and neuronal signaling pathways. ATP6V1G2 and CHST6 emerged as core MGAT3-associated biomarkers with a strong antagonistic relationship (r = −0.63). Both demonstrated robust standalone diagnostic efficacy (AUC > 0.72) across diverse patient subgroups in the training set, with the CHST6 + ATP6V1G2 + PCSK1 combination achieving AUC 0.755. Validation confirmed ATP6V1G2 as the consistently top single biomarker (AUC = 0.761), while the ATP6V1G2 + ENO2 + SEZ6L2 panel reached peak AUC 0.764. ATP6V1G2 was significantly downregulated and CHST6 markedly upregulated in AD cohorts, with clinical utility validated via nomograms for risk stratification, decision curve analysis (net benefit 0.20–0.25), and >80% reliable correspondence between high-risk individuals and confirmed AD cases.
Conclusions
MGAT3-associated biomarkers (ATP6V1G2, CHST6) and their combinations offer promising AD diagnostic potential, advancing insights into N-glycosylation-mediated disease mechanisms.
Keywords
Introduction
Alzheimer's disease (AD) is a progressive neurodegenerative disorder that poses a significant global health challenge due to its profound impact on cognitive function and memory. 1 The burden of AD extends beyond the affected individuals, placing immense emotional and financial strain on caregivers and healthcare systems, with global costs projected to reach hundreds of billions annually. 2 As the disease progresses, patients exhibit a range of cognitive impairments that complicate diagnosis and treatment. 3 Current diagnostic methodologies primarily rely on clinical assessments and established biomarkers, such as amyloid-β and tau protein levels. However, these approaches often lack the precision needed for early detection, leading to a critical gap in effective intervention strategies. 4
Despite extensive research efforts, therapeutic options for AD remain limited, primarily focusing on symptomatic relief rather than addressing the underlying disease mechanisms. 5 The need for innovative biomarkers and therapeutic targets has never been more urgent, especially considering the rising prevalence of AD in an aging population. 6 Recent studies have indicated that post-translational modifications, particularly N-glycosylation, may play a pivotal role in the pathogenesis of AD. 7 N-glycosylation has been linked to various neurodegenerative processes, such as amyloid-beta aggregation and tau hyperphosphorylation, suggesting that dysregulated glycosylation could contribute to disease progression by altering protein function and neuronal signaling pathways. 8
In light of these findings, our study aims to explore the role of N-glycosylation in AD pathogenesis through a comprehensive multi-disciplinary approach. By integrating bibliometric analysis, bioinformatics, and machine learning, we seek to identify and characterize N-glycosylation-associated biomarkers that could enhance diagnostic accuracy and therapeutic efficacy. 9 Utilizing large-scale gene expression datasets, such as those from the Gene Expression Omnibus (GEO), we will employ advanced computational techniques, including Weighted Gene Co-expression Network Analysis (WGCNA) and penalized regression methods, to analyze complex biological data.
Our research objectives are threefold: first, to map the existing research landscape surrounding N-glycosylation in AD, thereby identifying key trends and gaps in the literature; second, to uncover co-expression networks and hub genes involved in AD pathology; and third, to validate specific biomarkers associated with the glycosylation enzyme MGAT3 for their potential diagnostic utility. 10 By addressing these objectives, we aim to pave the way for future studies focused on the mechanistic understanding of N-glycosylation in AD and its potential as a target for therapeutic intervention.
In conclusion, this study seeks to fill critical gaps in the understanding of AD by integrating multiple omics approaches to elucidate the role of N-glycosylation in disease progression. Our findings may contribute to the development of novel diagnostic and therapeutic strategies, ultimately improving outcomes for individuals affected by AD.
Methods
Data sources, retrieval strategy, and analytical approaches for N-glycosylation and Alzheimer's disease bibliometrics
To conduct a comprehensive bibliometric analysis of the N-glycosylation-AD intersection, we used the Web of Science Core Collection—a premier bibliometric database—and retrieved English research articles/reviews published 2001–2025 with the query: TS = ((“N-glycosylat*” OR “N-linked glycosylat*” OR “N-glycan*” OR “N-glycoprotein*” OR “protein N-glycosylat*” OR “n-glycoform*”) AND (Alzheimer* OR “Alzheimer's disease” OR “Alzheimer's” OR “Alzheimer disease” OR “Alzheimer's-like” OR “Alzheimer pathology” OR “Alzheimer pathogenesis” OR “tau protein” OR “beta-amyloid”)). This yielded 239 documents, filtered to retain only research articles/reviews (complete list and detailed results in Supplemental Table 1). For analysis, VOSviewer (v1.6.20) visualized bibliometric networks (author, institutional, and national collaborations; co-cited author, citation, and keyword clusters; timeline graphs), 11 CiteSpace (v6.4 R1) identified emerging trends, keyword bursts, and temporal theme evolution, 12 and R software (v4.3.3) assessed publication/citation trends, author productivity, journal distributions, and institutional contributions—with plotting and citation fitting. The analysis covered general trends, countries, institutions, journals, authors, documents, topics, keywords, and bioinformatics of AD- and N-glycosylation-related genes to ensure a systematic field overview.
Preprocessing, batch effect correction, and correlation analysis of Alzheimer's disease GEO datasets
All analyses were performed in R version 4.3.2: Two AD-related datasets—GSE118553 (401 samples) and GSE132903 (195 samples)—were designated as the training set, while three additional datasets—GSE28146 (30 samples), GSE48350 (253 samples), and GSE5281 (161 samples)—served as the validation set. Raw gene expression data for all datasets were retrieved via the GEOquery package, with probe-level data aggregated to the gene level; missing values were imputed using impute.knn () (impute package), and quantile normalization was performed with normalize.quantiles() (preprocessCore package). Batch effects were corrected using removeBatchEffect () (limma package) with a batch factor corresponding to each dataset, followed by calculation of batch separation metrics (Within Mean/SD, Between Mean/SD, Separation Ratio) via grouped summarization (dplyr package). Pairwise Pearson correlations between samples within each set (training and validation) were computed using base R's cor () function, and correlation plots were generated with ggplot2 following journal formatting standards.
WGCNA-mediated co-expression network analysis and hub gene identification in Alzheimer's disease
All WGCNA analyses and visualizations were performed in R version 4.3.2 using the WGCNA (v1.72-5), ggplot2 (v3.4.4), and dplyr (v1.1.3) packages, utilizing the training set as the data source. The training set's gene expression matrix was preprocessed to filter out genes with >20% missing values and low-expression genes (average expression < 0.5). Optimal soft threshold power was selected via pickSoftThreshold() based on the scale-free topology criterion (R2 > 0.85). Weighted co-expression networks were constructed using blockwiseModules() (minModuleSize = 30, mergeCutHeight = 0.25, networkType = “signed”, which clustered genes into color-labeled modules. Module-trait correlation was analyzed by calculating Pearson coefficients between module eigengenes and AD trait, with statistical significance assessed via corPvalueStudent(). Hub genes were identified based on high module membership (MM > 0.8) and gene significance (GS > 0.6). Visualizations—including soft threshold plots, module clustering trees, module-trait bubble plots, and hub gene mountain plots—were generated using built-in WGCNA functions and ggplot().
Differential expression, gene intersection, and machine learning for MGAT3-associated gene identification
Differential expression analysis of the training set was performed in R (v4.3.2) using the limma package: raw expression data were quantile-normalized, after which linear models were fitted to compute log2 fold change (log2FC) and Benjamini-Hochberg-adjusted p-values (FDR). Volcano plots were generated in R using the ggplot2 and ggrepel packages, with log2FC plotted on the x-axis and -log10(adjusted p-value) on the y-axis; genes were color-coded by expression status (light blue for downregulated, red for upregulated, gray for non-significant), and key genes were labeled without overlap. Overlaps between three gene sets—differentially expressed genes (DEGs), N-glycosylation genes, and WGCNA-derived genes—were quantified and visualized as a Venn diagram in R via the VennDiagram package.
Machine learning analysis of 336 candidate genes—intersecting DEGs, N-glycosylation-related genes, and WGCNA-identified AD-associated genes—was performed in R (v4.3.2) to identify MGAT3-associated molecules. Four algorithms (random forest, XGBoost, glmnet, PLS) were implemented via respective R packages, with model robustness enhanced by 5-fold cross-validation, sample size-adapted parameter tuning (tuneLength = 2–5), and parallel computing to mitigate overfitting. Gene overlap across algorithms was visualized using VennDiagram, while Pearson correlation analysis between candidate genes and MGAT3 was conducted with Benjamini-Hochberg-adjusted p-values. Feature importance ranking stability was assessed via 100 bootstrap resamples, with consistency measured by Spearman's rank correlation and variance analysis across runs. A combined ranking integrated correlation strength, log2 fold change, and scaled machine learning scores (0–1), with visualizations generated using ggplot2, ggrepel, ggraph, and tidygraph; statistical significance was denoted using standard notation (***, **, *, NS) corresponding to p-adjusted < 0.001, < 0.01, < 0.05, and non-significant, respectively, with color-coding for correlation direction (positive: red/orange; negative: blue/green) and strength (strong: |r| > 0.7; moderate: 0.5 < |r| ≤ 0.7; weak: |r| ≤ 0.5).
Penalized regression for core biomarker identification among MGAT3-associated molecules
Data preprocessing and statistical analyses were conducted in R (v4.3.2) using packages including tidyverse, glmnet, and ggplot2. All analyses were performed on training set data, focusing on the TOP 20 molecules closely associated with MGAT3; raw data were cleaned by removing missing values, converting the outcome variable to a binary factor (Control/Case), and generating a predictor matrix excluding the outcome. To further narrow down the scope and identify key molecules, three penalized regression approaches were implemented for feature selection: Lasso (alpha = 1), ElasticNet (alpha = 0.5), and AdaptiveLasso (alpha = 1), with the latter incorporating penalty factors derived from Ridge regression coefficients to prioritize biologically relevant features. To mitigate overfitting, 5-fold cross-validation was employed to select the optimal regularization parameter (λ.min) based on deviance minimization, and parallel computing (via the doParallel package) was enabled to improve computational efficiency. Model robustness was corroborated by intersecting results across the three methods to identify core biomarkers, while the stability of feature importance rankings was assessed by comparing the order of absolute standardized coefficients across all approaches. Visualizations comprised feature importance plots—displaying standardized coefficients, effect directions, and classifications as core or method-specific—and coefficient path plots illustrating how coefficients change along the regularization path, with the optimal λ.min marked for reference.
Machine learning-based diagnostic model development with MGAT3-associated candidates in the training set
Data preprocessing and statistical analyses were performed in R (v4.3.2) using packages including tidyverse, randomForest, xgboost, glmnet, pROC, rmda, and ggplot2. The 10 candidate molecules—including MGAT3—were derived from prior feature selection via Lasso, ElasticNet, and AdaptiveLasso algorithms, which narrowed down MGAT3-associated top molecules for subsequent validation in the training set. The dataset was split into training (70%) and test (30%) subsets using createDataPartition to ensure external validity. Three machine learning models were constructed: logistic regression, random forest (ntree=100), and XGBoost (nrounds=50). Overfitting was mitigated through 5-fold cross-validation, hyperparameter tuning, and retention of biologically relevant features based on coefficient magnitude and Gini impurity reduction. Feature importance stability was confirmed by cross-validating normalized scores across all three models, while biological effect directions were determined via logistic regression coefficients and Spearman's correlation with disease status. Visualizations—including feature importance plots, ROC curves, molecular interaction networks, decision curves, clinical impact curves, nomograms, and expression boxplots—were generated to quantify diagnostic performance, molecular interactions, and clinical utility.
MGAT3-associated candidate assessment with consistent methodologies in the validation set
The present analysis utilized data from the validation set, with candidate molecules selected based on their robust association with MGAT3—derived by narrowing down the top MGAT3-related molecules via three algorithms in the prior step—including MGAT3 itself. All statistical analyses, differential expression assessments, diagnostic efficacy evaluations, and visualizations were performed in R (v4.3.2) following the identical methodologies described previously.
Our study's complete workflow encompasses bibliometrics and bioinformatics, as illustrated in Figure 1, with detailed information for the used GEO datasets provided in Figure 2.


Results
Research landscape: publication trends, collaboration, and citations in N-glycosylation and Alzheimer's disease
Supplemental Figure 1A presents key field metrics: annual publication trends (mixed overall with a −1.94% annual growth rate yet recent increases), high collaborative engagement (21.76% international co-authorship, 6.92 average authors per document), a mean citation count of 41.69 per document, and data on author numbers, references, and average document age; Supplemental Figure 1B details 2001–2024 annual publication volumes and mean total citations per article: 8 articles launched the field in 2001 (3.94 mean annual citations), with a 2010 citation peak of 16.93, a 2013 decline to 1.14, a 2023 publication high of 24 articles (3.56 mean citations), and 2024 citations dropping to 0.84 due to recency and inherent citation lag.
Country-level contributions and collaborative networks in N-glycosylation and Alzheimer's disease research
Supplemental Figure 2A outlines global publication outputs: the United States led with 190 articles in 2025, followed by Japan (171), China (118, with rapid growth since 2010), Germany (82), and Sweden (48, steady increases since 2004); Supplemental Figure 2B details corresponding author distributions, with the U.S. (51 articles, 19.6% multiple-country publications [MCP]), Japan (43 articles, 9.3% MCP), China (38 articles, 5.3% MCP), Germany (28 articles, 42.9% MCP), and Sweden (16 articles); Supplemental Figure 2C depicts the global collaboration network, with the U.S. as the central hub (notably partnering with Germany), Germany connecting to France and the UK, and the UK collaborating with Sweden and the U.S.; Supplemental Figure 2D presents the country co-authorship network, where circle size denotes publication volume (U.S., Japan, Germany, and China as top contributors), TLS values reflect collaboration depth (high for the U.S., Germany, and Japan), and a color gradient indicates publication recency (China's warmer tone signaling a recent surge in contributions).
Institutional contributions, temporal trends, and collaborative networks: N-glycosylation and Alzheimer's disease
Supplemental Figure 3A shows institutional temporal publication trends: Karolinska Institutet (steady growth since 2005, recent peak), RIKEN (consistent upward trajectory since 2006), University of California System (modest 2012 start, dramatic 2023 + surge), Egyptian Knowledge Bank (EKB, late entry, significant growth since 2019); Supplemental Figure 3B presents organizational co-authorship networks—larger circles denote high publication counts (e.g., RIKEN, Osaka University), with Cluster 1 encompassing Gifu University, Hiroshima University, and Osaka University (tight-knit collaborations); Supplemental Figure 3C highlights total link strength and publication recency: node size reflects collaboration levels, warmer colors indicate recent contributions, with RIKEN and Karolinska Institutet showing extensive collaborative networks, and Gifu University/Hiroshima University standing out with warmer nodes (active recent involvement).
Journal-level analysis: publication trends, citation networks, and interdisciplinarity in N-glycosylation and Alzheimer's disease
Supplemental Figure 4A shows that 2001–2025 publication volumes have grown across key journals, with the Journal of Biological Chemistry maintaining sustained high output and recent increases, while the Journal of Proteome Research and PLoS One exhibit significant growth in recent years; Supplemental Figure 4B depicts the co-citation network, where node size corresponds to citation count—with the Journal of Biological Chemistry as a central node and Nature, Science, and PNAS featuring large nodes—and line thickness reflects total link strength, with PNAS demonstrating substantial interconnectivity; Supplemental Figure 4C illustrates journal coupling, where node size aligns with publication volume and line thickness indicates co-citation link strength, highlighting the Journal of Biological Chemistry, Analytical Chemistry, Journal of Neurochemistry, and Glycobiology as prominent contributors; Supplemental Figure 4D presents the dual-map overlay, with the left side highlighting primary research areas (molecular biology, immunology, clinical medicine) and the right side showcasing foundational literature emphasizing molecular genetics, alongside cross-disciplinary connections and node size, color, and clustering patterns that reflect the dynamic evolution of research themes.
Author-level contributions: publication trends, co-citation, and co-authorship networks in N-glycosylation and Alzheimer's disease
Supplemental Figure 5A illustrates authors’ annual publication trends, with circle size representing yearly output—Taniguchi N and Kizuka Y maintaining consistent contributions (larger circles in recent years) and Schedin-Weiss S emerging as a prominent new contributor; Supplemental Figure 5B depicts the co-citation network, where node size reflects citation count, with key clusters centered on Kizuka Y and Taniguchi N, and Schedin-Weiss S and Hoffmann A holding central positions due to significant citation impact; Supplemental Figure 5C presents the co-authorship network, where circle size denotes document count and line thickness indicates total link strength—Taniguchi Naoyuki and Kizuka Yasuhiko stand as central figures with substantial output and collaborative impact, followed by Hashimoto Yasuhiro and Yamaguchi Yoshiki with notable contributions; Supplemental Figure 5D adds a temporal dimension to co-authorship (2010–2020 color gradient), highlighting Hashimoto Yasuhiro (total link strength [TLS] = 41) and Arai Hajime (TLS = 30) for enduring influence, while Kizuka Yasuhiko and Taniguchi Naoyuki exhibit more pronounced recent activity.
Document-level impact analysis: citation metrics, bursts, and interconnectedness in N-glycosylation and Alzheimer's disease
Figure 3A highlights the most cited documents: Lee JH (2010, Cell) 13 with 926 total citations and 57.88 citations per year, Zielinska DF (2010, Cell) 14 with 747 citations, and Franceschi C (2018, Front Med-Lausanne) 15 with 557 citations and 69.63 citations per year; Figure 3B identifies citation bursts across time periods: the early 2000s (Vassar R, 1999, Science, 16 burst strength 4.78; Selkoe DJ, 2001, 17 Physiological Reviews, 3.65), 2005–2009 (Francis R, 2002, 18 Developmental Cell, 3.35), 2010–2014 (Futakawa S, 2012, 19 Neurobiology of Aging, 3.33), 2015–2019 (Schedin-Weiss S, 2014, 20 FEBS Journal, 12.17; Kizuka Y, 2015, 21 EMBO Molecular Medicine, 7.29), and 2020–2024 (Zhang Q, 2020, 22 Science Advances, 6.21); Figure 3C depicts the four-cluster co-citation network, where red cluster features Akasaka-Manya K 23 (2010, Glycobiology, 33 citations, total link strength 293), blue cluster includes Gizaw ST 24 (2016, 19 citations), green cluster comprises Vassar R (1999, Science, 23 citations, 68 total link strength) and Liu F (2002, Neuroscience, 14 citations, 146 total link strength), and yellow cluster contains Kizuka Y 25 (2017, BBA-Gen Subjects, 31 citations); Figure 3D visualizes bibliographic coupling (node size = citations, line thickness = total link strength) with Zielinska (2010, 747 citations, 21 total link strength), Lee 13 (2010, 926 citations, 18), Franceschi (2018, >500 citations, 9), and Zhang (2020, >500 citations, 33) as central nodes, while Figure 3E (node size = total link strength, color = publication year) highlights Kizuka (2017), Schedin-Weiss (2014), Zielinska (2010), and Zielinska (2012) 26 as key interconnected documents.

Thematic evolution and research trends: keyword clustering and temporal shifts in N-glycosylation and Alzheimer's disease
Figure 4A (CiteSpace timeline view) depicts thematic progression in N-glycosylation-AD research 2015–2025: 2015–2017 was dominated by foundational themes (#0 cerebrospinal fluid, #1 neurodegenerative diseases), 2018–2020 saw focused research on #3 GnT-III activity (prominent 2017–2019), and 2021–2025 expanded to #5 amyloid pathology and #7 ribosome complex; Figure 4B (Keyword Timezone map) outlines trends across three periods: 2001–2010 top keywords by centrality were “expression” (0.25), “glycosylation” (0.22), “amyloid precursor protein” (0.20), “cerebrospinal fluid” (0.15), and “identification” (0.12); 2011–2020 shifted to “bisecting GlcNAc”, “mass spectrometry”, “endoplasmic reticulum stress”, “N-glycosylation”, and “brain”; 2021–2024 featured emerging themes including “protein glycosylation” (0.12), “pathology”, “GNT-III (MGAT3)”, and “aggregation”.

Batch variation mitigation and dataset compatibility validation in Alzheimer's disease gene expression
Prior to batch correction, GSE118553 and GSE132903 exhibited high Separation Ratios (8.44 and 52.90, respectively) and a Pearson correlation coefficient below 0.9 (Figure 5A); post-correction, Separation Ratios were substantially reduced (1.01 and 1.15), the correlation coefficient increased above 0.9 (Figure 5B), and batch-induced divergence was effectively mitigated, confirming the datasets’ compatibility for integration into a training set. Prior to batch correction, GSE28146, GSE48350, and GSE5281 exhibited high Separation Ratios (17.75, 29.10, and 13.21, respectively) and low inter-dataset Pearson correlations (Figure 5C); post-correction, Separation Ratios were markedly reduced (2.06, 0.81, and 1.43, respectively), and inter-dataset correlations were improved (Figure 5D), demonstrating effective mitigation of batch-induced variation and confirming the three datasets are compatible for integration.

Identification of AD-associated co-expression modules and hub genes via WGCNA
WGCNA was conducted on the training set to identify AD-associated co-expression modules, with key findings visualized in five figures. Figure 6A illustrates the scale-free topology fit index across soft threshold powers (1–20), while Figure 6B presents the corresponding average connectivity; this analysis led to the selection of an optimal soft threshold power of 6 (R2 = 0.88), ensuring the network adhered to scale-free properties. Figure 6C displays the gene clustering tree, with 26 distinct co-expression modules assigned unique colors and numerical identifiers (0–25). Figure 6D is a module-trait bubble plot where the x-axis represents individual modules, the y-axis denotes correlation coefficients with AD, and bubble size reflects statistical significance; the green module (Module 5) exhibited the strongest association with AD, marked by the highest correlation coefficient and a significantly small p-value (<0.001). Figure 6E is a mountain plot of the top 30 hub genes in the green module, showing a continuous trend in expression intensity as genes are ordered by their relevance to the module, with consistent peaks and troughs that reflect coordinated co-expression patterns among these core genes.

Identification and characterization of MGAT3-associated genes via multi-step bioinformatic analysis
In the training set, differential expression analysis yielded 8397 molecules with an adjusted p-value < 0.05, which are visualized in the volcano plot (Figure 7A): this plot displays log2FC on the x-axis and -log10(adjusted p-value) on the y-axis, with molecules grouped by expression status (light blue for downregulated, red for upregulated, gray for non-significant) and key genes (e.g., NEFM, GFAP) annotated. Figure 7B presents a Venn diagram of these 8397 DEGs, 6458 N-glycosylation genes, and 1129 WGCNA-derived AD-associated genes, which identifies a 336-gene overlap across all three sets; the full list of these overlapping genes is included in Supplemental Table 2. Counts of unique and pairwise-overlapping genes in each segment are also marked on the diagram.

Analysis of 336 candidate genes via four machine learning algorithms yielded consistent MGAT3-associated molecules: Figure 7C illustrates gene overlap across the four algorithms, reflecting cross-analytical consensus in candidate prioritization. Figure 7D presents feature importance scores (0–100), with core molecules including BSN, SCN2B, and STMN3 distinguished by scores exceeding 70—indicating strong algorithmic agreement—while most other genes show lower importance, and low bootstrap variance confirms stable performance. Figure 7E (multi-dimensional landscape) maps these key genes’ correlation coefficients (−0.5 to 0.5) to -log10(adjusted p-value), all achieving the most stringent statistical significance and displaying either moderate positive (BSN, SCN2B, STMN3) or weak negative (CHST6) relationships with MGAT3. Figure 7F (ranked bar plot) features these top genes by combined scores (0.46–0.74), led by BSN (rank 1) and SCN2B (rank 2) as leading moderate positive correlates and CHST6 as the top weak negative hit. Figure 7G (integrated feature importance) details their multi-dimensional profiles: BSN exhibits the strongest correlation with MGAT3 (0.64) and highest scaled machine learning importance (1.0), STMN3 balances moderate correlation (0.58) and strong algorithmic relevance (0.72), and CHST6 combines weak negative correlation (−0.33) with moderate machine learning importance (0.39), with consistent log2 fold change values (0.02–0.15) across all high-priority genes. Figure 7H (interaction network) visualizes 15 significant MGAT3 associations, with core nodes BSN, SEZ6L2, and ADCY1 forming strong positive connections (weights > 0.60) and CHST6 establishing a key weak negative interaction (weight = 0.50), aligning with their correlation and importance profiles; the complete list of the top 20 MGAT3-associated molecules is provided in Supplemental Table 2.
Identifying core MGAT3-associated biomarkers: coefficient and regularization path profiles
Figure 8A presents Lasso feature importance, with core biomarkers—consistently identified by all three methods—exhibiting distinct standardized coefficients: ATP6V1G2 (−2.051), SEZ6L2 (−1.342), SNAP25 (1.094), CHST6 (0.999), FBXW7 (0.692), ENO2 (0.498), CPLX1 (−0.474), ADCY1 (0.443), PCSK1 (−0.415), and CKMT1B (0.345), alongside method-specific features (e.g., SCN2B, SYT1) that had smaller absolute effects, with positive and negative effect directions clearly distinguished. Figure 8B shows ElasticNet feature importance, where core biomarkers share similar effect directions to Lasso but with marginally adjusted coefficients (e.g., ATP6V1G2: −1.831, SNAP25: 0.972, CHST6: 0.976) and method-specific features (e.g., SV2B, SLC12A5) exhibit modest absolute effects. Figure 8C displays AdaptiveLasso feature importance, with core biomarkers maintaining consistent effect directions (e.g., SNAP25: 1.393, ATP6V1G2: −2.016, CHST6: 1.002) and fewer method-specific features relative to the other two approaches. Figures 8D–F illustrate the regularization paths for Lasso, ElasticNet, and AdaptiveLasso, respectively, with dashed vertical lines marking the optimal λ.min; core biomarkers are represented by solid lines and method-specific features by dashed lines, demonstrating how coefficients shrink toward zero with increasing log(λ) and core biomarkers retaining non-zero coefficients across a broader range of λ values.

MGAT3-associated core biomarker identification, combination efficacy, and clinical utility in the training set
Figure 9A depicts the intersection of molecules selected by Lasso, ElasticNet, and AdaptiveLasso, with 10 core molecules (ADCY1, ATP6V1G2, CHST6, CKMT1B, CPLX1, ENO2, FBXW7, PCSK1, SEZ6L2, and SNAP25) identified as consistent candidates across all three methods. Figure 9B identifies ATP6V1G2 as the top-ranked biomarker (normalized mean importance = 1.000) with a disease-inhibitory effect, and CHST6 as the second-most critical (0.579). Figure 9C shows CHST6 and ATP6V1G2 form a key antagonistic pair (r = −0.63). Figure 9D confirms CHST6 (AUC = 0.723) and ATP6V1G2 (AUC = 0.720) as robust single-molecule diagnostic biomarkers with AUC values exceeding 0.7. Figure 9E demonstrates progressive diagnostic improvement with 1–5 molecule combinations: CHST6 alone achieves an AUC of 0.723, CHST6 + ATP6V1G2 increases to 0.753, and adding PCSK1 further elevates it to 0.755. Figure 9F validates the model's population-level utility at a clinically relevant 0.3 high-risk threshold, identifying ∼500 high-risk individuals per 1000 subjects with ∼85% corresponding to actual disease events. Figure 9G shows the model delivers maximum net benefit (0.20–0.25) across a 0.2–0.5 threshold range, outperforming both “treat all” and “treat none” strategies. Figure 9H presents a clinical nomogram where total scores (20–260 points) directly map to disease risk—100 points correspond to ∼0.3 risk, 180 points to ∼0.7 risk, and 220 points to ≥0.9 risk—with ATP6V1G2 contributing the highest point weights for precise risk stratification. Figure 9I confirms Disease groups exhibit significantly reduced expression of ATP6V1G2 alongside elevated CHST6 expression, compared to Control groups.

Diagnostic efficacy, clinical utility, and differential expression of MGAT3-associated biomarkers in the validation set
Figure 10A identifies ATP6V1G2 as the most influential biomarker with a normalized mean importance score of 1.000 across integrated models, followed by SNAP25 (0.518) and ENO2 (0.414). Figure 10B highlights ATP6V1G2's key molecular interactions, showing its strongest positive correlations with SNAP25 (r = 0.76) and ENO2 (r = 0.72). Figure 10C confirms ATP6V1G2 as the top-performing standalone diagnostic biomarker, achieving an AUC of 0.761 that outperforms all other individual molecules. Figure 10D demonstrates that the optimal three-marker combination—ATP6V1G2 + ENO2 + SEZ6L2—yields the highest diagnostic efficacy (AUC = 0.764), representing the peak performance without incremental gains from additional molecules. Figure 10E (Clinical Impact Curve) shows that at clinically relevant high-risk thresholds (0.2–0.6), the model identifies 500–600 high-risk individuals per 1000 subjects, with over 80% of these cases corresponding to confirmed disease. Figure 10F (Decision Curve Analysis) reveals the model's maximum net benefit (0.20–0.25) across a threshold range of 0.1–0.7, significantly outperforming both “treat all” and “treat none” strategies. Figure 10G (Nomogram) presents a clinical prediction tool where total scores directly map to disease risk, with ATP6V1G2 contributing the highest weight to risk stratification and a total score of ∼150 corresponding to a ∼0.5 disease risk. Figure 10H (Signature Boxplot) validates significant differential expression of ATP6V1G2, ENO2, and SEZ6L2 between disease and control groups, with ATP6V1G2 showing notably reduced expression in the disease cohort, further reinforcing their utility as a cohesive diagnostic signature.

Discussion
AD is a complex and progressive neurodegenerative disorder that significantly impacts cognitive function, leading to memory loss and a decline in daily living activities. 27 Given its rising prevalence, AD is considered a major public health challenge, imposing substantial economic burdens on healthcare systems globally, with annual costs reaching hundreds of billions of dollars. 28 The pathophysiology of AD involves a multifaceted interplay of genetic, environmental, and lifestyle factors, resulting in neuronal degeneration and synaptic loss. 29 Current diagnostic approaches primarily rely on clinical assessments and biomarkers such as amyloid-beta and tau proteins, which have shown limitations in early detection and specificity, underscoring the urgent need for innovative diagnostic tools and therapeutic strategies. 30
This study aims to explore the role of N-glycosylation, a crucial post-translational modification, in the context of AD pathogenesis. 31 Prior research has established connections between N-glycosylation alterations and neurodegenerative processes, including amyloid-beta aggregation and tau hyperphosphorylation, suggesting that dysregulated glycosylation may contribute to the progression of AD by modulating protein functions and neuronal signaling. 32 Our integrative approach utilizes advanced bioinformatics and machine learning techniques to identify N-glycosylation-associated biomarkers in AD, focusing on the enzyme MGAT3, which has emerged as a potential therapeutic target. 33 By mapping the research landscape, identifying co-expression networks, and validating core biomarkers for diagnostic utility, this study seeks to provide critical insights into the molecular mechanisms of AD and propose actionable strategies for early diagnosis and intervention.
The investigation into the role of N-glycosylation in AD has revealed significant insights into potential biomarkers and therapeutic targets. 34 Our findings highlight MGAT3 as a pivotal enzyme within the N-glycosylation pathway, suggesting its critical involvement in the disease's progression. 35 The identification of hub genes, such as ATP6V1G2 36 and CHST6, 37 further underscores the intricate network surrounding N-glycosylation modifications in neuronal function and signaling pathways. Notably, these genes exhibited robust correlations with MGAT3, suggesting that alterations in glycosylation patterns may influence AD pathology through mechanisms that warrant further exploration.
Additionally, our multi-omics integration approach has successfully identified a set of biomarkers with diagnostic potential, including ATP6V1G2 and CHST6. The application of machine learning algorithms, particularly penalized regression techniques, has enabled the selection of these core biomarkers, which demonstrate significant associations with AD risk. The ability to develop a multi-marker panel enhances the potential for early diagnosis and stratification of patient risk profiles, addressing a critical need in clinical practice. 38 The clinical utility of these biomarkers, validated through rigorous statistical analyses, supports their roles as promising candidates for future diagnostic assays and therapeutic interventions.
Furthermore, the validation of the diagnostic model utilizing these biomarkers demonstrates an improved area under the curve (AUC) in distinguishing AD patients from healthy controls. This model's performance indicates its potential applicability in clinical settings, allowing for better stratification of individuals at risk for AD based on their biomarker profiles. The findings not only advance our understanding of the molecular mechanisms underlying AD but also provide actionable insights for developing targeted therapies aimed at modifying disease progression through glycosylation pathways. As such, future research should focus on elucidating the functional roles of these identified biomarkers and their interactions within the broader context of neurodegenerative processes.
The limitations of this study primarily stem from the lack of wet-lab validation, which is crucial for confirming the biological relevance of the identified biomarkers. Experimental models, including MGAT3-knockout systems, are essential for elucidating the functional roles of the prioritized candidates in AD pathology. Additionally, while batch effects were addressed through comprehensive correction methods, residual heterogeneity may still exist across the utilized GEO datasets, potentially influencing the robustness of our findings. Furthermore, the clinical translation of our identified biomarkers necessitates prospective validation in diverse cohorts to ensure their applicability across different populations and settings. These limitations underscore the need for further investigations to establish a definitive link between our computational predictions and biological mechanisms.
In conclusion, this integrative study successfully maps the role of N-glycosylation in AD, highlighting ATP6V1G2 and CHST6 as promising biomarkers for diagnostic purposes. By employing a multi-disciplinary approach that combines bibliometric analysis, bioinformatics, and machine learning, we have identified core biomarkers and co-expression networks that may enhance early diagnosis and patient stratification. Our findings not only advance the understanding of N-glycosylation's involvement in AD but also pave the way for future research aimed at developing targeted therapies and clinical assays. Ultimately, this work contributes to the growing body of evidence that emphasizes the importance of novel biomarkers in addressing the pressing challenges posed by AD.
Conclusion
This integrative study combining bibliometrics and bioinformatics delineates the N-glycosylation-AD research landscape—featuring global contributions, evolving themes around MGAT3 and amyloid pathology—and identifies ATP6V1G2 and CHST6 as key MGAT3-associated biomarkers. Their antagonistic relationship, robust diagnostic potential, and validated clinical utility advance understanding of N-glycosylation-mediated AD mechanisms, providing promising tools for early diagnosis and risk stratification.
Supplemental Material
sj-xlsx-1-alr-10.1177_25424823251412909 - Supplemental material for N-Glycosylation and Alzheimer's disease: A 2001–2025 global bibliometric landscape revealing emerging diagnostic trends
Supplemental material, sj-xlsx-1-alr-10.1177_25424823251412909 for N-Glycosylation and Alzheimer's disease: A 2001–2025 global bibliometric landscape revealing emerging diagnostic trends by GuoXun Shi, Feng Ju, QiGang Dong and Ye Fang in Journal of Alzheimer's Disease Reports
Supplemental Material
sj-xlsx-2-alr-10.1177_25424823251412909 - Supplemental material for N-Glycosylation and Alzheimer's disease: A 2001–2025 global bibliometric landscape revealing emerging diagnostic trends
Supplemental material, sj-xlsx-2-alr-10.1177_25424823251412909 for N-Glycosylation and Alzheimer's disease: A 2001–2025 global bibliometric landscape revealing emerging diagnostic trends by GuoXun Shi, Feng Ju, QiGang Dong and Ye Fang in Journal of Alzheimer's Disease Reports
Supplemental Material
sj-pdf-3-alr-10.1177_25424823251412909 - Supplemental material for N-Glycosylation and Alzheimer's disease: A 2001–2025 global bibliometric landscape revealing emerging diagnostic trends
Supplemental material, sj-pdf-3-alr-10.1177_25424823251412909 for N-Glycosylation and Alzheimer's disease: A 2001–2025 global bibliometric landscape revealing emerging diagnostic trends by GuoXun Shi, Feng Ju, QiGang Dong and Ye Fang in Journal of Alzheimer's Disease Reports
Footnotes
Acknowledgements
We are deeply grateful to the authors of the 239 research articles and reviews incorporated into this bibliometric analysis.
Ethical considerations
This study uses de-identified public GEO data; original studies approved by respective IRBs. Analyses adhere to the Declaration of Helsinki and public dataset secondary use ethics. No additional ethical approval needed.
Consent to participate
Not applicable
Consent for publication
Not applicable
Author contribution(s)
Funding
This work was supported by four grants awarded to Hao Zhang: the Wuxi Science and Technology Development Fund (Award Number: Y20232005); the 2025 Open Project of the Provincial Key Laboratory for Integrative Chinese-Western Prevention and Treatment of Geriatric Diseases at Yangzhou University (Award Number: 202528); the 2025 Wuxi Aging Research Project (Award Number: WXLN25-A-24); and the 2025 Annual Scientific Research Project of the Jiangsu Provincial Association of Geriatrics (Award Number: JGS2025ZDM010).
Declaration of conflicting interests
The authors declared no potential conflicts of interest with respect to the research, authorship, and/or publication of this article.
Data availability statement
All supporting data are publicly available. Supplemental Table 1 lists retrieved bibliometric documents and detailed bibliometric results. Supplemental Table 2 includes DEGs, N-glycosylation-related genes, WGCNA-derived AD-associated genes, 336 overlapping candidates, and top 20 MGAT3-associated molecules. Supplemental Figures 1–5 present publication trends, collaborative networks, institutional/journal/author contributions, thematic evolution, and other key findings. Gene expression data (training set: GSE118553, GSE132903; validation set: GSE28146, GSE48350, GSE5281) are accessible via the GEO database.
Supplemental material
Supplemental material for this article is available online.
References
Supplementary Material
Please find the following supplemental material available below.
For Open Access articles published under a Creative Commons License, all supplemental material carries the same license as the article it is associated with.
For non-Open Access articles published, all supplemental material carries a non-exclusive license, and permission requests for re-use of supplemental material or any part of supplemental material shall be sent directly to the copyright owner as specified in the copyright notice associated with the article.
