Abstract
To overcome the issues of inadequate multi-scale feature representation and limited model interpretability in equipment fault diagnosis under complex operating conditions, this paper proposes a large language model enhanced dual-domain transformer for intelligent fault diagnosis, which integrates time–frequency features with text semantic enhancement. Initially, the raw vibration signals are transformed using the short-time Fourier transform, converting one-dimensional time-domain signals into two-dimensional time–frequency representations. Time-domain features are extracted using a Time Series Transformer, while frequency-domain features are obtained through a Vision Transformer. A dynamic adaptive weight learning-based feature fusion and alignment mechanism is then introduced to facilitate deep cross-domain modeling and enhance information complementarity between heterogeneous time–frequency features. In addition, domain knowledge pertinent to fault diagnosis is incorporated into a large language model through low-rank adaptation fine-tuning, thereby establishing a multimodal diagnostic framework with semantic reasoning capabilities. This approach effectively mitigates the pronounced “black-box” nature and limited interpretability associated with traditional deep learning-based methods. Experimental results indicate that the proposed method achieves diagnostic accuracies of 98.7 and 98.4% on the Case Western Reserve University and Northeast Forestry University datasets, respectively. Ablation studies further confirm the efficacy of each key module in enhancing diagnostic performance.
Keywords
Get full access to this article
View all access options for this article.
