
Announcement
Select search scope: search across all journals or within the current journal

High-throughput chromosome conformation capture (Hi-C) technology captures spatial interactions of DNA sequences into matrices, and software tools are developed to identify topologically associating domains (TADs) from the Hi-C matrices. With structural information theory, SuperTAD adopted a dynamic programming approach to find the TAD hierarchy with minimal structural entropy. However, the algorithm suffers from high time complexity. To accelerate this algorithm, we design and implement an approximation algorithm with a theoretical performance guarantee. We implemented a package, SuperTAD-Fast. Using Hi-C matrices and simulated data, we demonstrated that SuperTAD-Fast achieved great runtime improvement compared with SuperTAD. SuperTAD-Fast shows high consistency and significant enrichment of structural proteins from Hi-C data of human cell lines in comparison with the existing six hierarchical TADs detecting methods.
The physiological activities within cells are mainly regulated through protein–protein interactions (PPI). Therefore, studying protein interactions has become an essential part of researching protein function and mechanisms. Traditional biological experiments required for PPI prediction are expensive and time consuming. For this reason, many methods based on predicting PPI from protein sequences have been proposed in recent years. However, existing computational methods usually require the combination of evolutionary feature information of proteins to predict PPI docking situations. Because different relevant features of selected proteins are chosen, there may be differences in the predicted results for PPI. This article proposes a PPI prediction method based on the pretrained protein sequence model ProtBert, combined with the Bidirectional Gated Recurrent Unit (BiGRU) and attention mechanism. Only using protein sequence information and leveraging ProtBert’s powerful ability to capture amino acid feature information, BiGRU is used for further feature extraction of the amino acid vectors output by ProtBert. The attention mechanism is then applied to enhance the focus on different amino acid features and improve the expression ability of protein sequence features, ultimately obtaining binary classification results for protein interactions. Experimental results show that our proposed ProtBert-BiGRU-Attention model has good predictive performance for PPI. Through relevant comparative experiments, it has been proven that our model performs well in protein binary prediction. Furthermore, through the ablation experiment of the model, different deep learning modules’ contributions to the prediction have been demonstrated.
Gene duplication has a central role in evolution; still, little is known on the fates of the duplicated copies, their relative frequency, and on how environmental conditions affect them. Moreover, the lack of rigorous definitions concerning the fate of duplicated genes hinders the development of a global vision of this process. In this paper, we present a new framework aiming at characterizing and formally differentiating the fate of duplicated genes. Our framework has been tested via simulations, where the evolution of populations has been simulated using aevol, an
Understanding the genetic regulation, for example, gene expressions (GEs) by copy number variations and methylations, is crucial to uncover the development and progression of complex diseases. Advancing from early studies that are mostly focused on homogeneous groups of patients, some recent studies have shifted their focus toward different patient groups, explored their commonalities and differences, and led to insightful findings. However, the analysis can be very challenging with one GE possibly regulated by multiple regulators and one regulator potentially regulating the expressions of multiple genes, leading to two distinct types of commonalities/differences in the patterns of genetic regulation. In addition, the high dimensionality of both sides of regulation poses challenges to computation. In this study, we develop a two-way fusion integrative analysis approach, which innovatively applies two fusion penalties to simultaneously identify commonalities/differences in the regulated pattern of GEs and regulating pattern of regulators, and adopt a Huber loss function to accommodate the possible data contamination. Moreover, a simple yet efficient iterative optimization algorithm is developed, which does not need to introduce any auxiliary variables and extra tuning parameters and is guaranteed to converge to a globally optimal solution. The advantages of the proposed approach are demonstrated in extensive simulations. The analysis of The Cancer Genome Atlas data on melanoma and lung cancer leads to interesting findings and satisfactory prediction performance.
Recent technological advancements have enabled spatially resolved transcriptomic profiling but at a multicellular resolution that is more cost-effective. The task of cell type deconvolution has been introduced to disentangle discrete cell types from such multicellular spots. However, existing benchmark datasets for cell type deconvolution are either generated from simulation or limited in scale, predominantly encompassing data on mice and are not designed for human immuno-oncology. To overcome these limitations and promote comprehensive investigation of cell type deconvolution for human immuno-oncology, we introduce a large-scale spatial transcriptomic deconvolution benchmark dataset named S
Small molecules (SMs) play a pivotal role in regulating microRNAs (miRNAs). Existing prediction methods for associations between SM–miRNA have overlooked crucial aspects: the incorporation of local topological features between nodes, which represent either SMs or miRNAs, and the effective fusion of node features with topological features. This study introduces a novel approach, termed high-order topological features for SM–miRNA association prediction (HTFSMMA), which specifically addresses these limitations. Initially, an association graph is formed by integrating SM–miRNA association data, SM similarity, and miRNA similarity. Subsequently, we focus on the local information of links and propose target neighborhood graph convolutional network for extracting local topological features. Then, HTFSMMA employs graph attention networks to amalgamate these local features, thereby establishing a platform for the acquisition of high-order features through random walks. Finally, the extracted features are integrated into the multilayer perceptron to derive the association prediction scores. To demonstrate the performance of HTFSMMA, we conducted comprehensive evaluations including five-fold cross-validation, leave-one-out cross-validation (LOOCV), SM-fixed local LOOCV, and miRNA-fixed local LOOCV. The area under receiver operating characteristic curve values were 0.9958 ± 0.0024 (0.8722 ± 0.0021), 0.9986 (0.9504), 0.9974 (0.9111), and 0.9977 (0.9074), respectively. Our findings demonstrate the superior performance of HTFSMMA over existing approaches. In addition, three case studies and the DeLong test have confirmed the effectiveness of the proposed method. These results collectively underscore the significance of HTFSMMA in facilitating the inference of associations between SMs and miRNAs.