PO.BCS01.02 · 生物信息与计算
通过元聚类方法探索胰腺导管腺癌的高分辨率亚型
Exploring high-resolution subtypes for pancreatic ductal adenocarcinoma via a meta-clustering approach
作者与单位 Authors & Affiliations
摘要 Abstract
中文摘要
作为美国癌症死亡的第三大原因,胰腺癌(PC)是一种5年生存率极低的恶性肿瘤。由于缺乏可靠的临床指标(如生物标志物),PC的早期诊断和治疗仍具挑战性。胰腺导管腺癌(PDAC)是最常见的PC亚型,占病例的90%以上。鉴定PDAC亚型对于下游风险分层和定制化治疗设计至关重要。PDAC传统上被分为四种分子亚型,即异常分化的内分泌外分泌型(ADEX)、免疫原性型、鳞状型和胰腺祖细胞型。然而,众多研究表明这4种分子亚型内部存在高度异质性,表明亟需探索PDAC的高分辨率亚型。分子分析和组织病理学研究等传统湿实验室技术耗时、昂贵且费力。为填补这些空白,我们开发了一种元聚类方法,利用转录组学数据探索高分辨率的PDAC亚型。具体而言,我们首先利用随机投影(RP)降低PDAC RNA-seq数据的维度,然后使用我们的基础聚类方法通过Leiden算法进行聚类。为获得稳健的结果,我们实施了15次基于RP的Leiden聚类。随后,我们采用源自我们此前发表的名为SHARP的单细胞分析方法的加权元聚类(wMetaC)架构,对这15个聚类结果进行元聚类。结果表明,我们提出的方法在PDAC亚型分型方面显著优于最先进的聚类方法。此外,我们的元聚类方法的表现大幅优于所有单个基础聚类方法。我们进一步进行了簇特异性差异基因表达分析、通路分析及基因-药物-疾病关联研究,表明我们新鉴定的PDAC亚簇在不同簇之间具有独特性。我们期望我们提出的方法将为高分辨率PC亚型表征提供一个稳健的框架,以实现准确的下游风险评估和个体化治疗设计。
查看英文原文 English abstract
As the third leading cause of cancer death in the United States, pancreatic cancer (PC) is a malignancy with a very low 5-year survival rate. Early diagnosis and treatment of PC remain challenging due to the lack of reliable clinical indicators such as biomarkers. Pancreatic ductal adenocarcinoma (PDAC) is the most common PC subtype, accounting for over 90% of cases. Identifying PDAC subtypes is essential for downstream risk stratification and tailored treatment design. PDAC is conventionally categorized into four molecular subtypes, i.e., aberrantly differentiated endocrine exocrine (ADEX), immunogenic, squamous, and pancreatic progenitor. However, numerous studies have demonstrated high heterogeneity within these 4 molecular subtypes, indicating that exploring high-resolution subtypes for PDAC is highly needed. Conventional wet-lab techniques like molecular profiling and histopathological studies are time-consuming, costly, and laborious. To fill these gaps, we developed a meta-clustering approach to leverage transcriptomics data to explore high-resolution PDAC subtypes. Specifically, we first leveraged random projection (RP) to reduce the dimensions of the PDAC RNA-seq data, which were then clustered by our base clustering method using the Leiden algorithm. To obtain robust results, we implemented 15 runs of RP-based Leiden clustering. Then, we performed meta-clustering of these 15 clustering results by adopting a weighted meta-clustering (wMetaC) architecture from our previous published single-cell analysis method named SHARP. Results suggested that our proposed approach significantly outperformed state-of-the-art clustering methods for PDAC subtyping. Moreover, our meta-clustering approach performed substantially better than all individual base clustering methods. We further performed cluster-specific differential gene expression analysis, pathway analysis and gene-drug-disease association studies, suggesting that our newly identified PDAC sub-clusters were distinctive among different clusters. We expect our proposed approach will provide a robust framework for high-resolution PC subtype characterization for accurate downstream risk assessment and personalized treatment design.
利益披露 Disclosure
N. B. Peterson, None..
S. Wan, None.