PO.BCS01.06 · 生物信息与计算

Path2Prot提供了一种基于AI推断的蛋白质组学生物标志物进行乳腺肿瘤亚型分型和治疗反应预测的新方法

Path2Prot offers a new way for breast tumor subtyping and treatment response prediction from AI-inferred proteomic biomarkers

海报缩略图:Path2Prot提供了一种基于AI推断的蛋白质组学生物标志物进行乳腺肿瘤亚型分型和治疗反应预测的新方法
编号 87 展板 18 时间 4/19 02:00–05:00 区域 Section 4 主讲 Saugato Rahman Dhruba, PhD
分会场 Digital Pathology 1
查看 PDF 下载 PDF 🔒 查看 / 下载完整 PDF 需登录并开通下载套餐 · 查看套餐 / 开通 AACR 官方页面

作者与单位 Authors & Affiliations

Saugato Rahman Dhruba1, Danh-Tai Hoang1, Sumit Mukherjee1, Amos Stemmer1, Eldad Shulman1, Ranjan Kumar Barman2, Sanna Madan2, Sanju Sinha3, Kenneth D. Aldape4, Eytan Ruppin5

1National Cancer Institute - Cancer Data Science Laboratory (CDSL), Bethesda, MD,2National Cancer Institute - Cancer Data Science Laboratory (CDSL), Rockville, MD,3Sanford Bernham Prebys, La Jolla, CA,4Professor, Dept. of Pathology, Chair, NCI-CCR, Bethesda, MD,5Cedars-Sinai Medical Center, Los Angeles, CA

摘要 Abstract

中文摘要
背景:AI的出现正在变革精准医学,包括数字病理学,其中大型基础模型(FM)被应用于从肿瘤全切片组织病理学图像(WSI)中便捷地提取基因组学/转录组学模式。相比之下,尝试从WSI中的肿瘤形态学通过蛋白质组学获得直接功能性见解的研究较少,部分原因在于数据稀缺。因此,我们提出一种名为Path2Prot的弱监督深度学习模型,从肿瘤H&E切片图像推断乳腺癌(BC)中413个临床相关蛋白质组学生物标志物的相对丰度。我们通过利用所推断的蛋白质组学标志物,展示了此类模型在肿瘤亚型分型和治疗反应预测中的临床效用。 方法:Path2Prot由两个阶段组成:首先,通过标准流程将每张WSI预处理为20倍放大下的一组512 x 512瓦片图像,输入基于transformer的FM以提取形态学特征;其次,利用这些特征结合匹配的患者级蛋白质组学,训练一个多层感知机来推断蛋白质组学标志物水平。为进行训练,我们使用了来自841例TCGA-BRCA患者的2,074张WSI以及413个蛋白(总量+翻译后修饰)的匹配反相蛋白阵列(RPPA)数据。我们利用两种可用的WSI类型构建了三个不同的模型:FFPE模型,在893张福尔马林固定石蜡包埋WSI(用于诊断)上训练;FF模型,在1,181张新鲜冷冻WSI(RNA质量更佳)上训练;以及Combo模型,结合两个模型的预测。 结果:我们以推断与实测蛋白质组学之间的Pearson相关性(R)评估模型性能,其中R ≥ 0.4的蛋白称为良好预测蛋白(WPP)。Combo模型表现最佳,在交叉验证中WPP占23.7%(平均R = 0.31),并在CPTAC-BRCA的外部验证中成功泛化至跨平台质谱蛋白质组学,WPP占27.1%(平均R = 0.28;WPP重叠率 = 71.8%)。我们进一步将推断的HER2和ER水平二分类,以识别其免疫组化状态,并在TCGA-BRCA(n = 733)、CPTAC-BRCA(n = 89)、TransNEO(n = 160)和IMPRESS(n = 126)中将患者肿瘤分配到具有临床可操作性的亚型(HER2+、ER+和TNBC)。该任务能够较好地完成,得出HER2+ = 0.69-0.72、ER+ = 0.72-0.77和TNBC = 0.83-0.88的曲线下面积(AUC)值。最后,推断的蛋白靶标成功估计了TransNEO(n = 60,AUC = 0.71)和CSHS-BRCA(n = 20,AUC = 1.0)队列中的抗HER2反应,以及CSHS-BRCA(n = 16,AUC = 0.78)中的抗PD1反应。 结论:我们的分析揭示了乳腺癌中一个具有重要临床意义的蛋白亚组可从常规WSI中稳健预测以用于临床应用。随着更大规模蛋白质组学数据集的出现,可望在这些结果基础上取得显著改进。
查看英文原文 English abstract
Background: The advent of AI is revolutionizing precision medicine, including digital pathology where large foundation models (FMs) are applied to readily extract genomic/transcriptomic patterns from tumor whole-slide histopathology images (WSIs). In contrast, fewer studies have attempted to derive direct functional insights via proteomics from tumor morphology in WSIs, partly due to data scarcity. Henceforth, we propose a weakly-supervised deep learning model called Path2Prot to infer the relative abundance of 413 clinically relevant proteomic biomarkers in breast cancer (BC) from tumor H&E slide images. We show the clinical utility of such models in tumor subtyping and treatment response prediction by leveraging the inferred proteomic markers. Methods: Path2Prot is composed of two stages: First , each WSI is preprocessed via a standard pipeline to a set of 512 x 512 tile images at 20x magnification, which are fed to a transformer-based FM to extract morphological features; Next , using these features with matched patient-level proteomics, a multilayer perceptron is trained to infer the proteomic marker levels. To train, we used 2,074 WSIs from 841 TCGA-BRCA patients and the matched reverse-phase protein array (RPPA) data for 413 proteins (total + post-translationally modified). We leveraged both WSI types available via building three distinct models: FFPE model , trained on 893 formalin-fixed paraffin embedded WSIs (used for diagnosis); FF model , trained on 1,181 fresh-frozen WSIs (better RNA quality); and Combo model , combining the predictions of both models. Results: We assessed model performance with Pearson correlation ( R ) between inferred and measured proteomics, where proteins with R ≥ 0.4 are referred as the well-predicted proteins (WPPs). The Combo model performed the best with 23.7% WPPs (mean R = 0.31) in cross-validation and successfully generalized to cross-platform mass spectrometry proteomics in external validation with CPTAC-BRCA with 27.1% WPPs (mean R = 0.28; Overlap-in-WPPs = 71.8%). We further dichotomized the inferred HER2 and ER levels to identify their immunohistochemistry status and assigned patient tumors to clinically actionable subtypes (HER2+, ER+ & TNBC) across TCGA-BRCA ( n = 733), CPTAC-BRCA ( n = 89), TransNEO ( n = 160) and IMPRESS ( n = 126). This task can be done fairly well, yielding area under the curve (AUC) values of HER2+ = 0.69-0.72, ER+ = 0.72-0.77 and TNBC = 0.83-0.88. Finally, the inferred protein targets successfully estimated anti-HER2 response in TransNEO ( n = 60, AUC = 0.71) and CSHS-BRCA ( n = 20, AUC = 1.0) cohorts, and anti-PD1 response in CSHS-BRCA ( n = 16, AUC = 0.78). Conclusion: Our analysis reveals a clinically important subset of proteins in breast cancer can be robustly predicted from routine WSIs for clinical application. One may expect to significantly improve upon these results with the advent of larger proteomics datasets.
利益披露 Disclosure
S. Dhruba, None.. D. Hoang, None.. S. Mukherjee, None.. A. Stemmer, None.. E. Shulman, None.. R. K. Barman, None.. S. Madan, None.. S. Sinha, None.. K. D. Aldape, None. E. Ruppin, Medaware Ltd. Other, Eytan Ruppin is a cofounder of Medaware Ltd.. Metabomed Other, Eytan Ruppin is a cofounder of Metabomed. Pangea Biomed Other, Eytan Ruppin is a cofounder (divested) and non-paid scientific consultant of Pangea Biomed. GSK Other, Eytan Ruppin is a scientific advisory board member of GSK Oncology. WIN Consortium Other, Eytan Ruppin is a scientific advisory board member of WIN consortium. ProCan Program Other, Eytan Ruppin is a scientific advisory board member of ProCan program, Other, Eytan Ruppin is a scientific advisory board member of ProCan program.

← 返回 AACR 2026 检索