PO.BCS01.07 · 生物信息与计算

从非小细胞肺癌的H&E全切片图像预测基因表达和分子通路活性

Prediction of gene expression and molecular pathway activity from H&E whole slide images in non-small cell lung cancer

海报缩略图:从非小细胞肺癌的H&E全切片图像预测基因表达和分子通路活性
编号 1457 展板 20 时间 4/20 09:00–12:00 区域 Section 4 主讲 Mina Khoshdeli
分会场 Digital Pathology 2
查看 PDF 下载 PDF 🔒 查看 / 下载完整 PDF 需登录并开通下载套餐 · 查看套餐 / 开通 AACR 官方页面

作者与单位 Authors & Affiliations

Mina Khoshdeli1, Muhammad Sohaib2, Mohammed Qutaish1, Mahesh Bachu1, Matthew Loya1, Prianka Chohan1, Omar Jabado1, Craig Thalhauser1, Mirna Lechpammer1, David Soong1

1Genmab, Princeton, NJ,2Electrical and Biomedical Engineering, University of Nevada, Reno, Reno, NV

摘要 Abstract

中文摘要
从苏木精和伊红(H&E)全切片图像(WSI)预测转录组谱仍然是一项具有挑战性但极具价值的任务。这主要由三个因素驱动:(1) H&E WSI在患者群体中广泛可得;(2) 它们比RNA测序更具成本效益;(3) 从H&E图像推断此类信息有助于为更关键的诊断和预后检测保留有限的组织。既往研究探索了多种模型来解决这一问题,并展示了预测关键癌症过程中基因的潜力。在本研究中,我们开发了一个两阶段框架,采用最先进的基础模型进行患者级特征提取,并利用稳健的机器学习方法进行高效的模型训练。我们在一个精心整理的内部数据集上对这些模型进行了全面评估,该数据集包含67个具有匹配RNA-seq的商业非小细胞肺癌(NSCLC)患者样本。 Gigapath是一个在超过170,000张WSI上预训练的大规模病理学基础模型,用于从H&E染色的NSCLC中提取切块级和患者级视觉嵌入。每张切片被分割为256×256像素的切块,Gigapath嵌入通过注意力机制(LongNet)聚合为切片级表示。我们评估了多个回归模型,以从这些切片级嵌入预测数值型基因表达和通路水平。真实的通路活性通过对RNA-seq数据进行单样本基因集富集分析计算50个Hallmark基因集的归一化富集分数得出。模型在来自癌症基因组图谱肺腺癌(TCGA-LUAD)队列的425张切片上训练,并在67张切片的独立数据集上评估。使用Spearman相关性评估模型性能。 Gigapath-随机森林回归模型在多个Hallmark通路中展现出最强的预测性能,例如未折叠蛋白反应(ρ = 0.70)、MTORC1信号通路(ρ = 0.69)和上皮-间质转化(EMT)(ρ = 0.67)。这些通路不仅对NSCLC生物学至关重要,而且表现出独特的组织学特征,如细胞质应激、纤维化重塑和细胞核形态改变。该框架进一步扩展,使用相同的切片级特征表示预测单个基因的表达水平。初步结果表明,在17,719个表达基因中,有2,223个可以以大于0.4的Spearman相关性进行预测。这些发现表明,病理学基础模型可以高效地与回归模型整合,直接从组织学成像特征中捕捉多样的转录组活性,为生物标志物发现和个性化癌症疗法的开发提供了一个可泛化的框架。
查看英文原文 English abstract
Predicting transcriptomic profiles from hematoxylin & eosin (H&E) whole slide images (WSIs) remains a challenging but highly desirable task. This is driven by three main factors: (1) H&E WSIs are widely available across patient populations; (2) they are significantly more cost-effective than RNA sequencing; and (3) inferring such information from H&E images can help preserve limited tissue for more critical diagnostic and prognostic tests. Prior studies have explored a variety of models to address this problem and demonstrated potential in predicting genes in key cancer processes. In this study, we developed a two-stage framework that employed state-of-the-art foundation models for patient-level feature extraction, and leveraged robust machine learning methods for efficient model training. Comprehensive evaluation of these models was performed on a carefully curated internal dataset of 67 commercial non-small cell lung cancer (NSCLC) patient samples with matched RNA-seq. Gigapath, a large-scale pathology foundation model pre-trained on over 170,000 WSIs, was used to extract patch and patient level visual embeddings from H&E-stained NSCLC. Each slide was divided into 256×256 pixel patches, with Gigapath embeddings aggregated via an attention mechanism (LongNet) into slide-level representations. Several regression models were evaluated to predict numerical gene expression and pathway levels from these slide-level embeddings. Ground truth pathway activity was computed using normalized enrichment scores for 50 Hallmark gene sets by single-sample gene set enrichment analysis on the RNA-seq data. Models were trained on 425 slides from The Cancer Genome Atlas Lung -Adenocarcinoma (TCGA-LUAD) cohort and evaluated on the independent dataset of 67 slides. Model performance was assessed using Spearman correlation. The Gigapath-Random Forest regressor model demonstrated strongest predictive performance across several Hallmark pathways, such as Unfolded Protein Response (ρ = 0.70), MTORC1 Signaling (ρ = 0.69), and Epithelial-Mesenchymal Transition (EMT) (ρ = 0.67). These pathways are not only critical to NSCLC biology but also exhibit distinct histological signatures such as cytoplasmic stress, fibrotic remodeling, and nuclear morphological alterations. The framework was further extended to predict expression levels of individual genes using the same slide-level feature representations. Preliminary results indicate that out of 17,719 expressed genes, 2,223 can be predicted with a Spearman correlation greater than 0.4. These findings demonstrated that pathology foundation models can be efficiently integrated with regression models to capture diverse transcriptomic activities directly from histological imaging features, providing a generalizable framework for biomarker discovery and the development of personalized cancer therapies.
利益披露 Disclosure
M. Khoshdeli, Genmab Employment, Stock. M. Sohaib, None. M. Qutaish, Genmab Employment, Stock. M. Bachu, Genmab Employment, Stock. M. Loya, Genmab Employment, Stock. P. Chohan, Genmab Employment, Stock. O. Jabado, Genmab Employment, Stock. C. Thalhauser, Genmab Employment, Stock. M. Lechpammer, Genmab Employment, Stock. D. Soong, Genmab Employment, Stock.

← 返回 AACR 2026 检索