PO.BCS01.10 · 生物信息与计算

用于预测前列腺癌具有临床意义表型的多模态深度学习框架

Multimodal deep learning framework for predicting clinically significant phenotypes in prostate cancer

海报缩略图:用于预测前列腺癌具有临床意义表型的多模态深度学习框架
编号 4184 展板 11 时间 4/21 09:00–12:00 区域 Section 4 主讲 Priyanka Vasanthakumari, PhD
分会场 Integrative Computational Approaches 2
查看 PDF 下载 PDF 🔒 查看 / 下载完整 PDF 需登录并开通下载套餐 · 查看套餐 / 开通 AACR 官方页面

作者与单位 Authors & Affiliations

Priyanka Vasanthakumari1, Mohamed Omar2

1Computational Biomedicine, Cedars-Sinai Medical Center, Los Angeles, CA,2Cedars-Sinai Medical Center, West Hollywood, CA

摘要 Abstract

中文摘要
前列腺癌是最常见的非皮肤恶性肿瘤,也是美国男性中癌症相关死亡的第二大原因。其临床与生物学异质性给准确的结局预测与治疗规划带来了重大挑战。传统的单模态模型往往无法准确捕获疾病进展的风险。人工智能的近期进展,尤其是基础模型,能够将异质数据整合为肿瘤生物学的统一表征,进而揭示出比单模态方式所能实现的更好的预测与预后特征。我们开发了一个多模态模型,整合组织病理学全切片图像(WSI)、转录组学与临床数据,以预测前列腺癌的分子与临床表型,同时提升预测准确性与可解释性。该框架为每种模态采用独立的编码器:一个改编自基础模型CONCHv1.5与TITAN的图像编码器,一个基于ClinicalBERT的临床数据文本编码器,以及一个将Gene2Vec嵌入与投影到共享潜在空间的表达水平特征相结合的转录组编码器。这些编码器通过成对对齐进行对比微调,基于配对数据的可得性,以WSI作为锚定模态。多模态表征在若干具有临床意义的下游预测任务上进行评估,包括Gleason分级、转移预测、生化复发,以及诸如TMPRSS2:ERG融合与PTEN缺失等分子改变状态。该框架还通过量化每种模态的贡献并可视化注意力分布来提供可解释性,凸显出对每项任务重要的联合特征表征。本研究使用来自632例患者的数据,配对数据来自公共与机构存储库。WSI与临床文本数据之间的跨模态微调产生了有前景的对齐结果,配对的图像-文本样本相比未配对样本获得了显著更高的相似性排名(平均p值≈10⁻²⁹)。Gleason分级多类别分类下游任务实现了81.4%的总体测试准确率与79.6%的加权F1分数。
查看英文原文 English abstract
Prostate cancer is the most common non-cutaneous malignancy and the second leading cause of cancer-related death among men in the United States. Its clinical and biological heterogeneity poses major challenges for accurate outcome prediction and treatment planning. Traditional single-modality models often fail to accurately capture the risk of disease progression. Recent advances in artificial intelligence, particularly foundation models, enable the integration of heterogeneous data into unified representations of tumor biology, which can in turn uncover better predictive and prognostic features than what is possible with unimodal modalities.We develop a multimodal model that integrates histopathology whole-slide images (WSIs), transcriptomics, and clinical data to predict prostate cancer molecular and clinical phenotypes, enhancing both predictive accuracy and interpretability. The framework employs separate encoders for each modality: an image encoder adapted from foundation models CONCHv1.5 and TITAN, a text encoder based on ClinicalBERT for clinical data, and a transcriptomic encoder combining Gene2Vec embeddings with expression-level features projected into a shared latent space. The encoders are contrastively fine-tuned using pairwise alignment, with WSIs serving as the anchor modality based on paired data availability.The multimodal representations are evaluated on several downstream clinically significant prediction tasks, including Gleason grading, metastasis prediction, biochemical recurrence, and molecular alteration status such as TMPRSS2:ERG fusion and PTEN loss. The framework also provides interpretability by quantifying the contribution of each modality and visualizing attention distributions, highlighting joint feature representations that are important to each task.This study uses data from 632 patients with paired data from public and institutional repositories. Inter-modality fine-tuning between WSI and clinical text data yielded promising alignment results, with paired image-text samples achieving significantly higher similarity ranks than unpaired samples (mean p-value ≈ 10⁻²⁹). Gleason grade multiclass classification downstream task achieved an overall test accuracy of 81.4% and a weighted F1 score of 79.6%.
利益披露 Disclosure
P. Vasanthakumari, None.

← 返回 AACR 2026 检索