PO.BCS01.07 · 生物信息与计算
整合全玻片图像与临床特征以实现更优的前列腺癌风险分层
Whole-slide image and clinical feature integration for superior prostate cancer risk stratification
作者与单位 Authors & Affiliations
摘要 Abstract
中文摘要
背景:前列腺癌仍是男性最常见的恶性肿瘤之一,其临床结局具有显著异质性。早期准确的诊断和风险分层对于有效治疗和改善结局至关重要。尽管Gleason评分是一个成熟的风险分层工具,但它并不能完全解释结局的变异性。我们提出了一个多模态深度学习框架,将基于注意力的全玻片图像 (WSI) 特征与一个临床变量(年龄)整合,以预测前列腺腺癌患者的生存。
设计:从TCGA-PRAD获取了500名前列腺腺癌患者的临床数据和WSI。临床数据被解析出复发或进展 (RoP) 事件和年龄。RoP用作结局标签,年龄用作患者特征。使用QuPath (Bankhead Sci Rep 2017) 对诊断性WSI肿瘤区域进行注释并切片为224×224像素。使用UNI2-h (Mahmood Nat Med 2025) 提取切片嵌入。开发了一个端到端的机器学习流程,使用聚类约束注意力多示例学习 (CLAM) 将切片嵌入聚合为玻片层面的嵌入并预测RoP (Lu Nat Biomed 2021)。第二个流程使用经Gleason评分训练的二元交叉熵 (BCE) 模型来预测相同患者的RoP。评估了三个模型:仅WSI的CLAM模型、WSI+年龄的CLAM模型和仅Gleason的BCE模型。所有模型均使用5折分层交叉验证进行训练和评估。每一折以独立的随机初始化重复3次。每个模型在所有折和运行中对AUROC取平均。同时计算了标准差。使用Friedman检验比较模型性能,并通过Wilcoxon符号秩检验进行事后两两比较。
结果:Gleason BCE模型表现欠佳 (AUROC 0.67 ± 0.07)。两个基于WSI的模型均优于Gleason BCE模型,达到AUROC 0.76 ± 0.02(无年龄)和AUROC 0.76 ± 0.01(有年龄)。经Friedman检验,各AUROC存在显著差异 (p = 0.0224),在WSI与基于Gleason方法之间的事后比较中显示出趋向显著的趋势 (p = 0.06)。
结论:使用WSI衍生特征的AI模型,无论是否结合基本临床背景,在前列腺癌复发风险预测方面均优于传统的Gleason评分。这些发现支持将数字病理学和AI整合到前列腺癌的常规预后评估中。未来的工作将侧重于纳入更多患者、临床变量和多类别预后标签,以将预后预测扩展到RoP之外。
查看英文原文 English abstract
Background: Prostate cancer remains one of the most common malignancies in men, with significant heterogeneity in clinical outcomes. Early and accurate diagnosis and risk stratification is crucial for effective treatment and improved outcomes. Although Gleason score is an established risk stratification tool, it does not fully explain outcome variability. We present a multimodal deep-learning framework that integrates attention-based whole-slide image (WSI) features with a clinical variable (Age) to predict survival in patients with prostate adenocarcinoma.
Design: Clinical data and WSI for 500 patients with prostate adenocarcinoma were obtained from TCGA-PRAD. Clinical data were parsed for recurrence or progression (RoP) events and age. RoP was used as the outcome label, and age was used as a patient feature. Diagnostic WSI tumor regions were annotated and tiled at 224×224 pixels using QuPath (Bankhead Sci Rep 2017). Tile embeddings were extracted with UNI2-h (Mahmood Nat Med 2025). An end-to-end machine learning pipeline was developed to aggregate tile embeddings into slide-level embeddings using Clustering-Constrained Attention Multiple Instance Learning (CLAM) and predict RoP (Lu Nat Biomed 2021). A second pipeline used a Gleason score-trained binary cross entropy (BCE) model to predict RoP in the same patients. Three models were evaluated: WSI-only CLAM model, WSI+age CLAM model, and Gleason-only BCE model. All models were trained and evaluated using 5-fold stratified cross-validation. Each fold was repeated 3 times with independent random initialization. AUROC was averaged across all folds and runs per model. Standard deviation was also computed. Model performance was compared using Friedman's test, with post hoc pairwise comparisons by Wilcoxon signed-rank test.
Results: The Gleason BCE model underperformed (AUROC 0.67 ± 0.07). Both WSI-based models outperformed the Gleason BCE model, achieving AUROC 0.76 ± 0.02 (without age), and AUROC 0.76 ± 0.01 (with age). AUROCs were significantly different by Friedman's test (p = 0.0224) and showed a trend toward significance in post-hoc comparison between WSI and Gleason-based approaches (p = 0.06).
Conclusion: AI models using WSI-derived features, with or without basic clinical context, outperformed traditional Gleason score for recurrence risk prediction in prostate cancer. These findings support the integration of digital pathology and AI into routine prognostic assessment for prostate cancer. Future work will focus on including additional patients, clinical variables, and multi-class prognostic labels to expand prognostic prediction beyond RoP.
利益披露 Disclosure
J. Johnson, None..
K. Ebare, None.