PO.BCS02.06 · 生物信息与计算

通过多模态协同训练与可扩展的多组学编码器进行组学感知的图块聚合,以跨肿瘤生物标志物面板实现切片级预测

Omics-aware patch aggregation via multimodal co-training with a scalable multi-omics encoder for slide-level prediction across an oncology biomarker panel

编号 4207 展板 3 时间 4/21 09:00–12:00 区域 Section 5 主讲 hwanil choi
分会场 Machine Learning Approaches for Cancer Prediction
该海报暂无可下载的资料 AACR 官方页面

作者与单位 Authors & Affiliations

Hwanil Choi1, Tae Hyun Hwang2, Soonyoung Lee1, Jongseong Jang1

1Bio Intelligence Lab, LG AI Research, Seoul, Korea, Republic of,2Department of Surgery, Vanderbilt University Medical Center, Nashville, TN

摘要 Abstract

中文摘要
背景:全切片图像(WSI)广泛可用,但匹配的多组学谱有限,尤其是在需要多种模态时。我们开发了一个多模态学习框架,将RNA表达和DNA突变与H&E WSI整合,显式建模部分观测的组学数据,以训练一个组学感知的图块聚合器用于切片级预测。 方法:使用切片级对比目标训练组学感知图块聚合器,该目标对齐多组学编码器(MOE)和切片编码器(SE)。MOE是一个共享的Transformer,它对每种组学模态进行标记化,将模态特异性标记与组学类型编码连接,并使用自注意力捕获跨组学相互作用,从而允许在不改变架构的情况下添加新的组学。SE处理每张切片数万个图块,并通过相对位置偏置纳入图块坐标,以强调空间上邻近的区域。多模态预训练使用了来自癌症基因组图谱(The Cancer Genome Atlas)和基因型-组织表达(Genotype-Tissue Expression)项目的约20,000张部分配对的WSI和多组学谱。 结果:经过多模态预训练后,组学感知聚合器支持跨肿瘤生物标志物面板的切片级预测。对于基因过表达,ROC曲线下面积(AUC)值为:LAG3 0.84、CLDN6 0.68、CD274 0.98、EGFR 0.74、ERBB2 0.72、ERBB3 0.69、CD276 0.82、VTCN1 0.72、TACSTD2 0.77、FOLR1 0.93和MET 0.82。按肿瘤类型和任务划分,肺腺癌的AUC分别为肿瘤突变负荷0.70、EGFR突变0.87、KRAS突变0.62;结直肠癌的微卫星不稳定性为0.99;乳腺癌的ER、PR和HER2蛋白亚型分型分别为0.95、0.88和0.81,TP53和PIK3CA突变分别为0.74和0.86;肾细胞癌的PBRM1和BAP1突变分别为0.60和0.74;结肠腺癌的KRAS和TP53分别为0.88和0.89。基因过表达任务使用癌症基因组图谱,采用五折交叉验证并跨四个种子;肺和结直肠模型使用三星医疗中心队列;乳腺亚型分型使用BCNB队列;肾和结肠任务使用临床蛋白质组学肿瘤分析联盟(Clinical Proteomic Tumor Analysis Consortium)数据。 结论:一个与可扩展多组学编码器和WSI协同训练的组学感知图块聚合框架,能够对多种生物标志物和肿瘤类型实现准确的切片级预测,并说明了部分配对的多组学数据如何能够增强数字病理学模型。AI仅用于语言编辑;作者对所有内容负责并批准了最终版本。
查看英文原文 English abstract
Background: Whole-slide images (WSIs) are widely available, but matched multi-omics profiles are limited, especially when multiple modalities are required. We developed a multimodal learning framework that integrates RNA expression and DNA mutation with H&E WSIs, explicitly modeling partially observed omics to train an omics-aware patch aggregator for slide-level prediction. Methods: An omics-aware patch aggregator was trained with a slide-level contrastive objective that aligns a Multi-Omics Encoder (MOE) and a Slide Encoder (SE). The MOE is a shared Transformer that tokenizes each omics modality, concatenates modality-specific tokens with omic-type encodings, and uses self-attention to capture cross-omic interactions, allowing new omics to be added without changing the architecture. The SE processes tens of thousands of patches per slide and incorporates patch coordinates via relative positional bias to emphasize spatially proximal regions. Multimodal pretraining used ~20,000 partially paired WSIs and multi-omics profiles from The Cancer Genome Atlas and Genotype-Tissue Expression projects. Results: After multimodal pretraining, the omics-aware aggregator supported slide-level prediction across an oncology biomarker panel. For gene overexpression, area under the ROC curve (AUC) values were: LAG3 0.84, CLDN6 0.68, CD274 0.98, EGFR 0.74, ERBB2 0.72, ERBB3 0.69, CD276 0.82, VTCN1 0.72, TACSTD2 0.77, FOLR1 0.93, and MET 0.82. By tumor type and task, lung adenocarcinoma achieved AUCs of 0.70 for tumor mutational burden, 0.87 for EGFR mutation, and 0.62 for KRAS mutation; colorectal cancer 0.99 for microsatellite instability; breast cancer 0.95, 0.88, and 0.81 for ER, PR, and HER2 protein subtyping and 0.74 and 0.86 for TP53 and PIK3CA mutations; renal cell carcinoma 0.60 for PBRM1 and 0.74 for BAP1 mutations; and colon adenocarcinoma 0.88 for KRAS and 0.89 for TP53 mutations. Gene overexpression tasks used The Cancer Genome Atlas with five-fold cross-validation across four seeds; lung and colorectal models used Samsung Medical Center cohorts; breast subtyping used the BCNB cohort; and renal and colon tasks used Clinical Proteomic Tumor Analysis Consortium data. Conclusions: An omics-aware patch aggregation framework co-trained with a scalable multi-omics encoder and WSIs enables accurate slide-level prediction for diverse biomarkers and tumor types and illustrates how partially paired multi-omics data can strengthen digital pathology models. AI was used for language editing only; authors are responsible for all content and approved the final version.
利益披露 Disclosure
H. Choi, None.. T. Hwang, None.. S. Lee, None.. J. Jang, None.

← 返回 AACR 2026 检索