PO.BCS02.03 · 生物信息与计算

可扩展且可解释的多模态AI:通过基于图像的编码整合影像与基因组学以进行癌症和疾病分类

Scalable and interpretable multimodal AI: Integrating imaging and genomics via image-based encodingfor cancer and disease classification

海报缩略图:可扩展且可解释的多模态AI:通过基于图像的编码整合影像与基因组学以进行癌症和疾病分类
编号 5490 展板 3 时间 4/21 02:00–05:00 区域 Section 3 主讲 Sakib Mostafa, BS;MS;PhD
分会场 Machine Learning for Image Analysis
查看 PDF 下载 PDF 🔒 查看 / 下载完整 PDF 需登录并开通下载套餐 · 查看套餐 / 开通 AACR 官方页面

作者与单位 Authors & Affiliations

Sakib Mostafa, Md. Tauhidul Islam

Radiation Oncology, Stanford University School of Medicine, Stanford, CA

摘要 Abstract

中文摘要
整合不同的数据模态(如医学影像和基因组学)是现代肿瘤学的基础,能够全面捕捉复杂疾病异质性的整体视图。破译这些多模态关系对精准医学至关重要,然而数据的多样化性质对当前的计算方法构成了重大挑战。传统的多模态融合深度学习方法虽然强大,但往往存在巨大的计算开销以及复杂、资源密集型的架构。为解决这一问题,我们提出了一种新颖的框架,利用最优传输(Optimal Transport)将高维表格组学数据转换为紧凑的二维图像表征。这一转换将组学数据重塑为一个额外的图像通道,使得能够使用单个卷积神经网络(CNN)同时处理两种数据流,从而克服了计算效率上的关键限制。在整合全切片图像(WSI)和空间转录组学(ST)的多模态癌症数据集上评估时,我们的方法达到了96%的准确率,优于仅使用WSI的94%基线。我们进一步在ADNI阿尔茨海默病数据集上证明了该框架的泛化能力,其达到了97.6%的准确率(相比之下仅MRI为93.4%)。因此,该框架为统一的多模态分析提供了一种可扩展、可解释且高效的方法,为癌症诊断和复杂生物系统研究提供了新的机遇。
查看英文原文 English abstract
The integration of disparate data modalities, such as medical imaging and genomics, is fundamental to modern oncology, capturing a holistic view of complex disease heterogeneity. Deciphering these multimodal relationships is critical for precision medicine, yet the diverse nature of the data represents a significant challenge for current computational methods. Traditional deep learning approaches for multimodal fusion, while powerful, often suffer from massive computational overhead and complex, resource-intensive architectures. To address this, we present a novel framework that transforms high-dimensional tabular omics data into a compact, two-dimensional image representation using Optimal Transport. This transformation recasts omics data as an additional image channel, enabling the use of a single convolutional neural network (CNN) to concurrently process both data streams, thereby overcoming critical limitations in computational efficiency. When evaluated on a multimodal cancer dataset integrating Whole Slide Images (WSI) and Spatial Transcriptomics (ST), our method achieved 96% accuracy, outperforming the 94% baseline using WSI alone. We further demonstrated the framework's generalizability on the ADNI Alzheimer's dataset, where it achieved 97.6% accuracy (vs. 93.4% for MRI-only). This framework thus provides a scalable, interpretable, and efficient approach for unified multimodal analysis, offering new opportunities for cancer diagnosis and the study of complex biological systems.<!--EndFragment-->
利益披露 Disclosure
S. Mostafa, None.. M. Islam, None.

← 返回 AACR 2026 检索