PO.BCS02.03 · 生物信息与计算

从癌症样本水平的组成数据推断组织元素身份

Inferring tissue element identities from sample-level compositional data in cancer

海报缩略图:从癌症样本水平的组成数据推断组织元素身份
编号 5488 展板 1 时间 4/21 02:00–05:00 区域 Section 3 主讲 Georgios Asimomitis, PhD
分会场 Machine Learning for Image Analysis
查看 PDF 下载 PDF 🔒 查看 / 下载完整 PDF 需登录并开通下载套餐 · 查看套餐 / 开通 AACR 官方页面

作者与单位 Authors & Affiliations

Georgios Asimomitis1, Kevin M. Boehm1, Konstantinos Liosis1, Armaan Kohli1, Tom Pollard1, Andrew Aukerman1, Arfath Pasha1, Anika Begum1, Lora H. Ellenson2, Jinru Shia2, Hong A. Zhang2, Nikolaus Schultz1, Sohrab P. Shah1, Francisco Sanchez-Vega1

1Halvorsen Center for Computational Oncology, Memorial Sloan Kettering Cancer Center, New York, NY,2Department of Pathology & Laboratory Medicine, Memorial Sloan Kettering Cancer Center, New York, NY

摘要 Abstract

中文摘要
癌组织中细胞和结构成分的相对丰度反映了疾病生物学特征和结局。高通量技术,例如数字病理学和单细胞转录组学(scRNA-seq),可对组织区域或细胞进行分析,揭示构成整体表型的异质性构建单元(元素)。然而,转化应用越来越要求模型不仅能推断样本水平的分布,还能提供元素水平的表征(例如特定的组织学区域、细胞),以解释肿瘤的宏观行为。在无需繁琐注释的情况下从样本水平组成数据中实现这种精细度仍然是一项挑战,需要能够将局部特性与全局表型联系起来的模型。 现有方法使用注意力机制来加权各元素对样本水平预测的贡献,但缺乏概率学基础和生物学可解释性,因此在临床环境中的应用价值有限。在此,我们利用最优传输(Optimal Transport,OT)显式地对样本水平的组成约束与元素水平的分配进行建模。具体而言,我们提出了Composer,一个与领域无关的双任务机器学习模型,在无需元素水平注释的情况下训练,用于:1)估计样本水平的组成;2)更重要的是,将元素标签推断为可解释的组成分配。 该模型的性能在不同数据模态、癌症类型和任务中进行了评估。我们重点介绍3个概念: a)全切片图像(WSIs)上的组织类型分类:在60张来自高级别浆液性卵巢癌(HGSOC,Boehm等,2022)的苏木精-伊红(H&E)染色WSIs上,Composer预测了整体WSI组织组成(Jensen-Shannon散度(JSD):0.26;平均绝对误差:0.11),并在未经区域注释训练的情况下,正确分类(AUROC 0.98)了WSIs内各个区域的组织类型(肿瘤、间质、坏死、脂肪、其他)。 b)WSIs上的肿瘤分割:在185张涵盖HGSOC、乳腺癌和结直肠癌的H&E染色WSIs(MSKCC)上,Composer估计了每张WSI的整体肿瘤比例(JSD 0.12;MAE 0.15),高效地区分了WSIs内的肿瘤区域与非肿瘤区域(AUROC 0.91)。 c)scRNA-seq中的单细胞类型注释:使用156个scRNA-seq HGSOC样本(Vazquez-Garcia等,2022),Composer不仅推断了整体细胞类型分布(JSD:0.15;MAE:0.06),还准确地将各个细胞分类(AUROC:0.97)到其指定类型(T细胞、单核细胞、成纤维细胞、癌细胞、其他)。 总之,所提出的基于OT的弱监督框架为在癌症中将元素水平表征与样本水平组成图谱联系起来提供了一种有效方法。其在各种数据分析中的适用性有助于肿瘤生态系统的空间和分子表征、生物标志物定量以及在精准肿瘤学中的潜在应用。
查看英文原文 English abstract
The relative abundance of cellular and structural components in cancer tissues reflects disease biology and outcomes. High-throughput technologies, e.g., digital pathology and single-cell transcriptomics (scRNA-seq), profile tissue regions or cells, revealing the heterogeneous building blocks (elements) that comprise bulk phenotypes. Yet, translational applications increasingly demand that models not only infer sample-level distributions but also provide element-level characterizations (e.g., of specific histology regions, cells) that explain tumor macroscopic behavior. Achieving such granularity from sample-level compositional data without laborious annotations remains a challenge, requiring models capable of connecting local properties to global phenotypes. Existing methods, using attention to weight element contributions to sample-level predictions, lack probabilistic grounding and biological interpretability; thus, have limited utility in clinical settings. Here, we explicitly model sample-level compositional constraints with element-level assignments using Optimal Transport (OT). Particularly, we introduce Composer , a domain-agnostic dual-task machine learning model trained without element-level annotations to 1) estimate sample-level compositions, and, 2) importantly infer element labels as interpretable compositional allocations. The model's performance was evaluated across data modalities, cancer types, and tasks. We highlight 3 concepts: a) Tissue type classification on whole-slide images (WSIs): On 60 hematoxylin & eosin (H&E) WSIs from high-grade serous ovarian cancer (HGSOC, Boehm et al. 2022), Composer predicted the overall WSI tissue composition (Jensen-Shannon divergence (JSD): 0.26; mean absolute error: 0.11), classifying correctly (AUROC 0.98) the tissue type (tumor, stroma, necrosis, adipose, other) of individual regions within WSIs without training on regional annotations. b) Tumor segmentation on WSIs: On 185 H&E WSIs spanning HGSOC, breast, and colorectal cancers (MSKCC), Composer estimated the overall tumor fraction per WSI (JSD 0.12; MAE 0.15), distinguishing efficiently tumor from non-tumor regions (AUROC 0.91) within WSIs. c) Single-cell type annotation in scRNA-seq: Using 156 scRNA-seq HGSOC samples (Vazquez-Garcia et al. 2022), Composer not only inferred bulk cell type distributions (JSD: 0.15; MAE: 0.06), but also accurately classified (AUROC: 0.97) individual cells to their designated type (T cells, monocytes, fibroblasts, cancer cells, other). In summary, the proposed OT-based weakly supervised framework provides an effective approach for linking element-level representations to sample-level compositional profiles in cancer. Its applicability across data analyses empowers spatial and molecular characterization of tumor ecosystems, biomarker quantification, and potential applications to precision oncology.
利益披露 Disclosure
G. Asimomitis, None. K. M. Boehm, Monograph Capital, LLC Independent Contractor. Japanese Society of Obstetrics and Gynecology Travel. Memorial Sloan Kettering Cancer Center Patent. K. Liosis, None.. A. Kohli, None.. T. Pollard, None.. A. Aukerman, None.. A. Pasha, None.. A. Begum, None.. L. H. Ellenson, None.. J. Shia, None.. H. A. Zhang, None. N. Schultz, Stand Up to Cancer Independent Contractor. Innovation in Cancer Informatics Independent Contractor. S. P. Shah, Bristol-Myers Squibb Independent Contractor. F. Sanchez-Vega, None.

← 返回 AACR 2026 检索