PO.BCS01.15 · 生物信息与计算

绘制癌症DNA甲基化组的可分类性图谱:一种数据学习得出的疾病层级

Mapping classifiability in the cancer DNA methylome: A data-learned disease hierarchy

海报缩略图:绘制癌症DNA甲基化组的可分类性图谱:一种数据学习得出的疾病层级
编号 1514 展板 21 时间 4/20 09:00–12:00 区域 Section 6 主讲 Hao Xu, BS
分会场 Sequence Analysis
查看 PDF 下载 PDF 🔒 查看 / 下载完整 PDF 需登录并开通下载套餐 · 查看套餐 / 开通 AACR 官方页面

作者与单位 Authors & Affiliations

Hao Xu, Jenny Z. Li, Wanding Zhou

Center for Computational and Genomic Medicine, The Children’s Hospital of Philadelphia, Philadelphia, PA

摘要 Abstract

中文摘要
基于DNA甲基化的肿瘤分类已成功应用于临床。传统分类器假设诊断标签是固定且相互排斥的,然而许多生物学实体的定义并不一致,或存在内在的重叠。我们提出了一个通用框架,用于量化和解读基于DNA甲基化的疾病分类中的可分类性。在此,我们将可分类性本身视为数据的一种经验属性,通过对来自TCGA和GEO、涵盖超过324种癌症类型的13,698个以上协调统一的甲基化组队列进行交叉验证来加以衡量。该方法揭示了哪些疾病或癌症类型在分子水平上可以稳定区分,哪些则会跨标签坍缩合并,从而提供了一种基于原理、由数据驱动的生物学边界视角。通过分析交叉验证的一致性和标签的可混淆性,我们重建了一个直接由甲基化组衍生的层级分类系统,其中实体之间的关系无需事先的人为定义即可涌现,既重现了已知的谱系关系,又揭示了新颖的跨实体邻近性。我们进一步将主要的公共数据集整合为一个泛疾病基础分类器,该分类器既报告预测结果,也报告可分类性感知的置信度评分,反映沿学习得到的层级结构上的可分离性。最后,我们证明该框架同样可用于评估新的或罕见的队列,检验一个拟议的实体究竟能否构成一个独立、可分类的单元,还是会与既有类型合并。总之,这些进展将甲基化分类从一项预测任务重塑为一项在人类表观基因组中发现可分类性结构本身的任务,为完善肿瘤分类系统和诊断标准提供了数据驱动的基础。
查看英文原文 English abstract
DNA methylation-based tumor classification has been successfully applied in clinical settings. Traditional classifiers assume that diagnostic labels are fixed and mutually exclusive, yet many biological entities are inconsistently defined or intrinsically overlapping. We propose a general framework for quantifying and interpreting classifiability in DNA methylation-based disease classification. Here, we treat classifiability itself as an empirical property of data, measured through cross-validation across more than 13,698 harmonized methylome cohorts drawn from TCGA and GEO, spanning over 324 cancer types. This approach reveals which disease or cancer types are stably separable at the molecular level and which collapse across labels, providing a principled, data-driven view of biological boundaries. By analyzing cross-validation consistency and label confusability, we reconstruct a hierarchical taxonomy derived directly from the methylome, in which relationships between entities emerge without prior human definitions, recapitulating known lineage relationships and uncovering novel cross-entity proximities. We further integrate major public datasets into a pan-disease foundation classifier that reports both predictions and classifiability-aware confidence scores, reflecting separability along the learned hierarchy. Finally, we demonstrate that the same framework can evaluate new or rare cohorts, testing whether a proposed entity forms a distinct, classifiable unit or merges with established types. Together, these advances recast methylation classification from a task of prediction into one of discovering the structure of classifiability itself in the human epigenome, offering a data-driven foundation for refining tumor taxonomies and diagnostic criteria.
利益披露 Disclosure
H. Xu, None.. J. Z. Li, None.. W. Zhou, None.

← 返回 AACR 2026 检索