PO.CL01.07 · 临床研究

通过多特征全基因组cfDNA分析进行癌症类型识别的机器学习分类器

Machine learning classifier for cancer type identification via multi-feature genome-wide cfDNA profiling

海报缩略图:通过多特征全基因组cfDNA分析进行癌症类型识别的机器学习分类器
编号 1123 展板 4 时间 4/19 02:00–05:00 区域 Section 44 主讲 Haimeng Tang, MS
分会场 Liquid Biopsies: Circulating Nucleic Acids 1
查看 PDF 下载 PDF 🔒 查看 / 下载完整 PDF 需登录并开通下载套餐 · 查看套餐 / 开通 AACR 官方页面

作者与单位 Authors & Affiliations

Yunjian Zhang1, Liang Liu2, Hua Bao3, Haimeng Tang3, Ke Xu3, Hao Zhang4, Song Wang3, Shuang Chang3, Dongqin Zhu3, Zongyao Huang5, Zheng Wang2, Liu Yang6, Bingzhong Zhang7, Ji Tao8, Wenhua Liang9, Jierong Chen10, Shanshan Yang3, Xue Wu3, Yang Shao3, Wenquan Wang2, Dongyuan Zhu11

1First Affiliated Hospital of Sun Yat-sen University, Guangzhou, China,2Zhongshan Hospital, Fudan University, Shanghai, China,3Geneseeq Techonolgy Inc., Toronto, ON, Canada,4Xuzhou Medical University, Affiliated Hospital of Xuzhou Medical University, Xuzhou, China,5University of Electronic Science and Technology of China, Sichuan Cancer Hospital & Institute, Chengdu, China,6Colorectal Center, Jiangsu Cancer Hospital, Nanjing, China,7Sun Yat-sen Memorial Hospital, Sun Yat-sen University, Guangzhou, China,8Harbin Medical University Cancer Hospital, Harbin, China,9The First Affiliated Hospital of Guangzhou Medical University, Guangzhou, China,10Guangdong Provincial People’s Hospital, Guangzhou, China,11Shandong First Medical University and Shandong Academy of Medical Sciences, Jinan, China

摘要 Abstract

中文摘要
背景:确定癌症的组织起源(TOO)对于恰当的临床管理和治疗选择至关重要。利用循环游离DNA(cfDNA)的液体活检为癌症检测和TOO预测提供了一种非侵入性方法。循环肿瘤DNA(ctDNA)是cfDNA中源自肿瘤的部分,携带反映其起源的基因组和表观基因组特征。机器学习的最新进展使得能够开发从cfDNA图谱预测TOO的模型。然而,现有方法表现不一,尤其在ctDNA分数较低的样本中(ctDNA<3%),且不同癌症类型间的准确性仍不一致。 方法:我们利用来自17种癌症类型、1814例患者的全基因组cfDNA图谱开发了一个TOO分类器。提取了多种不同的cfDNA特征以揭示多样的癌症相关改变,包括拷贝数变异、重复元件、与DNA甲基化相关的片段末端基序、片段大小分布和覆盖度、微卫星不稳定性、突变特征、核小体占位、组织特异性片段化模式以及癌症相关病毒DNA的存在。模型性能在一个独立的外部队列(1221例患者)中进行了评估。还在原发灶不明癌(CUP)和多原发癌(MPC)患者队列中进行了额外测试。 结果:我们的癌症分类器在训练队列中实现了总体top 1准确率78%和top 2准确率89%,各癌症类型准确率一致较高。在独立验证队列中,模型保持稳健性能,top 1和top 2准确率分别为80%和90%。灵敏度随癌症分期升高而增加,从I期的66.8%提高到IV期的86.2%。在612份低ctDNA样本中,435例(71.1%)被正确分类。该分类器在CUP中也显示出强大潜力,15例中有11例(73.3%)与临床疑似原发部位相符。此外,在20例具有两个原发部位的MPC病例中,9例的两个部位均在前三个预测中被正确识别。在3例具有三个原发部位的MPC患者中,其中2例的三个部位中有两个在前三个预测中被准确捕获。 结论:我们基于cfDNA的机器学习分类器为准确识别癌症组织起源提供了一种稳健、非侵入性的方法。通过整合11种不同的cfDNA衍生的片段组学、基因组、表观基因组和微生物组特征,该模型在多种癌症类型中实现了高准确率,并在低ctDNA样本中保持强劲性能。其在CUP和MPC中的可喜结果进一步凸显了其在解决诊断困难病例和指导精准肿瘤学应用方面的潜在临床效用。
查看英文原文 English abstract
Background: Determination of the tissue of origin (TOO) of cancer is essential for appropriate clinical management and treatment selection. Liquid biopsy using circulating cell-free DNA (cfDNA) offers a non-invasive approach for cancer detection and TOO prediction. Circulating tumor DNA (ctDNA), a tumor-derived fraction of cfDNA, carries genomic and epigenomic signatures reflective of its origin. Recent advances in machine learning have enable the development of models to predict TOO from cfDNA profiles. However, current methods show variable performance, particularly in samples with low ctDNA fractions (ctDNA < 3%), and accuracy remains inconsistent across different cancer types. Methods: We developed a TOO classifier using whole genome cfDNA profiles from 1814 patients across 17 cancer types. Multiple distinct cfDNA features were extracted to reveal diverse cancer-associated alterations, including copy number variations, repeat elements, fragment end motifs associated with DNA methylation, fragment size distribution and coverage, microsatellite instability, mutational signatures, nucleosome occupancy, tissue-specific fragmentation patterns, and the presence of cancer-associated viral DNA. Model performance was evaluated in an independent external cohort of 1221 patients. Additional tests were conducted in cohorts of patients with cancers of unknown primary (CUP) and multiple primary cancers (MPC). Results: Our cancer classifier achieved an overall top 1 accuracy of 78% and top 2 accuracy of 89% in the training cohort, with consistently high accuracy across all cancer types. In the independent validation cohort, the model maintained robust performance, with top 1 and top 2 accuracies of 80% and 90%, respectively. Sensitivity increased with the advancing cancer stage, improving from 66.8% in stage I to 86.2% in stage IV. Among 612 low-ctDNA samples, 435 cases (71.1%) were correctly classified. The classifier also showed strong potential in CUP, with 11 of 15 cases (73.3%) aligning with the clinically suspected primary site. Furthermore, among 20 MPC cases with two primary sites, both were correctly identified within the top three predictions in 9 cases. In 3 MPC patients with three primary sites, two of the three sites were accurately captured among the top three predictions. Conclusion: Our cfDNA-based machine learning classifier provides a robust, non-invasive approach for accurate cancer tissue-of-origin identification. Integrating 11 distinct cfDNA-derived fragmentomic, genomic, epigenomic, and microbiomic features, the model achieved high accuracy across multiple cancer types and maintained strong performance in low-ctDNA samples. Its promising results in CUP and MPC further highlight its potential clinical utility in resolving diagnostically challenging cases and guiding precision oncology applications.
利益披露 Disclosure
Y. Zhang, None.. L. Liu, None. H. Bao, Geneseeq Technology Inc. Employment. H. Tang, Geneseeq Technology Inc. Employment. K. Xu, Geneseeq Technology Inc. Employment. H. Zhang, None. S. Wang, Geneseeq Technology Inc. Employment. S. Chang, Geneseeq Technology Inc. Employment. D. Zhu, Geneseeq Technology Inc. Employment. Z. Huang, None.. Z. Wang, None.. L. Yang, None.. B. Zhang, None.. J. Tao, None.. W. Liang, None.. J. Chen, None. S. Yang, Geneseeq Technology Inc. Employment. X. Wu, Geneseeq Technology Inc. Employment. Y. Shao, Geneseeq Technology Inc. Employment. W. Wang, None.. D. Zhu, None.

← 返回 AACR 2026 检索