PO.MCB08.05 · 分子与细胞生物学
利用癌症基因组图谱(The Cancer Genome Atlas)全基因组测序数据对癌症患者进行AI驱动的分层
AI-Driven stratification of cancer patients using The Cancer Genome Atlas whole-genome sequencing data
作者与单位 Authors & Affiliations
摘要 Abstract
中文摘要
近期的基因组基础模型推进了DNA序列的解读,但大多数模型仍局限于局部序列模式,无法产生临床决策所需的患者层面见解。为解决这一局限,我们开发了一个框架,该框架超越了序列层面的推断,通过癌症基础模型(Cancer Foundation Model)实现稳健的患者分层。我们的方法始于"DNAChunker",它采用基于动态H-net的分词策略,将基因组划分为可变长度的片段,在调控区和编码区保留高分辨率细节,同时高效压缩重复序列。在Nucleotide Transformer和Genomic Benchmarks上评估时,DNAChunker仅使用1.56亿参数,就实现了与最先进的GENERator(12亿参数)相当的性能。为将这些基因组嵌入转化为患者层面的见解,我们实现了一个基于transformer的癌症聚合模型(Cancer Aggregation Model),它将突变嵌入与体细胞拷贝数改变(SCNA)特征整合。该框架在大型全基因组测序(WGS)队列上进行了评估,包括PCAWG(n=2,040)和CUBRICS乳腺癌样本(n=1,053),并以TCGA-BRCA(乳腺癌;n=920)作为外部验证队列。该模型有效地按癌症类型(准确率96.89%)、同源重组缺陷(HRD;准确率92.83%)和PAM50亚型(准确率84.05%)对患者进行了分层。值得注意的是,它仅使用DNA层面的信息即可对PAM50内在亚型进行分类,摆脱了传统上对基于RNA表达谱分析的依赖。该癌症基础模型证明,从全基因组数据进行患者层面的表征学习能够在多种肿瘤类型中实现具有临床意义的分层。通过弥合基因组序列解读与可操作表型分类之间的鸿沟,该框架为基于AI的精准肿瘤学奠定了基础。经过进一步验证,它将有助于直接从WGS数据进行生物标志物发现和临床试验中的患者分层。
查看英文原文 English abstract
Recent genomic foundation models have advanced DNA sequence interpretation, yet most remain constrained to local sequence patterns and fail to produce the patient-level insights required for clinical decision-making. To address this limitation, we developed a framework that extends beyond sequence-level inference, enabling robust patient stratification through a Cancer Foundation Model. Our approach begins with “DNAChunker”, which employs a dynamic H-net-based tokenization strategy that divides the genome into variable-length segments, preserving high-resolution detail in regulatory and coding regions while efficiently compressing repetitive sequences. When evaluated on the Nucleotide Transformer and Genomic Benchmarks, DNAChunker achieved performance comparable to the state-of-the-art GENERator (1.2 billion parameters) while using only 156 million parameters. To translate these genomic embeddings into patient-level insights, we implemented a transformer-based Cancer Aggregation Model that integrates mutation embeddings with somatic copy-number alteration (SCNA) features. The framework was evaluated on large whole-genome sequencing (WGS) cohorts, including PCAWG (n=2,040) and CUBRICS breast cancer samples (n=1,053), with TCGA-BRCA (breast cancer; n=920) serving as an external validation cohort. The model effectively stratified patients by cancer type (accuracy, 96.89%), homologous recombination deficiency (HRD; accuracy, 92.83%), and PAM50 subtype (accuracy, 84.05%). Notably, it classified PAM50 intrinsic subtypes using only DNA-level information, eliminating the conventional reliance on RNA-based expression profiling. The Cancer Foundation Model demonstrates that patient-level representation learning from whole-genome data can achieve clinically meaningful stratification across diverse tumor types. By bridging the gap between genomic sequence interpretation and actionable phenotypic classification, this framework establishes a foundation for AI-based precision oncology. With further validation, it will facilitate biomarker discovery and patient stratification in clinical trials directly from WGS data.
利益披露 Disclosure
J. Lee, None..
C. Bao, None..
H. Park, None..
G. Lee, None..
Y. Lee, None..
B. Lee, None..
D. Lehotzky, None..
R. Solan, None..
A. Kowalewski, None..
X. Loinaz, None..
V. Narasimha Swamy, None..
D. I. Heiman, None..
S. Van Seters, None..
S. Belkin, None..
S. Wiseman, None.
A. D. Cherniack,
Bayer ).
L. Corchete Sanchez, None..
B. Danysh, None..
Z. Everton, None..
C. Stewart, None..
H. Tomono, None..
G. Wang, None.
E. Rheinbay,
Inocras ), E.R. receives research funding from Inocras, Inc.
G. Getz,
IBM ).
Pharmacyclics/Abbvie ).
Bayer ).
Genentech ).
Calico ).
Ultima Genomics ).
Inocras ).
Google ).
Kite ).
Novartis ).
Scorpion Therapeutics Stock Option, He is a founder, consultant, and holds privately held equity in Scorpion Therapeutics.
Predicta Biosciences Stock Option, He is also a founder of, and holds privately held equity in, Predicta Biosciences.
Antares Therapeutics Stock Option.
Y. Ju, None..
W. Lee, None..
R. Kim, None.