PO.BCS02.05 · 生物信息与计算
通过训练可扩展的深度神经网络从游离DNA中提取肿瘤信号以改善癌症早期检测
Improving early cancer detection by training scalable deep neural networks to extract tumor signal from cell-free DNA
作者与单位 Authors & Affiliations
摘要 Abstract
中文摘要
引言:基于血液的液体活检为无创癌症筛查提供了潜力。然而,由于循环肿瘤生物标志物水平低以及来自正常细胞的背景噪声,早期疾病的检测变得复杂。为解决这一问题,我们开发了一种新型深度学习框架,可在单条DNA读取的分辨率上检测癌症信号。将该方法应用于经亚硫酸氢盐转化的游离DNA(cfDNA)样本后,我们的方法显著提高了早期癌症的灵敏度。
方法:我们设计了一种大规模并行的二维卷积神经网络架构,通过学习数千个基因组区域的局部甲基化模式,来区分cfDNA中的癌症信号与非癌症信号。该模型以下一代测序(NGS)数据作为输入,将基因组窗口内的比对序列编码为图像,并输出用于分类的信息性特征向量。然而,训练受到两个现实世界数据限制的困扰:(1)疾病样本包含来自正常细胞和患病细胞的未标记片段的混合物;(2)获取足够的早期疾病数据成本高、负担重且耗时。为解决这些问题,我们设计了一种新型数据生成技术:(1)通过体外掺入肿瘤活检读取,为读取组分配阳性标签;(2)通过对非癌症cfDNA读取进行精细调整的体外混合,生成大规模、多样化的数据集。我们模型的架构具有关键优势:紧凑的输入编码、可解释的显著性图以及可扩展的并行架构:我们在一天之内对横跨1000万个基因组碱基的720TB数据进行了训练,凸显了我们框架的效率。
结果:我们使用来自CORE-HH临床研究(NCT05435066,N=1229例非癌症,N=1118例癌症,其中包括N=599例I/II期)的靶向亚硫酸氢盐测序数据验证了我们的方法。我们在一组预留的非癌症血浆(N=174)和肿瘤组织活检(N=505)上生成的13亿个训练样本上对模型进行了预训练。对临床样本数据的预测产生了特征向量,显著性图证实该模型突出了从活检中学习到的模式。在10x5交叉验证中,与在区域整体平均甲基化值上训练的分类器相比,在这些特征向量上训练的分类器在98.5%特异性下将总体灵敏度提高了9.9个百分点(I期:+6.5个百分点,II期:+17.6个百分点,III期:+14.3个百分点,IV期:+9.5个百分点)。这些性能提升,加上该框架的可扩展性,凸显了其作为癌症早期诊断变革性工具的潜力,并为在其他液体活检检测的NGS数据上训练模型奠定了基础。未来的工作将研究在不牺牲可扩展性的前提下,纳入来自基于DNA的大语言模型的逐读取嵌入的潜力。
查看英文原文 English abstract
Introduction: Blood-based liquid biopsies offer potential for non-invasive cancer screening. However, detecting early-stage disease is complicated by low levels of circulating tumor biomarkers and background noise from normal cells. To address this, we developed a novel deep-learning framework to detect cancer signal at the resolution of single DNA reads. Applied to bisulfite-converted cell-free DNA (cfDNA) samples, our method significantly improves early-stage cancer sensitivity.
Methods: We designed a massively parallel 2-D convolutional neural network architecture that differentiates cancer and non-cancer signal in cfDNA by learning local methylation patterns at thousands of genomic regions. The model takes next generation sequencing (NGS) data as input, encodes aligned sequences within genomic windows as images, and outputs informative feature vectors for classification. However, training is complicated by two real-world data limitations: (1) disease samples contain a mix of unlabeled fragments from normal and diseased cells, and (2) acquiring sufficient early-stage disease data is costly, burdensome, and time-intensive. To address these, we designed a novel data generation technique that (1) assigns positive labels for groups of reads via in silico spike-in of tumor biopsy reads and (2) generates large, diverse datasets via fine-tuned in silico mixing of non-cancer cfDNA reads. Our model's architecture has key advantages: compact input encoding, interpretable saliency maps, and scalable parallel architecture: we trained on 720TB of data across 10 million genomic bases in a single day, highlighting our framework's efficiency.
Results: We validated our method with targeted bisulfite sequencing data from the CORE-HH clinical study (NCT05435066, N=1229 non-cancers, N=1118 cancers, including N=599 Stage I/II). We pretrained the model on 1.3 billion training examples generated using a held-out set of non-cancer plasma (N=174) and tumor tissue biopsies (N=505). Predictions on data from clinical samples yielded feature vectors, with saliency maps confirming the model highlights biopsy-learned patterns. In a 10x5 cross-validation, classifiers trained on these feature vectors improved overall sensitivity by 9.9 points (Stage I: +6.5 pts, II: +17.6 pts, III: +14.3 pts, IV: +9.5 pts) at 98.5% specificity, compared to classifiers trained on region-wide average methylation values. These performance improvements, coupled with the scalability of the framework, underscore its potential as a transformative tool in the early diagnosis of cancer and establish a foundation for training models on NGS data in other liquid biopsy assays. Future work will investigate the potential to incorporate per-read embeddings from DNA-based large language models, without sacrificing scalability.
利益披露 Disclosure
J. A. Killian,
Harbinger Health Employment, Stock, Stock Option.
K. Pettie,
Harbinger Health Employment, Stock, Stock Option.
K. Gowen,
Harbinger Health Employment, Stock, Stock Option.
Precede Biosciences Employment.
S. Farashahi,
Harbinger Health Employment, Stock, Stock Option.
E. Brown,
Harbinger Health Employment, Stock, Stock Option.
Nexus Medical Labs Employment.
F. Hantash,
Harbinger Health Employment, Stock, Stock Option.
J. Charlton,
Harbinger Health Employment, Stock, Stock Option.
F. Michor,
Harbinger Health g., Board of Directors, non-salaried role), Stock, Patent, Other, Co-founder.
Zephyr AI Stock.
K. I. Chacko,
Harbinger Health Employment, g., Board of Directors, non-salaried role), Stock, Stock Option, Other, Co-founder.
Ambrosia Biosciences Employment, g., Board of Directors, non-salaried role), Stock, Other, Co-founder.
Flagship Pioneering Employment.
D. Kashef,
Harbinger Health Employment, Stock, Stock Option.