PO.BCS01.11 · 生物信息与计算
片段组学助力液体活检中体细胞突变分类的改进
Fragmentomics powers improved classification of somatic mutations in liquid biopsy
作者与单位 Authors & Affiliations
摘要 Abstract
中文摘要
引言:在游离DNA(cfDNA)中准确区分体细胞变异与胚系变异对精准肿瘤学至关重要。液体活检中的体细胞变异判读可能具有挑战性,尤其是对于高肿瘤分数样本中具有高突变等位基因频率(MAF)的变异。我们开发了一个基于片段组学的机器学习模型,以改进液体活检中的体细胞变异分类。
方法:我们在Guardant360 Liquid(Guardant Health,加州Palo Alto)上处理的、来自3,313名独特患者的4,250份临床样本上对模型进行了基准测试,其中变异的胚系状态可从纵向数据中可靠推导。训练数据包含5,253个独特变异;独立测试数据包含11,612个变异。我们工程化构建了55个片段组学特征并评估了多个分类器;使用片段长度和相对VAF特征的L2惩罚逻辑回归表现出最佳性能。我们选择了一个分类评分的临界阈值,以将模型的假阳性率维持在10%以下。该模型作为一个额外的校正步骤被整合进来,用于重新评估被基线判读算法分类为胚系的变异。
结果:从一个包含来自2,874份独特样本、26,900个变异的真值集出发(这些变异的体细胞或胚系来源已从纵向数据中获知),其中22,687个为真正的体细胞来源,4,213个为胚系。我们的基线判读算法将770/22,678个体细胞变异判定为胚系状态,错误判定的可能性与高ctDNA肿瘤分数相关。片段组学辅助方法能够“挽救”558个被错误判定的变异,其中包括226个具有直接临床可干预性的变异。就敏感性(真体细胞变异/(真体细胞变异+被判为胚系的体细胞变异))和特异性(真胚系变异/(真胚系变异+被判为体细胞的胚系变异))而言,我们的方法将敏感性从基线判读算法的96.61%提高到99.07%,同时特异性从99.00%下降到93.38%。尽管特异性下降,准确率和F1评分分别提高了1.18%和0.76%。逻辑回归模型在各癌症类型和不同片段谱中均表现出稳健性,最具预测性的特征是双核小体片段长度以及180至220个碱基范围内片段的相对等位基因频率。
结论:在cfDNA胚系-体细胞区分判读器中纳入基于模型的片段组学信号,可显著提高体细胞变异检测的敏感性,同时保持高准确性,从而能够可靠报告更多临床可干预的改变。考虑到该检测用于检测可靶向体细胞改变的预期用途,以特异性为代价换取敏感性的提升在临床上是可接受的。
查看英文原文 English abstract
Introduction: Accurate discrimination between somatic and germline variants in cell-free DNA (cfDNA) is critical for precision oncology. Somatic variant calling in liquid biopsy can be challenging, particularly for variants with high mutant allele frequency (MAF) in high tumor fraction samples. We developed a fragmentomics-based machine learning model to improve somatic variant classification in liquid biopsy.
Methods: We benchmarked our model on 4,250 clinical samples from 3,313 unique patients processed on Guardant360 Liquid (Guardant Health, Palo Alto, CA), where variant germline status could be confidently derived from longitudinal data. Training data comprised 5,253 unique variants; independent test data included 11,612 variants. We engineered 55 fragmentomics features and multiple classifiers were evaluated; optimal performance was demonstrated with logistic regression with L2 penalty using fragment length and relative VAF features. We chose a cutoff threshold for classification scores to maintain less than 10% false positives by the model. The model was integrated as an additional correction step to re-evaluate variants classified as germline by a baseline calling algorithm.
Results: Starting with a truth set of 26,900 variants from 2,874 unique samples whose somatic or germline origin were known from longitudinal data, 22,687 had a truly somatic origin and 4,213 were germline. Our baseline calling algorithm assigned germline status to 770/22,678 somatic variants and likelihood of mis-assignment was associated with a high ctDNA tumor fraction. The fragmentome-assisted approach was able to 'rescue' 558 of the mis-assigned variants including 226 with direct clinical actionability. In terms of sensitivity (true somatic variant/true somatic variant+somatic variant called germline) and specificity (true germline variant/true germline variant+germline variant called somatic), our approach improved sensitivity compared to the baseline calling algorithm from 96.61% to 99.07%, while specificity decreased from 99.00% to 93.38%. Despite the drop in specificity, accuracy and F1 score increased by 1.18% and 0.76%, respectively. The logistic regression model demonstrated robustness across cancer types and varying fragment profiles, with the most predictive features being di-nucleosome fragment lengths and relative allele frequencies over fragments ranging between 180 to 220 bases.
Conclusions: Inclusion of model-based fragmentomics signals in a cfDNA germline-somatic discrimination caller significantly improves somatic variant detection sensitivity while maintaining high accuracy, enabling confident reporting of additional clinically actionable alterations. The boost in sensitivity at the expense of specificity is clinically acceptable given the intended use of the test for detection of targetable somatic alterations.
利益披露 Disclosure
M. Stamboulian, None..
A. Valouev, None..
T. Jiang, None..
J. Hutchins, None.