PO.BCS01.15 · 生物信息与计算
一种基于深度学习的新型过滤工具,用于增强全基因组测序数据中技术性假象的检测
A novel, deep learning -based filtering tool for enhanced detection of technical artifacts in whole-genome sequencing data
作者与单位 Authors & Affiliations
摘要 Abstract
中文摘要
假象变异通过在这种低信噪比条件下模拟低变异等位基因分数(VAF)突变的某些特质,使本已困难的体细胞变异检出任务更加复杂。统计学检出工具无法覆盖假象的全部图景,留下了一个关键空白。虽然用于变异检出的深度学习工具已经存在,但它们通常消耗大量标注的训练数据、运行成本高昂,且可能对单一测序技术过拟合。在此,我们提出并基准测试了Permutect,一种轻量级的基于深度学习的变异检出工具,用于跨测序技术和基因组识别假象变异。Permutect结合了用于技术背景的假象模型和用于生物学背景的后验模型,使其不仅能学习技术特征,还能学习特定实验室环境的特殊性。该工具利用一组读段固有的无序性(从而具备排列不变性),以赋予其作为技术假象的概率。因此,Permutect是一种轻量级深度学习模型,训练时最少只需一个(未标注的)基因组,从而弥补了现有统计学和深度学习检出工具的不足。我们通过在四个Genome-in-a-Bottle样本上训练并将其与Mutect2、Strelka2和DeepSomatic在成熟的检出集(ICGC Dream Challenge Sets 1-4以及测序质量控制第二阶段(SEQC2)HiSeq和NovaSeq重复样本)上进行基准比较,测试了这一新型工具。Permutect是一种有效且精确的过滤工具,尤其在配对的肿瘤-正常(TN)设置中。它在Dream TN队列中取得了最高的平均Precision得分(0.928)和最高的平均F1得分(0.899),相对于其传统竞争者展现出更优的整体平衡性能。在SEQC2数据中性能保持高且一致(例如HiSeq TN上的F1得分为0.941);我们排除了DeepSomatic的得分,因为它是在同一数据上训练的。虽然所有工具的精度在仅肿瘤(TO)设置中都有所下降,但Permutect在Dream TO中保持了强劲的平均Recall(0.810),与Mutect2(0.811)相当。尽管所有检出工具在TO背景下都表现吃力,但我们的发现表明Permutect的性能与成熟方法相比具有竞争力且往往更优,凸显了其解决假象变异这一普遍问题的潜力。我们正在进一步开展工作,将Permutect的领域扩展至全外显子组和长读长测序,以及诸如纳入变异的三核苷酸背景等其他改进,这将进一步优化该工具以弥补在TO设置中所见的不足。通过这些努力,我们预期Permutect能够增强下游基因组分析,其应用范围涵盖从临床分子病理学到个性化免疫疗法的设计。
查看英文原文 English abstract
Artifactual variants obfuscate the already difficult task of somatic variant calling by mimicking some qualities of low variant allele fraction (VAF) mutations in this low signal-to-noise ratio regime. Statistical callers cannot cover the entire landscape of artifacts, leaving a critical gap. While deep learning tools for variant calling exist, they generally consume enormous amounts of labeled training data, are expensive to run, and can be overfit to a single sequencing technology. Here we present and benchmark Permutect, a lightweight deep learning-based variant caller for identifying artifactual variants across sequencing technologies and genomes. Permutect combines an artifact model for technological context and posterior model for biological context, respectively, allowing it to learn the characteristics not only for the technology but additionally the particularities of a given lab environment. The tool leverages the inherent non-orderedness of a set of reads (and thus their permutation invariance) to ascribe a probability for being a technical artifact. Permutect is therefore a lightweight deep learning model which can use as little as a single (unlabeled) genome for training thereby addressing the shortcomings of existing statistical and deep-learning callers. We tested this novel tool by training on four Genome-in-a-Bottle samples and benchmarking it against Mutect2, Strelka2, and DeepSomatic across well-established callsets: the ICGC Dream Challenge Sets 1-4 and SEquencing Quality Control Phase 2 (SEQC2) HiSeq and NovaSeq replicates. Permutect is an effective and precise tool for filtering, particularly in the paired Tumor-Normal (TN) setting. It achieved the highest mean Precision score (0.928) and the highest mean F1 score (0.899) in the Dream TN cohort, demonstrating superior overall balanced performance relative to its traditional competitors. Performance remained high and consistent across SeqC2 data (e.g., F1 score of 0.941 on HiSeq TN); we exclude scores from DeepSomatic, which was trained on that same data. While the precision of all tools are depressed in the Tumor-Only (TO) setting, it maintains a strong mean Recall (0.810) comparable to that of Mutect2 (0.811) in Dream TO. Though all callers struggled in the TO context, our findings show that Permutect's performance is competitive, and often superior, to established methods, highlighting its potential to address the pervasive issue of artifactual variants. Further work is being done to extend the domain of Permutect to whole-exome and long read sequencing and other improvements such as including the tri-nucleotide context about a variant will further refine the tool to address the shortcomings seen in the TO setting. In doing so, we expect Permutect to enhance downstream genomic analysis with applications ranging from clinical molecular pathology to the design of personalized immunotherapies.
利益披露 Disclosure
J. Gascoyne, None..
D. Benjamin, None..
J. Gallegos, None.
S. Shukla,
Agenus Inc. Stock.
Agios Pharmeceuticals Stock.
Breakbio Corp. Stock.
Bristol-Meyers Squibb Stock.
Imunon Stock, Other, Advisory/Consulting Role.
Jivanu Therapeutics Stock, Other, Advisory/Consulting Role.
Lumos Pharma Stock.
NeuroDiscovery AI Stock, Other, Advisory/Consulting Role.