PO.BCS01.11 · 生物信息与计算
跨两种全基因组测序平台评估健康血浆样本中的测序错误率
Assessment of sequencing error rates in healthy plasma samples across two whole genome sequencing platforms
作者与单位 Authors & Affiliations
摘要 Abstract
中文摘要
在健康样本中准确检测错误信号对于改善ctDNA分析的检测特异性和灵敏度至关重要。源于PCR、氧化损伤或克隆性造血变异的测序伪影和生物学噪声,可能掩盖低频肿瘤体细胞突变。本研究旨在系统性表征健康血浆样本的错误谱,并开发情境感知的过滤方法和模型,以区分真实变异与背景噪声。使用两种测序平台(测序仪A:N=52;测序仪B:N=28)从80份健康血浆样本生成全基因组测序数据(pWGS)。使用一个定制的pWGS分析模块,整合片段水平和位置水平的信息,跨测序片段计算突变特异性特征。我们比较了一个概率分类器(模型1)和一个先进的基于深度学习的模型(模型2),二者在同一特征集上训练,纳入了我们所识别的测序错误关键预测因子特征。在血浆样本中,于对应从26份组织样本(乳腺N=4,肺N=5,卵巢N=10,膀胱N=7)识别出的肿瘤突变靶标集的基因组位置上,对错误率进行了量化。使用测序平台A时,错误率强烈依赖于片段水平和序列情境特征。错误率在各变异类型间分布不均。错误率在片段末端附近以及GC富集区(GC>55%)内升高。跨癌种的中位原始错误率范围为每百万分之19至59(ppm)。应用模型1可将背景错误率降低最多65%,而模型2实现了最多90%的降低,最低错误率接近约5 ppm。来自平台B、以碱基质量>50进行过滤处理的测序数据实现了相当的覆盖度(约65×)和相似的错误率(约5 ppm)。在两种测序平台上,通过纳入更多片段水平特征和采用更先进的建模方法,均有可能进一步降低错误率。本研究建立了pWGS错误表征的基础框架,并证明了基于片段和情境的建模在降低测序噪声方面的有效性。
查看英文原文 English abstract
Accurate detection of error signals in healthy samples is critical for improving assay specificity and sensitivity in ctDNA analysis. Sequencing artifacts and biological noise, arising from PCR, oxidative damage, or clonal hematopoiesis variants, can obscure low-frequency tumor somatic mutations. This study aims to systematically characterize error profiles in healthy plasma samples and develop context-aware filtering and models to distinguish true variants from background noise. Whole-genome sequencing data were generated from 80 healthy plasma samples (pWGS) using two sequencing platforms (sequencer A: N=52; sequencer B: N=28). Mutation-specific features were computed across sequencing fragments using a custom pWGS analysis module that incorporated information both at the fragment- and the position-level. We compared a probabilistic classifier (Model 1) and an advanced deep learning based model (Model 2) trained on the same feature set, incorporating features that we identified as key predictors of sequencing error. Error rates within the plasma samples were quantified at genomic positions corresponding to tumor mutation target sets identified from 26 tissue samples (breast N=4, lung N=5, ovarian N=10, N=7 bladder). Using sequencing platform A, error rates showed strong dependence on fragment-level and sequence-context features. Error rates were not evenly distributed among variant types. Error rates were elevated near fragment ends and within GC-rich regions (>55% GC). Median raw error rates across cancer indications ranged from 19 to 59 parts per million (ppm). Application of Model 1 reduced background error rates by up to 65%, while Model 2 achieved up to a 90% reduction, with minimum error rates approaching ~5 ppm. Sequencing data from platform B that were processed with a base quality > 50 filter achieved comparable coverage (~65×) and similar error rates (~5 ppm). Across both sequencing platforms, further reductions are possible by incorporating additional fragment-level features and leveraging more advanced modeling approaches. This work establishes a foundational framework for pWGS error characterization and demonstrates the effectiveness of fragment- and context-based modeling in reducing sequencing noise.
利益披露 Disclosure
A. Shahpurwalla,
Natera, Inc. Employment, Stock, Stock Option.
Z. Montague,
Natera, Inc. Employment, Stock, Stock Option.
F. Lu,
Natera, Inc. Employment, Stock, Stock Option.
S. Alexander,
Natera, Inc. Employment, Stock, Stock Option.
G. Goyal,
Natera, Inc. Employment, Stock, Stock Option.
D. Hafez,
Natera, Inc. Employment, Stock, Stock Option.
M. Rabinowitz,
Natera, Inc. Employment, g., Board of Directors, non-salaried role), Stock, Stock Option, ), Travel, Patent, Consulting/Advisory Role.
MyOme Employment, g., Board of Directors, non-salaried role), Stock, Stock Option, ), Travel, Patent, Consulting/Advisory Role.
Marble Therapeutics Employment, g., Board of Directors, non-salaried role), Stock, Stock Option, Consulting/Advisory Role.
E. Kirkizlar,
Natera, Inc. Employment, Stock, Stock Option.
A. Zehir,
Natera, Inc. Employment, Stock, Stock Option.