PO.BCS01.13 · 生物信息与计算
Wakhan:利用长读长测序重建肿瘤基因组染色体尺度的拷贝数图谱
Wakhan: Reconstruction of chromosome-scale copy number profiles of tumor genomes with long-read sequencing
作者与单位 Authors & Affiliations
摘要 Abstract
中文摘要
引言。拷贝数改变(CNA)是癌症演化过程中的一种现象,其中基因组的某些区域可能被扩增或缺失。这导致了异质性的癌细胞集合。CNA图谱的分析和分类在理解癌症异质性和演化以更好地指导诊断和治疗方面发挥着至关重要的作用。目前有多种基于短读长的单倍型特异性CNA分析工具,但短读长提供的定相范围有限。长读长有助于将基因组变异直接定相为兆碱基尺度的单倍型,从而支持重建更长的、直至染色体尺度的CNA图谱。在此,我们提出了Wakhan,一种利用长读长分析单倍型特异性染色体尺度体细胞拷贝数畸变的工具。利用高质量的基因组组装覆盖度图谱,我们表明Wakhan在实现染色体级别CNA一致性方面显著优于其他常见的短读长和长读长CNA检测工具。
方法。Wakhan使用肿瘤-正常长读长BAM文件和定相的种系SNP检测结果作为输入。它首先通过利用单倍型覆盖度不平衡将输入定相扩展到染色体尺度。Wakhan检测那些相位切换区域,并通过考虑单倍型特异性覆盖度的变化来校正它们。接下来,Severus利用这种增强的定相来生成定相的结构变异(SV)检测结果。最后,Wakhan集成的CNA算法使用SV检测结果作为边界,并采用单倍型覆盖度模型为所得的CNA区域分配整数拷贝数状态。
https://github.com/KolmogorovLab/Wakhan
结果。我们试图将Wakhan的性能与几种最先进的单倍型特异性CNA检测工具进行比较。所选用于短读长分析的工具包括:Purple、Hatchet、Battenberg,用于长读长分析的工具包括Purple和Savana。由于有小变异和SV检测的基准,但没有针对体细胞CNA检测的类似基准。我们设计了一个基于CASTLE面板的CNA检测基准,包含6对用多种短读长和长读长测序技术测序的肿瘤/正常细胞系。我们将片段误差(SE)定义为:对每个CNA片段,我们计算杂合SNP处预期覆盖度与参考覆盖度之间的单倍型特异性均方距离。然后用它来计算加权的染色体平均值,并按肿瘤单倍型的平均覆盖度进行归一化。类似地,对于染色体误差(CE),将整条染色体的相位与参考覆盖度进行比较。在五个CASTLE数据集中,Wakhan和PURPLE具有最低的SE50和SE75,表明在重建单个CNA片段方面具有高准确性。我们还在一个仅肿瘤数据集上评估了Wakhan。Wakhan和PURPLE都很好地处理了正常样本的缺失,并准确反映了预期的肿瘤/正常图谱。
查看英文原文 English abstract
Introduction. Copy number alterations (CNA) is a phenomenon during cancer evolution where some regions of the genome may be amplified or deleted. This results in heterogeneous collections of cancer cells. Profiling and classification of CNA profiles play a vital role in understanding the cancer heterogeneity and evolution to better inform diagnosis and treatment. There are several short-reads haplotype-specific CNA profiling tools but short reads provide a limited phasing range. Long-reads facilitate the direct phasing of genomic variants into megabase-scale haplotypes, which supports the reconstruction of longer, up to chromosome-scale, CNA profiles. Here we present Wakhan, a tool to analyze haplotype-specific chromosome-scale somatic copy number aberrations using long reads. Leveraging high-quality genome assembly coverage profiles, we show that Wakhan significantly outperforms other common short- and long-read CNA callers in achieving chromosome-level CNA consistency.
Methods. Wakhan uses tumor-normal long-read BAMs and phased germline SNP calls as input. It first extends the input phasing to be chromosome-scale by exploiting haplotype coverage imbalance. Wakhan detects those phase switch regions and corrects them by taking into consideration the changes in haplotype-specific coverage. Next, Severus utilizes this enhanced phasing to generate phased structural variant (SV) calls. Finally, Wakhan's integrated CNA algorithm uses the SV calls as boundaries and employs a haplotype coverage model to assign integer copy-number states to the resultant CNA regions.
https://github.com/KolmogorovLab/Wakhan
Results. We sought to compare Wakhan's performance against several state-of-the-art haplotype-specific CNA calling tools. The tools selected for short-read analysis included: Purple, Hatchet, Battenberg and for long-read analysis Purple and Savana are included. As benchmarks for small variants and SV calling are available but no similar benchmarks for somatic CNA calls are available. We designed a CASTLE panel based CNA calling benchmark, consisting of 6 pairs of tumor/normal cell lines sequenced with multiple short- and long-read sequencing technologies. We define segment error (SE) as for each CNA segment, we calculate the haplotype-specific mean squared distance between expected and reference coverage at heterozygous SNPs. This is then used to compute a weighted chromosomal average, normalized by the tumor haplotype's mean coverage. Similarly, for chromosome error (CE), compare the phase of the whole chromosome against the reference coverage. In the five CASTLE datasets, Wakhan and PURPLE had the lowest SE50 and SE75, indicating high accuracy in reconstructing individual CNA segments. We also evaluated Wakhan on a tumor-only dataset. Both Wakhan and PURPLE handled the absence of normal samples well and accurately reflected the expected tumor/normal profiles.
利益披露 Disclosure
T. Ahmad, None..
M. Kolmogorov, None..
S. C. Sahinalp, None..
B. Paten, None..
S. Malikić, None..
Y. Liu, None..
A. Ataberk Donmez, None..
A. Goretsky, None.