PO.BCS01.14 · 生物信息与计算
Quartet:肿瘤细胞系稳健体细胞突变数据库
Quartet: A database of robust somatic mutations in tumor cell lines
作者与单位 Authors & Affiliations
摘要 Abstract
中文摘要
引言:当前的肿瘤细胞系集合及其基因组、蛋白质组和表型参考数据对癌症研究极为重要。鉴于这些数据集的核心地位,确保其达到最高质量至关重要。在此我们表明,这些集合中所识别突变的高达20%是错误的,原因在于测序伪影、错误比对的读段以及其他误差来源。为了给癌症细胞系百科全书(CCLE)测序的全部329个肿瘤细胞系获得一套高质量的突变谱,我们使用四种互补突变识别工具的共识处理了所有基因组,产生了一个我们称之为Quartet的在线资源。我们证明,与当前使用的突变谱相比,在多种任务中使用Quartet具有可观的益处。
方法:我们检索了癌症细胞系百科全书先前生成的原始序列数据。我们过滤掉不成比例地导致突变谱噪声的短序列,使用4种突变识别工具识别候选突变,然后将这些候选突变过滤为至少3种工具一致认可的子集。使用以超高深度测序的HCC1395癌细胞系作为阳性对照,我们将CCLE最初生成突变的约1/5归因于技术伪影,并证明Quartet中真阳性突变的富集。这些伪影突变在基因水平持续存在,我们观察到我们真实标准预测的变异效应与CCLE中所表达的变异效应之间存在分歧,而这一分歧通过Quartet得到显著纠正。当跨所有细胞系聚合并就原始CCLE突变中的变异等位基因分数(VAF)进行观察时,我们观察到变异的双峰分布,其较小的峰富集了Quartet有效过滤的技术伪影。我们证明,这些伪影突变无法通过VAF阈值或变异效应预测事后去除,也无法通过使用DepMap等替代数据库去除。随后我们利用这一召回的CCLE数据集训练药物基因组学模型,并显示在10种抗癌药物中预测性能的显著改善。
结论:Quartet提供了一个使用突变检出工具共识支持的癌症细胞系公共资源。Quartet证明了伪影突变及其在多种下游任务中有害效应的减少。这些结果凸显了通过Quartet增强癌症生物学发现的潜力,我们预期随着更多细胞系的纳入,这些益处将会累积。Quartet数据通过zenodo公开提供,源代码可在https://github.com/digitaltumors/quartet获取。
查看英文原文 English abstract
Introduction: Current collections of tumor cell lines, and their genomic, proteomic, and phenotypic reference data, have been vastly important to cancer research. Given the centrality of these datasets, it is paramount that they be of the highest quality. Here we show that up to 20% of mutations identified within these collections are erroneous, owing to sequencing artifacts, misaligned reads, in addition to other sources of error. To achieve a high-quality set of mutation profiles for all 329 tumor cell lines sequenced by the Cancer Cell Line Encyclopedia (CCLE), we processed all genomes using the consensus of four complementary mutation identification tools, yielding an online resource we call Quartet. We demonstrate considerable benefit from its use compared to currently used mutation profiles across a variety of tasks.
Methods: We retrieved raw sequence data previously generated by the Cancer Cell Line Encyclopedia. We filter for short sequences that disproportionately contribute to noise in mutation profiles, identify candidate mutations using 4 mutation identification tools, and then filter those candidate mutations to the subset with agreement from at least 3 tools. Using a HCC1395 cancer cell-line sequenced at ultra-high depths as a positive control, we attribute roughly 1/5 of mutations originally generated by CCLE to technical artifacts and demonstrate enrichment of True-Positive mutations from Quartet. These artifactual mutations persist at the gene-level, where we observe divergence between predicted variant effects from our ground truth and those expressed in CCLE, a divergence which is significantly recovered through Quartet. When aggregated across all cell lines and viewed with respect to Variant Allele Fraction (VAF) in original CCLE mutations, we observe a bimodal distribution of variants whose lesser mode is enriched for technical artifacts which Quartet effectively filters. We demonstrate these artifactual mutations cannot be removed post-hoc through either VAF thresholds or Variant Effect Prediction, nor through use of alternate databases such as DepMap. We then leverage this recalled CCLE dataset to train pharmacogenomics models and show significant improvement to predictive performance across 10 anti-cancer drugs.
Conclusion: Quartet presents a public resource of cancer cell lines which use consensus support of mutation callers. Quartet demonstrates depletion of artifactual mutations and their deleterious effects in a variety of downstream tasks. These results highlight the potential to enhance cancer biological discovery through Quartet, and we expect these benefits to compound as additional cell lines are incorporated. Quartet data are made publicly available through zenodo and source code is made available at https://github.com/digitaltumors/quartet.
利益披露 Disclosure
D. Halmos, None.