PO.BCS01.15 · 生物信息与计算
一种可泛化的超快速序列分析软件框架及其在实现脑癌精准肿瘤学1日深度多组学数据分析周转中的应用
A generalizable software framework for ultra-rapid sequence analysis and its application in enabling 1 day deep multi-omic data analysis turnaround for brain cancer precision oncology
作者与单位 Authors & Affiliations
摘要 Abstract
中文摘要
背景:复发的儿童脑癌患者缺乏标准治疗方案且临床预后不良;而成人胶质母细胞瘤既是最常诊断的中枢神经系统恶性肿瘤,也是最致命的。虽然他们最有可能从由多组学肿瘤分析支持的精准指导、个性化治疗选择中获益,但他们术后识别和获取可能有效药物的机会窗口极为短暂。这一短暂的临床时间框架使得利用新一代测序(NGS)肿瘤分析技术的临床试验无法设计,因为肿瘤-正常配对全基因组DNA测序和肿瘤单细胞RNA测序所产生的海量数据无法用当前主流方法及时分析。
方法:我们着眼于图形处理单元(GPU)加速计算以显著加快序列分析。我们既利用了现有的加速成果(最著名的是NVIDIA Parabrick软件包中GPU版本的BWA-MEM和STAR比对算法),也利用了我们自行开发的、目前尚无GPU版本的软件的加速版本(FreeBayes短变异和FACETS拷贝数变异检出算法)。我们既关注与非加速版本的结果等效性,也关注代码可复用性,以便他人能够轻松地将更多软件移植到GPU平台。
结果:以Genome In A Bottle的HG008肿瘤-正常配对作为测试案例(2×150bp Illumina NovaSeq 6000,正常样本150X和肿瘤样本190X标称覆盖度),使用现成的计算机硬件,我们的流程在序列比对上耗时<2小时,在变异检出和注释上耗时<30分钟。因此,整个序列分析工作流程可在1个工作日内完成,并生成可供专家审阅、分子肿瘤委员会讨论和治疗选择的已注释变异检出集。我们进一步在一个由16例患者、包含DNA和scRNA测序数据的院内儿童脑癌数据集上测试了我们的方法,并观察到相似结果。
结论:我们开发了一种将GPU加速计算应用于序列分析的可泛化方法,以及它们如何能够通过提供1日分析周转,显著扩展将基于NGS的精准肿瘤学方法应用于成人/儿童脑癌的可能性。结合测序机构的物流优化(例如使用专用流动槽和优先队列),我们已实现<1周的手术到洞见周转。我们的工作可作为设计依赖NGS为脑癌患者优先排序治疗方案的下一代精准肿瘤学临床试验的构建模块。
查看英文原文 English abstract
BACKGROUND: Pediatric brain cancer patients in relapse lack standard-of-care options and experience poor clinical outcome; and adult glioblastoma is simultaneously the most commonly diagnosed central nervous system malignancy and the most deadly. While they stand to benefit the most from precision-guided, personalized therapy selection supported by multi-omic tumor profiling, they also experience an extremely short post-surgery window-of-opportunity to identify and acquire likely effective drugs. This short clinical timeframe precludes clinical trials utilizing next generation sequencing (NGS) tumor profiling techniques from being designed, as the vast amount of data produced by tumor-normal pair whole genome DNA sequencing and tumor single cell RNA sequencing cannot be analyzed in time using current prevailing methods.
METHODS: We look towards graphics processing unit (GPU) accelerated computation to significantly speed up sequence analysis. We utilize both existing acceleration efforts (most notably the GPU version of BWA-MEM and STAR alignment algorithms as part of the NVIDIA Parabrick package) as well as accelerated versions of software we developed for which no GPU versions currently exist (FreeBayes short variant and FACETS copy number variant calling algorithms). We focus both on result equivalency to the unaccelerated versions, as well as code reusability so that additional software can be ported to GPU platforms by others with ease.
RESULTS: Using the HG008 tumor normal pair from Genome In A Bottle as a test case (2x150bp Illumina NovaSeq 6000, 150X normal and 190X tumor nominal coverage), our pipeline spent < 2 hours in sequence alignment, and < 30 minutes in variant calling and annotation using readily available computer hardware. As a result, the entire sequence analysis workflow can be finished within 1 working day, and produces an annotated variant call set ready for expert review, molecular tumor board discussion, and therapy selection. We further tested our approach on an in-house pediatric brain dataset consisting of 16 patients and both DNA and scRNA sequencing data, and observed similar results.
CONCLUSION: We have developed a generalizable approach to adopting GPU accelerated computation to sequence analysis, and how they can significantly expand the possibilities of applying NGS-based precision oncology approaches to adult / pediatric brain cancer by offering 1 day analysis turnaround. Coupled with logistic optimization at sequencing facilities (e.g. using dedicated flow cells and priority queues), we have achieved <1 week surgery-to-insight turnaround. Our work serves as building blocks for designing next general precision oncology clinical trials that rely on NGS to prioritize treatment options for brain cancer patients.
利益披露 Disclosure
A. Pitman, None..
D. Bean, None..
Y. Qiao, None.