PO.BCS01.12 · 生物信息与计算
Exacto:利用整合性长读长测序准确识别突变蛋白型和新抗原
Exacto: Accurate identification of mutant proteoforms and neoantigens using integrative long-read sequencing
作者与单位 Authors & Affiliations
摘要 Abstract
中文摘要
新抗原发现对于个性化免疫治疗至关重要,但目前的方法受限于对可通过短读长测序识别的小体细胞变异的关注。这些变异往往产生与自身抗原相似且免疫原性较弱的肽段。长读长测序能够更灵敏地检测大型结构变异和全长转录本。大型突变可产生与自身差异更大、因而免疫原性更强的新抗原。然而,目前尚无方法能够识别全面的体细胞DNA和RNA变异集、对其进行整合,并利用长读长数据对每个氨基酸进行情境化解读。为填补这一空白,我们开发了Exacto,一款公开可用的软件程序,利用长读长测序数据准确表征肿瘤基因组、转录组和突变蛋白型。Exacto执行三项主要功能。首先,它对参考基因组比对的长读长进行分析,以识别主要类型的肿瘤特异性DNA变异(SNV、多核苷酸变异、插入、缺失、重复、倒位和易位)和RNA变异(SNV、多核苷酸变异、插入、缺失、隐性外显子、内含子保留、外显子跳跃/截短、融合基因、环状RNA和未注释的基因间异构体)。其次,它整合RNA和DNA变异以预测体细胞突变的剪接后果。第三,它翻译全长RNA序列,并用潜在的RNA和DNA变异注释每个氨基酸。我们还在Exacto中开发了基因组和转录组变异图构建器,用于生成合成肿瘤及匹配的正常基因组以及肿瘤转录组。为对突变肽段识别进行全面的基准研究,我们开发了VSTOL和Nexus。VSTOL引入了一个称为奥卡姆变异语法(Occam's Variant Grammar)的新框架,以统一表示法为现有变异检测工具表征DNA和RNA变异。Nexus是一个运行超过50个新抗原发现工具的Nextflow套件。使用从变异图生成的合成样本,我们对Exacto、ClairS、Nanomonsv、Savana、Severus和SVision-pro的体细胞DNA变异检测进行了基准测试。Exacto在所有模拟变异类型上均达到最高或并列最高的召回率(SNV、缺失、插入 = 1.000;易位和倒位 = 0.983)。它以显著优势超越了次优工具:对于易位和倒位,Exacto达到0.688的精确率,而Savana为0.205(两者召回率均为0.983);对于缺失,Exacto的精确率为0.945,而Savana为0.761(两种方法召回率均为1.000);对于插入,Exacto的召回率为1.000,而Nanomonsv为0.901(两者精确率均为1.000)。鉴于这一性能,我们预期Exacto将成为利用长读长DNA和RNA测序发现免疫原性新抗原的标准方法。
查看英文原文 English abstract
Neoantigen discovery is essential for personalized immunotherapy, but current approaches are limited by a focus on small somatic variants identifiable by short-read sequencing. These variants often produce peptides that resemble self-antigens and are weakly immunogenic. Long-read sequencing enables more sensitive detection of large structural variants and full-length transcripts. Large mutations can generate neoantigens that are more dissimilar to self and thus more immunogenic. However, no method exists to identify a comprehensive set of somatic DNA and RNA variants, integrate these, and contextualize each amino acid using long-read data. To address this gap, we developed Exacto, a publicly available software program that uses long-read sequencing data to accurately characterize tumor genomes, transcriptomes, and mutant proteoforms. Exacto performs three main functions. First, it profiles reference-genome aligned long reads to identify major types of tumor-specific DNA variants (SNV, multi-nucleotide variant, insertion, deletion, duplication, inversion, and translocation) and RNA variants (SNV, multi-nucleotide variant, insertion, deletion, cryptic exon, intron retention, exon skipping / truncation, fusion gene, circular RNA, and unannotated intergenic isoforms). Second, it integrates RNA and DNA variants to predict the splicing consequences of somatic mutations. Third, it translates full-length RNA sequences and annotates each amino acid with underlying RNA and DNA variants. We have also developed a genome and transcriptome variation graph builder in Exacto to generate synthetic tumor and matched normal genomes as well as tumor transcriptomes. To perform a comprehensive benchmark study for mutant peptide identification, we developed VSTOL and Nexus. VSTOL introduces a new framework, called Occam's Variant Grammar, to characterize DNA and RNA variants in a unified representation for existing variant callers. Nexus is a Nextflow suite that runs over 50 tools for neoantigen discovery. Using synthetic samples generated from the variation graphs, we benchmarked somatic DNA variant calling with Exacto, ClairS, Nanomonsv, Savana, Severus, and SVision-pro. Exacto achieved the highest or tied-highest recall for all simulated variant types (SNV, deletion, insertion = 1.000; translocation and inversion = 0.983). It outperformed the next-best tools by substantial margins: for translocations and inversions, Exacto achieved 0.688 precision versus 0.205 for Savana (both with 0.983 recall); for deletions, 0.945 (Exacto) precision compared to 0.761 (Savana) with both methods obtaining 1.000 recall; and for insertions, 1.000 (Exacto) recall versus 0.901 (Nanomonsv) with both delivering 1.000 precision. Given this performance, we expect Exacto will become the standard method for discovery of immunogenic neoantigens using long-read DNA and RNA sequencing.
利益披露 Disclosure
J. Lee, None..
M. J. Sambade, None.
J. Wang,
Oxford Nanopore Technologies Travel.
A. Rubinsteyn,
Pathfinder Oncology Other, Consulting.
Decade Bio Other, Consulting.
B. G. Vincent,
Pathfinder Oncology Other, Consulting.