PO.BCS01.01 · 生物信息与计算
使用PEPMatch和CEDAR鉴定癌症新表位的非经典来源
Identifying noncanonical sources of cancer neoepitopes using PEPMatch and CEDAR
作者与单位 Authors & Affiliations
摘要 Abstract
中文摘要
新出现的证据表明,免疫原性癌症抗原可能来源于"暗基因组",其包含非编码区、替代开放阅读框(ORF)和非翻译区(UTR),这些区域传统上被大多数新表位预测流程排除在外。因此,对这些非经典抗原进行系统性鉴定和验证在技术上一直具有挑战性。为了弥补这一空白,我们利用PEPMatch(一种高通量肽段搜索工具),并将其与癌症表位数据库和分析资源(CEDAR)相整合,以系统性地鉴定经实验验证的新表位的非经典来源。
我们汇编了七个全面的人类数据库,涵盖经典和非经典蛋白来源:UniProtKB(含异构体的人类参考蛋白质组)、UniParc(UniProt蛋白档案)、Ensembl经典蛋白(ENSP)、经验证的非经典ORF(ncORF)、互补DNA(cDNA)和非编码RNA(ncRNA)的三框翻译,以及包括UTR和内含子的完整基因序列(ENSG)的六框翻译。使用PEPMatch的精确匹配算法,我们针对所有七个数据库对在40,800个T细胞测定中测试的28,601个CEDAR新肽段进行了搜索。
为了验证我们的流程,我们首先分析了通过质谱发现的840个MHC I类隐性新肽段,使用序贯归属搜索实现了98.33%的总映射成功率,该搜索揭示了非经典来源的富集(ncORF:14.76%,cDNA:41.55%,ncRNA:4.76%,ENSG:1.67%)。在这些隐性新表位中,5.60%在非经典来源之前于人类参考蛋白质组中被发现,28.93%在UniParc数据库中被发现。将此方法应用于所有28,601个CEDAR新肽段,揭示出1.91%来源于非经典来源。随后我们聚焦于6,394个经T细胞测定确认具有免疫原性的阳性新表位的相关子集。值得注意的是,与阴性肽段相比,这些具有免疫原性的阳性新表位中有显著比例映射到非经典数据库(ncORF:0.1% vs 0.0%,cDNA:8.8% vs 2.6%,ncRNA:1.4% vs 0.6%,ENSG:3.4% vs 1.5%),提示"暗基因组"是可靶向新表位的来源。使用随机打乱的肽段对照验证了这种映射方法的特异性,其产生的虚假匹配<1%。
这些发现表明,只要有经过适当整理的数据库来源,PEPMatch就能大规模成功鉴定癌症新表位的非经典基因组起源。CEDAR正在实施这一注释流程,为研究人员提供暗基因组抗原的映射,从而能够验证和发现用于免疫治疗开发的非常规靶标。这项工作通过系统性地编目来自此前被忽视的基因组区域的表位,扩展了癌症免疫治疗的可靶向图景。
查看英文原文 English abstract
Emerging evidence suggests that immunogenic cancer antigens may originate from the "dark genome", which comprises non-coding regions, alternative open reading frames (ORF), and untranslated regions (UTR), traditionally excluded from most neoepitope prediction pipelines. Consequently, systematic identification and validation of these noncanonical antigens has been technically challenging. To address this gap, we leveraged PEPMatch, a high-throughput peptide search tool, integrated with the Cancer Epitope Database and Analysis Resource (CEDAR) to systematically identify noncanonical sources of experimentally validated neoepitopes.
We compiled seven comprehensive human databases encompassing both canonical and noncanonical protein sources: UniProtKB (human reference proteome with isoforms), UniParc (UniProt protein archive), Ensembl canonical proteins (ENSP), validated non-canonical ORFs (ncORF), and three-frame translations of complementary DNAs (cDNA) and non-coding RNAs (ncRNA), and six-frame translations of full gene sequences, including UTRs and introns (ENSG). Using PEPMatch's exact matching algorithm, we performed searches of 28,601 CEDAR neopeptides tested in 40,800 T cell assays against all seven databases.
To validate our pipeline, we first analyzed 840 MHC class I cryptic neopeptides found by mass spectrometry, achieving a total mapping success of 98.33% using a sequential attribution search that revealed an enrichment in noncanonical sources (ncORF: 14.76%, cDNA: 41.55%, ncRNA: 4.76%, ENSG: 1.67%). Out of these cryptic neoepitopes, 5.60% were found in the human reference proteome and 28.93% in the UniParc database prior to the noncanonical sources. Applying this methodology to all 28,601 CEDAR's neopeptides revealed that 1.91% originated from noncanonical sources. We then focused on the relevant subset of 6,394 positive neoepitopes with confirmed immunogenicity from T cell assays. Notably, a significant percentage of these immunogenic positive neoepitopes mapped to noncanonical databases compared with the negative peptides (ncORF: 0.1% vs 0.0%, cDNA: 8.8% vs 2.6%, ncRNA: 1.4% vs 0.6%, ENSG: 3.4% vs 1.5%), suggesting that the 'dark genome' is a source of targetable neoepitopes. The specificity of this mapping approach was validated using randomly shuffled peptide controls, which yielded <1% spurious matches.
These findings demonstrate that PEPMatch successfully identifies noncanonical genomic origins of cancer neoepitopes at scale given the appropriately curated database sources. CEDAR is implementing this annotation pipeline to provide researchers with mappings for dark genome antigens, enabling validation and discovery of unconventional targets for immunotherapy development. This work expands the targetable landscape of cancer immunotherapy by systematically cataloging epitopes from previously overlooked genomic regions.
利益披露 Disclosure
D. Marrama, None..
I. Carri, None..
N. Blazeska, None..
R. Vita, None..
M. Nielsen, None..
A. Sette, None..
B. Peters, None.