PO.BCS01.10 · 生物信息与计算
hla2vec:利用概率性交叉反应组构建HLA嵌入空间
hla2vec: Developing a HLA embedding space using probabilistic cross-reactive groups
作者与单位 Authors & Affiliations
摘要 Abstract
中文摘要
主要组织相容性复合体(MHC)中编码人类白细胞抗原(HLA)蛋白复合物的基因是基因组中最具多样性的基因之一。虽然HLA-A*02:01等某些等位基因在近25%的人群中出现,但IMGT/HLA等数据库报告了近2万个变体。在癌症人群中,这产生了在采样人群中具有独特性的个体的长尾分布。The Cancer Genome Atlas(TCGA)的HLA I类分型覆盖了33种癌症类型的近1万名个体,产生了超过300种HLA等位基因。在该队列中,HLA-A*03、HLA-A*33和HLA-A*2等等位基因已在多种癌症中与预后和免疫治疗反应相关联。并非所有HLA等位基因都不同,一些等位基因表现出相似的结合模式,可将其归入交叉反应组(CREGs)。虽然这一概念传统上被记录为一组离散的已记载等位基因,但从大规模实验中,我们可以开始将其视为代表概率性CREGs(pCREGs)的概率空间。一个pCREG中的两个等位基因可能仅结合相同肽段的某一比例(例如90%)。以此为基础,我们可以构建一个嵌入空间,其中不同HLA等位基因之间的距离相对于它们的pCREG关联性。我们开发了一种名为HLA2Vec的方法,该方法在BigMHC数据集中累积的1500万次HLA/肽结合实验上进行了训练。此外,我们开发了一个基准来描述新模型识别CREGs的能力。我们将该方法与其他蛋白质建模语言(PML)——如Evolutionary Scale Modeling(ESM)——在识别可能属于同一pCREG的HLA对方面的能力进行了比较,发现HLA2Vec在使用少得多的参数的同时,准确度要高得多。将HLA等位基因置于该嵌入空间后,我们能够基于pCREG状态将拥有近乎独特HLA等位基因的较小个体群聚为更大的亚型,从而将可能具有相似免疫HLA结合模式的患者汇集在一起以进行统计分析。这一分析已应用于TCGA数据,将罕见等位基因归入更大的反应组。这项研究的下一步是探究这种HLA嵌入方法是否也能为构建T细胞受体(TCR)的类似嵌入空间提供洞见,并改进HLA pCREG条件结合预测。
查看英文原文 English abstract
The genes in the major histocompatibility complex (MHC) that code for the Human Leukocyte Antigen (HLA) protein complex are some of the most diverse genes in the genome. And while some of the alleles such as HLA-A*02:01 are seen in almost 25% of the population, databases such as IMGT/HLA report nearly 20 thousand variants. In cancer populations, this produces a long tail of individuals who are unique within a sampled population. HLA class I typing for The Cancer Genome Atlas (TCGA) covered almost 10K individuals across 33 cancer types, and resulted in over 300 HLA alleles. Within this cohort alleles such as HLA-A*03, HLA-A*33, and HLA-A*2 have been associated with prognosis and immunotherapy response in various cancers. Not all HLA alleles are different, with some alleles observing similar binding patterns that can place them in cross-reactive groups (CREGs). And while this concept has traditionally been documented as a discrete set of documented alleles, from large scale experiments we can start to think of this as a probabilistic space representing probabilistic CREGs (pCREGs). Two alleles in a pCREG may only bind to some fraction, for example 90%, of the same peptides. Using this as a basis we can develop an embedding space where the distance between different HLA alleles is relative to their pCREG association. We have developed a method called HLA2Vec, which has been trained on the 15 million HLA/peptide binding experiments accumulated in the BigMHC dataset. In addition, we developed a benchmark to describe a new model's ability to identify CREGs. We have compared this method to other Protein Modelling Languages (PML), such as Evolutionary Scale Modeling (ESM) in their ability to identify pairs of HLAs that are likely to be in the same pCREG, and found the HLA2Vec to be much more accurate, while using many fewer parameters. With HLA alleles placed in this embedding space, we are able to cluster together smaller groups of individuals with near unique HLA alleles into larger subtypes, based on pCREG status, bringing together patients likely to have similar immune HLA binding patterns for statistical analysis. This analysis has been applied to TCGA data to group rare alleles into larger response groups. The next steps of this research are to see if this method for HLA embedding may also provide insights for developing similar embedding spaces for T Cell Receptor (TCR), and improve HLA pCREG conditional binding prediction.
利益披露 Disclosure
I. Quesada, None..
J. Tagle, None..
K. Ellrott, None.