PO.BCS02.06 · 生物信息与计算

解读 PLM 用于癌症发现:高注意力热点预测致病突变位置和新型药物结合位点

Interpreting PLMs for cancer discovery: High attention hotspots predict pathogenic mutation positions and novel drug binding sites

海报缩略图:解读 PLM 用于癌症发现:高注意力热点预测致病突变位置和新型药物结合位点
编号 4216 展板 12 时间 4/21 09:00–12:00 区域 Section 5 主讲 Sophia Pribus, BS
分会场 Machine Learning Approaches for Cancer Prediction
查看 PDF 下载 PDF 🔒 查看 / 下载完整 PDF 需登录并开通下载套餐 · 查看套餐 / 开通 AACR 官方页面

作者与单位 Authors & Affiliations

Sophia J. Pribus1, Gowri Nayar2, Russ Altman3

1Stanford University School of Medicine, Stanford, CA,2Biomedical Informatics, Stanford University, Stanford, CA,3Bioengineering, Stanford University, Stanford, CA

摘要 Abstract

中文摘要
计算蛋白质组学已彻底革新了癌症研究,为有针对性的实验探索提供指导,以加速基于蛋白质的机制发现。蛋白质语言模型(PLM)实现了可扩展、资源高效的研究;通过仅在初级蛋白质序列上进行大规模训练,PLM 生成蛋白质结构的向量表示,已被证明可捕捉生化和结构特性。PLM 的核心组成部分是注意力机制,它专门在注意力矩阵中捕捉蛋白质序列中的长程相互作用。利用 Evolutionary Scale Modelling 2(ESM-2)PLM 生成的、此前未被探索的注意力矩阵,我们开发了一种新方法来识别高注意力(HA)残基,即 ESM-2 在训练早期赋予最多注意力的特定残基。我们发现,HA 残基在整个人类蛋白质组中与生物学功能存在可解释的联系,包括与活性位点的邻近性以及在蛋白质家族中的保守性。我们进一步使用 AlphaMissense 致病性预测和 TCGA 标记的致病变异位置,确定 HA 残基可预测具有高致病风险的蛋白质区域。最后,我们探索了 HA 残基在新型结合位点发现方面的效用,这是癌症研究中的一个开放性挑战。使用 Uniprot 和 Biolip 注释,我们确认 HA 残基在空间上邻近于此前已知的结合位点。随后,我们使用 SiteMap 预测已注释和未注释蛋白质中 HA 残基区域的可结合性。我们识别出多个癌症蛋白质实例,其中 HA 残基识别出此前未被发现的具有高可结合性、从而具有潜在新型治疗价值的区域。总之,我们的工作证明了 PLM 表示的生物学可解释性,并提供了一种有价值的方法,用于为有针对性的生物医学研究优先排序功能相关的蛋白质残基。
查看英文原文 English abstract
Computational proteomics has revolutionized cancer research, guiding targeted experimental exploration to accelerate protein-based mechanistic discovery. Protein Language Models (PLMs) enable scalable, resource-efficient study; through large-scale training on only primary protein sequences, PLMs generate vector representations of protein structure that have been shown to capture biochemical and structural properties. A core component of PLMs is the attention mechanism, which specifically captures long-range interactions across a protein sequence in attention matrices. Using the previously unexplored attention matrices generated by the Evolutionary Scale Modelling 2 (ESM-2) PLM, we developed a novel method to identify High Attention (HA) residues, the specific residues that ESM-2 assigns the most attention to early in training. We found that HA residues had interpretable links to biological function across the human proteome, including proximity to active sites and conservation across protein families. We further used AlphaMissense pathogenicity predictions and TCGA-labeled pathogenic variant positions to determine that HA residues predict protein regions with high pathogenic risk. Finally, we explored the utility of HA residues for novel binding site discovery, an open challenge in cancer research. Using Uniprot and Biolip annotations, we confirmed that HA residues were spatially proximal to previously-known binding sites. We then used SiteMap to predict the bindability of HA residue regions in both annotated and unannotated proteins. We identified multiple cancer protein examples where HA residues identified regions with previously undiscovered high bindability and thus potential novel therapeutic utility. In summary, our work demonstrates the biological interpretability of PLM representations and offers a valuable method to prioritize functionally relevant protein residues for targeted biomedical research.
利益披露 Disclosure
S. J. Pribus, None.. G. Nayar, None.. R. Altman, None.

← 返回 AACR 2026 检索