PO.BCS01.09 · 生物信息与计算
CAB:一种置信度感知的序列模型,用于高通量预测抗体类抗癌治疗药物中突变驱动的结合亲和力变化
CAB: A confidence-aware sequence model enabling high-throughput prediction of mutation-driven binding affinity change in antibody-based cancer therapeutics
该海报暂无可下载的资料
AACR 官方页面
作者与单位 Authors & Affiliations
摘要 Abstract
中文摘要
背景:抗体-抗原(Ab-Ag)亲和力的计算预测对于设计肿瘤靶向治疗药物、实现快速的计算机模拟亲和力成熟以及优化双特异性抗体和ADC等形式的分子日益重要。然而,现有模型几乎完全在具有可测量亲和力的阳性结合案例上训练,从而产生阳性案例偏倚,限制了它们识别真正结合丧失事件的能力,并常常对实验上不结合的Ab-Ag对给出不可靠的亲和力预测。
方法:为避免预测突变蛋白准确结构的困难,我们采用了序列到功能的框架,并对仅基于序列的蛋白质语言模型(如ESM2、DPLM)进行微调,用于突变水平的亲和力预测。我们纳入了AbAgym数据集(约335k条测量数据),以解决弱结合和非结合案例代表性不足的问题。为更好地区分结合与非结合状态,我们增加了一个序列对对比学习阶段,将已验证的结合体作为阳性样本,将随机采样的非结合体作为阴性样本。在对比预训练之后,我们联合优化了一个亲和力变化回归头和一个受AlphaFold启发的置信度评分头,以捕捉预测不确定性并标记可能的非结合体。这种结合对比学习与多任务的策略提高了突变水平的灵敏度,并增强了对结合丧失事件的检测能力,而这对抗肿瘤抗体发现至关重要。
结果:CAB实现了对抗体突变文库的高通量计算机模拟筛选,并在病毒和肿瘤相关抗原系统上进行了评估。对于SARS-CoV-2 RBD,CAB快速优先排序了数千个CDR变体,并识别出预测可带来>10倍亲和力提升的高置信度突变,成功重现了已知来自深度突变扫描的重设计。对于HER2和Plexin-B2等肿瘤靶点,CAB同样选出了具有增强预测亲和力的顶级高置信度突变,与实验工程化抗体一致。
讨论:与无论可靠性如何都输出单一亲和力估计值的传统模型不同,CAB的置信度头能够明确识别低置信度、可能不结合的变体,解决了当前预测框架的一个关键局限。这提高了虚拟筛选的可信度,并减少了假阳性的重设计。CAB为在癌症免疫治疗中针对肿瘤相关抗原识别有前景的变体提供了一个高效的计算引擎。
查看英文原文 English abstract
Background: Computational prediction of antibody-antigen (Ab-Ag) affinity is increasinglyimportant for designing tumor-targeting therapeutics, enabling rapid in-silico affinity maturation,and optimization of formats such as bispecifics and ADC. However, existing models are trainedalmost exclusively on positive binding cases with measurable affinity, creating a positive-casebias that limits their ability to recognize true loss-of-binding events and often yields unreliableaffinity predictions for Ab-Ag pairs that are experimentally non-binding.
Method: To avoid challenges in predicting accurate structures for mutated proteins, we adopteda sequence-to-function framework and fine-tuned sequence-only protein language models (e.g.,ESM2, DPLM) for mutation-level affinity prediction. We incorporated the AbAgym dataset(~335k measurements) to address the under-representation of weak and non-binding cases. Tobetter separate binding from non-binding regimes, we added a sequence-pair contrastivelearning stage using validated binders as positives and randomly sampled non-binders asnegatives. After contrastive pretraining, we jointly optimized an affinity-change regression headand an AlphaFold-inspired confidence-score head to capture prediction uncertainty and flaglikely non-binders. This combined contrastive and multi-task strategy improves mutation-levelsensitivity and strengthens detection of loss-of-binding events essential for anti-tumor antibodydiscovery.
Result: CAB enabled high-throughput in-silico screening of antibody mutation libraries and wasevaluated on both viral and tumor-associated antigen systems. For SARS-CoV-2 RBD, CABrapidly prioritized thousands of CDR variants and identified high-confidence mutations predictedto yield >10-fold affinity improvements, successfully recovering redesigns known from deepmutational scanning. For tumor targets such as HER2 and Plexin-B2, CAB similarly selected tophigh-confidence mutations with enhanced predicted affinity, consistent with experimentallyengineered antibodies.
Discussion: Unlike conventional models that output a single affinity estimate regardless ofreliability, CAB's confidence head enables explicit identification of low-confidence, likely non-binding variants, addressing a critical limitation of current prediction frameworks. This improvesthe trustworthiness of virtual screening and reduces false-positive redesigns. CAB provides anefficient computational engine for identification of promising variants against tumor-associatedantigens in cancer immunotherapy.
利益披露 Disclosure
Y. Sun, None..
B. Jiang, None.