PO.BCS01.16 · 生物信息与计算
一个全面的、由LLM赋能的药效学生物标志物资源库,用于加速癌症药物研发
A comprehensive LLM-enabled pharmacodynamic biomarker resource to accelerate cancer drug development
作者与单位 Authors & Affiliations
摘要 Abstract
中文摘要
引言:肿瘤药物研发中的一大挑战是确认治疗药物在患者肿瘤中的靶点结合。药效学(PD)生物标志物提供了这一关联,可实现对靶点有效性、药物活性和合理剂量选择的早期评估。然而,将药物靶点与经验证的PD生物标志物相连接的系统性资源仍然有限。为此,我们开发了一个全面的数据集和分析框架,以识别并优先排序涉及癌症生物学的九大主要靶点类别中的靶点特异性PD生物标志物。
材料与方法:我们从多个基因组学和药理学资源中整理了九大靶点类别的候选生物标志物:转录因子/辅因子、激酶、磷酸酶、泛素连接酶、去泛素化酶、乙酰转移酶、去乙酰化酶、甲基转移酶和去甲基化酶。源数据库对转录因子等功能类别往往采用宽泛的定义。为提高准确性并减少注释假象,我们交叉参考PFAM、酶分类(Enzyme Classification)和PDB数据库以完善蛋白分类。利用canSAR相互作用组,我们识别了有实验证据支持的直接靶点-生物标志物相互作用。为捕获情境特异性的转录生物标志物,我们使用TCGA、TARGET和GTEx数据集计算了队列特异性相关性。最后,我们采用基于LLM的事实核查代理,从抗体注册库(Antibody Registry)中提取并统一抗体注释,重点关注具有可测量底物修饰的酶靶点。
结果:从2,900个靶点和100,000个相互作用中,我们提出了涉及2,100多个潜在药物靶点的73,000个高置信度靶点-生物标志物关系。其中,67%代表转录因子-基因相互作用,33%代表酶-底物相互作用。为2,800多个候选生物标志物识别了商品化抗体,为实验验证提供了支持。所得数据集涵盖了19种癌症类型中排名前20的预测靶点的60%以上。我们在canSAR平台中提供所有数据,并可从canSAR-PD下载。
讨论:该资源为肿瘤学中PD生物标志物的发现提供了一个系统性框架。通过整合经整理的分子相互作用与LLM衍生的抗体注释,它能够对靶点结合和药物活性进行稳健的评估。该数据集为开发PD生物标志物以指导剂量选择、监测反应和加速癌症药物研发奠定了基础。
查看英文原文 English abstract
Introduction: A major challenge in oncology drug development is confirming target engagement of therapeutics in patient tumors. Pharmacodynamic (PD) biomarkers provide this link, enabling early assessment of target validity, drug activity and rational dose selection. However, systematic resources connecting drug targets to validated PD biomarkers remain limited. To address this, we developed a comprehensive dataset and analytic framework to identify and prioritize target-specific PD biomarkers across nine major target classes implicated in cancer biology.
Materials and Methods: We curated biomarker candidates from multiple genomic and pharmacologic resources for nine target classes: transcription factors/cofactors, kinases, phosphatases, ubiquitin ligases, deubiquitinases, acetyltransferases, deacetylases, methyltransferases, and demethylases. Source databases often apply broad definitions to functional classes such as transcription factors. To improve accuracy and reduce annotation artifacts, we cross-referenced PFAM, Enzyme Classification, and PDB databases to refine protein classifications. Using the canSAR interactome, we identified direct target-biomarker interactions supported by experimental evidence. To capture context-specific transcriptional biomarkers, we computed cohort-specific correlations using TCGA, TARGET, and GTEx datasets. Finally, we employ LLM-based fact-checking agent to extract and harmonize antibody annotations from the Antibody Registry, focusing on enzyme targets with measurable substrate modifications.
Results: From 2,900 targets and 100,000 interactions, we propose 73,000 high-confidence target-biomarker relationships involving over 2,100 potential drug targets. Of these, 67% represent transcription factor-gene interactions and 33% enzyme-substrate interactions. Commercial antibodies were identified for over 2,800 biomarker candidates, supporting experimental validation. The resulting dataset covers more than 60% of the top 20 predicted targets across 19 cancer types. We provide all the data in our canSAR platform and as a download from canSAR-PD.
Discussion: This resource provides a systematic framework for PD biomarker discovery in oncology. By integrating curated molecular interactions with LLM-derived antibody annotations, it enables robust evaluation of target engagement and drug activity. The dataset establishes a foundation for developing PD biomarkers to guide dose selection, monitor response, and accelerate cancer drug development.
利益披露 Disclosure
Y. Yang, None..
L. Zhao, None..
S. Orouji, None..
Y. Zhu, None..
R. Johnson, None..
D. Maxwell, None..
K. Brickey, None..
B. Al-Lazikani, None.