PO.BCS01.09 · 生物信息与计算
统一分子结构与细胞形态以增强癌症中的药物-靶点相互作用建模
Unifying molecular structure and cellular morphology to enhance drug-target interaction modeling in cancer
作者与单位 Authors & Affiliations
摘要 Abstract
中文摘要
准确的药物-靶点相互作用(DTI)预测对于加速药物发现和揭示治疗机理至关重要。然而,化合物和蛋白质靶点的庞大组合空间,加上支配其相互作用的复杂非线性关系,带来了重大的实验挑战。现有计算方法严重依赖分子特征,往往忽视扰动诱导的形态学表型,且缺乏将分子性质与响应相连接的统一表征。为解决这些局限,我们提出了一个两阶段对比学习框架,将药物和蛋白质的结构信息与来自Cell Painting高内涵筛选的细胞形态学谱整合到一个统一的多模态嵌入空间中。在第一阶段,独立训练了两个模态特异性对比学习模型:一个基于结构的模型,将药物结构的嵌入与蛋白质序列对齐;一个基于图像的模型,将来自药物和基因敲除扰动的Cell Painting检测的形态学嵌入对齐。在第二阶段,这些模态特异性嵌入通过跨模态对比学习进一步整合,构建一个共享嵌入空间,联合编码药物结构、蛋白质序列及其对应的细胞形态学谱。第一阶段的两个嵌入空间(基于结构和基于图像的模型)有效地使已注释的药物-靶点对更紧密地对齐,同时将非相互作用对推开。这些模型分别产生了0.4和0.3的相关性差异(以余弦相似度衡量)。在最终的统一嵌入空间中,已知DTI对的药物、靶点及其相关形态学嵌入之间出现了连贯的簇,证明了成功的跨模态对齐。在排名靠前的DTI对中,gedatolisib和PIK3CB显示出0.95的相似度,这与该药物作为PI3K/mTOR通路抑制剂的已知活性一致。总体而言,高相似度的DTI对通常涉及芳香族和杂环化合物,其物理化学性质与G蛋白偶联受体(GPCR)(例如CHRM4)和激酶(例如PIK3CB)等靶点的结合偏好密切匹配。这些模式与显示GPCR和激酶配体共享有利于结合的特征性结构特征的研究相一致。相比之下,低相似度的DTI对涉及的分子,其大小、极性或几何形状对代谢酶和血红素结合蛋白等靶点构成挑战。总体而言,本研究凸显了通过对比学习整合分子和形态学表征的前景,为推进DTI建模和精准癌症治疗发现提供了强大的框架。
查看英文原文 English abstract
Accurate drug-target interaction (DTI) prediction is critical for accelerating drug discovery and uncovering therapeutic mechanisms. However, the vast combinatorial space of chemical compounds and protein targets, coupled with the complex nonlinear relationships that govern their interactions, presents significant experimental challenges. Existing computational approaches rely heavily on molecular features, often overlooking perturbation-induced morphological phenotypes and lacking a unified representation that connects molecular properties to responses. To address these limitations, we propose a two-stage contrastive learning framework that integrates structural information of drugs and proteins with cellular morphological profiles derived from the Cell Painting high-content screening into a unified multimodal embedding space. In stage one, two modality-specific contrastive learning models were trained independently: a structure-based model that aligns embeddings of drug structures with protein sequences, and an image-based model that aligns morphological embeddings derived from Cell Painting assays of drug and gene knockout perturbations. In stage two, these modality-specific embeddings were further integrated through cross-modal contrastive learning to construct a shared embedding space that jointly encodes drug structures, protein sequences, and their corresponding cellular morphological profiles. Both embedding spaces from stage 1 (structure- and image-based models) effectively aligned annotated drug-target pairs more closely while pushing non-interacting pairs apart. These models yielded correlation differences of 0.4 and 0.3, respectively, as measured by cosine similarity. In the final unified embedding space, coherent clusters emerged among drugs, targets, and their associated morphological embeddings for known DTI pairs, demonstrating successful cross-modal alignment. Among the top-ranking DTI pairs, gedatolisib and PIK3CB showed a similarity of 0.95, consistent with the drug's known activity as a PI3K/mTOR pathway inhibitor. Overall, high-similarity DTI pairs often involved aromatic and heterocyclic compounds whose physicochemical properties closely matched the binding preferences of targets like G protein-coupled receptors (GPCRs) (e.g., CHRM4 ) and kinases (e.g., PIK3CB ). These patterns align with studies showing that GPCR and kinase ligands share characteristic structural features that facilitate binding. In contrast, low-similarity ones involved molecules whose size, polarity, or geometry posed challenges for targets such as metabolic enzymes and heme-binding proteins. Overall, this study highlights the promise of integrating molecular and morphological representations via contrastive learning, providing a powerful framework for advancing DTI modeling and precision cancer therapeutic discovery.
利益披露 Disclosure
Y. Lai, None.