LBPO.BCS02 · 生物信息与计算 · Late-Breaking
RNA1-DA:一种用于正向与反向转化的领域自适应RNA基础模型
RNA1-DA: A domain-adaptive RNA foundation model for forward and reverse translation
作者与单位 Authors & Affiliations
摘要 Abstract
中文摘要
引言。临床前癌症模型与患者肿瘤在细胞组成、分子特征和环境背景上存在差异。为实现临床与临床前系统之间的转化,我们开发了一种新的基础模型RNA1-DA,可在肿瘤、细胞系、类器官和异种移植样本之间实现领域自适应。我们证明RNA1-DA能够支持关键的转化研究任务,包括分子亚型迁移、临床前模型选择和药物反应预测。
方法。我们此前描述了RNA1——一种基于transformer的RNA表达基础模型,采用自监督和多任务训练,在182,383份批量RNA-seq癌症样本上进行训练。此处我们开发了RNA1-DA,它扩展了RNA1,通过以下方式实现临床与临床前样本的联合整合:(a) 一个从肿瘤样本中反卷积癌细胞表达的层,以及 (b) 一个使用对抗性自编码器整合跨样本类型RNA1嵌入的领域自适应层。我们建立了一个系统化的评估框架,通过测量疾病身份、分子亚型、癌症驱动基因生物学和药物反应在各系统间的保持与迁移,来评估RNA1-DA嵌入中的临床-临床前对齐。
结果。使用RNA1-DA,我们整合了30,810份肿瘤组织、94,973份细胞系、714份类器官、1290份细胞系来源异种移植和2526份患者来源异种移植样本。RNA1-DA在各系统间对齐了关键的生物学结构,包括疾病身份、分子亚型和驱动基因改变。基于与临床肿瘤的接近程度,临床前样本实现了62-88%的准确疾病分类,优于对照方法。13种TCGA和61种RNA1衍生的临床分子分型被系统地从临床样本迁移至临床前样本,并显示出与经典标志物和遗传依赖性的一致性;例如,所分配的乳腺癌亚型与经典细胞系亚型注释相符(p=1.8e-8),并与已知遗传依赖性相符(如ESR1、ERBB2、CDK4)。转化模型选择进一步得到一种新颖的转录组-基因组邻居重叠度量的支持,该度量在所评估的全部15种癌症中均显示RNA1-DA嵌入邻居与驱动基因改变邻居之间存在显著对应关系(p<0.05)。此外,RNA1-DA通过在CTRP筛选上的多任务微调,改善了细胞系药物反应预测,其表现显著高于基线方法(Spearman相关中位数0.60对0.35)。总之,这些结果证明了RNA1-DA在统一框架内支持关键转化研究任务的效用,包括分子亚型迁移、临床前模型选择和药物反应预测。
查看英文原文 English abstract
Introduction. Preclinical cancer models and patient tumors differ in cellular composition, molecular profiles, and environmental contexts. To enable translation between clinical and preclinical systems, we developed a new foundation model, RNA1-DA, with domain adaptation between tumor, cell line, organoid, and xenograft samples. We demonstrate that RNA1-DA enables key translational research tasks, including molecular subtype transfer, preclinical model selection, and drug response prediction.
Methods. We previously described RNA1 - a transformer-based RNA expression foundation model trained on 182,383 bulk RNA-seq cancer samples with self-supervised and multi-task training. Here we develop RNA1-DA, which extends RNA1 to enable the joint integration of clinical and preclinical samples using (a) a layer to deconvolve cancer cell expression from tumor samples, and (b) a domain adaptation layer using an adversarial autoencoder to integrate RNA1 embeddings across sample types. We developed a systematic evaluation framework to assess clinical-preclinical alignment in RNA1-DA embeddings by measuring the preservation and transfer of disease identity, molecular subtypes, cancer driver gene biology, and drug response across systems.
Results. Using RNA1-DA, we integrated 30,810 tumor tissue, 94,973 cell line, 714 organoid, 1290 cell line-derived xenograft, and 2526 patient-derived xenograft samples. RNA1-DA aligned key biological structure across systems, including disease identity, molecular subtypes, and driver gene alterations. Preclinical samples achieved 62-88% accurate disease classification based on proximity to clinical tumors, outperforming comparator methods. Thirteen TCGA and 61 RNA1-derived clinical molecular subtypings were systematically transferred from clinical to preclinical samples and showed concordance with canonical markers and genetic dependencies; for example, assigned breast cancer subtypes matched canonical cell line subtype annotations (p=1.8e-8) and known genetic dependencies (e.g., ESR1, ERBB2, CDK4). Translational model selection was further supported by a novel transcriptomic-genomic neighbor overlap metric, which demonstrated significant correspondence (p<0.05) between RNA1-DA embedding neighbors and driver gene alteration neighbors across all 15 cancers evaluated. In addition, RNA1-DA enabled improved cell line drug response prediction through multi-task fine-tuning on CTRP screens, achieving substantially higher performance than baseline methods (median Spearman correlation 0.60 vs 0.35). Together, these results demonstrate the utility of RNA1-DA in supporting key translational research tasks within a unified framework, including molecular subtype transfer, preclinical model selection, and drug response prediction.
利益披露 Disclosure
E. O'Brien,
Seres Therapeutics Employment, Stock, Patent.
M. Kukiełka, None..
A. Cupriak, None..
R. Ronen, None..
J. Dutkowski, None.