PO.BCS02.06 · 生物信息与计算

生成式基因组学准确预测癌症基因表达

Generative genomics accurately predicts cancer gene expression

海报缩略图:生成式基因组学准确预测癌症基因表达
编号 4213 展板 9 时间 4/21 09:00–12:00 区域 Section 5 主讲 Alexander Abbas, PhD
分会场 Machine Learning Approaches for Cancer Prediction
查看 PDF 下载 PDF 🔒 查看 / 下载完整 PDF 需登录并开通下载套餐 · 查看套餐 / 开通 AACR 官方页面

作者与单位 Authors & Affiliations

Gregory Koytiger1, Alice M. Walsh1, Vaishali Marar1, Kayla A. Johnson1, Max Highsmith1, Alexander R. Abbas1, Andrew Stirn2, Ariel Brumbaugh1, Alex David1, Darren Hui1, Jeffrey Kahn1, Sheng-Yong Niu1, Liza J. Ray1, Candace Savonen1, Stein Setvik1, Jeffrey T. Leek1, Robert K. Bradley1

1Synthesize Bio, Seattle, WA,2Variational Bio, Seattle, WA

摘要 Abstract

中文摘要
能够预测实验结果的 AI 模型可以通过规避实验室实验和临床试验的根本性限制来加速生物医学研究。我们开发了 GEM-1(Generate Expression Model-1),一种生成式基因组学框架,可对真实世界基因表达实验的多样性进行建模,并准确预测未来的实验结果。 我们使用来自 NCBI Sequence Read Archive 中 24,715 个数据集的 470,691 个 bulk RNA-seq 样本训练 GEM-1,涵盖多种组织、疾病和超过 18,000 种不同的扰动。一个自动化的元数据代理使用大型语言模型对碎片化的实验描述进行了统一。GEM-1 采用一种深度潜变量模型,将实验元数据划分为生物学、技术和扰动成分,利用预训练的基础模型嵌入以实现对新型扰动的泛化。 在训练后存入的留出数据上进行测试,GEM-1 对先前观察到的情境达到了伪重复水平的准确度(样本间基因排名的 pearson 相关系数 r_gene 为 0.65-0.75),并对完全新型的遗传扰动(r_gene = 0.58-0.63)和化学扰动(r_gene = 0.52-0.68)保持了强劲的性能。该模型正确预测了新型遗传扰动和化学扰动中分别 63% 和 70% 的富集基因集。我们将 GEM-1 扩展到单细胞数据,使用 4,150 万个细胞,在细胞类型注释方面达到了与已有模型相当的性能,同时实现了可解释的生物学特征推断。 我们通过生成准确重现关键生物学现象的合成队列,展示了其临床实用性:SYNTH-TEx(5,300 个健康组织样本,匹配 GTEx 模式)、SYNTH-interferon(200 个样本,正确建模狼疮干扰素失调)以及 SYNTH-cancer(10,523 个样本,展现出已知的癌症分子特征)。我们展示了 GEM-1 能够对特定患者样本模拟新型扰动,我们将这一能力称为“参考条件化”(reference conditioning)。 这一方法代表了朝向能够在进行物理实验之前预测实验结果的 AI 系统的重大进步,可绕过实验速度和临床试验招募方面的根本性限制,并有可能彻底革新药物开发和个性化医疗。
查看英文原文 English abstract
AI models capable of predicting experimental outcomes could accelerate biomedical research by circumventing fundamental constraints of laboratory experimentation and clinical trials. We developed GEM-1 (Generate Expression Model-1), a generative genomics framework that models the diversity of real-world gene expression experiments and accurately predicts future experimental results. We trained GEM-1 using 470,691 bulk RNA-seq samples from 24,715 datasets in the NCBI Sequence Read Archive, spanning diverse tissues, diseases, and over 18,000 distinct perturbations. An automated metadata agent harmonized fragmented experimental descriptions using large language models. GEM-1 employs a deep latent variable model that partitions experimental metadata into biological, technical, and perturbational components, using pretrained foundation model embeddings to enable generalization to novel perturbations. Testing on holdout data deposited after training, GEM-1 achieved pseudoreplicate-level accuracy (pearson correlation of gene rank across samples, r_gene, of 0.65-0.75) for previously observed contexts and maintained strong performance for completely novel genetic (r_gene = 0.58-0.63) and chemical perturbations (r_gene = 0.52-0.68). The model correctly predicted 63% and 70% of enriched gene sets for novel genetic and chemical perturbations, respectively. We extended GEM-1 to single-cell data using 41.5 million cells, achieving comparable performance to established models for cell type annotation while enabling interpretable biological feature inference. We demonstrated clinical utility by generating synthetic cohorts that accurately recapitulated key biological phenomena: SYNTH-TEx (5,300 healthy tissue samples matching GTEx patterns), SYNTH-interferon (200 samples correctly modeling lupus interferon dysregulation), and SYNTH-cancer (10,523 samples exhibiting known molecular features of cancer). We show that GEM-1 can simulate novel perturbations on specific patient samples, an ability we term "reference conditioning". This approach represents a significant advance toward AI systems that can predict experimental outcomes before physical experiments are conducted, shortcutting fundamental limitations in experimental speed and clinical trial recruitment and potentially revolutionizing drug development and personalized medicine.
利益披露 Disclosure
G. Koytiger, None.. A. M. Walsh, None.. V. Marar, None.. K. A. Johnson, None.. M. Highsmith, None.. A. R. Abbas, None.. A. Stirn, None.. A. Brumbaugh, None.. A. David, None.. D. Hui, None.. J. Kahn, None.. S. Niu, None.. L. J. Ray, None.. C. Savonen, None.. S. Setvik, None.. J. T. Leek, None.. R. K. Bradley, None.

← 返回 AACR 2026 检索