PO.BCS02.02 · 生物信息与计算

VGL:整合组织病理学与基因表达的视觉-基因-语言多模态LLM用于肺癌细胞类型分类

VGL: Vision-Gene-Language multimodal LLM integrating histopathology and gene expression for cell type classification in lung cancer

海报缩略图:VGL:整合组织病理学与基因表达的视觉-基因-语言多模态LLM用于肺癌细胞类型分类
编号 2754 展板 18 时间 4/20 02:00–05:00 区域 Section 3 主讲 Haenara Shin
分会场 Large Language Models in the Clinic
查看 PDF 下载 PDF 🔒 查看 / 下载完整 PDF 需登录并开通下载套餐 · 查看套餐 / 开通 AACR 官方页面

作者与单位 Authors & Affiliations

Haenara Shin, Dongjoo Lee, Jeongbin Park, Hongyoon Choi

Portrai, Inc., Seoul, Korea, Republic of

摘要 Abstract

中文摘要
背景 理解肿瘤微环境需要能够跨分子和空间模态解析细胞异质性的模型。随着空间转录组学、单细胞RNA-seq和高分辨率组织病理学成像的扩展,人们需要一个能够联合解读基因表达、空间背景和视觉组织特征的统一基础模型。我们开发了一个多模态大语言模型(LLM),将这些模态整合到单一自适应框架中,能够处理异质输入——包括基因表达概况、空间转录组学spot、单细胞测量和组织学切片——同时生成协调统一的输出,如基因、细胞类型和图像衍生的描述符。 方法 我们在视觉-基因-语言(VGL)框架内构建了一个多模态LLM,整合了基因表达、组织学图像和生物学语言表征。该模型基于MedGemma-4b-it,并使用QLoRA进行微调以实现参数高效的训练。训练使用了520万个多模态样本,即来自非小细胞肺癌的H&E切片与高变基因表达概况配对,跨scRNA-seq和空间转录组学平台(Visium、Xenium)共计1,745,240个细胞和空间spot。该模型采用多任务学习进行训练,涵盖image-to-gene、gene-to-cell type和cell type-to-gene目标的五项标准任务,并在4块H100 GPU上采用任务特异性比例调度。我们评估了从基因表达概况对12种主要免疫和基质细胞类型进行细胞类型分类的性能。 结果 训练后的VGL模型在留出测试集(n=20,764)上从基因表达概况预测细胞类型达到70.07%的准确率,相比之下朴素预训练模型的准确率为16.32%,提升了4.3倍。验证性能相似(69.85%准确率,n=41,529),表明具有稳健的泛化能力。这些提升证明了联合利用单细胞、空间转录组学和组织学信息的多模态LLM的价值。通过跨模态学习和掩码,该模型学习到的基因和细胞类型嵌入能够跨数据平台和空间背景泛化,即使仅有部分模态可用时也能产生生物学上一致的输出。 结论 我们引入了一个基于多模态LLM的空间基础模型VGL,它将单细胞RNA-seq、空间转录组学和组织病理学成像统一到一个模态无关的框架中。细胞类型分类的改进凸显了该模型捕捉和推理跨模态生物学结构的能力。该框架为解读异质分子和成像数据、并实现可扩展肿瘤微环境分析的空间AI系统奠定了基础。
查看英文原文 English abstract
Background Understanding the tumor microenvironment requires models that resolve cellular heterogeneity across molecular and spatial modalities. With the expansion of spatial transcriptomics, single-cell RNA-seq, and high-resolution histopathology imaging, there is a need for a unified foundation model that jointly interprets gene expression, spatial context, and visual tissue features. We developed a multimodal large language model (LLM) that integrates these modalities into a single adaptive framework handling heterogeneous inputs-including gene expression profiles, spatial transcriptomics spots, single-cell measurements, and histology patches-while generating harmonized outputs such as genes, cell types, and image-derived descriptors. Method We built a multimodal LLM within a Vision-Gene-Language (VGL) framework that integrates gene expression, histology images, and biological language representations. The model is based on MedGemma-4b-it and was fine-tuned using QLoRA for parameter-efficient training. Training used 5.2 million multimodal samples of H&E patches paired with highly variable gene expression profiles from non-small cell lung cancer, totaling 1,745,240 cells and spatial spots across scRNA-seq and spatial transcriptomics platforms (Visium, Xenium). The model was trained using multi-task learning across five canonical tasks spanning image-to-gene, gene-to-cell type, and cell type-to-gene objectives with task-specific ratio scheduling on 4 x H100 GPUs. We evaluated performance on cell type classification from gene expression profiles across 12 major immune and stromal cell types. Results The trained VGL model achieved 70.07% accuracy on the held-out test set (n=20,764) for predicting cell types from gene expression profiles, compared to 16.32% accuracy for the naive pre-trained model, a 4.3-fold improvement. Validation performance was similar (69.85% accuracy, n=41,529), indicating robust generalization. These gains demonstrate the value of a multimodal LLM that jointly leverages single-cell, spatial transcriptomics, and histology information. Through cross-modal learning and masking, the model learned gene and cell type embeddings that generalized across data platforms and spatial contexts and produced biologically consistent outputs even when only a subset of modalities was available. Conclusion We introduce a multimodal LLM-based spatial foundation model, VGL, that unifies single-cell RNA-seq, spatial transcriptomics, and histopathology imaging into a modality-agnostic framework. The improvement in cell type classification highlights the model's ability to capture and reason over cross-modal biological structure. This framework lays the groundwork for spatial AI systems that interpret heterogeneous molecular and imaging data and enable scalable tumor microenvironment profiling.
利益披露 Disclosure
H. Shin, Portrai, Inc. Employment. D. Lee, Portrai, Inc. Employment. J. Park, Portrai, Inc. Employment. H. Choi, Portrai, Inc. Stock. Institute of Radiation Medicine, Medical Research Center, Seoul National University, Seoul, Republic of Korea Employment. Department of Nuclear Medicine, Seoul National University Hospital, Seoul, Republic of Korea Employment. Department of Nuclear Medicine, Seoul National University College of Medicine, Seoul, Republic of Korea Employment.

← 返回 AACR 2026 检索