PO.BCS02.06 · 生物信息与计算
基线外周血 scRNA-seq AI 估计器框架通过分子基础模型和细胞到患者学习预测实体瘤反应和不良事件
Baseline peripheral blood scRNA-seq AI estimator framework predicts solid-tumor response and adverse events via molecular foundation models and cell-to-patient learning
该海报暂无可下载的资料
AACR 官方页面
作者与单位 Authors & Affiliations
摘要 Abstract
中文摘要
引言:从基线血液样本中准确识别可能从治疗中获益以及有不良事件风险的患者,是肿瘤学中一项尚未满足的需求。我们开发了一个转化型 AI 估计器框架,可从实体瘤患者的基线 PBMC 单细胞 RNA-seq 中预测治疗反应、不良事件(AE)和分子特征。我们使用基线 PBMC scRNA-seq 计数(可选细胞类型注释)、预训练的分子基础模型(FM)(scGPT v1.0 [1]、scFoundation [2])、基于 RECIST 标签的细胞到患者多示例学习(MIL)[3](CR/PR 为反应者;SD/PD 为无反应者)来预测反应。我们对下游任务进行微调,应用基于 scVI 的数据增强 [4] 以提高稳定性和泛化能力,并设置了一个将预测与细胞类型和基因程序相关联的可解释性层。
结果:使用接受免疫治疗的实体瘤患者的基线 PBMC scRNA-seq(103 例患者,约 12,000 个细胞/患者),该估计器可跨 FM 主干网络从基线血液中预测治疗反应,并识别出最优的下游架构。我们使用了带有和不带分层注意力的 scFoundation 以及 scGPT。对于不带分层注意力的 scFoundation,我们观察到伴随显著损失下降的学习过程,类别区分的 AUC 略低于 0.8,F1 评分 ≈ 0.78。对于带分层注意力的 scFoundation,我们观察到相似的性能,但准确率略差,AUC ≈ 0.75,F1 ≈ 0.70。对于 scGPT,我们观察到学习过程(损失下降),良好的类别区分,AUC ≈ 0.75,F1 评分略高于 0.7。所有情况下验证评分均略差。我们迄今报告的结果反映了我们的框架的泛化能力,展示出具有竞争力的性能,其差异可归因于分层与非分层聚合。用于提高训练稳定性的数据增强以及初步的 AE 风险建模(优先考虑 ≥3 级事件)显示出令人鼓舞的结果。分子特征用于支持机制性见解并作为解释预测的生物标志物。
结论:基线血液单细胞信号结合 FM 嵌入和细胞到患者学习,可预测实体瘤的治疗反应,并为 AE 分层和特征发现提供路径。通过更大的数据集可以取得改进,该框架也可促进可解释的生物学依据。
参考文献:1.Cui H. 等. Nature Methods, 21:1470-1480 (2024). 2.Hao M. 等. Nature Methods, 21:1481-1491 (2024). 3.Do C 等. Bioinformatics, 41: i96-i104 (2025). 4.Lopez R. 等. Nature Methods, 15:1053-1058 (2018)。
查看英文原文 English abstract
Introduction: Accurately identifying patients likely to benefit from therapy and at risk of adverse events from baseline blood samples, is an unmet need in oncology. We developed a translational AI estimator framework that predicts treatment response, adverse events (AEs), and molecular signatures from baseline PBMC single-cell RNA-seq, for solid tumors. We used baseline PBMC scRNA-seq counts (with optional cell-type annotations), pre-trained molecular foundation models (FMs) (scGPT v1.0 [1], scFoundation [2]), cell-to-patient Multi-Instance Learning (MIL) [3] with RECIST labels (CR/PR Responders; SD/PD Non-Responders), to predict responses. We fine-tuned downstream tasks, applied scVI-based data augmentation [4] to improve stability and generalization, with an interpretability layer linking predictions to cell types and gene programs.
Results: Using baseline PBMC scRNA-seq from patients with solid tumor treated with immunotherapy (103 patients, ~12,000 cells/patient), the estimator predicts treatment response from baseline blood across FM backbones and identifies an optimal downstream architecture. We used scFoundation with and without hierarchical attention and scGPT. For scFoundation without hierarchical attention, we observe learning with substantial loss reduction, discrimination of classes with AUC just below 0.8 and F1 scores ≈ 0.78. For scFoundation with hierarchical attention we observe similar performance with slightly worse accuracy with AUC ≈ 0.75 and F1 ≈ 0.70 . For scGPT, we observe learning (loss reduction), good discrimination of classes with AUC ≈ 0.75 and F1 scores just over 0.7. Validation scores are slightly worst in all cases. Results we reported so far are reflecting the ability of our framework to generalize, show competitive performance with differences attributable to hierarchical vs non-hierarchical aggregation. Data augmentation to improve training stability and preliminary AE risk modeling, prioritizing grade ≥3 events, are showing encouraging results. Molecular signatures are used to support mechanistic insight and biomarker explaining predictions.
Conclusions: Baseline blood single-cell signals combined with FM embeddings and cell-to-patient learning can predict treatment response in solid tumors and provide a path to AE stratification and signature discovery. Improvements can be made with larger dataset and interpretable biological rationales can be facilitated by this framework.
References: 1.Cui H. et al. Nature Methods, 21:1470-1480 (2024).2.Hao M. et al. Nature Methods, 21:1481-1491 (2024).3.Do C et al. Bioinformatics, 41: i96-i104 (2025).4.Lopez R. et al. Nature Methods, 15:1053-1058 (2018).
利益披露 Disclosure
M. Milo, None..
A. Proutski, None..
V. Savova, None.