PO.BCS01.09 · 生物信息与计算
在多模态肿瘤学基础模型上采用组相对策略优化进行机制感知的癌症治疗规划
Mechanism-aware cancer therapy planning with group-relative policy optimization on a multimodal oncology foundation model
作者与单位 Authors & Affiliations
摘要 Abstract
中文摘要
现代多模态肿瘤学基础模型(OFM)能够从临床、基因组、转录组和组织学数据预测患者轨迹并模拟治疗效果,但它们在很大程度上仍是黑箱:它们很少解释是哪些机制驱动了风险,或如何通过可行的干预来调控这些机制。
我们在超过120万名癌症患者的纵向临床记录、肿瘤DNA和RNA测序以及H&E成像数据上训练了一个多模态OFM。该模型学习每位患者不断演变的疾病状态的联合潜在表征,并能够预测结局和模拟反事实的治疗轨迹。随后我们增加了一个机制感知的推理层,将OFM从一个预测器转变为基于机制的治疗设计引擎。
我们的关键创新是组相对策略优化(GRPO),这是一个强化学习框架,它连接了由OFM推断的机制状态(例如通路激活模式和耐药程序)、由精选的药物-靶点关系和扰动特征构建的药物干预空间,以及一个明确的临床奖励。我们并非为单个患者优化一个抽象目标,而是将奖励定义为对共享相似疾病状态和驱动机制的患者队列的预测结局的改善。对于每个疾病背景,我们鉴定与不良结局相关的机制,将它们与候选药物及真实世界数据中观测到的结局相关联,然后使用GRPO评估策略(机制层面干预的集合),方法是询问将某个策略应用于该队列相较于匹配的标准治疗对照,是否能改善预测生存或延缓进展。这种组相对奖励稳定了学习,避免了对特异离群值的过拟合,并使习得的策略与临床医生自然地对"类似这些的患者"进行推理的方式相一致。
为使输出在生物学和临床上可解释,我们将所提出的机制转变映射到现有药物及组合、映射到其结局由相似潜在程序驱动的机制定义患者簇,以及映射到从头机制机会(在这些机会中最优策略改善了结局,但目前没有药物能完全解释该效应),从而标记出潜在的发现靶点。
在多个适应症中,GRPO恢复了已知的机制-治疗关系,鉴定出当所接受治疗与GRPO建议的机制策略一致时其结局相较接受不一致治疗的匹配患者得到改善的患者亚组,并提出了涉及当前尚未被联合靶向的程序的组合或序贯调控的新颖机制-治疗假设。这将大规模OFM转变为一个机制感知的决策引擎,能够在用于结局预测的同一潜在空间中对耐药和治疗机会进行推理。
查看英文原文 English abstract
Modern multimodal oncology foundation models (OFMs) can predict patient trajectories and simulate treatment effects from clinical, genomic, transcriptomic, and histologic data, but they remain largely black boxes: they rarely explain which mechanisms drive risk or how to modulate those mechanisms with feasible interventions.
We trained a multimodal OFM on over 1.2 million cancer patients with longitudinal clinical records, tumor DNA and RNA sequencing, and H&E imaging. The model learns a joint latent representation of each patient's evolving disease state and can forecast outcomes and simulate counterfactual treatment trajectories. We then added a mechanism-aware reasoning layer that turns the OFM from a predictor into an engine for mechanism-based therapy design.
Our key innovation is group-relative policy optimization (GRPO), a reinforcement-learning framework that links mechanism states inferred by the OFM (e.g., pathway activation patterns and resistance programs), a drug intervention space built from curated drug-target relationships and perturbation signatures, and an explicit clinical reward. Rather than optimizing an abstract objective for a single patient, we define the reward as improvement in predicted outcomes for cohorts of patients who share similar disease state and driver mechanisms. For each disease context, we identify mechanisms associated with poor outcome, link them to candidate drugs and observed outcomes in real-world data, and then use GRPO to evaluate policies (sets of mechanism-level interventions) by asking whether applying a policy to that cohort improves predicted survival or delays progression relative to matched standard-of-care controls. This group-relative reward stabilizes learning, avoids overfitting to idiosyncratic outliers, and aligns the learned policies with how clinicians naturally reason about “patients like these.”
To make outputs biologically and clinically interpretable, we map proposed mechanism shifts to existing drugs and combinations, to mechanism-defined patient clusters whose outcomes are driven by similar latent programs, and to de novo mechanism opportunities where the optimal policy improves outcomes but no current drug fully explains the effect, flagging potential targets for discovery.
Across multiple indications, GRPO recovers known mechanism-therapy relationships, identifies patient subgroups whose outcomes are improved when therapies they receive align with GRPO-suggested mechanism policies compared with matched patients receiving discordant therapies, and proposes novel mechanism-therapy hypotheses involving combined or sequential modulation of programs that are not jointly targeted today. This turns a large-scale OFM into a mechanism-aware decision engine that reasons about resistance and treatment opportunities in the same latent space used for outcome prediction.
利益披露 Disclosure
E. E. Schadt,
Pathos AI Employment, Stock.
J. Stokes,
Pathos AI Independent Contractor.
L. Zhao,
Pathos AI Employment, Stock.
J. Chin,
Pathos AI Employment, Stock.
J. Tyler,
Pathos AI Employment, Stock.
L. Sun,
Pathos AI Employment, Stock.
L. Beck,
Pathos AI Employment, Stock.
D. Xu,
Pathos AI Employment, Stock.
A. G. Beckmann,
Pathos AI Employment, Stock.
I. Huerga,
Pathos AI Employment, Stock.