PO.BCS02.02 · 生物信息与计算
面向癌症免疫治疗中治疗相关不良事件的基础模型
Towards a foundation model for treatment-related adverse events in cancer immunotherapy
该海报暂无可下载的资料
AACR 官方页面
作者与单位 Authors & Affiliations
摘要 Abstract
中文摘要
背景:免疫相关不良事件(irAEs)为检查点阻断治疗期间的免疫激活提供了重要线索,然而真实世界的不良事件(AE)报告往往稀疏且零散。另一个挑战在于,FAERS提供的是非结构化的病例级报告,而ClinicalTrials.gov报告的是组别级AE表格,其完整性和详细程度在各试验间差异很大。这些差异使得难以以统一的方式跨数据源研究毒性模式。为弥补这一空白,我们开发了一个去噪AE基础模型,可重建缺失的首选术语(PTs)、学习潜在的AE共现结构,并生成具有生物学可解释性的AE嵌入。
方法:我们收集了与免疫检查点抑制剂相关的十年FAERS报告(2015Q1-2025Q2),并整理了包含200个治疗相关PT的词汇表。该模型将每个AE概况视为部分观测,并在BioBERT主干上执行PT感知的掩码、病例级PT丢弃和多标签重建。采用候选集验证来评估对真正未观测到的PT的恢复能力,并应用中间层池化以提高在不同AE概况间的稳定性。对于ClinicalTrials.gov,我们通过标准化报告深度和AE粒度的差异,协调统一了组别级AE表格。随后我们为每个治疗组生成汇总的AE特征,并评估这些特征能否区分已知临床结局存在差异的免疫治疗组。
结果:该模型以强劲的性能恢复了刻意隐藏的AE(Recall@5 = 0.38;Recall@10 = 0.51)。学习到的嵌入空间将知名的irAEs与非免疫性毒性区分开来,并形成了与MedDRA SOC模式一致的聚类。当应用于ClinicalTrials.gov治疗组时,AE特征反映了癌症类型特异性模式,包括黑色素瘤中的皮肤事件,以及肝细胞癌和消化道癌症中的肝脏或胃肠道模式。组别级嵌入概况还在多项ICI试验中区分出了结局更优的治疗组,表明该表征空间捕捉到了与治疗反应相关的有意义的生物学信息。
结论:我们提出了一个整合FAERS和ClinicalTrials.gov AE数据的基础模型,用于去噪AE概况、学习具有生物学基础的毒性嵌入,并揭示与临床获益相关的AE模式。该框架为研究AE生物学提供了一种可扩展的方法,并可能支持基于AE的生物标志物开发以及免疫治疗疗效的组别级预测。
查看英文原文 English abstract
Background: Immune-related adverse events (irAEs) provide important clues about immune activation during checkpoint blockade, yet real-world AE reporting is often sparse and fragmented. A further challenge is that FAERS provides unstructured case-level reports, while ClinicalTrials.gov reports arm-level AE tables that differ greatly in completeness and level of detail across trials. These differences make it difficult to study toxicity patterns across data sources in a unified way. To address this gap, we developed a denoising AE foundation model that reconstructs missing Preferred Terms (PTs), learns latent AE co-occurrence structure, and generates biologically interpretable AE embeddings.
Methods: We collected ten years of FAERS reports related to immune checkpoint inhibitors (2015Q1-2025Q2) and curated a vocabulary of 200 treatment-related PTs. The model treats each AE profile as partially observed and performs PT-aware masking, case-level PT dropping, and multi-label reconstruction on a BioBERT backbone. Candidate-set validation was used to evaluate recovery of truly unobserved PTs, and mid-layer pooling was applied to improve stability across diverse AE profiles. For ClinicalTrials.gov, we harmonized arm-level AE tables by standardizing differences in reporting depth and AE granularity. We then generated pooled AE signatures for each treatment arm and evaluated whether these signatures could distinguish immunotherapy arms with known differences in clinical outcomes.
Results: The model recovered intentionally hidden AEs with strong performance (Recall@5 = 0.38; Recall@10 = 0.51). The learned embedding space separated well-known irAEs from non-immune toxicities and formed clusters that aligned with MedDRA SOC patterns. When applied to ClinicalTrials.gov treatment arms, the AE signatures reflected cancer-type-specific patterns, including dermatologic events in melanoma and hepatic or gastrointestinal patterns in hepatocellular carcinoma and GI cancers. Arm-level embedding profiles also differentiated treatment arms with superior outcomes across several ICI trials, suggesting that the representation space captures meaningful biology related to treatment response.
Conclusions: We present a foundation model that integrates FAERS and ClinicalTrials.gov AE data to denoise AE profiles, learn biologically grounded toxicity embeddings, and reveal AE patterns associated with clinical benefit. This framework provides a scalable approach for studying AE biology and may support AE-based biomarker development and arm-level prediction of immunotherapy efficacy.
利益披露 Disclosure
B. Baek, None..
T. Li, None..
B. Cao, None..
X. Yu, None..
X. Wang, None.