PO.BCS02.02 · 生物信息与计算

基于LLM的免疫治疗毒性提取揭示严重程度依赖性的总生存期效应

LLM-based extraction of immunotherapy toxicities reveals severity-dependent effects on overall survival

海报缩略图:基于LLM的免疫治疗毒性提取揭示严重程度依赖性的总生存期效应
编号 2762 展板 26 时间 4/20 02:00–05:00 区域 Section 3 主讲 Zeyun Lu, PhD
分会场 Large Language Models in the Clinic
查看 PDF 下载 PDF 🔒 查看 / 下载完整 PDF 需登录并开通下载套餐 · 查看套餐 / 开通 AACR 官方页面

作者与单位 Authors & Affiliations

Zeyun Lu1, Mustafa Saleh1, Charles Lu2, Intae Moon1, Razane El Hajj Chehade1, Elio Ibrahim1, Yevgeniy R. Semenov2, Toni K. Choueiri1, Alexander Gusev1

1Dana-Farber Cancer Institute, Boston, MA,2Massachusetts General Hospital, Boston, MA

摘要 Abstract

中文摘要
在接受免疫治疗的癌症患者中刻画免疫相关不良事件(irAE)仍具有挑战性。传统基于ICD的识别常常遗漏或错误分类irAE,而人工病历审查耗费人力、易出错且不可扩展。现有基于大型语言模型(LLM)的方法常常未能充分捕捉irAE的发生日期和严重程度信息,经常遗漏轻度事件,限制了其在时间敏感和按严重程度分层分析中的效用。我们开发了一种两阶段提示策略,从接受免疫治疗患者的大量非结构化临床病历中提取irAE类型、严重程度和发生日期。在第一阶段,模型从冗长且异质的临床文档中总结irAE相关信息。在第二阶段,它提取结构化的irAE细节。该方法提高了提取准确性,并支持更可靠地识别irAE类型和发生时间。 使用专家整理的金标准,我们评估了多个可靠且日益广泛使用的大型语言模型,包括ChatGPT 4o、Llama 3.1(70B和450B)、Llama 3.3(70B)和DeepSeek R1。在各模型中,与基于ICD的提取相比,基于LLM的方法取得了更高的灵敏度(0.95对0.57)、更高的精确率(0.69对0.64)和更高的F1评分(0.77对0.55)。对于LLM和ICD均识别出irAE的病例,LLM更准确地捕捉了事件类型(83%对38%),并一致地识别出更早的发生日期(平均早27天)。该策略与模型无关,并将随着底层LLM模型的改进而持续提高提取准确性。 我们将整合了ChatGPT 4o的流程应用于Dana-Farber癌症研究所Profile队列中8,768例接受免疫治疗患者的340,277份病历,涵盖21种癌症类型。87%的患者在治疗开始后两年内发生了至少轻度的irAE。与先前研究一致,使用纳入irAE发生时间(此类信息在早期研究中通常无法获得)的Cox模型,CTLA-4抑制剂与较高的irAE发生率相关(HR=3.93;P=1.94e-19)。按时间依赖性irAE严重程度分层的生存分析显示,轻度irAE与死亡风险降低相关(HR=0.86;P=5.64e-92),而重度irAE与死亡风险增加相关(HR=1.04;P=1.15e-3)。我们还观察到与潜在癌症类型相符的系统特异性重度irAE;例如,与其他癌症相比,非小细胞肺癌患者发生重度呼吸系统irAE的瞬时风险更高(HR=2.37;P=1.29e-9)。这些发现表明,基于LLM的临床文本提取能够实现对irAE可扩展且准确的刻画,为理解其预后意义提供了更深入的见解。
查看英文原文 English abstract
Characterizing immune-related adverse events (irAEs) among patients with cancer receiving immunotherapy remains challenging. Traditional ICD-based identification often misses or misclassifies irAEs, and manual chart review is labor-intensive, error-prone, and not scalable. Existing large language model (LLM)-based approaches often do not fully capture irAE onset dates and severity information, frequently missing mild events and limiting their utility for time-sensitive and severity-stratified analyses.We developed a two-stage prompting strategy to extract irAE type, severity, and onset date from large volumes of unstructured clinical notes for patients receiving immunotherapy. In the first stage, the model summarizes irAE-related information from long and heterogeneous clinical documentation. In the second stage, it extracts structured irAE details. This approach improves extraction accuracy and supports more reliable identification of irAE types and onset timing. Using expert-curated ground truth, we evaluated multiple reliable and increasingly used large language models, including ChatGPT 4o, Llama 3.1 with 70B and 450B, Llama 3.3 with 70B, and DeepSeek R1. Across models, LLM-based methods achieved higher sensitivity (0.95 vs. 0.57), higher precision (0.69 vs. 0.64), and higher F1 scores (0.77 vs. 0.55) compared with ICD-based extraction. For cases in which both LLM and ICD identified an irAE, LLMs more accurately captured event types (83% vs. 38%) and consistently identified earlier onset dates (on average 27 days earlier). This strategy is model-agnostic and will continue to improve extraction accuracy as the underlying LLM models improve. We applied our pipeline, which incorporates ChatGPT 4o, to 340,277 notes from 8,768 patients treated with immunotherapy in the Dana-Farber Cancer Institute Profile cohort, covering 21 cancer types. 87% of patients developed at least mild irAEs within two years of treatment initiation. Consistent with prior work, CTLA-4 inhibitors were associated with higher irAE incidence (HR = 3.93; P = 1.94e-19), using Cox models that incorporate irAE onset timing, information typically unavailable in earlier studies. Survival analyses stratified by time-dependent irAE severity showed that mild irAEs were associated with reduced hazard of death (HR = 0.86; P = 5.64e-92), whereas severe irAEs were associated with an increased hazard of death (HR = 1.04; P = 1.15e-3). We also observed system-specific severe irAEs that aligned with underlying cancer type; for example, patients with non-small cell lung cancer had a higher instantaneous risk of severe respiratory irAEs (HR = 2.37; P = 1.29e-9) compared with other cancers. These findings demonstrate that LLM-based clinical text extraction enables scalable and accurate characterization of irAEs, providing deeper insights into their prognostic implications.
利益披露 Disclosure
Z. Lu, None.. M. Saleh, None.. C. Lu, None.. I. Moon, None.. R. El Hajj Chehade, None.. E. Ibrahim, None.. Y. R. Semenov, None.

← 返回 AACR 2026 检索