PO.BCS02.02 · 生物信息与计算

大型语言模型大规模提取不良事件揭示治疗毒性的潜在结构

Large scale extraction of adverse events by large language models uncovers latent structure of treatment toxicity

海报缩略图:大型语言模型大规模提取不良事件揭示治疗毒性的潜在结构
编号 2755 展板 19 时间 4/20 02:00–05:00 区域 Section 3 主讲 John Lazar, MD;PhD
分会场 Large Language Models in the Clinic
查看 PDF 下载 PDF 🔒 查看 / 下载完整 PDF 需登录并开通下载套餐 · 查看套餐 / 开通 AACR 官方页面

作者与单位 Authors & Affiliations

John Lazar1, Divneet Mandair2, Catherine C. Smith2, Travis Zack3

1UCSF - University of California San Francisco, San Francisco, CA,2UCSF, San Francisco, CA,3UCSF School of Medicine, San Francisco, CA

摘要 Abstract

中文摘要
管理治疗毒性是肿瘤学中的一大挑战。关于治疗相关不良事件(AE)的大规模数据可帮助肿瘤科医生更好地理解和预测毒性。然而,由于传统上记录AE需要劳动密集的人工报告,大多数AE数据集局限于数量较少或类型受限的事件。从电子健康记录(EHR)中自动提取AE有望实现大规模分析,但这是一项具有挑战性的任务,因为关于大多数AE的信息分散于临床病历等非结构化数据中。因此,我们开发并验证了一个基于智能体的大型语言模型(LLM)流程AE-Extract,用于从肿瘤科医生的病程记录中提取AE并对其分级。与人工标注相比,该流程灵敏度很高,对同一器官系统内事件的召回率>95%,精确率为67%。相对于召回率而言精确率较低,反映了将事件归因于治疗效应时的不确定性,以及优先保证高召回率而非高精确率的设计选择。为了全面刻画癌症治疗期间的AE,我们随后将AE-Extract应用于某学术性癌症中心的84,684份病程记录。我们共识别出279,457例AE,平均每份记录3.3例。为识别治疗毒性的反复出现模式,我们对AE数据应用了非负矩阵分解(NMF),识别出85个潜在因子,捕捉了数据集中的毒性模式。这些潜在模式既捕捉了机制相关事件的共现(如神经病变与跌倒),也捕捉了不同器官系统中无已知机制关联事件之间的关联,例如接受免疫治疗患者的瘙痒与结肠炎,或接受细胞毒性化疗患者的神经病变与指甲变化。由于潜在因子量化了AE的核心模式,我们接下来研究了如何将其用作分析治疗毒性的稳健特征集。首先,我们发现在患者治疗过程前段识别出的潜在模式可以预测新AE的发生以及已有AE严重程度的加重。我们还发现这些治疗毒性模式与患者人口学特征及合并症相关:我们既重现了已知关联,如性别与癌症所致恶心/呕吐之间的联系,也发现了新的关系,如缺血性心脏病与包括耳鸣和听力受损在内的前庭蜗症状之间的联系。总体而言,我们的工作展示了一种从临床病历中准确、高通量提取AE的新方法。我们展示了对治疗毒性的全面评估如何能更好地刻画AE模式,识别治疗毒性与患者特征之间的稳健关联,并可作为构建个体化模型以预测治疗毒性的起点。
查看英文原文 English abstract
Managing treatment toxicity is a major challenge in oncology. Large-scale data on treatment related adverse events (AEs) could help oncologists better understand and predict toxicity. However, most AE datasets are limited to small numbers or restricted types of events due to the labor-intensive manual reporting traditionally required to document AEs. Automated extraction of AEs from the electronic health record (EHR) could enable large-scale analysis but is a challenging task as information about most AEs is in diffuse, unstructured data such as clinical notes.We therefore developed and validated an agentic large language model (LLM) pipeline, AE-Extract, to extract and grade AEs from oncologists' progress notes. The pipeline is highly sensitive with a recall of >95% for events in the same organ system and a precision of 67% compared to manual annotation. Lower precision relative to recall reflects uncertainty in the attribution of events to treatment effects as well as design choices to prioritize higher recall over precision. To comprehensively characterize AEs during cancer therapy, we subsequently applied AE-Extract to a set of 84,684 progress notes from an academic cancer center. We identified 279,457 total AEs, a mean of 3.3 per note. To identify recurrent patterns of treatment toxicity, we applied non-negative matrix factorization (NMF) to the AE data and identified 85 latent factors that capture patterns of toxicities in the dataset. These latent patterns capture co-occurrence of mechanistically related events such as neuropathy and falls as well as associations in different organ systems without known mechanistic connection such as pruritus and colitis in patients receiving immunotherapy or neuropathy and nail changes in patients receiving cytotoxic chemotherapy. As the latent factors quantify core patterns of AEs, we next studied how they could be used as robust feature sets for analysis of treatment toxicity. First, we found that the latent patterns identified in the first part of a patient's treatment course can predict development of new AEs as well as worsening severity of existing AEs. We also found these patterns of treatment toxicity correlated with patient demographics and comorbidities: we recapitulated known associations such as the link between gender and cancer induced nausea/vomiting as well as novel relationships such as a link between ischemic heart disease and vestibulocochlear symptoms including tinnitus and hearing impairment.Overall, our work demonstrates a new method for accurate, high-throughput extraction of AEs from clinical notes. We show how comprehensive evaluation of treatment toxicity allows for better characterization of patterns of AEs, identifies robust associations between treatment toxicity and patient characteristics, and may serve as a starting point for building personalized models to predict treatment toxicity.
利益披露 Disclosure
J. Lazar, None.. D. Mandair, None. C. C. Smith, Daiichi Sankyo Independent Contractor. Revolution Medicines ). Genentech Independent Contractor. Abbvie Independent Contractor. Syndax ). Biomea Independent Contractor, ), Clinical trial funding. T. Zack, Open Evidence Employment, Stock.

← 返回 AACR 2026 检索