PO.BCS02.02 · 生物信息与计算
利用Google云计算将大语言模型应用于CAR-T细胞疗法临床数据
Applications of large language models to CAR-T cell therapy clinical data using Google cloud computing
作者与单位 Authors & Affiliations
摘要 Abstract
中文摘要
大语言模型(LLM)因其分析和总结大量文本数据的出色能力,正在医学领域被广泛采用。这些模型使临床医生和研究人员能够从复杂数据集中提取有意义的洞见,并可能协助决策。在此,我们展示了将LLM应用于CAR-T细胞疗法相关临床数据解读与总结的工作流程与应用。
在Google云计算(GCP)环境中使用一个LLM(Gemini 2.5 pro),开发了两个应用来分析CAR-T细胞疗法临床数据:1)提取并总结CRS和ICANS事件相关数据,以简化合规团队的工作流程;2)识别CAR-T输注时可获得的、能够将患者在CAR-T后14天分类为高监测或低监测需求的特征。患者数据(生命体征、化验、血液学记录和EKG)通过Google BigQuery从电子病历(EMR)提取到GCP中的SQL表。对每个应用,检索相关数据字段,格式化为JSON对象,并嵌入LLM提示词中以进行上下文感知处理。两个应用在将LLM输出与EMR中的真实数据进行分析后,均经过了迭代式的提示词工程优化。
对于提取和总结CRS及ICANS事件的应用,LLM在Mayo Clinic Rochester的数据上进行了优化。与IEC合规数据库相比,LLM对CRS事件实现了100%的准确率和F1评分,对ICANS事件实现了96%的准确率和82%的F1评分。CRS和ICANS分级的匹配率分别为83%和89%。我们将同一LLM应用于Mayo Clinic Arizona(MCA)和Mayo Clinic Florida(MCF)的病例(这些机构对临床记录和毒性流程表的记录方式不同),对CRS实现了准确率(MCA:93%,MCF:93%)和F1评分(MCA:96%,MCF:96%),对ICANS实现了准确率(MCA:81%,MCF:82%)和F1评分(MCA:76%,MCF:75%)。LLM能够捕获人工审查遗漏的事件。对于大多数需要合规团队最终裁定的差异,LLM将更新为标记差异以供合规团队审查。
对于CAR-T后14天监测需求相关的应用,LLM识别出5类可预测CAR-T输注后高或低监测需求的类别(疾病状态、炎症和肿瘤负荷标志物、血液学状态、肾功能和体能状态)。与LLM从EMR提取的数据相比,我们的模型在第一队列中实现了85.7%的敏感性、23.8%的特异性和65.5%的F1评分。对于人口统计学特征与队列1统计学上相似且使用相同5类类别的第二队列,LLM实现了83.3%的敏感性、28.6%的特异性和65.4%的F1评分。我们的工作流程和LLM应用为有兴趣将LLM应用于临床和研究场景的其他人提供了范例和指导。
查看英文原文 English abstract
Large Language Models (LLM) are being widely adopted into the medical field for their impressive ability to analyze and summarize large amounts of text data. These models enable clinicians and researchers to extract meaningful insights from complex datasets and may assist with decision making. Here we present our workflows and application of LLM for the interpretation and summarization of clinical data related to CAR-T cell therapy.
Using an LLM (Gemini 2.5 pro) in the Google Cloud Computing (GCP) environment, two applications were developed to analyze CAR-T cell therapy clinical data: 1) extracting and summarizing CRS and ICANS event-related data to streamline the compliance team workflow, and 2) identifying features available at time of CAR-T infusion able to classify patients into high- or low-monitoring needs 14 days post CAR-T. Patient data (vitals, labs, hematology notes, and EKGs) were extracted from the electronic medical record (EMR) using Google BigQuery into SQL tables in GCP. For each application, relevant data fields were retrieved, formatted into JSON objects, and embedded in the LLM prompt for context-aware processing. Both applications have undergone iterative prompt engineering after analyzing the LLM output against ground truth data in the EMR.
For the application extracting and summarizing CRS and ICANS events, the LLM was optimized on Mayo Clinic Rochester data. When compared to the IEC compliance database, the LLM achieved 100% accuracy and F1 score for CRS events, and 96% accuracy and 82% F1 score for ICANS events. The match rate for CRS and ICANS grades were 83% and 89% respectively. We applied the same LLM to Mayo Clinic Arizona (MCA) and Mayo Clinic Florida (MCF) cases, who document clinical notes and toxicity flowsheet differently, and achieved an accuracy (MCA: 93%, MCF:93%) and F1 score (MCA: 96%, MCF: 96%) for CRS and accuracy (MCA: 81%, MCF:82%) and F1 score (MCA: 76%, MCF: 75%) for ICANS. LLM was able to capture events missed by manual reviews. For most of the discrepancies where compliance team final adjudication is needed, LLM will be updated to flag discrepancies for review by the compliance team.
For the application related to the monitoring needs 14 days post CAR-T, the LLM identified 5 categories predictive of high or low monitoring needs post CAR-T infusion (disease status, inflammatory and tumor burden markers, hematologic status, renal function and performance status). Our model achieved a sensitivity of 85.7%, specificity of 23.8%, and F1 score of 65.5% for our first cohort, compared with data extracted by the LLM from the EMR. For our second cohort, with demographics statistically similar to cohort 1 and using the same 5 categories, the LLM achieved a sensitivity of 83.3%, a specificity of 28.6%, and F1 score of 65.4%. Our workflow and applications of LLM's provide examples and guidance to others interested in applying LLM's for clinical and research applications.
利益披露 Disclosure
E. Contreras Guzman, None..
M. Jankowski, None..
A. De Menezes Silva Corraes, None..
M. Gupta, None..
M. L. Shaw, None..
S. Chhabra, None..
J. K. Mclean, None..
K. R. Riester, None..
K. Joseph, None..
H. Ross, None..
C. Barnett, None..
S. Carter, None..
S. Girmay, None..
R. Wolan, None..
M. Ramsey, None..
C. Downhour, None..
K. Morgan, None..
S. Sibley, None..
E. Rushing, None..
L. Holmes, None..
A. Burgstahler, None..
H. Alkhateeb, None..
M. Hathcock, None..
R. Bruno, None..
A. C. Rosenthal, None..
H. Murthy, None..
Y. Lin, None.