PO.BCS02.02 · 生物信息与计算

基于大语言模型的血液/肿瘤科患者消息分诊:性能、安全性与临床意义

Large language model-based triage of Hematology/Oncology patient messages: Performance, safety, and clinical implications

海报缩略图:基于大语言模型的血液/肿瘤科患者消息分诊:性能、安全性与临床意义
编号 2741 展板 5 时间 4/20 02:00–05:00 区域 Section 3 主讲 Dinh Nguyen, MD
分会场 Large Language Models in the Clinic
查看 PDF 下载 PDF 🔒 查看 / 下载完整 PDF 需登录并开通下载套餐 · 查看套餐 / 开通 AACR 官方页面

作者与单位 Authors & Affiliations

Dinh Nguyen, Sinjin Lee, Brett Anwar, Mason Kellogg, Ronil Synghal, Khang Nguyen

SCPMG, Pasadena, CA

摘要 Abstract

中文摘要
背景:患者门户消息是血液/肿瘤科实时症状报告的重要来源,但人工审查劳动强度大,可能延迟对紧急问题的识别。大语言模型(LLM)能够大规模处理自由文本,但其在癌症诊疗中用于分诊的安全性和性能尚不确定。 方法:我们对2024年11月至2025年11月期间在南加州Permanente医疗集团(Kaiser Permanente的诊疗实施分支)某单一中心接受诊疗的血液系统恶性肿瘤成年患者进行了一项回顾性研究。从电子健康记录门户中提取了所有患者向血液/肿瘤科临床团队发起的消息。临床医生对分层抽样的消息进行了紧急和非紧急类别的标注。标记为紧急的消息反映了需在24-48小时内优先由临床医生审查的症状。我们开发了一个基于LLM的分类器,采用少样本学习(few-shot learning)生成标签,填充到接收临床医生收件箱中的“主题”列。主要结局包括判别性能指标,如分类紧急消息的AUROC、敏感性、特异性和F1评分。 结果:我们的模型共处理了来自2,287名独立患者的10,683条消息;433条消息被标记为紧急。对于紧急消息分类,少样本LLM分别实现了91.05%(95% CI,85.69%-96.42%)的AUROC、97.62%(95% CI,92.50%-99.89%)的敏感性、84.48%(95% CI,74.60%-93.34%)的特异性和89.13%(95% CI,81.82%-95.74%)的F1评分。 结论:基于LLM的分诊系统可靠地将患者消息分类为紧急和非紧急类别,有望促进血液/肿瘤科对紧急症状采取更及时的行动。该系统的实用性凸显了其对于寻求利用NLP改善癌症诊疗中消息管理的科室的相关意义。
查看英文原文 English abstract
Background: Patient portal messages are a major source of real-time symptom reports in Hematology/Oncology, but manual review is labor-intensive and may delay identification of urgent issues. Large language models (LLMs) can process free text at scale, yet their safety and performance for triage in cancer care are uncertain. Methods: We conducted a retrospective study of adult patients with hematologic malignancies receiving care at a single center within Southern California Permanente Medical Group, the care delivery arm of Kaiser Permanente, from November 2024 to November 2025. All patient-initiated messages to the Hematology/Oncology clinical team were extracted from the electronic health record portal. A stratified sample of messages was annotated by clinicians for urgent and non-urgent categories. Messages labeled as urgent reflected symptoms warranting prioritized clinician review within 24-48 hours. We developed an LLM-based classifier using few-shot learning that generates a label populating the “Topic” column in the receiving clinician's inbox. Primary outcomes included discriminative performance metrics such AUROC, sensitivity, specificity, and F1 score for classifying urgent messages. Results: A total of 10,683 messages from 2,287 unique patients were processed by our model; 433 messages were labeled as urgent. For urgent message classification, the few-shot LLM achieved an AUROC, sensitivity, specificity, and F1 score of 91.05% (95% CI, 85.69%-96.42%), 97.62% (95% CI, 92.50%-99.89%), 84.48% (95% CI, 74.60%-93.34%), and 89.13% (95% CI, 81.82%-95.74%) respectively. Conclusions: An LLM-based triage system reliably categorized patient messages into urgent and non-urgent categories and has the potential to facilitate timelier action for urgent symptoms within a Hematology/Oncology department. The system's practical utility highlights its relevance for departments seeking to leverage NLP to improve message management in cancer care.
利益披露 Disclosure
D. Nguyen, None.. S. Lee, None.. B. Anwar, None.. M. Kellogg, None.. R. Synghal, None.. K. Nguyen, None.

← 返回 AACR 2026 检索