PO.BCS02.02 · 生物信息与计算

来源规范至关重要:以指南为锚定的大语言模型在急性白血病决策支持中优于Open Evidence

Source discipline matters: Guideline anchored large language model outperforms Open Evidence for decision support in acute leukemias

海报缩略图:来源规范至关重要:以指南为锚定的大语言模型在急性白血病决策支持中优于Open Evidence
编号 2747 展板 11 时间 4/20 02:00–05:00 区域 Section 3 主讲 Peter Palumbo, BA
分会场 Large Language Models in the Clinic
查看 PDF 下载 PDF 🔒 查看 / 下载完整 PDF 需登录并开通下载套餐 · 查看套餐 / 开通 AACR 官方页面

作者与单位 Authors & Affiliations

Peter Palumbo1, Connor Yost2, Emilio Del Toro1, Demetrios Garbis1, Peter Odutola3, Yash Kumar4, Arturo Loaiza5, Matthew Sullivan6

1Dartmouth Geisel School of Medicine, Hanover, NH,2Department of Internal Medicine, Creighton University School of Medicine, Phoenix, AZ,3Department of Molecular Biology, Harvard University, Cambridge, MA,4Institutional Liaquat National Medical College Hospital, Karachi, Pakistan,5Department of Hematology and Oncology, St. Luke's University Health Network, Bethlehem, PA,6Department of Hematology and Oncology, Dartmouth Hitchcock Medical Center, Lebanon, NH

摘要 Abstract

中文摘要
背景 急性白血病是血液肿瘤学中最复杂、发展最迅速的领域之一,其治疗选择取决于分子亚型和体能状态等多种因素。美国国家综合癌症网络(NCCN)为急性髓系白血病(AML)和急性淋巴细胞白血病(ALL)提供了不断更新的、基于细胞谱系的特异性算法,但这些指南内容密集且频繁修订。大语言模型(LLM)可能有助于临床医生整合这些数据,但其输出的可靠性在很大程度上取决于其证据来源。本研究比较了以NCCN为锚定的检索增强模型(RAG GPT-5)与Open Evidence(OE,一种关联NEJM和JAMA等期刊来源的模型),以评估其在急性白血病决策支持中的准确性、安全性和指南一致性。 方法 由两个模型独立评估40例去标识化的AML和ALL临床情景:Open Evidence(O₁)和以NCCN为锚定的检索增强GPT-5模型(O₂)。评审者对模型身份不知情,并使用改良的生成式性能评分(mGPS = 指南一致性 − 幻觉惩罚;范围 −1.0 至 +1.0)对每个回答进行评分。统计比较采用独立样本t检验。 结果 与Open Evidence(O₁,均值 = 0.70,SD = 0.32)相比,RAG模型(O₂)表现出显著更高的整体性能(均值 = 0.84,SD = 0.25);t(≈78) = −2.17,p = 0.033。定性评审揭示了临床推理中的关键区别:Open Evidence经常虚构药物(如ipilimumab),遗漏既往治疗背景,并且在化疗前未能针对感染恢复或心脏风险进行调整。RAG GPT-5仅引用NCCN推荐意见,仅存在轻微的取整误差(如ATRA剂量),偶尔会默认采用保守但仍与指南一致的给药方案(如daunorubicin)。两个模型均未能充分处理双原发肿瘤或BCR-ABL阳性情景,且均低估了近期更新,例如用于MLL重排AML的menin抑制剂,这类药物正在兴起但尚未被NCCN列入。RAG系统的方差较小,表明其在各病例间的表现更为一致。 结论 在急性白血病中,证据来源实质性地改变了LLM的行为和可靠性。相比OE,以指南为锚定的检索产生了显著更符合NCCN、且幻觉更少的推荐意见。虽然两个系统偶尔会遗漏细微的治疗史或近期的研究性药物,但只有OE提出了临床上不安全的建议。这些发现支持将以NCCN为锚定的RAG作为急性白血病中基于LLM的决策支持的更安全、更一致的基础,因为在该领域精确性和患者背景至关重要。未来的工作应扩展至复发和移植情景,并进行前瞻性的临床医生验证。
查看英文原文 English abstract
Background Acute leukemia is one of the most complex and rapidly evolving domains in hematologic oncology, where treatment selection depends on a variety of factors such as molecular subtype and performance status. The National Comprehensive Cancer Network (NCCN) provides updated, lineage-specific algorithms for Acute Myeloid Leukemia (AML) and Acute Lymphoblastic Leukemia (ALL), yet these guidelines are dense and frequently revised. Large language models (LLMs) may assist clinicians in synthesizing this data, but the reliability of their outputs depends critically on their evidence sources. This study compared an NCCN-anchored retrieval-augmented model (RAG GPT-5) with Open Evidence (OE), a model linked to journal-based sources such as NEJM and JAMA , to assess accuracy, safety, and guideline concordance in acute leukemia decision support. Methods Forty de-identified AML and ALL vignettes were independently evaluated by two models: Open Evidence (O₁) and an NCCN-anchored retrieval-augmented GPT-5 model (O₂). ). Reviewers were blinded to model identity and rated each response using a modified Generative Performance Score (mGPS = Guideline Concordance - Hallucination Penalty; range −1.0 to + 1.0). Statistical comparison used independent-samples t-tests. Results The RAG model (O₂) demonstrated significantly higher overall performance (mean = 0.84, SD = 0.25) compared with Open Evidence (O₁, mean = 0.70, SD = 0.32); t (≈78) = −2.17, p = 0.033. Qualitative review revealed key distinctions in clinical reasoning: Open Evidence frequently hallucinated agents (e.g., ipilimumab), omitted prior therapy context, and failed to adjust for infection recovery or cardiac risk before chemotherapy. RAG GPT-5 exclusively cited NCCN recommendations, with minor rounding errors (e.g., ATRA dose), and occasionally defaulted to conservative but still guideline-concordant dosing (e.g., daunorubicin). Neither model fully addressed dual-tumor or BCR-ABL-positive scenarios, and both under-recognized recent updates such as menin inhibitors for MLL-rearranged AML, which are emerging but not yet NCCN-listed. Variance was smaller for the RAG system, indicating more consistent performance across cases. Conclusions In acute leukemias, evidence source materially alters LLM behavior and reliability. Guideline-anchored retrieval produced significantly more NCCN-concordant recommendations and fewer hallucinations than OE. While both systems occasionally missed nuanced treatment history or recent investigational agents, only OE introduced clinically unsafe suggestions. These findings support NCCN-anchored RAG as the safer and more consistent foundation for LLM-based decision support in acute leukemias, where precision and patient context are paramount. Future work should expand to relapse and transplant scenarios with prospective clinician validation.
利益披露 Disclosure
P. Palumbo, None.. C. Yost, None.. E. Del Toro, None.. D. Garbis, None.. P. Odutola, None.. Y. Kumar, None.. A. Loaiza, None.. M. Sullivan, None.

← 返回 AACR 2026 检索