PO.BCS02.02 · 生物信息与计算

自动化乳腺癌患者的临床和病理分期

Automating clinical and pathological staging for breast cancer patients

海报缩略图:自动化乳腺癌患者的临床和病理分期
编号 2744 展板 8 时间 4/20 02:00–05:00 区域 Section 3 主讲 Arshad Mohammed, BA
分会场 Large Language Models in the Clinic
查看 PDF 下载 PDF 🔒 查看 / 下载完整 PDF 需登录并开通下载套餐 · 查看套餐 / 开通 AACR 官方页面

作者与单位 Authors & Affiliations

Arshad Mohammed1, Umair Ayub2, Pooja Advani3, Shakeela W. Bahadur2, Amye J. Tevaarwerk4, Tufia C. Haddad4, Elisabeth I. Heath4, Brenda J. Ernst2, Ben Zhou5, Cui Tao6, Sara J. Holton4, Karthik V. Giridhar4, Irbaz B. Riaz2

1Mayo Clinic Alix School of Medicine, Phoenix, AZ,2Division of Hematology/Oncology, Mayo Clinic, Phoenix, AZ,3Division of Hematology/Oncology, Mayo Clinic, Jacksonville, FL,4Division of Hematology/Oncology, Mayo Clinic, Rochester, MN,5Department of Computing and Augmented Intelligence ASU, Tempe, AZ,6Department of Artificial Intelligence and Informatics, Mayo Clinic, Jacksonville, FL

摘要 Abstract

中文摘要
背景:准确的分期对治疗选择、预后评估和试验入组资格至关重要。虽然既往研究尝试自动化病理(p)分期,但很少涉及临床(c)分期,且尚无研究将LLM衍生的分期与临床就诊时的临床医生分期及回顾性肿瘤登记分期进行比较。 方法:我们使用Gemini-2.0-Flash-001和GPT-5开发了一个多智能体框架,用于(1)识别报告,(2)提取关键数据,以及(3)应用AJCC第8版分期标准。在乳腺癌患者中,将LLM分期与临床医生记录和登记分期进行比较。临床医生-登记一致性作为参考标准。两名独立的乳腺肿瘤学专家审查了不一致的病例。 结果:我们分析了2018-2023年间横跨所有三个Mayo Clinic站点的122名随机选择的乳腺癌患者。LLM在病理分期上的性能达到或超过了人类评审者间的一致性(95-99.2% vs. 95-98.4%)。临床分期更具挑战性,LLM对cT的一致性为73-77.9%,对cN为87-89.3%,而临床医生-登记一致性分别为87.8%和91.0%。在89名临床医生-登记一致的患者中,LLM对pT/pN的一致性为98.9%,对cT为79.8%,对cN为91.0%。在cT不一致病例中专家倾向于LLM分期的比例为27.8%(5/18),在cN不一致病例中为25.0%(2/8)。误差分析揭示LLM在处理非肿块强化(6例)、选择放射学测量值(2例)、识别离散肿块(2例)以及其他方面(3例)存在挑战。 结论:LLM在病理分期上达到了人类水平的性能(98%),支持人在环路的部署。对于临床分期(cT为79.8%,cN为91.0%),未来的工作必须增强多模态推理并整合体格检查数据。需要前瞻性验证以评估真实世界的影响。通过有针对性的优化和监督,自动化分期能够变革登记工作流程,同时增强临床决策。 临床-病理分期一致性比较 pT pN cT cN LLM vs. 登记 96.7% 99.2% 73.8% 87.8% LLM vs. 临床医生 95.9% 96.7% 77.9% 89.3% 临床医生 vs. 登记 95.1% 98.4% 87.8% 91.0%
查看英文原文 English abstract
Background: Accurate staging is essential for treatment selection, prognosis assessment, and trial eligibility. While prior studies attempted automating pathological (p) staging, few address clinical (c) staging and no studies have compared LLM-derived staging with clinician staging during clinical visit and retrospective cancer registry staging. Methods: We developed a multi-agent framework using Gemini-2.0-Flash-001 and GPT-5 to (1) identify reports (2) extract key data, and (3) apply AJCC 8 th edition staging criteria. LLM staging was compared with clinician documentation and registry staging in breast cancer patients. Clinician-registry agreement served as the reference standard. Two independent breast oncologists reviewed discordant cases. Results: We analyzed 122 randomly selected breast cancer patients across all three Mayo Clinic sites from 2018 - 2023. LLM performance matched or exceeded human inter-rater agreement for pathological staging (95-99.2% vs. 95-98.4%). Clinical staging was more challenging, with LLM concordance of 73-77.9% for cT and 87-89.3% for cN versus 87.8% and 91.0% clinician-registry agreement. Among 89 patients with clinician-registry concordance, LLM concordance was 98.9% for pT/pN, 79.8% for cT, and 91.0% for cN. Experts favored LLM staging in 27.8% (5/18) of cT and 25.0% (2/8) of cN discordances. Error analysis revealed LLM challenges in handling non-mass enhancements (6 cases), selecting radiologic measurements (2 cases), identifying discrete masses (2 cases), and miscellaneous (3 cases). Conclusion: LLMs achieved human-level performance for pathological staging (98%), supporting human-in-the-loop deployment. For clinical staging (79.8% cT, 91.0% cN), future work must enhance multimodal reasoning and integrate physical examination data. Prospective validation is needed to assess real-world impact. With targeted refinements and oversight, automated staging can transform registry workflows while augmenting clinical decision-making. Clinical-Pathological Staging Concordances Comparison pT pN cT cN LLM vs. Registry 96.7% 99.2% 73.8% 87.8% LLM vs. Clinician 95.9% 96.7% 77.9% 89.3% Clinician vs. Registry 95.1% 98.4% 87.8% 91.0%
利益披露 Disclosure
A. Mohammed, None.. U. Ayub, None.. P. Advani, None.. S. W. Bahadur, None.. A. J. Tevaarwerk, None.. T. C. Haddad, None.. E. I. Heath, None.. B. J. Ernst, None.. B. Zhou, None.. C. Tao, None.. S. J. Holton, None.. K. V. Giridhar, None.. I. B. Riaz, None.

← 返回 AACR 2026 检索