LBPO.BCS02 · 生物信息与计算 · Late-Breaking

AI辅助的乳腺癌病理报告信息提取

AI-assisted pathology report abstraction for breast cancer

海报缩略图:AI辅助的乳腺癌病理报告信息提取
编号 LB450 展板 18 时间 4/22 09:00–12:00 区域 Section 52 主讲 Sarah Van Alsten, BS;MPH;PhD
分会场 Late-Breaking Research: Bioinformatics, Computational Biology, Systems Biology, and Convergent Science 2
查看 PDF 下载 PDF 🔒 查看 / 下载完整 PDF 需登录并开通下载套餐 · 查看套餐 / 开通 AACR 官方页面

作者与单位 Authors & Affiliations

Sarah C. Van Alsten1, Saianand Balu1, Isaiah W. Zipple1, Georgia C. Mudd1, Nader Mehri1, Daniel Fabbri2, Melissa A. Troester1

1UNC Lineberger Comprehensive Cancer Center, Chapel Hill, NC,2Vanderbilt University, Nashville, TN

摘要 Abstract

中文摘要
背景:临床癌症研究常涉及从病理报告中人工提取肿瘤信息,耗时费力。基于大语言模型(LLM)的信息提取提供了一种高效的替代方案,但由于LLM的“黑箱”性质而缺乏透明度。我们评估了一种使用BRIM(一种LLM辅助的人在环临床信息提取工具)的组合方法的速度和准确性。 方法:我们转录了828份PDF乳腺肿瘤病理报告(其中399份此前由受过训练的肿瘤登记员提取),涉及卡罗来纳乳腺癌研究第4阶段(CBCS4)的588名参与者的文本。我们设计了提示词,用于从报告中提取肿瘤大小、分级、雌激素和孕激素受体(ER、PR)以及人表皮生长因子受体2(HER2)信息,包括阳性百分比、染色强度、免疫组化(IHC)评分(视情况而定)。提示词通过在BRIM中结合GPT-OSS-20B(一种开源LLM)进行迭代测试来调优,提取人员审查文本内证据和模型推理,提供实时反馈以改进后续输出。性能指标包括与金标准人工提取相比的提取速度和准确性。仅纳入LLM和登记员均有可用数据的病例。 结果:对828份记录的LLM辅助提取耗时11.5小时,而我们的认证肿瘤登记员则需100小时(平均每份报告7分钟),提取时间减少90%。如表所示,提取准确性范围从85%(PR阳性百分比)到98%(肿瘤分级、ER强度和ER分类解释),大多数准确性超过90%。 结论:从病理报告中对肿瘤变量进行LLM辅助提取是可行的、准确的,并可能大幅减轻人工提取的负担。纳入额外复杂性(如附录、多灶性)的进一步工作对于更贴近地模拟登记员工作流程将非常重要。 LLM辅助提取的准确性 范围 准确性(正确标签/所有生成标签) ER状态 阳性百分比 每份记录 89%(184/207) 强度 每份记录 98%(145/148) 分类/文本 每份记录 98%(196/200) PR状态 阳性百分比 每份记录 85%(168/198) 强度 每份记录 92%(126/137) 分类/文本 每份记录 89%(170/191) HER2状态 强度 每份记录 97%(148/153) ISH 每份记录 97%(30/31) HER2/CEP比值 每份记录 94%(31/33) 肿瘤分级 每份记录 98%(336/344) 大小(数值) 每份记录 94%(263/279) ER状态汇总 每例患者 97%(308/316) PR状态汇总 每例患者 95%(299/313)
查看英文原文 English abstract
Background: Clinical cancer research often involves time-consuming manual abstraction of tumor information from pathology reports. Large language model (LLM)-based abstraction offers an efficient alternative but lacks transparency due to the “black-box” nature of LLMs. We evaluated the speed and accuracy of a combined approach using BRIM, an LLM-assisted human-in-the-loop clinical abstraction tool. Methods: We transcribed text from 828 PDF breast tumor pathology reports (399 previously abstracted by a trained tumor registrar) for 588 participants in the Carolina Breast Cancer Study, Phase 4 (CBCS4). We designed prompts to abstract tumor size, grade, estrogen and progesterone receptor (ER, PR) and human epidermal growth factor receptor 2 (HER2) information from reports, including percent positivity, staining intensity, immunohistochemistry (IHC) scores, as appropriate. Prompts were tuned through iterative testing in BRIM in conjunction with GPT-OSS-20B (an open-source LLM), where abstractors reviewed in-text evidence and model reasoning, providing real-time feedback to improve subsequent outputs. Performance metrics included abstraction speed and accuracy compared to gold-standard manual abstraction . Only cases with available data from both the LLM and registrar were included. Results: LLM-assisted abstraction of the 828 notes took 11.5 hours compared to 100 hours by our certified tumor registrar (average 7 min/report), representing a 90% reduction in abstraction time. As shown in Table, abstraction accuracies ranged from 85% (% positivity for PR) to 98% (tumor grade, ER intensity, and ER categorical interpretation), with most accuracies exceeding 90%. Conclusion: LLM-assisted abstraction of tumor variables from pathology reports is feasible, accurate, and may substantially reduce the burden of manual abstraction. Further work incorporating additional complexity (e.g. addenda, multifocality) will be important for more closely mimicking registrar workflows. Accuracy of LLM-based Abstraction Scope Accuracy (Correct Label/All Generated Labels) ER Status % Positive Per Note 89% (184/207) Intensity Per Note 98% (145/148) Category/Text Per Note 98% (196/200) PR Status % Positive Per Note 85% (168/198) Intensity Per Note 92% (126/137) Category/Text Per Note 89% (170/191) HER2 Status Intensity Per Note 97% (148/153) ISH Per Note 97% (30/31) HER2/CEP Ratio Per Note 94% (31/33) Tumor Grade Per Note 98% (336/344) Size (Numeric) Per Note 94% (263/279) ER Status Summary Per Patient 97% (308/316) PR Status Summary Per Patient 95% (299/313)
利益披露 Disclosure
S. C. Van Alsten, None.. S. Balu, None.. I. W. Zipple, None.. G. C. Mudd, None.. N. Mehri, None. D. Fabbri, Brim Analytics Other Business Ownership. M. A. Troester, None.

← 返回 AACR 2026 检索