PO.BCS02.02 · 生物信息与计算

一种领域专业化、可扩展且具临床相关性的人工智能流程,用于自动化卵巢癌影像报告的标准化O-RADS分层

An artificial intelligence domain-specialized scalable and clinically relevant pipeline to automate standardized O-RADS stratification for imaging reports in ovarian cancer

编号 2742 展板 6 时间 4/20 02:00–05:00 区域 Section 3 主讲 Asmi Agarwal
分会场 Large Language Models in the Clinic
该海报暂无可下载的资料 AACR 官方页面

作者与单位 Authors & Affiliations

Asmi Agarwal1, Min Ren2, Jingjing Gong3, Richard Selinfreund4, Ruchika Goel5, Yanhui Guo6

1SIU School of Medicine, Springfield, IL,2Ultrasound Department, Shanghai First Maternity and Infant Hospital, Shanghai, China,3Shanghai Changning Maternity and Infant Health Hospital, Shanghai, China,4Department of Pathology, SIU School of Medicine, Springfield, IL,5Department of Hematology and Oncology, Johns Hopkins University and SIU School of Medicine, Springfield, IL,6Department of Computer Science, University of Illinois Springfield, Springfield, IL

摘要 Abstract

中文摘要
目的:尽管治疗手段不断进步,卵巢恶性肿瘤在全球女性中仍承担着不成比例的死亡负担。及时准确地评估附件病变对改善预后至关重要,但由于早期表现细微或被误判,该疾病常在晚期才被诊断。卵巢-附件报告和数据系统(O-RADS)为恶性肿瘤风险分层提供了标准化框架;然而在实践中,其手动应用耗时且易受观察者间变异性的影响,为一致性及纳入临床工作流程造成了障碍。为满足这一紧迫的临床需求,我们开发了一个人工智能(AI)流程,可直接从自由文本盆腔超声报告中自动化O-RADS分类。 程序/方法:通过将Lingshu(一个用于医学报告推理的多模态领域专业化大语言模型LLM)与传统机器学习分类器相结合,我们的系统将非结构化的放射学叙述转化为结构化、高保真的恶性肿瘤风险评估,无需人工评分。我们还将此框架的性能与使用MedGemma的等效流程进行了比较。 数据/结果:我们分析了413份去标识化的盆腔超声报告,并使用Lingshu提取语义嵌入。这些代表临床有意义语言模式的嵌入随后通过5折交叉验证用于训练机器学习分类器。Lingshu与逻辑回归模型表现最佳,实现了0.773 ± 0.031的平均准确率、0.777 ± 0.029的加权精确度、0.767 ± 0.028的召回率、0.765 ± 0.032的F1评分以及0.929 ± 0.019的宏平均AUROC。值得注意的是,这一模型实现了较低的0.923 ± 0.027的AUROC。 结论:像Lingshu这样的基础模型展现出卓越的语义理解能力。我们表明,该模型可被安全有效地用于直接从非结构化超声报告中标准化O-RADS风险评估,弥合AI能力与真实世界放射学实践之间的差距。我们的方法建立了一条可扩展且具临床影响力的路径,用于将AI整合到妇科肿瘤学工作流程中,为LLM驱动工具在卵巢癌早期检测和风险分层中的更广泛应用铺平道路。重要的是,这一AI驱动框架并非旨在取代放射科专家的判断,而是通过实现O-RADS标准的一致应用、支持诊断信心并减少不同医生和机构间的变异性来加以增强。如此,它有望加速高危附件病变的早期识别,并最终改善临床决策和患者预后。
查看英文原文 English abstract
Purpose: Despite therapeutic advances, ovarian malignancies continue to carry a disproportionate mortality burden among women worldwide. Timely and accurate assessment of adnexal lesions is critical for improving outcomes, yet the disease is often diagnosed at an advanced stage due to subtle or misinterpreted early findings. The Ovarian-Adnexal Reporting and Data System (O-RADS) provides a standardized framework for malignancy risk stratification; however, in practice, its manual application can be time-consuming and prone to inter-observer variability, creating barriers to consistency and incorporation within clinical workflows. To address this urgent clinical need, we developed an artificial intelligence (AI) pipeline that automates O-RADS classification directly from free-text pelvic ultrasound reports. Procedures/Methods: By integrating Lingshu , a multimodal Domain-Specialized large language model (LLM) for medical reports reasoning, with traditional machine learning classifiers, our system transforms unstructured radiology narratives into structured, high-fidelity malignancy risk assessments, eliminating the need for manual scoring. We also compared the performance of this framework with that of an equivalent pipeline using MedGemma. Data/Results: We analyzed 413 de-identified pelvic ultrasound reports and extracted semantic embeddings using Lingshu. These embeddings, representing clinically meaningful linguistic patterns, were used to then train machine learning classifiers via a 5-fold cross-validation. Lingshu and a logistic regression model performed the best, achieving a mean accuracy of 0.773 ± 0.031, weighted precision of 0.777 ± 0.029, recall of 0.767 ± 0.028, F1-score of 0.765 ± 0.032, and a macro-averaged AUROC of 0.929 ± 0.019. Notably, this achieved a lower AUROC of 0.923 ± 0.027. Conclusion: A foundation model such as Lingshu demonstrates remarkable semantic understanding. We show that this model can be safely and effectively leveraged to standardize O-RADS risk assessment directly from unstructured ultrasound reports, bridging the gap between AI capability and real-world radiology practice. Our approach establishes a scalable and clinically impactful pathway for integrating AI into gynecologic oncology workflows, paving the way for broader adoption of LLM-driven tools in early ovarian cancer detection and risk stratification. Importantly, this AI-driven framework is not intended to replace expert radiologic judgment but to augment it by enabling the consistent application of O-RADS criteria, supporting diagnostic confidence, and reducing variability across providers and institutions. In doing so, it has the potential to expedite early identification of high-risk adnexal lesions and ultimately improve clinical decision-making and patient outcomes.
利益披露 Disclosure
A. Agarwal, None.. M. Ren, None.. J. Gong, None.. R. Selinfreund, None.. R. Goel, None.. Y. Guo, None.

← 返回 AACR 2026 检索