PO.BCS02.01 · 生物信息与计算

一种从非结构化患者记录中自动化、高保真地整理癌症诊断和分期的智能体式AI工作流程

An agentic AI workflow for automated, high-fidelity curation of cancer diagnosis and staging from unstructured patient records

海报缩略图:一种从非结构化患者记录中自动化、高保真地整理癌症诊断和分期的智能体式AI工作流程
编号 20 展板 5 时间 4/19 02:00–05:00 区域 Section 2 主讲 Tian Kang
分会场 Agentic AI in Cancer
查看 PDF 下载 PDF 🔒 查看 / 下载完整 PDF 需登录并开通下载套餐 · 查看套餐 / 开通 AACR 官方页面

作者与单位 Authors & Affiliations

Tian Kang, Elizabeth Dougherty, Shivam Mishra, Ryan Godart, Arpita Saha, Jonathan Wills, Victoria L. Chiou, Maria A. Berezina, Kunal Nagpal

Tempus AI, Inc., Chicago, IL

摘要 Abstract

中文摘要
目的 使用大语言模型(LLM)的人工智能(AI)驱动的临床文档摘录(CDA)正在改变肿瘤学数据管理,它能从非结构化的记录中提取关键数据,例如分期和分子谱。我们验证了一个能够自主将癌症诊断摘录为结构化、高质量真实世界证据(RWE)的智能体,该智能体加速研究,并以临床级准确性、大规模处理能力和运营可持续性支持及时的循证治疗决策。 方法 我们的AI-CDA工作流程采用三阶段混合多智能体AI设计,以实现高准确性、高成本效益的数据摘录: 1. 预筛选:两个在临床记录上微调过的专用自然语言处理(NLP)模型扫描整份患者病历,对文档进行分类并识别关键的癌症诊断事件,从而筛选出最多100份相关文档,降低LLM推理成本。 2. 提取器:一个非推理型LLM(GPT-4.1)从所选文档中提取所有相关的诊断信息。 3. 结构化与标准化:一次GPT-4.1调用综合诊断摘要,根据预定义的数据模型生成结构化的临床字段(例如诊断日期、分期、组织学)。最后,结构化字段由o3-mini标准化为固定的分类体系。 系统性能在从 Tempus RW 数据库中抽样的1497例患者(499例乳腺癌、499例肺癌、499例泛癌种)队列上,与人工摘录进行了基准比较。 结果 该自动化工作流程相较于人工摘录表现优异。原发组织摘录取得了0.94的微观F1分数,在常见癌症上表现出色(例如肺癌:0.98,乳腺癌:0.98,结肠癌:0.97)。转移状态检测取得了0.92的F1分数,组织学为0.91,总体分期为0.83。初始诊断日期准确性为0.90。随后一次对不一致之处进行裁定的再评估显示,通过纠正最初的人工标注错误,得出了更高的真实世界性能。组织学F1分数升至0.96(+0.05),总体分期升至>0.91(提升超过10%)。 在成本效益方面:使用两个专用NLP模型进行预筛选,将需要LLM审阅的平均文档数量减少了>90%(从每位患者266份减至24份)。这降低了LLM成本,缓解了与较长上下文相关的性能下降,并仅优先处理最相关的信息。 结论 本研究证明了一种智能体式AI-CDA工作流程用于高准确性癌症数据摘录的可行性和运营可持续性。这一混合解决方案满足了三项关键需求:通过自动化、高保真的数据摘录扩展临床研究;支持及时的循证治疗决策;以及通过混合NLP/LLM工作流程降低推理成本和资源消耗,从而实现运营可持续性。
查看英文原文 English abstract
Purpose Artificial intelligence (AI)-driven clinical document abstraction (CDA) using large language models (LLMs) is transforming oncology data management by extracting crucial data, such as staging and molecular profiles, from unstructured notes. We validated an agent that autonomously abstracts cancer diagnoses into structured, high-quality real-world evidence (RWE), accelerating research and supporting timely evidence-based treatment decisions with clinical-grade accuracy at scale and operational sustainability. Methods Our AI-CDA workflow uses a three-stage hybrid multi-agent AI design for high-accuracy, cost-efficient data abstraction: 1. Pre-screening: two specialized natural language processing (NLP) models fine-tuned on clinical notes scan the entire patient chart, classifying documents and identifying key cancer diagnosis events to select up to 100 relevant documents, reducing the LLM inference cost. 2. Extractor: a non-reasoning LLM (GPT-4.1) extracts all relevant diagnostic information from selected documents. 3. Structuring and normalization: a GPT-4.1 call synthesizes diagnosis summaries to produce structured clinical fields (e.g., diagnosis date, stage, histology) according to a predefined data model. Finally, structured fields are normalized to a fixed taxonomy by o3-mini. System performance was benchmarked against manual abstraction on a cohort of 1497 patients (499 breast, 499 lung, 499 pan-cancer) sampled from the Tempus RW database. Results The automated workflow performed strongly against manual abstraction. Tissue of origin abstraction achieved a micro F1-score of 0.94, with excellent performance on common cancers (e.g., lung: 0.98, breast: 0.98, colon: 0.97). Metastasis status detection achieved 0.92 F1-score, histology 0.91, and overall stage 0.83. Initial diagnosis date accuracy was 0.90. A subsequent, adjudicated re-evaluation of discordance revealed an even higher real-world performance by correcting initial manual labeling errors. Histology F1-score rose to 0.96 (+0.05) and overall staging rose to > 0.91 (a lift of over 10%). With regard to cost efficiency: pre-screening with two specialized NLP models reduced the average number of documents requiring LLM review >90% (from 266 to 24 per patient). This lowers LLM cost, mitigates performance degradation associated with longer context, and prioritizes only the most pertinent information. Conclusions This study demonstrates the feasibility and operational sustainability of an agentic AI-CDA workflow for highly accurate cancer data abstraction. This hybrid solution addresses three critical needs: scaling clinical research through automated, high-fidelity data abstraction; supporting timely evidence-based treatment decisions; and achieving operational sustainability by reducing inference cost and resourcing through a hybrid NLP/LLM workflow.
利益披露 Disclosure
T. Kang, Tempus AI Employment, Stock. E. Dougherty, Tempus AI Employment, Stock. S. Mishra, Tempus AI Employment, Stock. R. Godart, Tempus AI Employment, Stock. A. Saha, Tempus AI Employment, Stock. J. Wills, Tempus AI Employment, Stock. V. L. Chiou, Tempus AI Employment, Stock. M. A. Berezina, Tempus AI Employment, Stock. K. Nagpal, Tempus AI Employment, Stock.

← 返回 AACR 2026 检索