PO.BCS02.01 · 生物信息与计算

推理引导的检索提升了从临床病历中匹配肿瘤学试验入组资格的能力

Reasoning‑guided retrieval improves oncology trial eligibility matching from clinical notes

海报缩略图:推理引导的检索提升了从临床病历中匹配肿瘤学试验入组资格的能力
编号 29 展板 14 时间 4/19 02:00–05:00 区域 Section 2 主讲 Patrycja Krawczuk
分会场 Agentic AI in Cancer
查看 PDF 下载 PDF 🔒 查看 / 下载完整 PDF 需登录并开通下载套餐 · 查看套餐 / 开通 AACR 官方页面

作者与单位 Authors & Affiliations

Vivek Shetye, Patrycja Nikol Krawczuk, Ryan Godart, Arpita Saha, Gabriel Altay, Chelsea Osterman, Gena Rangel, Samantha Garrett, Victoria L. Chiou, Kunal Nagpal

Tempus AI, Inc., Chicago, IL

摘要 Abstract

中文摘要
目的:肿瘤学试验入组不足因需要人工筛查非结构化临床记录以对照复杂的入组标准而加剧。经检索机制增强的大语言模型(LLM)在自动化该流程方面展现出前景,但往往无法捕捉复杂的多标准入组决策所需的分散证据。我们假设,一种赋予LLM能力使其自主、迭代地搜索、综合并验证相关信息的智能体(agentic)检索策略,能够提升入组分类中证据的完整性与准确性。 方法:评估涵盖两类任务:(i) 复杂入组推理和 (ii) 生物标志物提取。复杂任务基于两项入组标准繁复的非小细胞肺癌(NSCLC)研究:III期不可切除NSCLC和转移性NSCLC。每个查询代表一次患者层面的入组资格评估(共 n=618 个查询),需要整合多种证据类型,包括分子检测结果、既往治疗和临床状态。生物标志物任务包括针对关键基因组改变(EGFR、ESR1、RAS)的148次评估。我们将一个检索固定的、基于相似度文本片段的检索增强生成(RAG)系统与一种执行多达8次自适应搜索的智能体检索方法进行比较。该智能体自主评估检索到的证据、重新表述其查询,并迭代扩展搜索范围以构建全面的上下文。两个系统均使用LLM(Gemini 2.5 Pro),每个查询审阅多达64个文本片段。模型输出以专家审校的金标准为基准,采用F1分数、召回率和准确率进行评价。 结果:智能体检索在复杂入组推理任务中提升了性能。在III期NSCLC中,准确率从68%升至80%(+12.7个百分点;95% CI:7.9–17.7),这得益于召回率提升15%(69%至84%)以及F1分数从79%提高9%至88%。在转移性NSCLC中,准确率从77%提高至84%(+6.3个百分点;95% CI:1.3–11.6),F1分数从86%提高至90%。相比之下,生物标志物提取任务两种方法表现相当(准确率95–96%,F1为96%),表明在检索复杂度较低时,智能体推理带来的收益极小。 结论:本研究证明了智能体检索在精准、可扩展的肿瘤学试验筛查中的转化价值。通过实现自适应证据综合而非静态检索,该系统提高了入组评估的完整性,并最大限度减少了人工审查工作量。
查看英文原文 English abstract
Purpose: Under-enrollment in oncology trials is exacerbated by the manual effort required to screen unstructured clinical records against complex eligibility criteria. Large language models (LLMs) augmented with retrieval mechanisms show promise for automating the process but often fail to capture dispersed evidence required for complex, multi-criterion eligibility decisions. We hypothesized that an agentic retrieval strategy, which empowers an LLM to autonomously and iteratively search, synthesize, and validate relevant information, would enhance evidence completeness and accuracy in eligibility classification. Methods: Evaluation encompassed two task types: (i) complex eligibility reasoning and (ii) biomarker extraction. The complex tasks drew on two non-small cell lung cancer (NSCLC) studies with intricate eligibility criteria: stage III unresectable NSCLC and metastatic NSCLC. Each query represented a patient-level eligibility assessment (total n=618 queries) that required integrating multiple evidence types, including molecular findings, prior therapies, and clinical status. The biomarker task included 148 evaluations targeting key genomic alterations (EGFR, ESR1, RAS).We compared a retrieval-augmented generation (RAG) system that retrieved fixed, similarity-based text chunks with an agentic retrieval approach performing up to 8 adaptive searches. The agent autonomously assessed retrieved evidence, reformulated its queries, and iteratively expanded the search scope to assemble a comprehensive context. Both systems, using LLM (Gemini 2.5 Pro), reviewed up to 64 text chunks per query. Model outputs were benchmarked against expert-curated ground truth using F1-score, recall, and accuracy. Results: Agentic retrieval improved performance across complex eligibility reasoning tasks. In the stage III NSCLC, accuracy rose from 68% to 80% (+12.7 pp; 95% CI: 7.9-17.7), driven by a 15% gain in recall (69% to 84%) and a 9% increase in F1-score from 79% to 88%. In metastatic NSCLC, accuracy improved from 77% to 84% (+6.3 pp; 95% CI: 1.3-11.6) and the F1-score from 86% to 90%. In contrast, the biomarker extraction task showed comparable performance (accuracy 95-96%, F1 96%) for both methods, indicating minimal benefit from agentic reasoning where retrieval complexity was low. Conclusions: This study demonstrates the translational value of agentic retrieval for precise, scalable oncology trial screening. By enabling adaptive evidence synthesis rather than static retrieval, the system improves the completeness of eligibility assessment and minimizes manual review effort.
利益披露 Disclosure
V. Shetye, Tempus AI, Inc. Tempus AI Consultant. P. N. Krawczuk, Tempus AI, Inc. Employment, Stock. R. Godart, Tempus AI, Inc. Employment, Stock. A. Saha, Tempus AI, Inc. Employment, Stock. G. Altay, Tempus AI, Inc. Employment, Stock. C. Osterman, Tempus AI, Inc. Employment, Stock. G. Rangel, Tempus AI, Inc. Employment, Stock. S. Garrett, Tempus AI, Inc. Employment, Stock. V. L. Chiou, Tempus AI, Inc. Employment, Stock. K. Nagpal, Tempus AI, Inc. Employment, Stock.

← 返回 AACR 2026 检索