PO.BCS02.02 · 生物信息与计算

评估大语言模型在肿瘤学临床试验自动匹配中的应用

Evaluation of large language models for automated clinical trial matching in oncology

海报缩略图:评估大语言模型在肿瘤学临床试验自动匹配中的应用
编号 2739 展板 3 时间 4/20 02:00–05:00 区域 Section 3 主讲 Aakash Desai, MD
分会场 Large Language Models in the Clinic
查看 PDF 下载 PDF 🔒 查看 / 下载完整 PDF 需登录并开通下载套餐 · 查看套餐 / 开通 AACR 官方页面

作者与单位 Authors & Affiliations

Aakash Desai, Ellen McNeeley, Sanad Alhuski, Maya Khalil, Matthew Might, Rebecca Arend, Andrew Crouse, Mehmet Akce

University of Alabama at Birmingham, Birmingham, AL

摘要 Abstract

中文摘要
背景:高效的患者-试验匹配仍是肿瘤学中的关键挑战,异质性的文档记录、缺失数据以及复杂的入组标准使其更加复杂。大语言模型(LLM)通过解读非结构化的临床记录和生物标志物数据,具有自动化入组资格筛查的潜力。 方法:我们评估了6个模型:llama3.2:3b、llama3.3:70b、medgemma_27b_text_it、deepseek-r1:8b、gpt-oss20b和gpt-oss120b,针对反映肿瘤学临床试验常见入组标准的19个关键问题进行临床试验入组资格判定。数据从具有已知试验匹配结果的患者病历中提取,并分析各模型的二元(是/否)回答、置信度评分和推理摘录。评估了模型间的一致性及输出的可解释性。 结果:gpt-oss20b和gpt-oss120b模型对记录完善的标准(如可测量疾病、ECOG状态、年龄和组织可获得性)的入组资格判定均表现出高度一致,置信度评分通常高于0.90。在需要推断或文档记录不完整的标准上出现差异;gpt-oss120b在模糊病例中表现出更高的置信度和更细致的推理。两个模型均能标记缺失或不明确的数据,提供支持临床审查的推理透明度。一致性指标显示对明确标准具有较强的可靠性(Cohen's kappa >0.8),有望显著减少人工筛查负担。其余模型总体上回答质量较差,且在要求以结构化格式作答时完全无法给出连贯的回应。 结论:LLM能够准确且透明地自动化肿瘤学试验入组资格筛查的关键环节,增强人工审查流程。模型在面对不确定数据时的置信度差异凸显了持续优化的必要性,并突出了可解释AI在临床决策支持中的价值。这些发现支持将LLM整合到临床试验匹配工作流程中,以改善试验可及性和入组效率。 影响:基于LLM的自动化、可解释的临床试验匹配代表了迈向精准肿瘤学的一项有前景的进展,通过扩大患者获得个性化治疗的机会并优化试验通量。
查看英文原文 English abstract
Background: Efficient patient-trial matching remains a critical challenge in oncology, complicated by heterogeneous documentation, missing data, and complex eligibility criteria. Large Language Models (LLMs) offer potential to automate eligibility screening by interpreting unstructured clinical notes and biomarker data. Methods: We evaluated 6 models: llama3.2:3b, llama3.3:70b, medgemma_27b_text_it, deepseek-r1:8b, gpt-oss20b and gpt-oss120b for clinical trial eligibility determination across 19 key questions reflecting common eligibility criteria from oncology clinical trials. Data were extracted from patient medical records with known trial matches, and models' binary (yes/no) responses, confidence scores, and reasoning excerpts were analyzed. Concordance between models and interpretability of outputs were assessed. Results: Both gpt-oss20b and gpt-oss120b models demonstrated high agreement on eligibility determinations for well-documented criteria such as measurable disease, ECOG status, age, and tissue availability, with confidence scores commonly above 0.90. Differences emerged in criteria requiring inference or where documentation was incomplete; gpt-oss120b showed greater confidence and nuanced reasoning in ambiguous cases. Both models flagged missing or unclear data, providing reasoning transparency that supports clinical review. Concordance metrics suggested strong reliability (Cohen's kappa >0.8) for explicit criteria, with potential to significantly reduce manual screening burden. The remaining models provided poorer quality responses in general and were unable to respond coherently at all if required to provide that response in a structured format. Conclusions: LLMs can accurately and transparently automate critical components of oncology trial eligibility screening, augmenting manual review processes. Differences in model confidence with uncertain data underscore the need for ongoing refinement and highlight the value of explainable AI in clinical decision support. These findings support integrating LLMs into clinical trial matching workflows to improve trial access and enrollment efficiency. Impact: Automated, interpretable LLM-based clinical trial matching represents a promising advancement toward precision oncology by scaling patient access to tailored therapies and optimizing trial throughput.
利益披露 Disclosure
A. Desai, None.. E. McNeeley, None.. S. Alhuski, None.. M. Khalil, None.. M. Might, None.. R. Arend, None.. A. Crouse, None.. M. Akce, None.

← 返回 AACR 2026 检索