PO.MD01.01 · 分子诊断与数据
评估一款用于 AACR GENIE BPC 数据临床基因组分析的智能体式 LLM 聊天机器人
Evaluation of an agentic LLM chatbot for clinico-genomic analysis of AACR GENIE BPC data
作者与单位 Authors & Affiliations
摘要 Abstract
中文摘要
目的
基因组和临床数据的庞大体量与复杂性可能阻碍基于临床基因组数据集的高效研究,需要人工投入和专业知识。智能体式大语言模型(LLM)工作流程或许有助于加速数据处理,但现有 LLM 在此任务上的表现尚未得到充分表征。
方法
开发了一款基于智能体式大语言模型的聊天机器人,利用 Gemini-2.5-pro LLM 来解读肿瘤学研究查询,并基于 AACR GENIE BPC NSCLC 队列(2.0 公开版)自主执行一系列分析任务。将该 LLM 的表现与一套经过审编的基准问题集进行对照评估,该基准集包含源自一项已发表研究(https://pubmed.ncbi.nlm.nih.gov/37223888/)的 125 个经专家审阅的临床和基因组问题,准确性定义为与稿件报告的参考值数值一致性在 ±10% 以内。
结果
使用该聊天机器人提出了从该出版物中人工提取的 118 个问题,包括大致归类为量化队列规模(n=92)或进行统计分析(n=26)的问题。总体准确率为 42.37%。对不准确的回答进行人工审阅,并归入以下类别:无明显误差或差异来源(33.8%),其中 39.1% 与参考值的偏差 < 20%;聊天机器人推理错误(36.8%);聊天机器人未能澄清用户问题中的某一概念(22.1%);参考出版物的分析描述不够详尽以致无法复现(7.4%);用户对聊天机器人的回应有误(4.4%);聊天机器人未按预期解读分析问题(1.5%);以及不明确/其他(1.5%)。
结论
智能体式 LLM 数据分析工作流程在自动化肿瘤学数据解读的部分环节方面具有潜力,但当前的性能局限——归因于推理不一致、临床概念澄清不完整,以及需要明确规范已发表的分析方案以确保可重现性和可评估性——凸显了在这些系统能够可靠地整合入真实世界临床研究流程之前,需要在这些具体方面进一步改进模型。
查看英文原文 English abstract
Purpose
The significant volume and complexity of genomic and clinical data can hinder efficient research based on clinico-genomic datasets, requiring manual effort and specialized expertise. Agentic large language model (LLM) workflows may help accelerate data processing, but the performance of existing LLMs for this task is not well-characterized.
Methods
An agentic large language model-based chatbot was developed to leverage the Gemini-2.5-pro LLM to interpret oncology research queries and autonomously execute sequential analytic tasks based on the AACR GENIE BPC NSCLC cohort (version 2.0 public). The LLM's performance was assessed against a curated benchmark set of 125 expert-reviewed clinical and genomic questions derived from a published study (https://pubmed.ncbi.nlm.nih.gov/37223888/), with accuracy defined as numerical concordance within ±10% of manuscript-reported reference values.
Results
The chatbot was used to ask 118 questions manually extracted from the publication, including questions broadly categorized as quantifying cohort sizes (n=92) or conducting statistical analyses (n=26). The overall accuracy rate was 42.37%. Inaccurate responses were manually reviewed and assigned to the following categories: no obvious source of error or discrepancy (33.8%), where 39.1% of these deviated < 20% from the reference value; chatbot reasoning faulty (36.8%); chatbot failed to clarify a concept in the user question (22.1%); reference publication analysis insufficiently specified to replicate (7.4%); user error in response to chatbot (4.4%); chatbot did not interpret analysis question as intended (1.5%); and unclear/other (1.5%).
Conclusion
Agentic LLM data analysis workflows hold potential for automating components of oncology data interpretation, but current performance limitations, attributable to inconsistent reasoning, incomplete clarification of clinical concepts, and a need for clear specification of published analysis plans for reproducibility and evaluation, highlight the need for further model refinement in these specific areas before these systems can be reliably integrated into real-world clinical research pipelines.
利益披露 Disclosure
L. Thiriveedi, None..
K. L. Kehl, None.