PO.BCS02.02 · 生物信息与计算
Arkangel AI、OpenEvidence、ChatGPT、Medisearch:它们是否客观上达到了医学标准?对医疗领域LLM的真实世界评估
Arkangel AI, OpenEvidence, ChatGPT, Medisearch: Are they objectively up to medical standards? A real-life assessment of LLMs in healthcare
该海报暂无可下载的资料
AACR 官方页面
作者与单位 Authors & Affiliations
摘要 Abstract
中文摘要
背景:大语言模型(LLM)在医疗领域的应用日益增多,但标准化基准测试未能反映其在真实世界情景中的有效性和安全性。评估其质量对于安全地融入实践至关重要。
方法:由独立的专科医生编写四个虚构的临床情景,并在四个对话式智能体中进行测试:ArkangelAI、OpenEvidence、ChatGPT和Medisearch。每个情景包含四个问题。由四名外部临床医生使用八项标准的Likert量表评估回答:1-2 = 不满意,3 = 中立,4-5 = 满意,6 = 不适用。评估标准考虑了正确性、共识性、偏倚、诊疗标准、信息更新程度、患者安全、参考文献中的真实来源以及情境感知能力。响应时间以中位数/四分位距(IQR)测量。结果以频率形式报告。应用了假设检验(alpha = 0.05)。
结果:共有128个问答对。ArkangelAI-Deep满意度最高(92.9%),其次是OpenEvidence(83.6%)、ChatGPT-Deep(80.5%)和Medisearch(71.1%)。大多数不满意集中在参考文献真实来源这一标准上:GPT-Personalized为75%,GPT-Regular为97%。相反,ArkangelAI-Deep、ChatGPT-Deep和OpenEvidence获得了100%的满意度。所有模型在正确性和与共识一致性方面表现良好。ChatGPT在无偏倚回答方面得分最低。对患者最安全的是GPT-Personalized,其次是ArkangelAI-Deep。Medisearch的响应时间最快(18秒),而GPT-Deep(13分钟)和ArkangelAI-Deep(7.4分钟)最慢,显示出深度与可用性之间的权衡。
结论:ArkangelAI-Deep和OpenEvidence始终优于其他模型,而Medisearch和GPT-Regular存在显著局限性。这些结果强调了需要建立标准化框架,以确保LLM在医疗领域的安全使用。
查看英文原文 English abstract
Background: Large language models (LLMs) are increasingly used in healthcare, but standardized benchmarks fail to capture their validity and safety in real-world scenarios. Evaluating their quality is critical for safe integration into practice.
Methods: Four fictitious clinical vignettes were developed by independent specialists and tested in four conversational agents: ArkangelAI, OpenEvidence, ChatGPT, and Medisearch. Each vignette included four questions. Responses were evaluated by four external clinicians using an eight-criterion Likert scale: 1-2 = dissatisfaction, 3 = neutral, 4-5 = satisfaction, 6 = not applicable. The criteria considered correctness, consensus, bias, standard of care, updated information, patient safety, real sources in references, and context-awareness. Response times were measured with medians/interquartile ranges (IQR). Results were reported as frequencies. Hypothesis tests were applied (alpha= 0.05).
Results: There were 128 Question-answer pairs. ArkangelAI-Deep had the highest satisfaction (92.9%), followed by OpenEvidence (83.6%), ChatGPT-Deep (80.5%), and Medisearch (71.1%). Most dissatisfaction was for the real-source-of-references criteria: GPT-Personalized 75%, GPT-Regular 97%. Conversely, ArkangelAI-Deep, ChatGPT-Deep, and OpenEvidence obtained 100% satisfaction. All performed well in correctness and agreement with the consensus. ChatGPT was the lowest-scoring in non-biased answers. The safest for patients was GPT-Personalized, followed by Arkagel AI-Deep. Medisearch had the fastest response time (18 s), while GPT-Deep (13 min) and ArkangelAI-Deep (7.4 min) were slowest, showing a trade-off between depth and usability.
Conclusions: ArkangelAI-Deep and OpenEvidence consistently outperformed others, while Medisearch and GPT-Regular had significant limitations. These results underscore the need for standardized frameworks to ensure safe use of LLMs in healthcare.
利益披露 Disclosure
N. Castano -Villegas, None..
M. Villa, None..
K. Monsalve, None..
I. Llano, None..
L. Velásquez, None..
J. Zea, None.