PO.BCS02.02 · 生物信息与计算

从混乱到列:使用CIDER进行高准确度临床数据提取

From chaos to columns: High-accuracy clinical data extraction with CIDER

海报缩略图:从混乱到列:使用CIDER进行高准确度临床数据提取
编号 2738 展板 2 时间 4/20 02:00–05:00 区域 Section 3 主讲 Balazs Gyorffy, MD;PhD
分会场 Large Language Models in the Clinic
查看 PDF 下载 PDF 🔒 查看 / 下载完整 PDF 需登录并开通下载套餐 · 查看套餐 / 开通 AACR 官方页面

作者与单位 Authors & Affiliations

Mate Posta, Aida Figler, Zsofia Dobolyi, Balazs Gyorffy

Semmelweis University, Budapest, Hungary

摘要 Abstract

中文摘要
非结构化医疗记录的分析是临床研究和医疗保健中的一项关键挑战。大语言模型(LLM)为从叙述性文本中提取结构化信息提供了变革性的机会;然而,它们在医疗环境中的使用受到安全性、伦理和可重复性问题的限制。在此我们介绍CIDER(临床数据提取器),一个本地部署的、开源的基于LLM的系统,专为医疗文档的安全分析而设计。 CIDER通过一个自动化流程运行,该流程整合了基于vLLM的推理、预定义的数据模式和提示工程的提取规则,将非结构化临床文本转换为结构化变量。该系统处理批量上传,使用微调模型解析报告,并生成标准化输出表格以供直接分析使用。我们评估了CIDER从真实世界匈牙利语病理和组织学记录中提取结构化临床数据的能力。以Qwen3-VL-32B-FP8模型为骨干,我们分析了2046份病理记录,并在六个关键临床变量上验证了模型的输出:性别、T分期、N分期、原发肿瘤器官、手术年份和肿瘤大小。所提取的数据与人工挖掘的数据进行了比较。 当有人工数据可用时,提取准确度在性别(99.4%,1971/1982一致)、T分期(95.34%,879/922)、N分期(92.19%,437/474)、手术年份(97.94%,1998/2040)和原发肿瘤器官(95.52%,1771/1854)方面均非常高。最大肿瘤大小达到77.05%的准确度(1333/1730一致)。值得注意的是,CIDER还能够在人工注释缺失的情况下检索出临床相关信息,识别出额外的实例:性别(n=64)、T分期(n=780)、N分期(n=213)、肿瘤大小(n=291)、手术年份(n=6)和原发肿瘤器官(n=15)。 总之,CIDER在所评估的参数上展示出强劲的性能。这些结果表明,一个本地部署的、开源的LLM系统可以在从复杂的非英语医疗文本中提取结构化数据方面达到接近专家水平的准确度。通过完全在机构基础设施内运行,CIDER确保了完整的数据主权,并为自动化医疗记录解读提供了可扩展的解决方案,支持多语言医疗环境中的研究、注册库开发和临床决策。CIDER平台可公开访问:https://llm.gyorffylab.com/cider。
查看英文原文 English abstract
The analysis of unstructured medical records represents a crucial challenge in clinical research and healthcare. Large Language Models (LLMs) offer a transformative opportunity to extract structured information from narrative text; however, their use in medical environments is limited by security, ethical, and reproducibility issues. Here we present CIDER (ClinIcal Data ExtractoR), a locally deployed, open-source LLM-based system designed for the secure analysis of medical documentation. CIDER operates through an automated pipeline integrating vLLM-based inference, predefined data schemas, and prompt-engineered extraction rules to convert unstructured clinical text into structured variables. The system processes batch uploads, parsed reports using a fine-tuned model, and generates standardized output tables for direct analytical use. We evaluated CIDER's ability to extract structured clinical data from real-world Hungarian-language pathology and histology records. Using the Qwen3-VL-32B-FP8 model as the backbone, we analyzed 2046 pathological records and validated the model's outputs across six key clinical variables: sex, T stage, N stage, primary tumor organ, year of surgery, and tumor size. The extracted data were compared with manually mined data. When manual data were available, extraction accuracy was very high for sex (99.4%, 1971/1982 identical), T stage (95.34%, 879/922), N stage (92.19%, 437/474), year of surgery (97.94%, 1998/2040), and primary tumor organ (95.52%, 1771/1854). The largest tumor size reached an accuracy of 77.05% (1333/1730 identical). Notably, CIDER was also capable of retrieving clinically relevant information in cases where manual annotations were missing, identifying additional instances for sex (n=64), T stage (n=780), N stage (n=213), tumor size (n=291), year of surgery (n=6), and primary tumor organ (n=15). In summary, CIDER demonstrated strong performance across the evaluated parameters. These results show that a locally deployed, open-source LLM system can achieve near-expert level accuracy in structured data extraction from complex, non-English medical texts. By operating entirely within institutional infrastructure, CIDER ensures full data sovereignty and provides a scalable solution for automated medical record interpretation, supporting research, registry development, and clinical decision-making in multilingual healthcare environments. The CIDER platform is publicly accessible at https://llm.gyorffylab.com/cider.
利益披露 Disclosure
M. Posta, None.. A. Figler, None.. Z. Dobolyi, None.. B. Gyorffy, None.

← 返回 AACR 2026 检索