PO.BCS01.12 · 生物信息与计算

一款用户友好、无需编程、符合HIPAA规范的大规模表格数据自动化分析应用

A user-friendly, no-code, application for HIPAA-compliant automated analysis of tabular data at scale

海报缩略图:一款用户友好、无需编程、符合HIPAA规范的大规模表格数据自动化分析应用
编号 5507 展板 12 时间 4/21 02:00–05:00 区域 Section 4 主讲 Paraic Kenny, PhD
分会场 New Software Tools for Data Analysis
查看 PDF 下载 PDF 🔒 查看 / 下载完整 PDF 需登录并开通下载套餐 · 查看套餐 / 开通 AACR 官方页面

作者与单位 Authors & Affiliations

Paraic A. Kenny

Gundersen Medical Foundation, La Crosse, WI

摘要 Abstract

中文摘要
临床环境中的大量研究涉及审阅电子病历中大量往往是非结构化的数据,这一过程耗时且需要一支受过训练的临床研究协调员、住院医师和/或专科培训医师队伍。诸如ChatGPT之类的大语言模型已展现出高效分析和总结大量文本的能力,有望大大加速临床环境中基于病历审阅的研究。对个人健康信息(PHI)泄露的担忧严重限制了将商用聊天机器人用于此目的的潜力。在治理良好的系统中,安全、隔离防护的内部LLM聊天机器人可以解决与PHI泄露相关的担忧。尽管以每次一个提示的方式处理大型研究数据相较于以往纯人工流程可能带来可观的效率提升,但通过自动化提示提交和结果检索过程可以实现远为更大的效率。 为解决这些问题,我们开发了一款用户友好的应用,利用在Emplify Health的Azure云环境中内部部署的ChatGPT 4o-mini,大规模处理任何形式的表格数据。该应用导入表格数据(Excel或文本格式),引导用户完成简单直接的提示工程流程,然后自动逐行将数据提交给LLM,并按用户指定的方式检索并制表返回的数据。 该应用最初的开发目标是快速推理数千份病理报告,以识别其中对应癌症病例的报告,并针对这些病例自动提取多个癌症相关参数。在这一早期成功之后,我们开发了一款完全通用、数据无关的应用,该应用在我们的研究机构内获得了广泛应用,被部署于癌症登记处、心脏病学CT报告、影像学叙述报告和临床就诊记录等多样化的数据源。通过这种方式,该应用大幅缩短了数据检索所需的时间,从而显著加速了研究项目,使我们的人员能够将更多时间投入到数据分析等更高价值的活动中。
查看英文原文 English abstract
Much research in clinical settings involves review of large quantities of often unstructured data in electronic medical records, which is time consuming and requires an educated workforce of clinical research coordinators, residents and/or fellows. Large language models, such as ChatGPT, have demonstrable capabilities for efficiently analyzing and summarizing large volumes of text, offering the potential to greatly accelerate chart review-based research in clinical settings. Concerns regarding leakage of personal health information (PHI) sharply limit the potential to utilize commercial chatbots for this purpose. Secure, fire-walled in-house LLM chatbots in a well-governed system can solve concerns related to PHI leakage. Although processing data for large studies one-prompt-at-a-time may offer substantial efficiency gains versus prior human-only processes, far greater efficiencies can be achieved by automating the prompt submission and retrieval process. To address these concerns, we developed a user-friendly application for processing any form of tabular data at scale using an in-house implementation of ChatGPT 4o-mini on Emplify Health's Azure cloud environment. The application imports tabular data (in excel or text format), guides the user through a straightforward prompt engineering process and then automatically submits the data, row-by-row, to the LLM and retrieves and tabulates the returning data in the manner specified by the user. The application was initially developed with the goal of rapidly reasoning through thousands of pathology reports to identify those which correspond to cancer cases and, for those cases, automatically extracting multiple cancer-related parameters. Following this early success, we developed a completely generalized data agnostic application which has found widespread utility within our research institute, being deployed on diverse data sources such as cancer registries, cardiology CT reports, imaging narratives and clinical encounter notes. In this way, the application has substantially accelerated research projects by dramatically reducing the time needed to retrieve data, enabling our personnel to spend more time on higher yield activities such as data analysis.
利益披露 Disclosure
P. A. Kenny, Tempus AI ).

← 返回 AACR 2026 检索