PO.BCS02.02 · 生物信息与计算
利用大型语言模型将病理报告分类到本体层级结构中
Leveraging large language models to classify pathology reports into ontological hierarchies
作者与单位 Authors & Affiliations
摘要 Abstract
中文摘要
自由文本病理报告在风格、术语和统一性方面各不相同,使得为研究队列开发进行一致的诊断提取变得困难。这一问题在大型档案库中被放大,其中不断演变的分类和有意的模糊性阻碍了标准化的分类诊断。诸如OncoTree之类的结构化本体通过强制使用编码来解决这一问题。虽然此处以OncoTree为例,但该方法可推广至任何层级本体,包括当前的WHO分类。
利用前沿的LLM,我们开发了OncoPath,一个在结构化本体内对病理报告进行分类的应用程序。我们通过结构化的树遍历应用OncoTree层级结构来测试OncoPath,从宽泛的解剖学类别逐步推进到具体的亚型。首先评估高层级类别,然后评估可能的子节点,每个节点的子节点都会显示出来,以便模型能够预判下游选项。对诊断路径进行汇总,模型以置信等级(<50%、50-90%、>90%)选择最佳编码。我们评估了LLaMA 3.3 70B Instruct的多种配置。外科病理和血液病理报告被分别处理,准确性通过病理学家审查针对10个最常见的分子编码进行验证,这些编码涵盖了49.6%的外科病例和84.9%的血液病理病例。
该分类器以高准确性将诊断映射到OncoTree编码。基于血液病理的编码在使用完整报告时效果良好,而基于外科的编码在使用有针对性的摘要时表现最佳,因为报告通常包含来自不同部位和诊断的多个部分。血液病理报告达到90-100%的准确性;外科编码达到80-100%,但HGSOC(高级别浆液性卵巢癌)仅为60%,这反映出即使对病理学家而言,确定肿瘤原发部位也存在固有困难。编码与原始报告一并存储,可在研究和运营中直接重复使用。该遍历方法提供了可审计性、置信水平,以及在诊断改变时自动重新评估的能力。基于LLM辅助的OncoTree遍历提供了一种可扩展的方法,将各种病理报告转换为结构化编码。通过在保留诊断细微差别的同时标准化报告,OncoPath实现了回顾性队列开发以及对发病率、进展和诊疗模式的分析。该框架可轻松扩展至其他结构化疾病本体,为将非结构化报告转化为标准化、可重用数据建立了一种可推广的方法。
查看英文原文 English abstract
Free-text pathology reports vary in style, terminology, and uniformity, making consistent diagnostic extraction difficult for research cohort development. This issue is amplified in large archives, where evolving classifications and intentional ambiguity hinder standardized, categorical diagnoses. Structured ontologies such as OncoTree address this by enforcing codes. While OncoTree is used here as an example, the approach generalizes to any hierarchical ontology, including current WHO classifications.
Using cutting-edge LLMs, we developed OncoPath, an application that classifies pathology reports within structured ontologies. We tested OncoPath by applying OncoTree hierarchy through a structured tree-walk, progressing from broad anatomic categories to specific subtypes. High-level categories are assessed first, then plausible child nodes, with each node's children displayed so the model can anticipate downstream options. Diagnostic paths are summarized, and the model selects the optimal code with a confidence tier (<50%, 50-90%, >90%). Multiple configurations of LLaMA 3.3 70B Instruct were evaluated. Surgical and hematopathology reports were processed separately, and accuracy was verified by pathologist review against the 10 most prevalent molecular codes, covering 49.6% of surgical and 84.9% of hematopathology cases.
The classifier mapped diagnoses to OncoTree codes with high accuracy. Heme-based coding worked well with the full report, whereas Surg-based coding performed best with a targeted summary because reports often contained multiple parts from different sites and diagnoses. Heme reports reached 90-100% accuracy; Surg codes reached 80-100%, except HGSOC (high grade serous ovarian carcinoma) at 60%, reflecting the inherent difficulty, even for pathologists, of determining the tumor's primary site of origin. Codes were stored with original reports, enabling direct reuse in research and operations. The traversal approach provided auditability, confidence levels, and automatic re-evaluation when diagnoses changed.An LLM-assisted OncoTree traversal provides a scalable method to convert diverse pathology reports into structured codes. By standardizing reports while preserving diagnostic nuance, OncoPath enables retrospective cohort development and analyses of incidence, progression, and practice patterns. This framework can readily extend to other structured disease ontologies, establishing a generalizable approach for transforming unstructured reports into standardized, reusable data.
利益披露 Disclosure
B. Fried, None..
A. Kamali, None..
C. Colorado-Jimenez, None..
M. Pulitzer, None..
D. Kim, None..
L. Boiocchi, None..
A. Chan, None..
M. Yabe, None..
M. Roshal, None..
S. Aijazuddin, None..
A. Dogan, None..
C. Vanderbilt, None..
K. H. Bilal, None..
G. Goldgof, None.