PO.BCS02.02 · 生物信息与计算
利用本地部署的peta级语言模型从病理报告中进行MSI分类以快速整理结肠癌队列
Rapid curation of colon cancer cohorts using on-site peta-scale language model for MSI classification from pathology reports
作者与单位 Authors & Affiliations
摘要 Abstract
中文摘要
背景:病理报告高度异质,使得难以提取队列整理所需的结构化信息。早期的机器学习模型必须针对每个特定任务单独训练,限制了灵活性。较新的大语言模型(LLM)能够遵循书面指令——“提示词”——直接从文本中识别或汇总信息,使单一模型无需重新训练即可执行多项任务。这些模型此前因体积过大而无法在医院系统内运行,但紧凑的GPU平台(如NVIDIA DGX)现已使包含数千亿参数的最先进LLM能够在本地部署成为可能。这使得灵活的信息提取工作流得以快速原型化。本研究以从结直肠病理报告中提取错配修复(MMR)状态为例,评估了本地部署的LLM在快速队列整理中的应用。
方法:从两个研究队列(ColoCare和SeroNet)中共收集了82份结直肠病理报告。大多数报告包含用于确定MSI(微卫星不稳定性)状态的MLH1、MSH2、MSH6和PMS2染色结果。两个本地部署的LLM(GPT-OSS 20B和120B)根据记录的MMR表达或任何直接的MSI陈述,将每个病例分类为MSI、MSS或不明确,其中缺乏明确MMR信息的病例被归入不明确类别。将模型预测与人工MSI注释进行比较。对于模型分类为不明确但注释者判断为MSI或MSS的报告,提示每个模型重述并论证其判断,以评估推理的一致性。
结果:两个模型在82份报告中均表现良好。20B模型的准确率为96.3%,而120B模型为95.1%。对于MSI,120B模型显示出更高的敏感性(87.5% vs. 77.8%)和F1分数(87.5% vs. 82.4%),两个模型均保持高精确度和特异性(分别≥ 87.5%和≥ 98.6%)。对于MSS,两个模型的表现均很强,120B模型的精确度(97.6% vs. 95.5%)和特异性(97.4% vs. 94.6%)略高。20B模型完美识别了所有不明确病例,而120B模型将两份MSS报告归入“无法评估”。在审查这些差异时,模型与其原始分类保持一致,但在如何解读“MSI IHC low prob”这一短语上存在差异:120B模型将其视为不确定并返回“未知”,而20B模型在MMR表达完整的背景下将同样的措辞解读为支持MSS。
结论:这些发现证明了在隐私受限的医院环境中利用最先进LLM的实际效用。进一步的工作将评估在更具技术挑战性场景中的性能,并阐明在将LLM整合到临床工作流时,模糊措辞如何影响质量控制。
查看英文原文 English abstract
Background: Pathology reports are highly heterogeneous, making it difficult to extract structured information needed for cohort curation. Earlier machine-learning models had to be trained separately for each specific task, limiting flexibility. Newer large language models (LLMs) can follow written instructions-“prompts”-to identify or summarize information directly from the text, allowing a single model to perform many tasks without retraining. These models were previously too large to run within hospital systems, but compact GPU platforms, such as NVIDIA DGX, now make state-of-the-art LLMs containing hundreds of billions of parameters feasible to deploy locally. This enables rapid prototyping of flexible extraction workflows. Here, we evaluate locally deployed LLMs for rapid cohort curation using mismatch repair (MMR) status extraction from colorectal pathology reports as an example.
Methods: A total of 82 colorectal pathology reports were collected across two study cohorts (ColoCare and SeroNet). Most reports included MLH1, MSH2, MSH6, and PMS2 staining results used to determine MSI (microsatellite instability) status. Two locally deployed LLMs (GPT-OSS 20B and 120B) classified each case as MSI, MSS, or ambiguous based on documented MMR expression or any direct MSI statement, with cases lacking clear MMR information assigned to an ambiguous category. Model predictions were compared with manual MSI annotations. For reports the model classified as ambiguous but annotators judged MSI or MSS, each model was prompted to restate and justify its decision to assess reasoning consistency.
Results: Both models performed well across the 82 reports. The 20B model achieved 96.3% accuracy, while the 120B model achieved 95.1%. For MSI, the 120B model showed higher sensitivity (87.5% vs. 77.8%) and F1 score (87.5% vs. 82.4%), with both models maintaining high precision and specificity (≥ 87.5% and ≥ 98.6%, respectively). For MSS, performance was strong for both models, with the 120B model showing slightly higher precision (97.6% vs. 95.5%) and specificity (97.4% vs. 94.6%). The 20B model perfectly identified all ambiguous cases, whereas the 120B model assigned two MSS reports to “unable to be assessed.” In reviewing these discrepancies, the models remained consistent with their original classifications but differed in how they interpreted the phrase “MSI IHC low prob”: the 120B model viewed it as indeterminate and returned “unknown,” while the 20B model interpreted the same wording as supporting MSS in the context of intact MMR expression.
Conclusion: These findings demonstrate the practical utility of leveraging state-of-the-art LLMs within privacy constrained hospital settings. Further work will assess performance in more technically challenging scenarios and clarify how ambiguous wording impacts quality control when integrating LLMs into clinical workflows.
利益披露 Disclosure
J. J. Levy, None..
N. Nguyen, None..
M. Le, None..
K. Yao, None..
J. Figueiredo, None.