PO.BCS02.02 · 生物信息与计算
用于分析异质性乳腺癌患者基因组报告的LLM框架的开发
Development of an LLM framework for analysis of heterogeneous breast cancer patients genomic reports
作者与单位 Authors & Affiliations
摘要 Abstract
中文摘要
背景:临床基因组报告的数量和多样性不断增加,给各医疗机构解读和恰当实施精准肿瘤学研究带来了重大挑战。来自多个不同供应商(如Invitae、Ambry Genetics、Foundation Medicine)的报告通常为PDF格式,且在panel设计、基因覆盖范围和报告标准方面存在显著异质性,这阻碍了在电子病历中高效开展回顾性数据挖掘和患者队列识别。我们试图开发一个框架来查询所有这些报告。
方法:我们提取了接受新辅助化疗的乳腺癌患者的基因组报告,并开发了MolHarmonizer——一个新颖、可扩展的框架,利用Python和Gemini LLM,旨在处理并协调来自不同多供应商报告的基因组数据。Gemini LLM被明确用于稳健地提取、标准化和结构化关键基因组特征,将非结构化数据转化为统一、可查询的数据集。
结果:我们的MolHarmonizer框架成功处理了来自1703例乳腺癌患者(2006-2023年)、来自23家不同公司的1,147份基因组报告,展现了提取和标准化关键可操作生物标志物的稳健能力。数据来源包括Invitae(n=554)、Ambry Genetics(n=189)、Natera(n=95)、Mayo Clinic(n=88)、Tempus(n=63)、Guardant Health(n=47)和Foundation Medicine(n=37),其余供应商各贡献少于20份报告。在这些样本中,827/1147(72.1%)为胚系(血液/唾液)样本。大多数患者因个人/家族史而接受panel检测(n=925)。总体而言,413/1147(36.0%)份报告至少识别出一个突变。就乳腺癌而言,75份报告显示BRCA1/2突变(37例BRCA1、37例BRCA2,一例患者同时携带BRCA1和2)。识别出的其他突变包括:PIK3CA(n=40)、TP53(n=96)、PTEN(n=21)、ESR1(n=13)和AKT1/2(n=8)。
结论:MolHarmonizer是一个利用Gemini LLM的强大框架,通过自动化生物标志物提取和协调,有效解决了基因组数据的异质性。这使得能够快速识别队列并进行深入的回顾性分析,以获得临床见解、发现生物标志物、理解疾病史、促进新模式的发现(例如从WSI预测BRCA1突变),并加速我们新辅助乳腺癌队列内的研究。未来计划包括扩展至涵盖超过20,000例乳腺癌患者、开发用户友好的聊天机器人,并确保跨机构对各种复杂疾病的适应性。
查看英文原文 English abstract
Background: The increasing volume and diversity of clinical genomic reports pose a significant challenge across healthcare institutions for the interpretation and proper implementation of precision oncology research. Reports from multiple different vendors (e.g., Invitae, Ambry Genetics, Foundation Medicine) are typically PDFs and exhibit substantial heterogeneity in panel design, gene coverage, and reporting standards, which hinders efficient retrospective data mining and patient cohort identification within the electronic medical records. We sought to develop a framework to interrogate all of these reports.
Methods: We extracted genomic reports from patients with breast cancer treated with neoadjuvant chemotherapy and developed MolHarmonizer, a novel, scalable framework leveraging Python and Gemini LLMs, designed to process and harmonize genomic data from disparate multi-vendor reports. Gemini LLMs are employed explicitly for robust information extraction, normalization, and structuring of key genomic features, transforming unstructured data into a unified, queryable dataset.
Results: Our MolHarmonizer framework successfully processed 1,147 genomic reports from 1703 breast cancer patients (2006-2023) from 23 different companies, demonstrating robust capability to extract and standardize critical actionable biomarkers. Data sources included Invitae (n=554), Ambry Genetics (n=189), Natera (n=95), Mayo Clinic (n=88), Tempus (n=63), Guardant Health (n=47), and Foundation Medicine (n=37), with others contributing less than 20 reports. Of the samples, 827/1147 (72.1%) were germline (blood/saliva). A majority of the patients were tested using the panels due to a personal/family history (n=925). Overall, 413/1147 (36.0%) reports identified at least one mutation. For breast cancer, 75 reports showed BRCA1/2 mutations (37 BRCA1, 37 BRCA2, and one patient with both BRCA1 and 2). Other mutations identified included: PIK3CA (n=40), TP53 (n=96), PTEN (n=21), ESR1 (n=13) and AKT1/2 (n=8).
Conclusion: MolHarmonizer, a powerful framework leveraging Gemini LLMs, effectively addresses genomic data heterogeneity by automating biomarker extraction and harmonization. This enables rapid cohort identification and deep retrospective analyses for clinical insights, biomarker discovery, understanding disease history, facilitating novel pattern discovery, e.g., predicting BRCA1 mutations from WSI, and accelerating research within our neoadjuvant BC cohort. Future plans include expanding to include over 20,000 breast cancer patients, developing a user-friendly chatbot, and ensuring inter-institutional adaptability for a variety of complex diseases.
利益披露 Disclosure
K. R. Kalari, None..
T. Boyapati, None..
T. L. Hoskin, None..
S. Myla, None.