PO.BCS02.02 · 生物信息与计算
从真实世界MTB数据中提取基于大语言模型的分子层面患者相似性
Large language model-derived molecular patient similarity from real-world MTB data
作者与单位 Authors & Affiliations
摘要 Abstract
中文摘要
背景:
分子肿瘤委员会(MTB)经常评估具有罕见分子特征的患者,而此类患者的前瞻性证据匮乏。为应对这一挑战,我们研究了大语言模型(LLM)能否将真实世界的MTB文档结构化为可分析的表征,以识别分子层面相似的患者。
方法:
使用NVIDIA Nemotron-49B对Charité MTB(2020-2025)讨论的2,788例患者的MTB文档和506例患者的病理报告进行结构化处理。提取了关于既往靶向治疗、免疫组化(IHC)和分子改变(MOL;CiViC加权)的文本摘要和结构化特征,共产生1,736个文本、二元和数值特征。LLM输出使用bge-multilingual-gemma2进行向量嵌入。患者相似性的计算采用嵌入的余弦距离,以及结构化特征的Jaccard(二元)和绝对误差(数值)。相似性排名通过倒数排名融合进行集成,并使用MRR和nDCG进行评估,与基于原始MTB文档的BM25和嵌入基线进行比较。作为真值的替代指标,将MTB推荐意见结构化为行动方案、药物、药物类别和临床试验要素,并用作相似性目标。
结果:
基于LLM的信息提取针对30份人工注释病理报告进行评估,在突变检测方面达到96%的F1分数,在IHC方面达到92%。集成相似性方法——通过倒数排名融合结合结构化特征相似性、摘要嵌入相似性和BM25——与MTB推荐意见的一致性最高(MRR@1000 = 25.8,nDCG@10 = 10.1),优于单独的BM25(22.8,8.6)或MTB文档的文本嵌入相似性(22.2,8.3),所有改进均具有统计学显著性(p < 0.01)。源自MTB文档的结构化特征(20.2,9.2)优于源自病理报告的特征(17.9,8.0)。总体而言,这些发现表明LLM衍生的表征相比基于文本和嵌入的基线,改善了与治疗一致的患者相似性。在所有病例中,共识别出486个独特的治疗实体;在290个药物实体中,163个被推荐给不止一名患者,65%的患者与另一病例至少共享一项推荐。
结论:
LLM衍生的临床-分子表征使得在真实世界MTB数据集中可扩展地检索分子匹配的病例成为可能。该方法支持机构病例库的构建和系统性病例系列的生成,增强了精准肿瘤学中的证据生成。
[O.S.和S.L.对本工作贡献相同。]
查看英文原文 English abstract
Background:
Molecular tumor boards (MTBs) frequently evaluate patients with rare molecular profiles where prospective evidence is scarce. To address this challenge, we investigated whether large language models (LLMs) can structure real-world MTB documentation into analyzable representations to identify molecularly similar patients.
Methods:
MTB documentation from 2,788 patients and pathology reports from 506 patients discussed at the Charité MTB (2020-2025) were structured using NVIDIA Nemotron-49B. Textual summaries and structured features on prior targeted therapies, immunohistochemistry (IHC), and molecular alterations (MOL; CiViC-weighted) were extracted, yielding 1,736 textual, binary and numeric features. The LLM output was vector embedded using bge-multilingual-gemma2. Patient similarity was computed using cosine distance for embeddings, and Jaccard (binary) and absolute error (numeric) for structured features. Similarity rankings were ensembled via reciprocal-rank fusion and evaluated using MRR and nDCG, against BM25 and embedding baselines on raw MTB documentation. As a proxy for ground truth, MTB recommendations were structured into course of action, drug, agent class, and clinical trial elements and used as similarity targets.
Results:
LLM-based information extraction achieved F1 scores of 96% for mutation detection and 92% for IHC, evaluated against 30 manually annotated pathology reports. The ensemble similarity method - combining structured-feature similarity, summary-embedding similarity, and BM25 via reciprocal-rank fusion - showed the highest alignment with MTB recommendations (MRR@1000 = 25.8, nDCG@10 = 10.1), outperforming BM25 alone (22.8, 8.6) or text-embedding similarity of MTB documentation (22.2, 8.3), with all improvements statistically significant (p < 0.01). Structured features derived from MTB documentation (20.2, 9.2) outperformed those derived from pathology reports (17.9, 8.0). Together, these findings indicate that LLM-derived representation improved therapy-aligned patient similarity over text- and embedding-based baselines. Across all cases, 486 unique therapeutic entities were identified; among 290 drug entities, 163 were recommended to more than one patient, and 65% of patients shared at least one recommendation with another case.
Conclusion:
LLM-derived clinical-molecular representations enabled scalable retrieval of molecularly matched cases in real-world MTB datasets. This approach supports institutional case library formation and systematic case series generation, enhancing evidence generation in precision oncology.
[O.S. and S.L. contributed equally to this work.]
利益披露 Disclosure
S. Lugani, None..
O. Serbetci, None..
A. Reinicke, None..
B. Körtum, None..
B. Özdin, None..
T. Debertshäuser, None.
D. Modest,
Servier ), Travel, Gift.
Amgen ), Travel, Gift.
Merck Gift.
Sanofi Gift.
BMS Gift.
MSD Gift.
AstraZeneca Gift.
Pierre Fabre Gift.
GSK Gift.
Seagen Gift.
G1 Gift.
Onkowissen Gift.
COR2ED Gift.
Taiho Gift.
Takeda Gift.
Incyte Gift.
Cureteq Gift.
IKF Gift.
AIO Studien gGmbH Gift.
Regeneron Gift.
U. Keilholz, None..
U. Leser, None..
M. Benary, None.