PO.CL05.13 · 临床研究

EpitopeMiner:用于证据驱动的个性化癌症疫苗设计的可扩展知识挖掘

EpitopeMiner: Scalable knowledge mining for evidence-driven personalized cancer vaccine design

海报缩略图:EpitopeMiner:用于证据驱动的个性化癌症疫苗设计的可扩展知识挖掘
编号 6698 展板 9 时间 4/21 02:00–05:00 区域 Section 49 主讲 Mai Chan Lau
分会场 Vaccines and Other Immunomodulatory Agents
查看 PDF 下载 PDF 🔒 查看 / 下载完整 PDF 需登录并开通下载套餐 · 查看套餐 / 开通 AACR 官方页面

作者与单位 Authors & Affiliations

Agamjyot Singh Chadha1, Isaac Jiasheng Cheong2, Marcia Zhang2, Wei Kit Tan2, Wei Lin Tang1, Jing Quan Lim3, Solomonraj Wilson1, Choon Kiat Ong3, Bernett Lee4, Chwee Ming Lim5, Olaf Rotzschke1, Mai Chan Lau1

1Singapore Immunology Network Lab, A*STAR - Agency for Science, Technology and Research, Singapore, Singapore,2Bioinformatics Institue, A*STAR - Agency for Science, Technology and Research, Singapore, Singapore,3Lymphoma Translational Research Laboratory, Division of Cellular and Molecular Research, National Cancer Centre Singapore, Singapore, Singapore,4Lee Kong Chian School of Medicine (LKCMedicine), Nanyang Technological University Singapore, Singapore, Singapore,5Surgery Academic Clinical Programme (Surgery ACP), Duke-NUS Medical School, Singapore, Singapore

摘要 Abstract

中文摘要
背景:个性化癌症疫苗通过引发肿瘤特异性免疫应答而具有巨大前景[1-3]。一个关键挑战是识别正确的靶标——呈递在肿瘤细胞上的免疫原性蛋白序列,即表位。虽然计算流程可从肿瘤测序中预测候选表位,但实验验证成本高昂且缓慢。利用文献和数据库知识可通过实现证据驱动的高置信度靶标选择来弥合这一差距,但受限于分散在各期刊和免疫学数据库中的碎片化信息[4-5]。我们推出EpitopeMiner,它将基于序列的候选筛选与证据驱动的知识检索相结合,用于表位优先级排序。 方法:使用标准工作流程,从九名患者(包括肺癌、肉瘤、NKTL、DLBCL)的肿瘤-PBMC配对全基因组测序中预测出共25,966个肿瘤特异性表位:HLA分型(OptiType)、变异检出(采用wANNOVAR的Strelka)、MHC结合预测(NetMHCpan)以及RNA支持的蛋白改变过滤。EpitopeMiner将OpenAI的大语言模型(LLM)与一个内部检索增强生成(RAG)数据库相结合,该数据库包含:(a)来自PMC、PLOS One和Europe PMC的78,461篇全文研究文章,以及(b)来自IEDB、dbPepNeo、SystemMHC、TANTIGEN和caAtlas数据库的约260万个独特表位。EpitopeMiner包括:(i)一个筛选模块,处理表位列表,检测来自内部数据库的精确或≥7个氨基酸的部分匹配;以及(ii)一个报告模块,分析每个排名靠前的命中(由最高序列相似性和证据密度定义),生成涵盖28个免疫学关键词并附引用的LLM响应。 结果:在从九名患者预测的25,966个表位中,EpitopeMiner发现7个精确匹配,17.6%具有≥6个部分匹配;每个表位的平均处理时间为0.97秒。在使用3个肺癌驱动基因表位(KITDFGRAK、ITDFGRAKL、TDFGRAKLL)进行基准测试时,EpitopeMiner的表现优于ChatGPT和Gemini,返回了最多的相关免疫学信息——概括为(总响应数,含证据的百分比)——分别为(18,100%)、(8,100%)和(2,100%),而ChatGPT为(10,70%)、(4,50%)、(6,33%),Gemini为(7,0%)、(1,0%)、(1,0%)。此外,EpitopeMiner为每个表位检索到≥10个部分匹配,而ChatGPT共检索到3个,Gemini则一个也没有。 结论:我们构建了EpitopeMiner,一个用于可持续文献和数据库整理的计算框架。在一个9名患者的数据集中,EpitopeMiner以人工分析无法企及的规模和速度检索到经实验和临床验证的表位证据。EpitopeMiner以带引用的响应优于通用LLM,在基准测试中实现100%的证据覆盖率,减少幻觉并提高可靠性。
查看英文原文 English abstract
Background: Personalized cancer vaccines hold great promise by eliciting tumor-specific immune responses [1-3]. A key challenge is identifying the right targets - immunogenic protein sequences, or epitopes, presented on tumor cells. While computational pipelines can predict epitope candidates from tumor sequencing, experimental validation is costly and slow. Leveraging literature and database knowledge could bridge this gap by enabling evidence-driven selection of high-confidence targets, but is constrained by fragmented information across journals and immunology databases [4-5]. We introduce EpitopeMiner, which integrates sequence-based candidate screening with evidence-driven knowledge retrieval for epitope prioritization. Methods: A total of 25,966 tumor-specific epitopes were predicted from whole-genome sequencing of tumor-PBMC pairs from nine patients (including lung, sarcoma, NKTL, DLBCL) using a standard workflow: HLA typing (OptiType), variant calling (Strelka with wANNOVAR), MHC binding prediction (NetMHCpan) and RNA-supported protein-altering filtering. EpitopeMiner combines OpenAI's Large Language Model (LLM) with an in-house Retrieval Augmented Generation (RAG) database comprising (a) 78,461 full-text research articles from PMC, PLOS One, and Europe PMC and, (b) ~2.6 million unique epitopes from IEDB, dbPepNeo, SystemMHC, TANTIGEN, and caAtlas database. EpitopeMiner includes: (i) a screening module that processes an epitope list, detecting exact or ≥ 7 amino acid partial matches from the in-house database, and (ii) a reporting module that analyses each top-ranked hits, defined by highest sequence similarity and evidence density, to generate an LLM response covering 28 immunology keywords with citations. Results: Among the 25,966 epitopes predicted from the nine patients, EpitopeMiner found 7 exact matches, and 17.6% had ≥ 6 partial matches; mean processing time per epitope was 0.97 seconds. In benchmarking with 3 lung cancer driver-gene epitopes (KITDFGRAK, ITDFGRAKL, TDFGRAKLL), EpitopeMiner outperformed ChatGPT and Gemini, returning the highest amount of relevant immunological information - summarized as (total responses, % with evidence) - (18, 100%), (8, 100%), and (2, 100%) respectively, compared to ChatGPT's (10, 70%), (4, 50%), (6, 33%) and Gemini's (7, 0%), (1, 0%), (1, 0%). In addition, EpitopeMiner retrieved ≥10 partial matches for each epitope, whereas ChatGPT retrieved total of 3 and Gemini none. Conclusion: We built EpitopeMiner, a computational framework for sustainable literature and database curation. In a 9-patient dataset, EpitopeMiner retrieved experimentally and clinically validated epitope evidence at a scale and speed infeasible with manual analysis. EpitopeMiner outperformed general-purpose LLMs with cited responses, achieving 100% evidence coverage on benchmarks, reducing hallucinations and improving reliability.
利益披露 Disclosure
A. Chadha, None.. I. Cheong, None.. M. Zhang, None.. W. Tan, None.. W. Tang, None.. J. Lim, None.. S. Wilson, None.. C. Ong, None.. B. Lee, None.. C. Lim, None.. O. Rotzschke, None.. M. Lau, None.

← 返回 AACR 2026 检索