PO.PR02.02 · 预防研究

使用早期可用的Illumina protein prep 6K检测和机器学习进行癌症早期检测

Early cancer detection using early-access Illumina protein prep 6K assay and machine learning

海报缩略图:使用早期可用的Illumina protein prep 6K检测和机器学习进行癌症早期检测
编号 7614 展板 1 时间 4/22 09:00–12:00 区域 Section 36 主讲 Mete Mulazimoglu, PhD
分会场 Cancer and Cancer Related Alterations, Detection Approaches, and Molecular Characterization
查看 PDF 下载 PDF 🔒 查看 / 下载完整 PDF 需登录并开通下载套餐 · 查看套餐 / 开通 AACR 官方页面

作者与单位 Authors & Affiliations

KAMEL LAHOUEL, Mete Mulazimoglu, Kameron Bates, Candice Wike, Kunjur Manasa Upadhyaya, Victoria Zismann, Kamawela Leka, Payton Smith, Gracyn Benck, Kianna Martos Rupp, Chaney Jambor, Matteo Munini, Sophie Pénisson, Stephanie Pond, Jeffrey Trent, Patrick Pirrotte, Cristian Tomasetti

TGen (The Translational Genomics Research Institute), Phoenix, AZ

摘要 Abstract

中文摘要
背景: 癌症早期检测研究揭示循环蛋白质是强有力且信息丰富的生物标志物。我们进一步旨在开发一个对独立处理数据集间批次效应稳健的机器学习框架,并确定一个能够在最早阶段检测癌症的蛋白质子集。在本研究中,我们试图评估早期可用的Illumina protein prep蛋白质组学检测,不仅评估其技术重现性,还评估其产生与癌症早期检测相关的生物学信息性蛋白质特征的能力。 实验流程: 分析了来自六块板的蛋白质组谱,涵盖约217份正常和206份癌症血浆样本。前四块板包含从东欧来源采集的样本,用于模型训练和特征选择,其余两块板包含在美国从不同血统背景人群采集的样本,作为外部测试集。这两块板包括来自膀胱癌、乳腺癌、胃癌和肺癌的样本,代表多样化的生物学和技术条件。 采用最小冗余-最大相关性(MRMR)方法进行特征选择,该方法通过最大化与癌症状态的互信息同时最小化冗余来对蛋白质进行排序。前200个蛋白质用于训练带有径向基核的支持向量机(SVM)分类器。 结果: 该模型在两块独立测试板上实现了0.89的平均AUC,展示了强大的跨批次和跨血统重现性,并证实可从Illumina平台中提取信息性、可推广的蛋白质特征。在95%特异性下,分类器实现了65%的总体敏感性(95% CI:51-77%),在II期癌症中表现尤为出色,达82%(95% CI:52-95%),凸显了其在早期检测中的潜在效用。各癌症类型的表现一致,肺癌中观察到最高的敏感性。重要的是,六折留一板交叉验证在99%特异性下产生了82%的平均敏感性,表明整合多样化数据源可能会增强模型的可推广性。 结论: 应用于大规模蛋白质组数据的机器学习框架确定了一个紧凑且具有生物学意义的、能够进行癌症早期检测的蛋白质子集。结果凸显了Illumina Protein Prep 6K检测的稳健性,以及开发适用于人群规模癌症筛查的批次不敏感蛋白质分类器的可行性。
查看英文原文 English abstract
Background: Studies in cancer early detection have revealed circulating proteins to be powerful and informative biomarkers. We further aimed to develop a machine-learning framework robust to batch effects across independently processed datasets and to identify a subset of proteins capable of detecting cancer at its earliest stages. In this study, we sought to evaluate the early-access Illumina protein prep proteomic assay not only for its technical reproducibility but also for its ability to yield biologically informative protein signatures relevant to cancer early detection. Experimental Procedures: Proteomic profiles from six plates were analyzed, encompassing approximately 217 normal and 206 cancer plasma samples. The first four plates contained samples collected from Eastern European sources and were used for model training and feature selection, while the remaining two plates contained samples collected in the United States from populations with diverse ancestry backgrounds and served as an external test set. These two plates included samples from bladder, breast, gastric, and lung cancers, representing diverse biological and technical conditions. Feature selection was performed using the Minimum Redundancy-Maximum Relevance (MRMR) method, which ranks proteins by maximizing mutual information with cancer status while minimizing redundancy. The top 200 proteins were used to train a Support Vector Machine (SVM) classifier with a radial-basis kernel. Results: The model achieved a mean AUC of 0.89 across the two independent test plates, demonstrating strong cross-batch and cross-ancestry reproducibility and confirming that informative, generalizable protein features can be extracted from the Illumina platform. At 95% specificity, the classifier achieved an overall sensitivity of 65% (95% CI: 51-77%), with particularly strong performance in Stage II cancers at 82% (95% CI: 52-95%), underscoring its potential utility for early detection. Performance was consistent across cancer types, with highest sensitivities observed in lung. Importantly, a six-fold leave-one-plate out cross-validation yielded an average sensitivity of 82% at 99% specificity, demonstrating that integrating diverse data sources will likely strengthen model generalizability. Conclusions: A machine-learning framework applied to large-scale proteomic data identifies a compact and biologically meaningful subset of proteins capable of early cancer detection. The results highlight the robustness of the Illumina Protein Prep, 6K assay, and the feasibility of developing batch-insensitive protein classifiers for population-scale cancer screening.
利益披露 Disclosure
K. Lahouel, None.. M. Mulazimoglu, None.. K. Bates, None.. C. Wike, None.. K. Upadhyaya, None.. V. Zismann, None.. K. Leka, None.. P. Smith, None.. G. Benck, None.. K. Martos Rupp, None.. C. Jambor, None.. M. Munini, None.. S. Pénisson, None.. S. Pond, None.. J. Trent, None.. P. Pirrotte, None.. C. Tomasetti, None.

← 返回 AACR 2026 检索