PO.PS01.09 · 人群科学
利用高通量蛋白质组学和机器学习识别预测5年胃癌风险的多蛋白特征
Identification of a multi-protein signature to predict 5-year stomach cancer risk using high-throughput proteomics and machine learning
该海报暂无可下载的资料
AACR 官方页面
作者与单位 Authors & Affiliations
摘要 Abstract
中文摘要
引言:胃癌是全球第五大常见癌症,也是癌症死亡的第四大原因,5年相对生存率平均仅为20%-30%。这一低生存率在很大程度上归因于晚期诊断,其往往发生在癌症已经扩散、更难治疗之后。美国预防服务工作组目前尚无关于胃癌筛查或幽门螺杆菌(H pylori)感染(胃癌的主要风险因素)筛查的指导意见。本研究探讨了是否可利用高通量蛋白质组学开发一种胃癌易感性模型,以帮助对个体风险进行分层并指导筛查程序。
方法:来自欧洲癌症与营养前瞻性研究的样本采用改良适配体蛋白质组学技术(SomaScan™ 7K检测)进行分析。经过数据质控后,14,787份柠檬酸盐血浆样本具有临床和蛋白质组学数据(总计约1.03亿次蛋白质测量),其中包括在5年内和20年以上随访期内分别诊断出胃癌的n=56例和n=219例。采用机器学习技术,使用80%的训练数据集,识别一个可预测采血后5年内诊断胃癌风险的模型。区分度通过C指数和5年AUC评估。在具有H pylori感染状态的个体(n=286)中进行了事后子集分析。
结果:检测到大量蛋白质组学信号,有922种蛋白质在FDR <0.1时与新发胃癌显著相关。识别出一个8蛋白加速失效时间模型,在留出验证数据集中准确预测胃癌风险,5年AUC为0.742,C指数为0.697。该蛋白质模型的表现高于用于预测终生胃癌风险的基于多基因风险评分模型的平均表现。该模型包含已知在胃癌中上调或下调的蛋白质(75%),尽管其中仅一部分蛋白质(37.5%)可预测胃癌风险,以及新型蛋白质(25%)。在H pylori感染阳性个体子集(n=214)中,胃癌风险模型的5年AUC = 0.698,C指数为0.599,提示该模型在H pylori感染个体中仍能准确预测风险。
结论:我们成功开发了一个多蛋白模型,可预测有和无H pylori感染个体的5年胃癌风险,凸显了高通量蛋白质组学作为评估癌症风险的新型筛查工具的潜在效用。
查看英文原文 English abstract
Introduction: Stomach cancer is the fifth most common cancer worldwide, and the fourth leading cause of cancer death, with an average 5-year relative survival rate of only 20%-30%. This low survival rate is largely due to late diagnosis, which often occurs after the cancer has spread and is more difficult to treat. There is currently no guidance from the US Preventive Services Task Force around screening for stomach cancer or Helicobacter pylori ( H pylori ) infection, a primary risk factor for stomach cancer. This study investigated whether high-throughput proteomics could be used to develop a stomach cancer susceptibility model that may help stratify individual risk and direct screening procedures.
Methods: Samples from the European Prospective Investigation into Cancer and Nutrition study were assayed using modified-aptamer proteomics technology (SomaScan™ 7K assay). Following data QC, 14,787 citrate plasma samples had clinical and proteomic data (totaling ~103 million protein measurements), which included n=56 and n=219 stomach cancer diagnosis within 5 years and over 20-year follow-up period, respectively. Machine learning techniques were used to identify a model that predicts risk of stomach cancer diagnosis within 5 years of blood draw using an 80% training split of the data. Discrimination was assessed via C-index and 5-year AUC. A post-hoc subset analysis was performed in individuals with H pylori infection status (n=286).
Results: An abundance of proteomic signal was detected with 922 proteins significantly associated with incident stomach cancer at FDR <0.1. An 8-protein accelerated failure time model was identified that accurately predicted stomach cancer risk with a 5-year AUC of 0.742 and C-Index of 0.697 in the hold-out validation dataset. Performance of the protein model was higher than the average performance of the Polygenic Risk Score-based models used to predict lifetime stomach cancer risk. The model included proteins known to be up or downregulated in stomach cancer (75%), although only a subset of these proteins (37.5%) are predictive of stomach cancer risk, and novel proteins (25%). In the subset of individuals positive for H pylori infection (n=214), the stomach cancer risk model had a 5-year AUC = 0.698 and C-Index of 0.599 suggesting the model can still accurately predict risk in individuals with H pylori infection.
Conclusions: We successfully developed a multi-protein model to predict 5-year risk of stomach cancer in individuals with and without H pylori infection, highlighting the potential utility of high-throughput proteomics as a novel screening tool for assessing cancer risk.
利益披露 Disclosure
J. Chadwick,
Illumina Employment.
C. Paterson,
Illumina Employment.
S. Shrestha,
Illumina Employment.
H. Biegel,
Illumina Employment.
J. Kuzma,
SomaLogic Employment.
S. Williams,
SomaLogic Employment.