PO.BCS02.01 · 生物信息与计算
人类不能仅靠人工智能(AI)而生存
Humans cannot live by artificial intelligence (AI) alone
作者与单位 Authors & Affiliations
摘要 Abstract
中文摘要
引言:差异表达(DE)分析是组学评价的基石。它用于识别癌症、治疗响应和药物诱导不良事件的生物标志物。DE方法使用AI/ML(机器学习)。如果多种DE方法能够识别出相同的生物标志物,这将强有力地支持将该生物标志物用作可供湿实验室验证研究的稳健候选物。目前尚不清楚顶级DE方法是否识别出相同的生物标志物(即共享的)。因此,我们评估了4种DE方法识别患有和不患有免疫检查点抑制剂诱导的甲状腺功能减退(TEAE ThyDis)的癌症患者中共享的血清自身抗体的能力。
方法:纳入接受durvalumab、ipilimumab、pembrolizumab、nivolumab或联合治疗、患有TEAE ThyDis(N=18)或无TEAE(N=15)的乳腺癌(N=8)或黑色素瘤(N=25)患者。在R中使用了4种DE方法(limma、DESeq2、edgeR、randomForest)。对每位患者评估了15,500种经和未经ComBat批次校正的治疗前自身抗体。
结果:在乳腺癌患者中,limma、DESeq2、edgeR和randomForest分别识别出201、109、158和472个生物标志物(表1)。然而,4种DE方法之间最多仅有53个生物标志物是共享的。使用limma或randomForest进行ComBat批次校正后分别识别出125和484个生物标志物,4种方法之间最多有114个共享生物标志物。在黑色素瘤患者中,limma、DESeq2、edgeR和randomForest分别识别出198、244、568和1042个生物标志物,最多有183个共享(表1)。使用limma或randomForest进行ComBat批次校正后分别识别出196和1088个生物标志物,最多有257个共享。没有任何生物标志物在所有方法中都是共享的。
结论:我们的数据表明,顶级AI/ML DE分析方法识别出不同的生物标志物。作为一个领域,现在是重新评估和革新这些工具、以及创建新工具以确保稳健、可重现的生物标志物识别的时候了。
查看英文原文 English abstract
INTRODUCTION: Differential expression (DE) analysis is the cornerstone of omics evaluation. It is used to identify biomarkers for cancer, therapeutic response, and drug-induced adverse events. DE methods use AI/ML (machine learning). If multiple DE methods could identify the same biomarkers, this would strongly support the biomarker's use as a robust candidate(s) for wet lab validation studies. It is unclear if the top DE methods identify the same biomarkers (i.e., shared). Therefore, we evaluated 4 DE methods for their ability to identify shared serum autoantibodies in cancer patients with and without immune-checkpoint inhibitor induced hypothyroidism (TEAE ThyDis).
METHODS: Patients with breast cancer (N=8) or melanoma (N=25) who were treated with durvalumab, ipilimumab, pembrolizumab, nivolumab, or combination who had TEAE ThyDis (N = 18) or No TEAE (N = 15) were included. Four DE methods (limma, DESeq2, edgeR, randomForest) were used in R. 15,500 pre-treatment autoantibodies with and without ComBat batch correction were evaluated for each patient.
RESULTS: In patients with breast cancer, limma, DESeq2, edgeR, and randomForest identified 201, 109, 158, and 472 biomarkers, respectively (Table 1). However, only up to 53 biomarkers were shared between the 4 DE methods. ComBat batch correction with limma or randomForest led to identification of 125 and 484 biomarkers respectively and up to 114 shared biomarkers between the 4 methods. In patients with melanoma, limma, DESeq2, edgeR, and randomForest identified 198, 244, 568, and 1042 biomarkers respectively with up to 183 biomarkers shared (Table 1). ComBat batch correction with limma or randomForest led to identification of 196 and 1088 biomarkers respectively and up to 257 shared. There was no biomarker that was shared in all methods.
CONCLUSIONS: Our data suggests that top AI/ML DE analysis methods identify different biomarkers. As a field, it is time to re-evaluate and re-vamp these tools as well as create new tools to ensure robust reproducible biomarker identifications.
利益披露 Disclosure
K. Blenman,
CareVive ).
O. Blaha, None..
S. Qiu, None..
K. Chen, None..
K. Spall, None..
Y. Liu, None..
M. Williams, None..
M. Villa, None..
V. Vankov, None..
K. Oteng Agyapong, None..
D. Li, None.