PO.BCS02.06 · 生物信息与计算
预测西班牙裔男性前列腺癌风险的机器学习模型
Machine learning models for predicting prostate cancer risk in Hispanic men
作者与单位 Authors & Affiliations
摘要 Abstract
中文摘要
引言:美国西班牙裔男性面临显著的前列腺癌(PCa)负担,有研究报告其在诊断时确诊为晚期疾病的比例高于非西班牙裔男性。然而,在当前的风险预测方法中,其遗传学和社会人口学特征仍未得到针对性关注。这些观察结果凸显了当前风险评估策略中的空白,以及对改进的、与人群相关的预测工具的需求。机器学习(ML)在生物医学研究中的应用已在癌症风险的早期检测和预测方面取得了宝贵的进展。本研究旨在利用来自 All of Us Research Program 的数据,开发一种用于预测西班牙裔男性 PCa 风险的 ML 模型。
方法:研究队列包括 20,172 名西班牙裔男性,其中 PCa 病例 220 例(1.1%),对照 19,952 例。预测变量包括年龄组(≤40、41-59、≥60 岁)、PCa 家族史、吸烟状况、饮酒频率以及与 PCa 相关的 rs10090154 基因型。数据按 80/20 拆分为训练集(n = 16,138)和测试集(n = 4,034)。为处理类别不平衡,应用了类别权重。训练并评估了 logistic 回归、随机森林和 XGBoost 模型。模型经过五折交叉验证并进行 sigmoid 校准,性能采用 ROC-AUC、PR-AUC、Brier 评分、精确率-召回率和 F1 评分进行评估。
结果:在三种模型中,logistic 回归模型在区分度和校准度之间取得了最佳平衡,平均 ROC-AUC 为 0.872(95% CI:0.856-0.885),PR-AUC 为 0.059(95% CI:0.045-0.074),Brier 评分为 0.0105,F1 评分为 0.066。该模型指标为:准确率 = 0.704,召回率 = 0.955,精确率 = 0.034。相应地,假阴性率(FNR)为 4.5%,表明 44 例真实 PCa 病例中仅有 2 例被错误分类为对照;而假阳性率(FPR)为 30.6%,意味着约三分之一的非 PCa 病例被错误地标记为阳性。特征重要性和比值比分析确定,年龄 ≥60 和 PCa 家族史是该模型最具影响力的预测变量,其次是 rs10090154 基因型和既往吸烟状况。相比之下,随机森林(ROC-AUC = 0.851)和 XGBoost(ROC-AUC = 0.858)模型表现出相似的区分度,但与 logistic 回归相比,其校准度和可解释性略低。
结论:一种经类别加权、sigmoid 校准的 logistic 回归模型实现了准确、可解释且校准良好的 PCa 风险预测。尽管 All of Us 队列中观察到的病例患病率较低,该模型仍表现出强大的敏感性和区分度,支持其在基于人群的精准健康研究中用于生物医学和临床风险分层的实用性。这些发现凸显了简单且可解释的 ML 模型在改善西班牙裔男性早期 PCa 风险评估方面的潜力。
查看英文原文 English abstract
Introduction: Hispanic men in the United States experience a notable prostate cancer (PCa) burden, with studies reporting higher rates of advanced-stage disease at diagnosis compared with non-Hispanic men. However, their genetic and sociodemographic profiles remain untargeted in current risk prediction approaches. These observations highlight gaps in current risk assessment strategies and the need for improved, population-relevant prediction tools. Machine learning (ML) applications in biomedical research have yielded valuable advances in the early detection and prediction of cancer risk. The aim of this study was to develop a ML model to predict PCa risk in Hispanic men leveraging data from the All of Us Research Program.
Methods: The study cohort included 20,172 Hispanic males, 220 PCa cases (1.1%) and 19,952 controls. Predictors included age group (≤40, 41-59, ≥60 years), family history of PCa, smoking status, drinking frequency and the rs10090154 genotype associated with PCa. Data were split 80/20 into training (n= 16, 138) and testing (n= 4,034) sets. To address class imbalance, class weights were applied. Logistic regression, random forest and XGBoost models were trained and evaluated. Models underwent a five-fold cross-validation with sigmoid calibration, and performance was assessed using ROC-AUC, PR-AUC, Brier score, precision recall and F1 score.
Results: Among the three models, the logistic regression model achieved the best balance of discrimination and calibration with a mean ROC-AUC of 0.872 (95% CI: 0.856-0.885), PR-AUC of 0.059 (95% CI: 0.045-0.074), Brier score of 0.0105 and an F1 score of 0.066. The model metrics were accuracy = 0.704, recall = 0.955 and precision = 0.034. Correspondingly, the false negative rate (FNR) was 4.5%, indicating that only 2 out of 44 true PCa cases were misclassified as controls, while the false positive rate (FPR) was 30.6%, meaning roughly one in three non PCa cases were incorrectly flagged as positive. Feature importance and odds ratio analyses identified Age ≥60 and family history of PCa as the most influential predictors for the model, followed by the rs10090154 genotype and former smoking status. Comparatively, the random forest (ROC-AUC= 0.851) and XGBoost (ROC- AUC= 0.858) models demonstrated similar discrimination but exhibited slightly lower calibration and interpretability compared with logistic regression.
Conclusions: A class-weighted, sigmoid-calibrated logistic regression model achieved accurate, interpretable and well calibrated PCa risk predictions. Despite the low observed case prevalence in the All of Us cohort, the model showed strong sensitivity and discrimination, supporting its utility for biomedical and clinical risk stratification in population-based precision health research. These findings highlight the potential of simple and interpretable ML models to improve early PCa risk assessment in Hispanic men.
利益披露 Disclosure
R. J. Rodríguez-Colón, None..
A. V. Varela-Parrilla, None..
Z. C. Medina-Nieves, None..
M. M. Sánchez-Vázquez, None..
M. Martínez-Ferrer, None.