PO.BCS02.06 · 生物信息与计算
基于机器学习、利用生活方式和临床风险因素对波多黎各女性乳腺癌进行分类
Machine learning based classification of breast cancer using lifestyle and clinical risk factors in Puerto Rican women
作者与单位 Authors & Affiliations
摘要 Abstract
中文摘要
在全球范围内,乳腺癌发病率上升超过20%,死亡率上升14%,使其成为重大的公共卫生挑战。有效的预防和筛查策略需要改进的早期风险评估工具,尤其是针对研究不足的人群。目前的风险模型主要基于欧洲血统人群进行训练,导致在混合血统人群中表现欠佳且不均衡。我们的目标是比较当前机器学习(ML)方法在波多黎各人群中对乳腺癌状态进行分类的表现。我们使用了来自一个包含1393名女性的波多黎各队列的结构化数据集,涵盖生活方式、人口统计学、生育以及已确立的流行病学风险因素变量。数据经过系统性的数据清洗、去除结局泄露特征、分类标签规范化、缺失值插补以及非数值变量编码进行预处理。最终数据集采用分层训练-测试策略进行划分,以保持病例-对照平衡,随后使用多种监督算法进行模型开发。训练并比较了Logistic回归、决策树、随机森林、梯度提升、XGBoost、LightGBM和CatBoost模型,采用分层10折交叉验证和标准化性能指标。经随机搜索进行超参数调优后,CatBoost分类器在留出测试集上取得了最佳性能,ROC-AUC为0.8009,PR-AUC为0.8285,F1分数为0.7239,平衡准确率为0.7342,优于传统模型和其他提升算法。使用SHapley Additive exPlanations(SHAP)和基于置换的特征重要性评估模型可解释性,二者一致地识别出有影响力的预测因子,包括绝经状态、吸烟史、多基因风险评分、年龄以及患乳腺癌的姐妹数量等风险因素。这些发现证实,梯度提升方法,尤其是CatBoost,能够有效捕捉多维健康数据中的非线性相互作用,同时保持可解释性。总体而言,这项工作证明了整合机器学习、利用非遗传和遗传风险因素预测波多黎各人群乳腺癌风险的可行性和实用性。
查看英文原文 English abstract
Globally, breast cancer incidence has increased by more than 20 percent and mortality has increased by 14 percent making it a significant public health challenge. Effective prevention and screening strategies require improved early risk assessment tools, especially for understudied populations. Current risk models are predominantly trained on European descent populations, resulting in suboptimal and uneven performance in admixed populations. Our goal was to compare how current machine learning (ML) methods perform to classify breast cancer status in a Puerto Rican population. We used a structured dataset of lifestyle, demographic, reproductive, and established epidemiological risk-factor variables from a Puerto Rican cohort of 1393 woman. Data was preprocessed using systematic data cleaning, removal of outcome-leaking features, normalization of categorical labels, imputation of missing values, and encoding of non-numeric variables. The final dataset was split using a stratified train-test strategy to preserve case-control balance, followed by model development using multiple supervised algorithms. Logistic Regression, Decision Tree, Random Forest, Gradient Boosting, XGBoost, LightGBM, and CatBoost models were trained and compared using stratified 10-fold cross-validation and standardized performance metrics. After hyperparameter tuning with randomized search, the CatBoost classifier achieved the strongest performance on the held-out test set, with a ROC-AUC of 0.8009, PR-AUC of 0.8285, F1-score of 0.7239, and balanced accuracy of 0.7342, outperforming both traditional models and other boosting algorithms. Model interpretability was assessed using SHapley Additive exPlanations (SHAP) and permutation-based feature importance, which consistently identified influential predictors including menopausal status, smoking history, polygenic risk score, age, and number of sisters with breast cancer, among other risk factors. These findings confirm that gradient-boosting approaches, particularly CatBoost, effectively capture nonlinear interactions in multidimensional health data while maintaining interpretability. Overall, this work demonstrates the feasibility and utility of integrating machine learning to predict breast cancer risk using non-genetic and genetic risk factors in a Puerto Rican population.
利益披露 Disclosure
J. E. Martínez-Jiménez, None..
S. V. Pérez-Mártir, None..
D. de León-Vázquez, None..
N. A. Arroyo, None..
J. Dutil, None.