PO.PS01.09 · 人群科学
结合电子健康记录、环境和基因组学数据用于南加州肺癌风险预测
Combining electronic health records, environmental, and genomics data for lung cancer risk prediction in Southern California
该海报暂无可下载的资料
AACR 官方页面
作者与单位 Authors & Affiliations
摘要 Abstract
中文摘要
背景:肺癌风险反映了相互交织的临床、环境和基因组因素,但这些数据类型很少在个体层面加以整合。我们在南加州组建了一个基于EHR的队列,以构建新发肺癌诊断的预测模型,并刻画驱动风险的临床、基因组和社区特征。
方法:我们从加州大学圣地亚哥分校健康系统的电子健康记录中构建了一个7,151名成年人的回顾性队列(50.4%为女性,65%≥65岁),并将约40项临床特征与来自CalEnviroScreen的普查区层面的环境和社会经济指标,以及ALK和EGFR的基因组突变状态相关联。我们采用分层五折交叉验证,比较了14种分类器(逻辑回归、随机森林、XGBoost和11种PyTorch神经网络)以预测肺癌诊断。表现最佳的模型的超参数采用Optuna中的贝叶斯搜索进行优化。模型表现采用AUROC、准确度、精确度、召回率和F1进行总结,特征重要性采用Shapley(SHAP)值评估。
结果:优化后的XGBoost取得了最佳的交叉验证区分度(AUROC 0.879),准确度为0.802,精确度为0.744,F1为0.694,优于线性和深度学习基线。按SHAP排名靠前的特征包括吸烟强度、心血管代谢共病和年龄,而社区失业率、农药负荷和臭氧水平贡献了额外但较小的预测信号,进一步强化了社区劣势和污染对该区域队列中肺癌易感性的贡献。在完成基因组分析的患者中,ALK阳性病例的诊断年龄显著小于ALK野生型病例(平均56岁对71岁,p=0.016),凸显了生物学上不同的疾病进程。
结论:一个整合EHR、环境和基因组数据的梯度提升模型能够在一个多样化的区域队列中有意义地对个体肺癌风险进行分层,并突出脆弱性的临床和社区驱动因素。这些发现支持利用常规收集的健康和环境数据来指导有针对性的肺癌筛查和预防工作,并促进未来关于外部验证、时变暴露以及跨种族和社会经济群体的明确公平性约束的研究。
查看英文原文 English abstract
Background: Lung cancer risk reflects intersecting clinical, environmental, and genomic factors, yet these data types are rarely integrated at the individual level. We assembled an EHR-based cohort in Southern California to build a predictive model of incident lung cancer diagnosis and to characterize the clinical, genomic, and neighborhood features that drive risk.
Methods: We constructed a retrospective cohort of 7,151 adults from UC San Diego Health electronic health records (50.4% female, 65% ≥ 65 years) and linked approximately 40 clinical features to census-tract-level environmental and socioeconomic indicators from CalEnviroScreen, as well as genomic mutation status for ALK and EGFR. We compared 14 classifiers (logistic regression, random forest, XGBoost, and 11 PyTorch neural networks) using stratified five-fold cross-validation to predict lung cancer diagnosis. Hyperparameters for top performing models were optimized using Bayesian search in Optuna. Model performance was summarized using AUROC, accuracy, precision, recall, and F1, and feature importance was assessed using Shapley (SHAP) values.
Results: Optimized XGBoost achieved the best cross-validated discrimination (AUROC 0.879), with accuracy 0.802, precision 0.744, and F1 0.694, outperforming linear and deep-learning baselines. Top-ranked features by SHAP included smoking intensity, cardiometabolic comorbidity, and age, with neighborhood unemployment, pesticide burden, and ozone levels contributing additional, though smaller, predictive signal, reinforcing the contribution of neighborhood disadvantage and pollution to lung cancer vulnerability in this regional cohort. Among genomically profiled patients, ALK-positive cases were diagnosed at a significantly younger age than ALK-wildtype cases (mean 56 vs 71 years, p=0.016), underscoring biologically distinct disease courses.
Conclusion: An integrated gradient-boosting model leveraging EHR, environmental, and genomic data can meaningfully stratify individual lung cancer risk in a diverse regional cohort and elevate both clinical and neighborhood drivers of vulnerability. These findings support the use of routinely collected health and environmental data to guide targeted lung cancer screening and prevention efforts and motivate future work on external validation, time-varying exposures, and explicit fairness constraints across racial and socioeconomic groups.
利益披露 Disclosure
G. Pucillo, None..
S. Naqvi, None..
A. Jue, None..
C. Law, None..
S. Patel, None..
U. Z. George, None.