PO.PS01.09 · 人群科学
利用机器学习对电子健康记录进行拉丁裔出生地细分:对结直肠癌筛查差异的洞见
Disaggregating Latino nativity using machine learning on electronic health records: Insights for colorectal cancer screening disparities
该海报暂无可下载的资料
AACR 官方页面
作者与单位 Authors & Affiliations
摘要 Abstract
中文摘要
背景:结直肠癌(CRC)预防方面的进展并不均衡,研究显示拉丁裔患者的CRC筛查率较低且诊断分期较晚。拉丁裔亚群在癌症相关危险因素方面存在差异,同时在社会人口学特征、迁移史、保险覆盖及可能影响癌症预防的医疗保健可及性模式方面也存在广泛差异。然而,大规模数据集很少包含研究拉丁裔亚群特异性癌症风险与结局差异所需的细粒度数据。本研究评估了一种旨在推断出生地和出生国、推进癌症预防研究以更好地评估拉丁裔健康公平性的机器学习方法。
方法:我们使用了来自28个州1,876家社区健康中心接受诊疗的1,500,191名拉丁裔患者的全面电子健康记录数据,连同按普查区级别地理编码的邻里构成数据和基于姓氏的数据,来开发拉丁裔亚群的机器学习模型。训练并测试了多种监督学习算法以预测出生地和出生国。使用受试者工作特征曲线下面积(AUC)评估模型的预测性能。作为一个案例,我们使用模型预测的拉丁裔亚群和出生地概率,按外国出生状态(已知和预测的)评估CRC筛查差异。
结果:在该网络的1,500,191名拉丁裔中,仅有173,278人(11.6%)由拉丁裔患者自我报告了出生国,凸显了使用现有EHR数据研究拉丁裔异质性的挑战。出生地预测模型在所有组别中均表现出优异的判别预测性能(美国出生 vs. 外国出生:AUC=0.90;墨西哥裔 vs. 非墨西哥裔:AUC=0.87;危地马拉裔 vs. 非危地马拉裔:AUC=0.84;古巴裔 vs. 非古巴裔:AUC=0.84)。在我们的案例中,使用拉丁裔患者已知的外国出生状态,我们观察到美国出生的拉丁裔接受CRC筛查的几率低于外国出生的拉丁裔(OR=0.55,95% CI=0.50—0.617)。我们观察到CRC筛查比值比的已知估计值与模型预测估计值之间具有高度一致性。
结论:包括针对拉丁裔在内的全国性数据细分呼吁面临诸多挑战。我们开发并验证了新颖的预测模型,用于推断拉丁裔的出生地和出生国,以应用于基于人群的癌症差异研究。这些方法为在未收集拉丁裔出生地的数据中评估癌症差异提供了机会。
查看英文原文 English abstract
Background: Advancements in colorectal cancer (CRC) prevention have not been equitable with studies showing reduced CRC screening rates and later-stage diagnoses among Latino patients. Latino subgroups vary in their cancer-related risk factors but also differ widely in their sociodemographic characteristics, migration histories, insurance coverage, and health care access patterns that may impact cancer prevention. However, large-scale datasets seldom contain granular data needed to study Latino subgroup-specific differences in cancer risk and outcomes. This study evaluated a machine learning approach designed to infer nativity and country of birth and advance cancer prevention research to better evaluate health equity among Latinos.
Methods: We used comprehensive electronic health record data from 1,500,191 Latino patients receiving care at 1,876 community health centers across 28 states, along with geocoded census-tract-level neighborhood composition data, and surname-based data to develop machine learning models of Latino subgroups. Multiple supervised learning algorithms were trained and tested to predict nativity and country of birth. Model predictive performance was evaluated using area under the receiver operating curve (AUC). As a case example, we used model-predicted probabilities of Latino subgroups and nativity to evaluate CRC screening disparities by foreign-born status, both known and predicted.
Results: Among 1,500,191 Latinos in the network, country of birth was self-reported by Latino patients for only 173,278 (11.6%), underscoring the challenges of using existing EHR data for studying Latino heterogeneity. Prediction models for nativity showed excellent discriminatory prediction performance across all groups (US-born vs. foreign-born: AUC=0.90; Mexican vs. non-Mexican: AUC=0.87; Guatemalan vs. non-Guatemalan: AUC=0.84; Cuban vs. non-Cuban: AUC=0.84). In our case example, using known foreign-born status of Latino patients, we observed that US-born Latinos had lower odds of CRC screening compared to foreign-born Latinos (OR=0.55, 95% CI=0.50-0.617). We observed high concordance between known and model-predicted estimates of CRC screening odds ratios.
Conclusion: National calls for data disaggregation, including among Latinos, have numerous challenges. We developed and validated novel prediction models to infer Latino nativity and country of birth for use in population-based cancer disparities research. These methods present an opportunity to evaluate cancer disparities in data where Latino nativity is not collected.
利益披露 Disclosure
M. Marino, None..
J. Hwang, None..
J. A. Lucas, None.