PO.PS01.09 · 人群科学
利用行政健康数据进行乳腺癌的个性化风险评估
Personalized risk assessment of breast cancer using administrative health data
作者与单位 Authors & Affiliations
摘要 Abstract
中文摘要
乳腺癌筛查项目采用“一刀切适配大多数”的方法,参与率欠佳。人群层面的行政健康数据库为构建可扩展、数据驱动的风险评估工具提供了独特机会,能够识别可能从更个性化筛查策略中获益的女性。我们汇集了近二十年的纵向健康数据,包括乳腺X线摄影筛查史、用药情况、就诊记录和出院摘要,涵盖不列颠哥伦比亚省的174万名女性,其中诊断出39,211例新发乳腺癌。我们团队正在开发新的乳腺癌风险评估模型,利用加拿大公共资助医疗系统的行政健康数据预测每位女性直至乳腺癌发病(BCo)的个体化时间。我们正在应用机器学习个体生存分布(ISD)模型,该模型为每个受试者x确定一个分布S(t | x),表示x直至BCo的时间至少还有t年的概率,适用于所有t > 0。随后,我们可利用这些模型估计每位女性直至BCo的预期时间以及她的风险评分。在使用25个与乳腺癌具有已知/疑似关联特征的初步模型中,随机生存森林(RSF)取得了最高的一致性指数(CI = 58.9%),而多任务逻辑回归(MTLR)取得了具有竞争力的5年Brier评分(BS = 0.0068)和较低的平均绝对误差(MAE = 30.4个月)。这些早期结果证明了利用行政健康数据进行个性化乳腺癌风险预测的可行性。后续工作将大幅扩展特征集,以改善模型区分度。
查看英文原文 English abstract
Breast cancer screening programs are one-size-fits-most approaches with suboptimal participation rates. Population-level administrative health databases provide a unique opportunity to build scalable, data-driven risk assessment tools capable of identifying women who may benefit from more personalized screening strategies. We assembled nearly two decades of longitudinal health data, including mammographic screening history, medication use, physician visits, and hospital discharge abstracts, for 1.74 million women in British Columbia, among whom 39,211 incident breast cancers were diagnosed. Our team is developing new breast cancer risk assessment models to predict each woman's individual time until Breast Cancer Onset (BCo) using administrative health data from Canada's publicly funded healthcare system. We are applying machine learning Individual Survival Distribution (ISD) models, which identify each subject x with a distribution S (t | x), showing the probability that x's time until BCo is at least t more years, for all t > 0. We can then use these models to estimate each woman's expected time until BCo, as well as her risk score. In preliminary models using 25 features with known/suspected links to breast cancer, random survival forest (RSF) achieved the highest concordance index (CI = 58.9%), while multitask logistic regression (MTLR) achieved a competitive 5-year Brier score (BS = 0.0068) and a low mean absolute error (MAE = 30.4 months). These early results demonstrate the feasibility of leveraging administrative health data for personalized breast cancer risk prediction. Ongoing work will substantially expand the feature sets to improve model discrimination.
利益披露 Disclosure
F. Mushashi, None..
S. Qi, None..
P. Bhatti, None..
A. Roth, None..
R. Greiner, None..
R. A. Murphy, None.