PO.PS01.06 · 人群科学

在一项大型队列研究中基于机器学习的膳食模式降维方法及其对癌症风险预测能力的比较

Comparison of machine learning-based dimensionality reduction methods for dietary patterns and their predictability of cancer risk in a large cohort study

编号 5053 展板 26 时间 4/21 09:00–12:00 区域 Section 35 主讲 Hyobin Lee, BS;MS
分会场 Diet, Alcohol, and Tobacco, and Other Lifestyle Factors
该海报暂无可下载的资料 AACR 官方页面

作者与单位 Authors & Affiliations

Hyobin Lee1, Dongseok Heo2, Sukhong Min1, Sinyoung Cho1, So-Yoon Lee1, Ji-Yeob Choi3, Bongwon Suh4, Daehee Kang1

1Department of Preventive Medicine, Seoul National University College of Medicine, Seoul, Korea, Republic of,2Integrated Major in Innovative Medical Science, Seoul National University Graduate School, Seoul, Korea, Republic of,3Department of Biomedical Sciences, Seoul National University Graduate School, Seoul, Korea, Republic of,4Department of Intelligence and Information, Seoul National University, Seoul, Korea, Republic of

摘要 Abstract

中文摘要
背景与目的:膳食模式分析在营养流行病学中至关重要,然而传统聚类方法可能因无法捕捉潜在膳食结构而受限。本研究比较了三种降维技术——主成分分析(PCA)、统一流形逼近与投影(UMAP)和自编码器(AE)——用于膳食模式的构建,并进一步考察了AE衍生的膳食模式与一项大型前瞻性队列研究中癌症发病的关联。方法:数据来自纳入Health Examinees-Gem(HEXA-G)研究(2004-2013年)的130,472名参与者,他们完成了经过验证的食物频率问卷。在k-means聚类之前分别应用PCA、UMAP和AE。使用轮廓系数评估聚类质量,使用SHAP值评估变量贡献。外部验证通过将HEXA训练的编码器应用于韩国国民健康与营养调查(KNHANES)进行。癌症发病通过与韩国中央癌症登记处的链接确定,截至2018年12月31日。多变量Cox比例风险模型估计了总癌症和特定部位癌症的风险比(HRs)和95%置信区间(CIs),重点关注韩国最常见的七种癌症。结果:在未进行降维的情况下,轮廓系数为0.05;PCA很少超过0.2,UMAP达到约0.4,AE达到>0.35,提供了具有竞争力的聚类质量和最均衡的变量贡献。识别出十种膳食模式:均衡、选择性、米饭、面包、蔬菜、乳制品、肉类、加工肉、面条和高盐。使用KNHANES的外部验证产生了相似的轮廓值(约0.36)并保留了质心位置,证实了可迁移性。在中位随访9.4年期间,发生7,390例癌症病例。总癌症未观察到显著关联;然而,特定部位分析显示,男性的加工肉模式与较高的结直肠癌风险相关(HR = 1.98,95% CI:1.12-3.49),选择性模式与较高的胃癌风险相关(HR = 1.32,95% CI:1.03-1.70),二者均相较于均衡模式。在女性中,面包模式与较低的胃癌风险相关(HR = 0.53,95% CI:0.32-0.89)。结论:在各种降维技术中,AE在聚类质量和变量贡献均衡性之间实现了最佳平衡,支持其在构建膳食模式方面的实用性。这些发现表明,基于机器学习的降维方法,特别是AE,能够强化膳食模式的构建并捕捉与癌症风险的有意义关联。
查看英文原文 English abstract
Background & Aims Dietary pattern analysis is essential in nutritional epidemiology, yet traditional clustering approaches may be limited by their inability to capture latent dietary structures. This study compared three dimensionality reduction techniques-Principal Component Analysis (PCA), Uniform Manifold Approximation and Projection (UMAP), and Autoencoders (AE)-for dietary pattern development, and further examined associations of AE-derived dietary patterns with cancer incidence in a large prospective cohort study. Methods Data were obtained from 130,472 participants enrolled in the Health Examinees-Gem (HEXA-G) study (2004-2013), who completed a validated food frequency questionnaire. PCA, UMAP, and AE were each applied prior to k-means clustering. Cluster quality was assessed using silhouette coefficients, and variable contributions were evaluated using SHAP values. External validation was conducted by applying the HEXA-trained encoder to the Korean National Health and Nutrition Examination Survey (KNHANES). Cancer incidence was ascertained through linkage with the Korea Central Cancer Registry up to December 31, 2018. Multivariable Cox proportional hazards models estimated hazard ratios (HRs) and 95% confidence intervals (CIs) for total and site-specific cancers, focusing on the seven most common cancers in Korea. Results Without dimensionality reduction, the silhouette coefficient was 0.05; PCA rarely exceeded 0.2, UMAP reached ~0.4, and AE achieved >0.35, providing competitive cluster quality with the most balanced variable contributions. Ten dietary patterns were identified: Balanced, Selective, Rice, Bread, Vegetables, Dairy, Meat, Processed meat, Noodles, and Salty. External validation using KNHANES produced similar silhouette values (~0.36) and preserved centroid positions, confirming transferability. Over a median follow-up of 9.4 years, 7,390 cancer cases occurred. No significant associations were observed for total cancer; however, site-specific analyses revealed that the Processed meat pattern in men was associated with higher colorectal cancer risk (HR = 1.98, 95% CI: 1.12-3.49), and the Selective pattern with higher gastric cancer risk (HR = 1.32, 95% CI: 1.03-1.70) compared to the Balanced pattern. In women, the Bread pattern was associated with lower gastric cancer risk (HR = 0.53, 95% CI: 0.32-0.89). Conclusion Among the dimensionality reduction techniques, AE achieved the most favorable balance of cluster quality and variable contribution balance, supporting its utility for developing dietary patterns. These findings demonstrate that machine learning-based dimensionality reduction methods, particularly AEs, can strengthen dietary pattern development and capture meaningful associations with cancer risk.
利益披露 Disclosure
H. Lee, None.. D. Heo, None.. S. Min, None.. S. Cho, None.. S. Lee, None.. J. Choi, None.. B. Suh, None.. D. Kang, None.

← 返回 AACR 2026 检索