PO.BCS01.16 · 生物信息与计算

对癌前诊断电子健康记录的全表型关联研究在All of Us研究计划中识别出风险因素和保护因素

Phenome-wide association study of pre-cancer diagnosis electronic health records identifies risk and protective factors in the All of Us Research Program

编号 2721 展板 14 时间 4/20 02:00–05:00 区域 Section 2 主讲 Christian Rich, BA
分会场 Integration of Clinical and Research Data
该海报暂无可下载的资料 AACR 官方页面

作者与单位 Authors & Affiliations

Charles C. D. Rich1, Alyssa B. Bair2, Britton E. Richardson1, Katelyn C. Forbes1, Blaine A. Bates3, Mary F. Davis4, Matthew H. Bailey5

1Biology, Brigham Young University, Provo, UT,2Data Science, Brigham Young University, Provo, UT,3Chemical Engineering, Brigham Young University, Provo, UT,4Microbiology and Molecular Biology, Brigham Young University, Provo, UT,5Simmons Center for Cancer Research, Brigham Young University, Provo, UT

摘要 Abstract

中文摘要
背景:All of Us研究计划是癌症流行病学研究的丰富资源,拥有40多万名参与者的全基因组序列并与电子健康记录(EHR)相关联。大型癌症数据集往往只关注病例而无对照,并忽视诊断前的医疗事件。在此,我们对癌症病例与匹配对照之间诊断前的EHR数据进行了全表型关联研究(PheWAS),揭示了可随后利用All of Us基因组数据进行深入研究的共现表型和互斥表型。 方法:使用SNOMED CT编码,我们在All of Us第8版中识别出跨23种癌症类型的48,000多例癌症病例。为在消除时间性确认偏倚的同时开展PheWAS,我们实施了一种匹配截断策略:对于在X岁诊断的每例癌症病例,我们按出生年份、性别和种族匹配对照个体,并在X岁截断其EHR数据,确保累积诊断的机会均等。我们使用逻辑回归检验了癌症诊断与约2,100种临床表型之间的关联,回归模型对年龄、性别以及包括EHR长度和ICD编码总数在内的EHR指标进行了校正,并采用Bonferroni校正进行多重检验。在本研究的编程部分,使用了Anthropic的Claude进行调试和故障排除。 结果:我们的分析证实了已确立的癌症风险因素,验证了All of Us作为癌症流行病学研究稳健平台的价值。值得注意的是,我们识别出与疼痛相关表型的意外负相关:慢性疼痛(P值=7.7×10⁻⁴⁷,OR=0.67)和一般性疼痛(P值=2.9×10⁻⁴⁴,OR=0.72)。其他负相关还包括睡眠障碍(P值=2.9×10⁻²⁷,OR=0.79)和情绪障碍(P值=5.2×10⁻³²,OR=0.77),提示可能存在值得进一步研究的保护性关系。 结论:这项对All of Us中诊断前EHR数据的全面PheWAS揭示了超越传统风险因素的、复杂的癌症相关表型全景。对潜在保护性表型的识别,结合我们消除时间偏倚的严谨方法,为精准癌症预防中的假设生成型研究奠定了基础。这些发现凸显了多样化生物样本库和严谨方法对于识别表型关系的重要性。使用了Anthropic的Claude对文本进行语法、清晰度和篇幅优化方面的修订。重要的是,所有科学内容、假设、分析和结论均为作者的原创工作。
查看英文原文 English abstract
Background: The All of Us Research Program represents a rich resource for cancer epidemiology research, with over 400,000 participants with whole genome sequences linked to electronic health records (EHRs). Large cancer datasets often focus exclusively on cases without controls and neglect pre-diagnosis healthcare occurrences. Here, we perform a phenome-wide association study (PheWAS) of pre-diagnosis EHR data between cancer cases and matched controls, revealing co-occurring and mutually exclusive phenotypes that can be subsequently investigated using All of Us genomic data. Methods: Using SNOMED CT codes, we identified 48,000+ cancer cases across 23 cancer types in All of Us version 8. To conduct PheWAS while eliminating temporal ascertainment bias, we implemented a matched truncation strategy: for each cancer case diagnosed at age X, we matched control individuals on birth year, sex, and race and truncated their EHR data at age X, ensuring equal opportunity to accrue diagnoses. We tested associations between cancer diagnosis and approximately 2,100 clinical phenotypes using logistic regression adjusted for age, sex, and EHR metrics including length of EHR and total ICD code count, with Bonferroni correction for multiple testing. For coding aspects of this study, Claude by Anthropic was used for debugging and troubleshooting. Results: Our analysis confirmed established cancer risk factors, validating All of Us as a robust platform for cancer epidemiology research. Notably, we identified unexpected inverse associations with pain-related phenotypes: chronic pain ( P -value=7.7×10⁻⁴⁷, OR=0.67) and general pain ( P -value=2.9×10⁻⁴⁴, OR=0.72). Additional inverse associations included sleep disorders ( P -value=2.9×10⁻²⁷, OR=0.79) and mood disorders ( P -value=5.2×10⁻³², OR=0.77), suggesting potentially protective relationships warranting further investigation. Conclusions: This comprehensive PheWAS of pre-diagnosis EHR data in All of Us reveals a complex landscape of cancer-associated phenotypes extending beyond traditional risk factors. The identification of potentially protective phenotypes, combined with our rigorous approach to temporal bias elimination, establishes a foundation for hypothesis-generating research in precision cancer prevention. These findings underscore the importance of diverse biobanks and rigorous methods for identifying phenotypic relationships. Claude by Anthropic was used to revise text for grammar, clarity, and length optimization. Importantly, all scientific content, hypotheses, analyses, and conclusions remain the original work of the authors.
利益披露 Disclosure
C. C. D. Rich, None.. A. B. Bair, None.. B. E. Richardson, None.. K. C. Forbes, None.. B. A. Bates, None.. M. F. Davis, None.. M. H. Bailey, None.

← 返回 AACR 2026 检索