PO.CL01.18 · 临床研究
一个可扩展的多模态框架用于跨多种癌症类型的无偏倚风险生物标志物发现
A scalable multimodal framework for unbiased risk biomarker discovery across multiple cancer types
作者与单位 Authors & Affiliations
摘要 Abstract
中文摘要
背景:大多数现有的癌症风险模型建立在单一模态和手工选择的特征之上。对种系遗传学、血浆蛋白质组学和深度临床表型进行系统、无偏倚的整合,有望在多种癌症类型中揭示新的风险生物标志物。
方法:我们开发了一个多模态生物标志物发现引擎,可用于发现风险、诊断、预后、预测和监测生物标志物。目前该框架可处理:种系遗传学和多基因风险评分;高维血浆蛋白质组学(Olink);纵向初级保健记录、住院事件、实验室结果、生活方式问卷以及癌症登记链接数据。关键设计特点包括模块化队列处理、自动化数据预处理,以及机器学习模型(包括:梯度提升和神经网络)。
应用:该平台目前部署于英国生物样本库(UK Biobank,n=502,505名参与者;涵盖22种癌症类型的>46,000例新发癌症),正在进行主动的模型训练和生物标志物发现。该架构与队列无关,可直接应用于新兴的大规模资源,包括Our Future Health和All of Us研究项目。
海报展示:我们将通过英国生物样本库中癌症风险建模的实例来展示该平台的可配置性,包括:(i)单个模态与多模态集成的性能比较,(ii)癌症特异性的模态贡献模式,以及(iii)时间窗口过滤对区分真实预测信号与现患疾病效应的影响。
结论:通过消除特征工程中的偏倚并支持多样化健康数据流的无缝整合,这一可扩展框架为数据驱动地发现多模态癌症风险生物标志物提供了稳健的基础,为下一代精准预防策略铺平了道路。
查看英文原文 English abstract
Background: Most existing cancer risk models are built on single modalities and hand-selected features. Systematic, unbiased integration of germline genetics, plasma proteomics, and deep clinical phenotyping holds promise for revealing novel risk biomarkers across diverse cancer types.
Methods: We developed a multi-modal biomarker discovery engine that can be used for discovering risk, diagnostic, prognostic, predictive and monitoring biomarkers. Currently the framework handles: Germline genetics and polygenic risk scores High-dimensional plasma proteomics (Olink) Longitudinal primary-care records, hospital episodes, laboratory results, lifestyle questionnaires, and cancer registry linkages Key design features include modular cohort handling, automated data preprocessing, and machine-learning models (including: gradient boosting and neural networks).
Application: The platform is currently deployed on the UK Biobank (n = 502,505 participants; >46,000 incident cancers across 22 cancer types) with active model training and biomarker discovery in progress. The architecture is cohort-agnostic and ready for direct application to emerging large-scale resources including Our Future Health and the All of Us Research Program.
Poster presentation: We will demonstrate the platform's configurability through examples of cancer-risk modelling in the UK Biobank, showcasing: (i) comparative performance of individual modalities versus multimodal ensembles, (ii) cancer-specific patterns of modality contribution, and (iii) the effect of time-window filtering on separating true predictive signals from prevalent disease effects.
Conclusions: By eliminating bias in feature engineering and supporting seamless integration of diverse health data streams, this scalable framework provides a robust foundation for data-driven discovery of multimodal cancer risk biomarkers, paving the way for next-generation precision prevention strategies.
利益披露 Disclosure
C. Petrescu,
Chronomics Limited Employment, Stock Option.
L. Schmunk,
Chronomics Limited Employment, Stock Option.
J. Monahan,
Chronomics Ltd. Employment, Stock Option.
A. Salami,
Chronomics Ltd. Employment, Stock Option.
T. M. Stubbs,
Chronomics Ltd Employment, g., Board of Directors, non-salaried role), Stock.