PO.BCS02.06 · 生物信息与计算
一种整合分期、组织学、分级和治疗以预测唾液腺癌死亡率的多变量机器学习模型:一项基于 SEER 2010-2021 人群的研究
A multi-variable machine learning model integrating stage, histology, grade, and treatment to predict mortality in salivary gland cancer: A SEER 2010-2021 population-based study
该海报暂无可下载的资料
AACR 官方页面
作者与单位 Authors & Affiliations
摘要 Abstract
中文摘要
背景:唾液腺癌是罕见、异质性强的肿瘤,临床结局差异极大。目前的预后工具依赖于有限的临床病理学特征,未能纳入肿瘤生物学与治疗模式之间的复杂相互作用。我们旨在利用一个大型美国人群队列,开发并验证一种整合分期、组织学、分级、肿瘤大小和治疗方式以预测全因死亡率的机器学习(ML)模型。
方法:我们在 SEER 数据库(2010-2021)中识别出原发性恶性唾液腺肿瘤患者。变量包括年龄、性别、种族、AJCC 分期、组织学亚型、肿瘤大小、分级、淋巴结状态、手术、放疗和化疗。缺失数据采用迭代多变量插补进行填补。评估的模型包括 logistic 回归、随机森林、梯度提升和 XGBoost。性能采用 70/30 训练-测试拆分以及时间验证(以 2010-2017 年的诊断作为训练集、2018-2021 年作为测试集)进行评估。模型可解释性采用 SHAP 值进行考察。
结果:共有 12,678 例唾液腺癌病例符合纳入标准。XGBoost 表现更优,测试集中 AUC 为 0.87,时间验证中为 0.85,优于 logistic 回归(AUC 0.76)。高风险分类的敏感性、特异性和 PPV 分别为 0.81、0.78 和 0.72。SHAP 分析确定分期、肿瘤大小、组织学亚型、分级和接受手术是死亡率最具影响力的预测变量。该模型揭示了肿瘤分级与淋巴结状态之间的非线性相互作用,而这些是传统回归所无法捕捉的。
结论:我们建立了一种稳健、可解释的机器学习模型,整合了关键的临床、病理和治疗变量,以高准确度预测唾液腺癌的死亡率。该工具为个体化风险分层提供了强有力的框架,并可能有助于关于治疗强度、监测和生存期护理的决策制定。有必要进行进一步的前瞻性验证以支持临床应用。
查看英文原文 English abstract
Background: Salivary gland cancers are rare, heterogeneous tumors with widely variable clinical outcomes. Current prognostic tools rely on limited clinicopathologic features and do not incorporate complex interactions between tumor biology and treatment patterns. We aimed to develop and validate a machine learning (ML) model integrating stage, histology, grade, tumor size, and treatment modalities to predict all-cause mortality using a large U.S. population-based cohort.
Methods: We identified patients with primary malignant salivary gland tumors in the SEER database (2010-2021). Variables included age, sex, race, AJCC stage, histologic subtype, tumor size, grade, nodal status, surgery, radiation, and chemotherapy. Missing data were imputed using iterative multivariate imputation. Models evaluated included logistic regression, random forest, gradient boosting, and XGBoost. Performance was assessed with 70/30 train-test split and temporal validation using diagnoses from 2010-2017 as training and 2018-2021 as testing. Model interpretability was examined using SHAP values.
Results: A total of 12,678 salivary gland cancer cases met inclusion criteria. XGBoost showed higher performance, with an AUC of 0.87 in the test set and 0.85 in temporal validation, outperforming logistic regression (AUC 0.76). Sensitivity, specificity, and PPV for high-risk classification were 0.81, 0.78, and 0.72, respectively. SHAP analysis identified stage, tumor size, histologic subtype, grade, and receipt of surgery as the most influential predictors of mortality. The model revealed nonlinear interactions between tumor grade and nodal status that were not captured by traditional regression.
Conclusion: We came up with a robust, interpretable machine learning model integrating key clinical, pathologic, and treatment variables to predict mortality in salivary gland cancer with high accuracy. This tool provides a strong framework for individualized risk stratification and may help decision-making regarding treatment intensity, surveillance, and survivorship care. Further prospective validation is warranted to support clinical adoption.
利益披露 Disclosure
C. Okoye, None..
C. M. Emeasoba, None..
C. Obialo-Ibeawuchi, None..
O. Ntukidem, None..
O. J. Ogedegbe, None..
U. Amaechi, None.